Media processing device, user terminal, and program

The media processing device and user terminal system addresses challenges in handling changing viewpoint information configurations by transmitting it as general-purpose metadata, ensuring consistent display of content with viewpoint freedom.

WO2026155170A1PCT designated stage Publication Date: 2026-07-23NIPPON HOSO KYOKAI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NIPPON HOSO KYOKAI
Filing Date
2026-01-14
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing mechanisms for transmitting content with a degree of freedom of viewpoint, such as 3DoF+ and 6DoF, face challenges in handling changes to viewpoint information configuration, leading to potential modifications in processing functions.

Method used

A media processing device and user terminal system that separates viewpoint information from specific content, transmitting it as general-purpose metadata, allowing appropriate display even with changing configurations.

Benefits of technology

Enables appropriate display of content with viewpoint freedom by handling viewpoint information separately, ensuring consistent functionality despite changes in configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2026000924_23072026_PF_FP_ABST
    Figure JP2026000924_23072026_PF_FP_ABST
Patent Text Reader

Abstract

A media processing device (200) comprises: a reception unit (210) that receives viewpoint information from a user terminal (300); a renderer (220) that generates, on the basis of the viewpoint information, specific content including at least content having a degree of freedom in viewpoint; and a transmission unit (230) that transmits the specific content generated by the renderer (220) to the user terminal (300). The transmission unit (230) transmits the viewpoint information used in the generation of the specific content to the user terminal (300) as general-purpose metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Media Processing Device, User Terminal, and Program ,

[0008] , ,

[0007] , ,

[0006] Cross - Reference to Related Applications

[0001] This application is based on Japanese Application No. 2025 - 006930 filed on January 17, 2025, the contents of which are incorporated herein by reference.

[0002] This disclosure relates to a media processing device, a user terminal, and a program. [[ID=IO]]

[0003] Conventionally, mechanisms for transmitting content such as 360° video and 3D objects have been proposed (for example, Non - Patent Document 1). Such mechanisms include 3DoF+ (Degree of Freedom) with viewpoint movement within the range where the user moves their head while sitting, and 6DoF with viewpoint movement within the range where the user can move freely. In such mechanisms, the positional relationship between the 360° video and the 3D object is shown by scene description.

[0004] 3GPP TR 26.928 V16.1.0 December 2020<000001OE]]

[0005] In the context described above, a case where specific content including content with a degree of freedom of viewpoint is generated by a media processing device and then the generated specific content is transmitted from the media processing device to a user terminal is conceivable.

[0006] As a result of intensive studies, the inventors have found that it is desirable to suppress modification of the function of processing specific content by handling viewpoint information separately from the specific content, assuming situations such as a situation where the configuration of the viewpoint information corresponding to the specific content can be changed.

[0007] Therefore, the present disclosure has been made to solve the above - described problems, and an object thereof is to provide a media processing device, a user terminal, and a program that enable appropriate display of specific content including content with a degree of freedom of viewpoint.

[0008] The disclosed aspect is a media processing device comprising: a receiving unit that receives viewpoint information from a user terminal; a renderer that generates specific content including at least content having degrees of freedom of viewpoint based on the viewpoint information; and a transmitting unit that transmits the specific content generated by the renderer to the user terminal, wherein the transmitting unit transmits the viewpoint information used in generating the specific content to the user terminal as general-purpose metadata.

[0009] The disclosed aspect is a user terminal comprising: a transmitting unit that transmits viewpoint information to a media processing device; and a receiving unit that receives specific content, which includes at least content having degrees of freedom of viewpoint, and which is generated by the media processing device based on the viewpoint information, wherein the receiving unit receives the viewpoint information used in generating the specific content from the media processing device as general-purpose metadata.

[0010] The disclosed aspect is a program that causes a computer to perform the following steps: step A, receiving viewpoint information from a user terminal; step B, generating specific content that includes at least content having degrees of freedom of viewpoint based on the viewpoint information; and step C, transmitting the specific content generated by step B to the user terminal, wherein step C includes transmitting the viewpoint information used in generating the specific content to the user terminal as general-purpose metadata.

[0011] The disclosed aspect is a program that causes a computer to perform the following steps: step A, which transmits viewpoint information to a media processing device; and step B, which receives specific content that includes at least content having degrees of freedom of viewpoint, and which is generated by the media processing device based on the viewpoint information, wherein step B includes receiving the viewpoint information used in generating the specific content from the media processing device as general-purpose metadata.

[0012] Figure 1 is a diagram showing a transmission system 10 according to an embodiment. Figure 2 is a block diagram showing a media processing device 200 and a user terminal 300 according to an embodiment. Figure 3 is a diagram for explaining a second content according to an embodiment. Figure 4 is a diagram showing a method for viewing specific content according to an embodiment. Figure 5 is a diagram for explaining operation example 1. Figure 6 is a diagram for explaining operation example 2. Figure 7 is a diagram for explaining operation example 2. Figure 8 is a diagram for explaining operation example 3. Figure 9 is a diagram for explaining operation example 3. Figure 10 is a diagram for explaining operation example 3. Figure 11 is a diagram for explaining operation example 3. Figure 12 is a diagram for explaining operation example 4. Figure 13 is a diagram for explaining operation example 4. Figure 14 is a diagram for explaining operation example 4. Figure 15 is a diagram for explaining operation example 5. Figure 16 is a diagram for explaining operation example 5. Figure 17 is a diagram for explaining operation example 5. Figure 18 is a diagram for explaining operation example 5. Figure 19 is a diagram for explaining the first method according to modification example 1. Figure 20 is a diagram illustrating the second method related to Modification Example 1. Figure 21 is a diagram illustrating the first method related to Modification Example 2. Figure 22 is a diagram illustrating the second method related to Modification Example 2. Figure 23 is a diagram illustrating viewpoint information related to Modification Example 2. Figure 24 is a diagram illustrating Modification Example 3. Figure 25 is a diagram illustrating Modification Example 4. Figure 26 is a diagram illustrating Modification Example 4. Figure 27 is a diagram illustrating Modification Example 4.

[0013] Next, embodiments of the present disclosure will be described. In the following drawings, identical or similar parts are denoted by the same or similar reference numerals. However, it should be noted that the drawings are schematic, and the proportions of the dimensions, etc., may differ from those of reality.

[0014] Therefore, specific dimensions should be determined by referring to the following explanation. Furthermore, it is important to note that there may be differences in the dimensional relationships and / or ratios between drawings.

[0015] [Summary of Disclosure] The summary of the disclosure is a media processing device comprising: a receiving unit that receives viewpoint information from a user terminal; a renderer that generates specific content including at least content having degrees of freedom of viewpoint based on the viewpoint information; and a transmitting unit that transmits the specific content generated by the renderer to the user terminal, wherein the transmitting unit transmits the viewpoint information used in generating the specific content to the user terminal as general-purpose metadata.

[0016] The disclosure outlines a user terminal comprising: a transmitting unit that transmits viewpoint information to a media processing device; and a receiving unit that receives specific content, which includes at least content having a degree of freedom of viewpoint, and which is generated by the media processing device based on the viewpoint information, wherein the receiving unit receives the viewpoint information used in generating the specific content from the media processing device as general-purpose metadata.

[0017] The disclosure outlines a program corresponding to the media processing device or user terminal described above.

[0018] In the disclosure summary, the viewpoint information used to generate specific content is transmitted from the media processing device to the user terminal as general-purpose metadata. With this configuration, even if a situation arises where the configuration of viewpoint information corresponding to specific content may change, it is possible to suppress modifications to the function that processes the specific content by handling the viewpoint information separately from the specific content. Consequently, since the user terminal can grasp the viewpoint information used to generate the specific content, even if the viewpoint information used when generating the specific content in the media processing device differs from the viewpoint information used when displaying the specific content in the user terminal, the specific content can be displayed appropriately.

[0019] [Embodiment] (Transmission System) A transmission system according to an embodiment will be described below. Figure 1 is a diagram showing a transmission system 10 according to an embodiment. As shown in Figure 1, the digital wireless transmission system comprises a transmitting device 100, a media processing device 200, and a user terminal 300.

[0020] In this embodiment, the transmitting device 100 transmits a first content that does not have freedom of viewpoint and a second content that does have freedom of viewpoint to the media processing device 200. Furthermore, the transmitting device 100 transmits first control information associated with the first content and second control information associated with the second content to the media processing device 200.

[0021] The first content may include at least one of 2D video and audio. The first content and the first control information may be transmitted using the first method. The first method may be a method compliant with ISO / IEC 23008-1 (hereinafter referred to as MMT (MPEG Media Transport)). In the following, an example will be given of the case where the first method is MMTP (MMT Protocol) compliant with MMT. The first control information may be referred to as MMT-SI (Signaling Information).

[0022] The second content may include 360° video and 3D objects. The second content and second control information may be transmitted using the second method, or they may be transmitted using a protocol such as HTTP (HyperText Transfer Protocol). The second content may conform to 3DoF+ (Degree of Freedom), which involves viewpoint movement within the range of head movement by the user while seated, or 6DoF, which involves viewpoint movement within the range of free movement by the user. Because the second content has degrees of freedom of viewpoint, it may include two or more 360° videos or two or more 3D objects at the same time (frame). The second control information may be called a scene description.

[0023] Here, the second control information may be transmitted using the first method described above. That is, the second control information may be transmitted using the same first method as the first control information (for example, MMTP). Alternatively, it may be transmitted using a protocol such as HTTP.

[0024] The transmission from the transmitting device 100 to the media processing device 200 is not particularly limited, but may be transmitted using satellite broadcasting, the Internet network, or a mobile communication network.

[0025] While not particularly limited, the transmission system may be a digital wireless transmission system. The digital wireless transmission system may be a system used for 4K or 8K satellite broadcasting.

[0026] The media processing device 200 generates specific content, which includes at least the second content described above, based on viewpoint information received from the user terminal 300, and transmits the generated specific content to the user terminal 300. Although not particularly limited, the transmission of the specific content may be via the Internet network or via a mobile communication network.

[0027] The user terminal 300 may be a smartphone, tablet, head-mounted display, or other user terminal. As shown in Figure 1, two or more user terminals 300 may be provided. In other words, two or more user terminals 300 may request the media processing device 200 to generate specific content. Each user terminal 300 may transmit different viewpoint information to the media processing device 200.

[0028] (Media Processing Device and User Terminal) The media processing device and user terminal according to the embodiment will be described below. Figure 2 is a block diagram showing the media processing device 200 and user terminal 300 according to the embodiment.

[0029] As shown in Figure 2, the media processing device 200 includes a receiving unit 210, a renderer 220, and an encoding processing unit 230.

[0030] The reception unit 210 receives viewpoint information. In this embodiment, the reception unit 210 constitutes a receiving unit that receives viewpoint information from the user terminal 300. The viewpoint information includes an information element indicating the viewpoint position of the user of the user terminal 300 and an information element indicating the direction of the user's line of sight of the user of the user terminal 300.

[0031] Renderer 220 generates specific content that includes at least the second content, based on viewpoint information. Since the specific content is generated based on viewpoint information, it may include one 360° video or one 3D object at the same time (frame). Below, we will illustrate the case where the specific content includes the first content in addition to the second content.

[0032] Firstly, the renderer 220 generates first content, including 2D video and audio, as part of specific content, based on first control information (MMT-SI). Viewpoint information is not required in the generation of first content.

[0033] Specifically, renderer 220 acquires 2D video, audio, and MMT-SI in the form of MMTP packets, which are packets containing 2D video, audio, and MMT-SI.

[0034] For example, an MMTP packet is stored in an IP (Internet Protocol) packet. The IP packet may be transmitted using UDP (User Datagram Protocol) or TCP (Transmission Control Protocol).

[0035] Here, the first content is processed in units divided into fixed time intervals (hereinafter referred to as MPU; Media Processing Unit). An MPU includes one or more access units. Access units are sometimes treated as MFUs (Media Fragment Units). An MFU related to 2D video may be called a NAL (Network Abstraction Layer) unit, and an MFU related to audio may be called an MHAS (MPEG-H 3D Audio Stream) packet.

[0036] The MMT-SI includes a PA (Package Access) message, which in turn includes an MPT (MMT Package Table) listing the first content. Furthermore, the MMT-SI includes an MPU timestamp descriptor indicating the presentation time of the first content. The MPU timestamp descriptor may represent the presentation time of the MPU, i.e., the time of the access unit that first presents the content in the MPU.

[0037] The MPU timestamp descriptor may be generated using UTC (Coordinated Universal Time) as the base time. The base time may be TAI (International Atomic Time) or the time provided by GPS (Global Positioning System). The base time may be the time provided by an NTP (Network Time Protocol) server or the time provided by a PTP (Precision Time Protocol) server.

[0038] Secondly, the renderer 220 generates second content, including 360° video and 3D objects, as part of specific content, based on the second control information (scene description). Viewpoint information is used in the generation of the second content.

[0039] Specifically, renderer 220 may obtain the scene description in the form of MMTP packets in which the scene description is packetized. The method for obtaining 360° video and 3D objects is not particularly limited.

[0040] 360° video may be converted to 2D video by a projection transformation such as ERP (Equirectangular projection) and / or cubemap. Metadata indicating the type of projection transformation applied to the 360° video may be added. 3D objects may be encoded in mesh format. ISO / IEC 14496-16 “Animation framework extension (AFX)” may be used for mesh format encoding. 3D objects may also be encoded in point cloud format. ISO / IEC 23090-5 “Video-based Point Cloud Compression” may be used for point cloud format encoding.

[0041] Here, the second content is compiled into a single file in units divided by a fixed time interval. This fixed time interval may be 500ms. For example, if the frame rate is 60fps (frames per second), one file will contain 30 frames.

[0042] A scene description is generated for each file and contains information for each frame that identifies the 360° video and 3D objects. For example, a scene description may include an information element indicating the name of the 3D object in the frame (object_name), an information element indicating the frame number (frame_number), an information element indicating the position of the 3D object in the frame (translation_object), an information element indicating the rotation of the 3D object in the frame (rotation_object), and an information element indicating the size of the 3D object in the frame (scale_object).

[0043] Thirdly, the renderer 220 outputs specific content, including the first content and the second content, to the encoding processing unit 230. The renderer 220 may also output the presentation time of the specific content to the encoding processing unit 230 along with the specific content.

[0044] Here, the presentation time of the specific content may be corrected based on the delay time between the media processing device 200 and the user terminal 300. Specifically, the renderer 220 may calculate the presentation time (T') of the specific content provided from the media processing device 200 to the user terminal 300 (T' = T + ΔT) based on the presentation time (T) of the specific content provided from the transmission device 100 to the media processing device 200 and the delay time (ΔT). The delay time (ΔT) may be a value predetermined in the media processing device 200, or may be a different value for each user terminal 300.

[0045] Fourthly, the renderer 220 may constitute a transmission unit that transmits the viewpoint information used for generating the specific content to the user terminal 300. The viewpoint information used for generating the specific content may be transmitted from the encoding processing unit 230 to the user terminal 300.

[0046] For example, the transmission method of the viewpoint information and the specific content may be MMTP or HTTP. When MMTP is used as the transmission method of the specific content, the viewpoint information may be stored as metadata in the OMAF (Omnidirectional Media Format) defined in ISO / IEC 23090-2.

[0047] The encoding processing unit 230 encodes the specific content generated by the renderer 220. In an embodiment, the encoding processing unit 230 may be an example of a transmission unit that transmits the specific content to the user terminal 300.<...>

[0048] Furthermore, the encoding processing unit 230 may encode the presentation time of the specific content. The encoding processing unit 230 may transmit an information element indicating the presentation time to the user terminal 300 together with the specific content.

[0049] Here, as the compression encoding method used by the encoding processing unit 230, any compression encoding method can be used. For example, the compression encoding method may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding).

[0050] As mentioned above, since the second content included in a specific piece of content is generated based on viewpoint information, the video included in that specific piece of content can be treated as a 2D video that does not have freedom of viewpoint.

[0051] For example, the transmission control method used to initiate and / or terminate viewing of specific content may include RTSP (Real Time Streaming Protocol). The transmission method may be MMTP or HTTP. If MMTP is used as the transmission method, the specific content may be stored in OMAF as defined in ISO / IEC 23090-2.

[0052] As shown in Figure 2, the user terminal 300 includes a detection unit 310, a decoding processing unit 320, and a renderer 330.

[0053] The detection unit 310 detects the user's viewpoint position and gaze direction. The detection unit 310 may include an acceleration sensor and a GPS (Global Positioning System) sensor. The detection unit 310 may include a user interface (e.g., touch sensor, keyboard, mouse, controller, etc.) that is manually input by the user. The detection unit 310 may transmit viewpoint information (viewpoint position and gaze direction) to the media processing device 200. The detection unit 310 may output viewpoint information (viewport) to the renderer 330.

[0054] The decoding processing unit 320 decodes specific content received from the media processing unit 200. The decoding processing unit 320 may also decode the presentation time received from the media processing unit 200. The decoding processing unit 320 may output the specific content to the renderer 330, or it may output the presentation time to the renderer 330.

[0055] The renderer 330 outputs specific content decoded by the decoding processing unit 320. The renderer 330 may also output specific content based on the presentation time decoded by the decoding processing unit 320. For example, the renderer 330 may output the video content included in the specific content to a display and the audio content included in the specific content to a speaker.

[0056] Here, the renderer 330 may generate specific content in which the viewpoint position and line of sight direction have been corrected based on the difference between the viewpoint information received from the media processing device 200 and the viewpoint information input from the detection unit 310.

[0057] (Second Content) The second content according to the embodiment will be described below. Here, the second content at t=0, t=1, and t=2 will be described. The time intervals at t=0, t=1, and t=2 are not particularly limited.

[0058] For example, as shown in Figure 3, at t=0, a 360° video may be displayed instead of the 3D object. The 360° video can be considered a background video for the 3D object. At t=1, the 3D object may be displayed superimposed on the 360° video. Furthermore, at t=1, the position and rotation of the 3D object superimposed on the 360° video may be changed.

[0059] The scene description described above includes information elements indicating the position, rotation, and size of the 3D object for each of t=0, t=1, and t=2, allowing the 3D object to be appropriately superimposed onto the 360° image.

[0060] (Viewing Method) The viewing method according to the embodiment will be described below. Here, we will illustrate the viewing of specific content, including the first content and the second content.

[0061] As shown in Figure 4, in step S11, the user terminal 300 sends an RTSP SETUP to the media processing unit. The RTSP SETUP is a message indicating that viewing of a specific content has begun.

[0062] Here, RTSP SETUP includes the IP address of the user terminal 300, the listening port number, and content identification information (content ID). RTSP SETUP may also include capability information for viewing specific content on the user terminal 300. Capability information may include frame rate, display resolution, etc. Display resolution may include field of view (FoV). Capability information may also include information elements indicating the encoding and compression methods supported by the user terminal 300.

[0063] Here, an example is given in which the capability information of the user terminal 300 is directly notified to the media processing device 200, but the embodiment is not limited to this. The capability information of the user terminal 300 may be notified to the transmitting device 100, and then notified from the transmitting device 100 to the media processing device 200.

[0064] In step S12, the media processing device 200 sends a response to the RTSP SETUP. Here, an ACK is sent as a response indicating that the RTSP SETUP has been received.

[0065] In step S21, the user terminal 300 transmits initial viewpoint information to the media processing device 200. The initial viewpoint information may be transmitted in MMT-SI format.

[0066] In step S22, the media processing device 200 generates initial specific content based on the initial viewpoint information (rendering process). For example, the media processing device 200 generates second content to be included in the initial specific content based on the initial viewpoint information and the scene description.

[0067] Here, the media processing device 200 may generate initial specific content using a viewport that is wider than the display resolution of the user terminal 300. For example, the wider range may be a range of display resolution + 20% in the horizontal direction and display resolution + 20% in the vertical direction.

[0068] The media processing device 200 applies a compression encoding scheme to the initial specified content. The compression encoding scheme may be HEVC or VVC, although it is not particularly limited.

[0069] In step S23, the media processing device 200 transmits initial specific content corresponding to the initial viewpoint information to the user terminal 300. The media processing device 200 transmits the presentation time of the initial specific content to the user terminal 300. As described above, the presentation time (T') provided to the user terminal 300 may be determined based on the delay time (ΔT).

[0070] Furthermore, if a different value for the delay time (ΔT) is used for each user terminal 300, the media processing device 200 can identify it by including the transmission time of the RTSP SETUP in the RTSP SETUP as described above.

[0071] The user terminal 300 outputs specific content based on the presentation time (T'). The user terminal 300 may also generate specific content with corrected viewpoint position and gaze direction based on the difference between viewpoint information received from the media processing device 200 and viewpoint information input from the detection unit 310.

[0072] In step S31, the user terminal 300 transmits viewpoint information to the media processing device 200. The viewpoint information may be transmitted in MMT-SI format. The user terminal 300 may transmit viewpoint information at predetermined intervals (e.g., 500ms), or it may transmit viewpoint information in response to a change in at least one of the viewpoint position and line of sight direction.

[0073] In step S32, the media processing device 200 generates specific content based on the viewpoint information received in step S31 (rendering process).

[0074] In step S33, the media processing device 200 transmits specific content corresponding to the viewpoint information received in step S31 to the user terminal 300.

[0075] The processing in steps S31 to S33 is the same as the processing in steps S21 to S23, except that the viewpoint information received in step S31 is used instead of the initial viewpoint information. Therefore, the details of the processing in steps S31 to S33 are omitted. The processing in steps S31 to S33 may be repeated at a predetermined interval, or it may be repeated each time the user's viewpoint position or line of sight changes.

[0076] In step S41, the user terminal 300 sends an RTSP TEARDOWN to the media processing unit. The RTSP TEARDOWN is a message indicating that viewing of a specific content has ended.

[0077] In step S42, the media processing device 200 sends a response to the RTSP TEARDOWN. Here, an ACK is sent as a response indicating that the RTSP TEARDOWN has been received.

[0078] Figure 4 illustrates a case where steps S11 and S12 are executed on an RTSP basis, but the embodiments are not limited to this. Steps S11 and S12 may also be executed on an MMTP basis or on an HTTP basis.

[0079] Similarly, while examples have been given of cases where steps S41 and S42 are performed on an RTSP basis, the embodiments are not limited thereto. Steps S41 and S42 may also be performed on an MMTP basis or on an HTTP basis.

[0080] Figure 4 illustrates a case where steps S31 to S33 are performed on an MMTP basis, but the embodiment is not limited to this. Steps S31 to S33 may also be performed on other methods (e.g., HTTP).

[0081] Similarly, while the examples illustrate the case where steps S41 to S42 are performed on an MMTP basis, the embodiments are not limited thereto. Steps S41 to S42 may also be performed on other methods (e.g., HTTP).

[0082] (Operation Example 1) The above-described embodiment may include Operation Example 1 shown below. In Operation Example 1, the media processing device 200 transmits the viewpoint information used in generating the specific content to the user terminal 300, in association with the sequence number attached to the specific content.

[0083] Specifically, the media processing device 200 (renderer 220), similar to the embodiment described above, generates second content, including 360° video and 3D objects, as part of specific content, based on second control information (scene description). In generating the second content, viewpoint information received from the user terminal 300 is used.

[0084] In Operation Example 1, the renderer 220 associates the viewpoint information used to generate a specific content (in this case, the second content) with a sequence number. The renderer 220 associates the viewpoint information used to generate the specific content with the sequence number attached to the specific content and sends it to the user terminal 300. The viewpoint information may be stored in a VP (View Port) message. The VP message may have the MMT-SI format specified in ISO / IEC 23008-1, ARIB STD-B60, ARIB TR-B39, etc. The VP message may be sent frame by frame. The VP message may also be sent to the user terminal 300 as an MMTP message (MMT-SI).

[0085] While not particularly limited, a VP message may have the data structure shown in Figure 5. As shown in Figure 5, a VP message may include message_id, version, length, fov, viewpoint_pos_x, viewpoint_pos_y, viewpoint_pos_z, viewpoint_yaw, viewpoint_pitch, viewpoint_roll, viewport_width, viewport_height, mpu_sequence_number_flag, mpu_sequence_number, etc.

[0086] The message_id is an identifier that identifies a VP message. The message_id may also be 0x0204.

[0087] The `version` field indicates the version of the MMTP protocol. The `version` field may also be 0x00.

[0088] The `length` parameter indicates the length of the VP message.

[0089] FOV is information that indicates the field of view.

[0090] `viewpoint_pos_x` is information indicating the x-coordinate of the viewpoint position. `viewpoint_pos_x` is an example of viewpoint information used in the generation of specific content.

[0091] `viewpoint_pos_y` is information indicating the y-coordinate of the viewpoint position. `viewpoint_pos_y` is an example of viewpoint information used in the generation of specific content.

[0092] `viewpoint_pos_z` is information indicating the z-coordinate of the viewpoint position. `viewpoint_pos_z` is an example of viewpoint information used in the generation of specific content.

[0093] `viewpoint_yaw` is information indicating the yaw of the viewpoint position. `viewpoint_yaw` is an example of viewpoint information used in the generation of specific content.

[0094] `viewpoint_pitch` is information indicating the pitch of the viewpoint position. `viewpoint_pitch` is an example of viewpoint information used in the generation of specific content.

[0095] `viewpoint_roll` is information indicating the view position roll. `viewpoint_roll` is an example of view information used in the generation of specific content.

[0096] `viewport_width` is information that indicates the width of the display area (specific content).

[0097] viewport_height is information that indicates the height of the display area (specific content).

[0098] The `mpu_sequence_number_flag` indicates whether or not the `mpu_sequence_number` field exists. For example, if `mpu_sequence_number_flag` is 1, the `mpu_sequence_number` field exists, but if `mpu_sequence_number_flag` is 0, the `mpu_sequence_number` field does not necessarily have to exist.

[0099] `mpu_sequence_number` is the MPU sequence number of the video corresponding to the specific content indicated in the VP message. `mpu_sequence_number` is an example of a sequence number assigned to specific content.

[0100] Here, `viewpoint_pos_x`, `viewpoint_pos_y`, and `viewpoint_pos_z` are examples of information elements that indicate the user's viewpoint position in the 3D space constructed by the scene description. `viewpoint_yaw`, `viewpoint_pitch`, and `viewpoint_roll` are examples of information elements that indicate the user's line of sight direction in the 3D space constructed by the scene description. `viewport_width` and `viewport_height` are examples of information elements that indicate the number of pixels in the video contained in specific content.

[0101] Under these conditions, the media processing device 200 and the user terminal 300 may perform the following operations.

[0102] Firstly, the media processing device 200 (renderer 220) may determine viewpoint_pos_x, viewpoint_pos_y, and viewpoint_pos_z based on the viewpoint position used to generate the specific content. The renderer 220 may also determine viewpoint_yaw, viewpoint_pitch, and viewpoint_roll based on the viewing direction used to generate the specific content. The renderer 220 may determine viewport_width based on the number of pixels in the horizontal direction of the specific content, and determine viewport_height based on the number of pixels in the vertical direction of the specific content.

[0103] The media processing unit 200 (encoding processing unit 230) may perform compression encoding of specific content generated by the renderer 220 and transmit the specific content. Here, the encoding processing unit 230 may transmit the VP message shown in Figure 5 to the user terminal 300. That is, the encoding processing unit 230 may transmit the viewpoint information used in generating the specific content to the user terminal 300, in association with the sequence number attached to the specific content.

[0104] Secondly, the user terminal 300 (decoding processing unit 320) may be configured as a receiving unit that receives viewpoint information used in the generation of specific content from the media processing unit 200, in association with the sequence number attached to the specific content. That is, the decoding processing unit 320 may receive the VP message (MMT-SI) shown in Figure 5 from the media processing unit 200.

[0105] The user terminal 300 (renderer 330) may associate the video constituting the decoded specific content with the viewpoint information contained in the VP message based on the sequence number (mpu_sequence_number). Based on the difference between the viewpoint information detected by the detection unit 310 and the viewpoint information contained in the VP message, the renderer 330 may identify the video within the range specified by the viewpoint information detected by the detection unit 310 from the display area defined by the information contained in the VP message (viewport_width and viewport_height). The renderer 330 may display the identified video.

[0106] (Operation Example 2) The above-described embodiment may also include Operation Example 2 shown below. In Operation Example 2, as shown in Figure 6, a case is assumed in which a first user terminal 400 that provides viewpoint information feedback and a second user terminal 500 that does not provide viewpoint information feedback are mixed. The first user terminal 400 may be a terminal such as a head-mounted display. The first user terminal may have the same functions as the user terminal 300 described above. The second user terminal 500 may be a terminal such as a volumetric display.

[0107] Specifically, in Operation Example 2, as shown in Figure 6, the media processing device 200 (renderer 220) may be configured as a receiving unit that receives content configuration and recommended viewport information from the transmitting device 100. The content configuration may be considered to include 2D video, audio, 360° video, and 3D objects. The content configuration may also be considered to include MMT-SI and scene description. Recommended viewport information may be considered as an example of specific viewpoint information. Recommended viewport information may also be information that defines (recommends) the position, direction, and field of view at which to view the video in the three-dimensional space (three-dimensional space constructed by the scene description) configured by the specific content. Recommended viewport information may include at least one information element that indicates the viewpoint position in the three-dimensional space configured by the specific content and an information element that indicates the direction of the line of sight in the three-dimensional space configured by the specific content. Recommended viewport information may be considered as viewpoint information mainly used by the second user terminal 500.

[0108] In the following, unless explicitly stated otherwise, the viewpoint information used in generating specific content may include specific viewpoint information (recommended viewport information) received from the transmitting device 100, and may also include viewpoint information received from the user terminal 300.

[0109] While not particularly limited, the recommended viewport information may be included in the scene description in the manner shown in Figure 7. As shown in Figure 7, the recommended viewport information may include camera_orientation, frame_number, translation, and yfov.

[0110] `camera_orientation` is information indicating the direction from which the image is viewed in the 3D space constructed by the scene description. `camera_orientation` can be considered synonymous with line of sight.

[0111] `frame_number` is information indicating the frame number of the video to which `camera_orientation`, `translation`, and `yfov` are applied. `camera_orientation` can be considered an example of specific viewpoint information.

[0112] Translation is information that indicates the viewing position of the image in the 3D space constructed by the scene description. Translation can be considered synonymous with viewpoint position. Translation can also be considered an example of specific viewpoint information. For example, Figure 7 illustrates the case where the viewpoint position is [0,0,-50] when the frame number is 0, and moves to [0,0,-75] when the frame number is 2505.

[0113] yfov is information that indicates the field of view in the 3D space constructed by the scene description.

[0114] While not strictly limited, camera_orientation and translation may be assigned by the content creator. Alternatively, if the content is captured by a camera, camera_orientation and translation may be automatically assigned by the GPS and sensors installed in the camera.

[0115] Here, `translation` is an example of an information element that indicates the viewpoint position in a 3D space composed of specific content. `camera_orientation` is an example of an information element that indicates the direction of the line of sight in a 3D space composed of specific content.

[0116] Under these conditions, the media processing device 200 and the user terminal 300 may perform the following operations.

[0117] Firstly, the media processing device 200 (renderer 220) may generate specific content based on specific viewpoint information (recommended viewport information). The media processing device 200 (encoding processing device 230) may perform compression encoding of the specific content generated by the renderer 220 and transmit the specific content. Here, the encoding processing device 230 may transmit the specific content generated based on the specific viewpoint information to a first user terminal 400, or it may transmit the specific content generated based on the specific viewpoint information to a second user terminal 500.

[0118] Secondly, the media processing device 200 (reception unit 210) may receive viewpoint information from the first user terminal 400 if the user terminal is the first user terminal 400 that provides viewpoint information as feedback. The media processing device (renderer 220) may generate specific content based on the viewpoint information received from the first user terminal 400. In such a case, the media processing device 200 (reception unit 210) may receive a reset signal from the first user terminal 400. The media processing device (renderer 220) may generate specific content based on specific viewpoint information (recommended viewport information) in response to the reset signal.

[0119] In such cases, the first user terminal 400 may have the same configuration as the user terminal 300. However, the first user terminal 400 (detection unit 310) may have a function to detect a reset signal. The detection unit 310 may detect a user operation that inputs a reset signal. The detection unit 310 may transmit the reset signal to the media processing device 200.

[0120] Thirdly, the media processing device 200 (renderer 220) may generate specific content based on specific viewpoint information (recommended viewport information) when the user terminal is a second user terminal 500 that does not provide viewpoint information feedback.

[0121] In such cases, the second user terminal 500 does not need to have a detection unit 310 for detecting viewpoint information. The second user terminal 500 does not need to have a renderer 330. The second user terminal 500 may have the same configuration as the user terminal 300, except that it does not have a detection unit 310 and a renderer 330.

[0122] (Operation Example 3) The embodiments described above may include Operation Example 3 shown below. In Operation Example 3, the media processing device 200 (for example, the selection unit 260 described later) may be configured as a receiving unit that receives quality information from the transmitting device 100 for each of two or more streams whose quality differs depending on the orientation of the 3D object based on viewpoint information, as quality information for the stream relating to the 3D object contained in the specific content.

[0123] Specifically, in Operation Example 3, as shown in Figure 8, the media processing device 200 has a selection unit 260 in addition to the configuration shown in Figure 2. The selection unit 260 receives the scene description and 3D objects from the transmitting device 100. The selection unit 260 inputs the selected stream (3D object) from two or more streams to the renderer 220. The selection unit 260 may also request the transmitting device 100 to transmit the selected stream. In the case where the media processing device 200 transmits specific content to multiple user terminals 300, the selection unit 260 may request the transmitting device 100 to transmit the streams required by each of the multiple user terminals 300, or it may request the transmitting device 100 to transmit all streams.

[0124] Here, the selection unit 260 receives quality information for each of two or more streams, each with a different quality depending on the orientation of the 3D object based on viewpoint information, from the transmission device 100 as quality information for the 3D object stream.

[0125] Quality information may also be information indicating the relative quality of each face that constitutes the bounding box for a 3D object. The bounding box may be represented by a three-dimensional rectangle onto which the 3D object is projected. For example, as shown in Figure 9, the bounding box may be defined by vertices A to H. In such a case, each face of the bounding box includes face #1 represented by vertices A, B, F, E; face #2 represented by vertices B, C, G, F; face #3 represented by vertices A, B, C, D; face #4 represented by vertices E, F, G, H; face #5 represented by vertices A, D, H, E; and face #6 represented by vertices D, C, G, H.

[0126] In such cases, assuming we view a 3D object in a three-dimensional space constructed by the scene description, it is assumed that three faces will be primarily observed. In other words, the remaining three faces are assumed to be observed less.

[0127] In example 3, two or more streams are prepared for 3D objects included in specific content, each with a different quality depending on the orientation of the 3D object based on viewpoint information.

[0128] While not particularly limited, quality information may be included in the scene description in the manner shown in Figure 10. In Figure 10, six streams are exemplified as streams with different orientations of 3D objects. Quality information may be represented in the format "quality" [#1,#2,#3,#4,#5,#6]. Within the brackets [ ], #1 to #6 represent the quality index of faces #1 to #6. The quality index may take values ​​in the range of 1 to 9. A higher quality index value may indicate higher quality. For example, in the stream identified by "id"="1", the quality of #1, #2, and #3 ("8") is high, and the quality of #4, #5, and #6 ("3") is low. In the stream identified by "id"="2", the quality of #1, #2, and #3 ("3") is low, and the quality of #4, #5, and #6 ("8") is high.

[0129] Under these conditions, the media processing device 200 may perform the operations shown below. The following description will mainly focus on the selection of a selected stream (3D object) from among two or more streams.

[0130] In Mode 1, as shown in the upper part of Figure 11, the media processing device 200 (selection unit 260) may first identify the vertex closest to the user's viewpoint (for example, vertex B), and then identify three faces with the closest vertex (for example, face #1 represented by vertices A, B, F, E; face #2 represented by vertices B, C, G, F; and face #3 represented by vertices A, B, C, D). The selection unit 260 may also select the stream that has the maximum sum of the quality indices of the three identified faces (in the example shown in Figure 10, the stream identified by "id" = "1").

[0131] In Mode 1, since the vertex closest to the user's viewpoint is identified based on viewpoint information, it can be considered that the selection unit 260 selects the stream to be transmitted to the user terminal 300 from among two or more streams based on viewpoint information and quality information.

[0132] Mode 2 may be the mode applied when a 3D object is to be displayed in a reduced or enlarged state. For example, as shown in the middle of Figure 11, the media processing device 200 (selection unit 260) may identify the vertex closest to the user's viewpoint (e.g., vertex B) and then identify three faces with the closest vertex (e.g., face #1 represented by vertices A, B, F, E; face #2 represented by vertices B, C, G, F; and face #3 represented by vertices A, B, C, D). When a 3D object is to be displayed in a reduced state, pixels of the 3D object are downsampled. Therefore, the selection unit 260 may select the stream that minimizes the sum of the quality indices of the three identified faces (in the example shown in Figure 10, the stream identified by "id"="2"). On the other hand, when a 3D object is to be displayed in an enlarged state, pixels of the 3D object are interpolated. Therefore, the selection unit 260 may select the stream that maximizes the sum of the quality indices of the three identified faces (in the example shown in Figure 10, the stream identified by "id"="1").

[0133] In Mode 2, since the vertex closest to the user's viewpoint is identified based on viewpoint information, it can be considered that the selection unit 260 selects the stream to be transmitted to the user terminal 300 from among two or more streams based on viewpoint information and quality information.

[0134] Mode 3 may be applied in cases where two 3D objects (3D object #1 and 3D object #2) overlap in the user's line of sight. Here, the selection of a stream related to 3D object #1 is described. For example, as shown in the lower part of Figure 11, the media processing device 200 (selection unit 260) may identify the vertex closest to the user's viewpoint (e.g., vertex B) and then identify three faces with the closest vertices (e.g., face #1 represented by vertices A, B, F, E; face #2 represented by vertices B, C, G, F; and face #3 represented by vertices A, B, C, D). Here, 3D object #2 overlaps with the line segment connecting the vertex closest to the user's viewpoint (e.g., vertex B) and the user's viewpoint, and the three identified faces are obscured by 3D object #2. Therefore, the selection unit 260 may select the stream that minimizes the sum of the quality indices of the three identified faces (in the example shown in Figure 10, the stream identified by "id" = "2").

[0135] In mode 3, since the vertex closest to the user's viewpoint is identified based on viewpoint information, the selection unit 260 can be thought of as selecting a stream to send to the user terminal 300 from among two or more streams based on viewpoint information and quality information. Furthermore, in mode 3, since the overlap of two 3D objects is identified based on viewpoint information and 3D object placement information, the selection unit 260 can be thought of as selecting a stream to send to the user terminal 300 from among two or more streams based on viewpoint information, quality information, and placement information. The 3D object placement information (for example, "rotation_object", "scale_object", and "translation_object" shown in Figure 10) may be included in the scene description. The 3D object placement information may be thought of as "link_area" shown in Figure 10. That is, the selection unit 260 may receive the 3D object placement information in three-dimensional space from the transmission device 100.

[0136] In the quality information shown in Figure 10, the sum of the quality indices for the six faces is equal in each stream. However, the embodiments are not limited to this. The sum of the quality indices for the six faces may differ between two or more streams.

[0137] (Operation Example 4) The embodiments described above may also include Operation Example 4 shown below. Here, Operation Example 4 includes the operations shown below in addition to Operation Example 3. In Operation Example 4, the media processing device 200 (for example, the selection unit 270 described later) may be configured as a receiving unit that receives importance information for each of two or more objects included in a specific content from the transmitting device 100.

[0138] Specifically, in Operation Example 4, as shown in Figure 12, the media processing device 200 has a selection unit 270 in addition to the configuration shown in Figure 2. The selection unit 270 receives the scene description, 3D objects, and 360° video from the transmission device 100. The selection unit 270 inputs the selected stream (3D object) from two or more streams to the renderer 220.

[0139] Here, the selection unit 270 receives importance information for each of two or more objects from the transmitting device 100. The objects may include 3D objects and 360° images.

[0140] For example, importance information may indicate the relative importance between two or more objects. For instance, consider the case shown in Figure 13, where object A (background), object B (person), and object C (dog) exist in a three-dimensional space constructed by a scene description. Object A (background) is an example of a 360° video, and object B (person) and object C (dog) are examples of 3D objects. In such a case, importance information may indicate the relative importance between each of object A (background), object B (person), and object C (dog).

[0141] While not particularly limited, importance information may be included in the scene description in the manner shown in Figure 14. In Figure 14, importance information may be represented by weight. Weight may take values ​​in the range of 1 to 9. A higher weight value may indicate higher importance. In Figure 14, an example is shown where object A (background) identified by "object_id"="0" has the highest weight ("9"), object B (person) identified by "object_id"="1" has the lowest weight ("3"), and object C (dog) identified by "object_id"="2" has a weight ("8") that is higher than object B (person) but lower than object A (background).

[0142] Under these conditions, the media processing device 200 may perform the operations shown below. The following description will mainly focus on the selection of a selected stream (3D object) from among two or more streams.

[0143] Firstly, the media processing unit 200 (selection unit 270) selects the stream with the highest quality for the 3D object with the highest importance. The method of selecting the stream may be the same as in operation example 3. For example, since the importance of object C (dog) is greater than that of object B (person), the selection unit 270 selects the stream for object C (dog) that maximizes the sum of the quality indices of the three faces having the vertices closest to the user's viewpoint.

[0144] Secondly, the media processing unit 200 (selection unit 270) selects the stream with the lowest quality for all 3D objects except the one with the highest importance. Subsequently, the selection unit 270 replaces the stream with the lowest quality with a stream with a higher quality for each 3D object, starting with the most important ones, within the range where specific conditions are met. The specific conditions may include a first condition that the bandwidth of the line from the transmitter 100 to the media processing unit 200 is below a threshold, and a second condition that the processing load of the media processing unit 200 is below a threshold. The specific conditions may also be defined by a combination of the first and second conditions. For example, since the importance of object B (person) is less than that of object C (dog), the selection unit 270 selects a stream with a higher quality for object B (person) within the range where the specific conditions are met.

[0145] As described above, the media processing device 200 (selection unit 270) may be thought to select a stream to be transmitted to the user terminal 300 from among two or more streams based on viewpoint information, quality information, and importance information. The media processing device 200 (selection unit 270) may be thought to select a stream to be transmitted to the user terminal 300 from among two or more streams based on viewpoint information, quality information, placement information, and importance information.

[0146] In Operation Example 4, we illustrated a case where there is one stream of 360° video. However, the embodiments are not limited to this. There may be two or more streams of 360° video with different qualities.

[0147] In Operation Example 4, we illustrated a case where, for a 3D object, there are two or more streams with different qualities depending on the orientation of the 3D object based on viewpoint information. However, the embodiments are not limited to this. For a 3D object, there may be two or more streams with different qualities regardless of the orientation of the 3D object.

[0148] (Operation Example 5) The embodiments described above may include Operation Example 5 shown below. In Operation Example 5, the media processing device 200 (for example, the renderer 220) may be configured as a receiving unit that receives information elements from the transmitting device 100 that define the range of movement of the user's viewpoint position in a three-dimensional space composed of specific content (a three-dimensional space constructed by a scene description).

[0149] Firstly, in Operation Example 5, the information element may include an information element (hereinafter referred to as the first information element) that restricts the movement of the user's viewpoint position inside a 3D object contained in specific content. For example, as shown in Figure 15, in the case where a 3D object is placed in a three-dimensional space constructed by a scene description, the movement of the viewpoint position inside the 3D object may be restricted. However, in cases such as when the 3D object is a building or when another scene exists inside the 3D object, the movement of the viewpoint position inside the 3D object may be permitted.

[0150] Secondly, in Operation Example 5, the information element may include an information element (hereinafter referred to as the second information element) that restricts the movement of the user's viewpoint outside the three-dimensional space. For example, as shown in Figure 16, the three-dimensional space may be defined by a combination of a cuboid and a spheroid. The number of cuboids defining the three-dimensional space may be two or more, and the number of spheroids defining the three-dimensional space may be two or more. However, there may be cases in which the movement of the user's viewpoint outside the three-dimensional space is permitted.

[0151] While not particularly limited, the first information element may be included in the scene description in the manner shown in Figure 17. In Figure 17, the first information element may be represented by viewing_inside_object_flag. viewing_inside_object_flag may be set for each 3D object. For example, if viewing_inside_object_flag is "0", movement of the viewpoint position within the 3D object may be restricted, and if viewing_inside_object_flag is "1", movement of the viewpoint position within the 3D object may be permitted.

[0152] While not particularly limited, the second information element may be included in the scene description in the manner shown in Figure 18. In Figure 18, the second information element may include information elements that define a cuboid that defines a three-dimensional space (cuboid_center_x, cuboid_center_y, cuboid_center_z, cuboid_size_x, cuboid_size_y, cuboid_size_z). cuboid_center_x, cuboid_center_y, and cuboid_center_z are information elements that indicate the center position of the cuboid, and cuboid_size_x, cuboid_size_y, and cuboid_size_z are information elements that indicate the size of the cuboid. The second information element may also include information elements that define a spheroid that defines a three-dimensional space (spheroid_center_x, spheroid_center_y, spheroid_center_z, spheroid_size_x, spheroid_size_y, spheroid_size_z). spheroid_center_x, spheroid_center_y, and spheroid_center_z are information elements indicating the center position of the ellipsoid, and spheroid_size_x, spheroid_size_y, and spheroid_size_z are information elements indicating the size of the ellipsoid. Note that cuboid_enable is an information element indicating whether or not a cuboid defines a 3D space, and spheroid_enable may be an information element indicating whether or not a 3D space defines a ellipsoid. Figure 18 illustrates a case in which a 3D space is defined by two cuboids and two ellipsoids.

[0153] Under these conditions, the media processing device 200 (renderer 220) may perform the following operations.

[0154] Firstly, when the user's viewpoint moves outside the movement range, the renderer 220 may generate specific content using the intersection of the user's viewpoint trajectory and the boundary of the movement range as the viewpoint position. In other words, the renderer 220 may fix the viewpoint position at the position (boundary position) at the time the viewpoint position attempted to move outside the movement range.

[0155] Secondly, the renderer 220 may notify the user that the movement of the viewpoint is restricted when the user's viewpoint moves outside the range of movement. For example, the renderer 220 may display a message such as "You cannot move beyond this point."

[0156] (Operation and Effects) In this embodiment, the media processing device 200 generates specific content based on viewpoint information and then transmits the specific content to the user terminal 300. With this configuration, there is no need for the user terminal 300 to generate specific content including second content with freedom of viewpoint, and the user terminal 300 can present the specific content by providing viewpoint information to the media processing device 200. Therefore, although a delay occurs between the media processing device 200 and the user terminal 300, the processing load on the user terminal 300 can be reduced.

[0157] In Operation Example 1, the media processing device 200 transmits viewpoint information used to generate the specific content to the user terminal 300, associated with the sequence number attached to the specific content. With this configuration, the user terminal 300 can understand the viewpoint information and sequence number used to generate the specific content. Therefore, even if the viewpoint information used by the media processing device 200 to generate the specific content differs from the viewpoint information used by the user terminal 300 to display the specific content, the specific content can be displayed appropriately.

[0158] In Operation Example 2, the media processing device 200 receives content configuration and specific viewpoint information (recommended viewport information) from the transmitting device 100. With this configuration, the media processing device 200 can generate specific content based on the specific viewpoint information, and can display the specific content appropriately even in cases where there is a mix of first user terminals 400 that provide viewpoint information feedback and second user terminals 500 that do not.

[0159] In Operation Example 2, the media processing device 200 generates specific content based on specific viewpoint information in response to a reset signal, even when the user terminal is the first user terminal 400. With this configuration, even in cases where the viewpoint position and line of sight direction become unknown to the first user terminal 400 in the three-dimensional space constructed by the scene description (cases where the user gets lost in the three-dimensional space), the reset signal allows the system to return to specific content based on specific viewpoint information.

[0160] In Operation Example 3, the media processing device 200 receives quality information from the transmitting device 100 for each of two or more streams, each with different quality depending on the orientation of the 3D object based on viewpoint information, as quality information for the streams relating to the 3D object contained in the specific content. With this configuration, based on the new insight that the quality of each face constituting the bounding box of the 3D object does not need to be uniform, it is possible to display the 3D object appropriately while suppressing transmission traffic.

[0161] In Operation Example 4, the media processing device 200 receives importance information for each of two or more objects included in a specific content from the transmitting device 100. With this configuration, by introducing a mechanism to set the importance of each object, such as 360° video and 3D objects, it is possible to appropriately display each object included in a specific content while suppressing transmission traffic.

[0162] In Operation Example 5, the media processing device 200 receives information elements from the transmitting device 100 that define the range of movement of the user's viewpoint position in a three-dimensional space composed of specific content. With this configuration, specific content, including content with freedom of viewpoint, can be displayed appropriately without causing any distortion of the specific content displayed on the user terminal 300.

[0163] [Modification Example 1] Below, we will describe Modification Example 1 of the embodiment. Below, we will mainly describe the differences from the embodiment.

[0164] Change Example 1 describes how to synchronize the first and second content when a specific piece of content includes both the first and second content.

[0165] In the following, synchronization means that the presentation times of the first content (e.g., MPU) and the second content (file) are properly aligned. Therefore, synchronization may include the presentation times of 2D video and 3D objects being aligned, or the presentation times of audio and 3D objects being aligned. Similarly, synchronization may include the presentation times of 2D video and 360° video being aligned, or the presentation times of audio and 360° video being aligned.

[0166] The first method describes a case in which the media processing device 200 synchronizes the first content and the second content based on the first control information (MMT-SI). The media processing device 200 uses the MMT-SI as an entry point to check for the presence or absence of a scene description (second content), and if a scene description exists, it uses the MPU timestamp descriptor to determine the presentation time of a specific content including the first content and the second content.

[0167] Specifically, as shown in Figure 19, the 2D video and audio are presented based on the MPU timestamp descriptor (simply called "timestamp" in Figure 19), thus enabling synchronization of the 2D video and audio.

[0168] On the other hand, the presentation time of the first frame included in the scene description is determined by referring to the MPU timestamp descriptor included in MMT-SI. The presentation times of the second and subsequent frames included in the scene description can be determined by the frame number included in the scene description and the frame rate of the second content. For example, in the case where the frame rate is 30fps, the presentation time of the nth frame is determined by adding 1 / 30 × n to the time determined by the MPU timestamp descriptor. However, the frame number of the first frame included in the scene description is "0".

[0169] In the first method, we exemplified a case where the scene description does not include the presentation time of the first frame included in the scene description, but the scene description may include the presentation time of the first frame included in the scene description.

[0170] The second method describes a case in which the media processing device 200 synchronizes the first content and the second content based on the second control information (scene description). The media processing device 200 uses the scene description as an entry point to check for the presence or absence of MMT-SI (first content), and if MMT-SI is present, it identifies the presentation time of the specific content, including the first content and the second content, based on the presentation time included in the scene description.

[0171] In such cases, the scene description includes absolute time information indicating the presentation time of the second content. The absolute time information may also be the presentation time of the first frame included in the scene description.

[0172] For example, absolute time information may be generated using UTC as the reference time. The reference time may be TAI, or the time provided by GPS. The reference time may be the time provided by an NTP server, or the time provided by a PTP server. Furthermore, absolute time information may be generated based on the same reference time as the MPU timestamp descriptor.

[0173] Furthermore, the scene description includes reference information to identify the first content. This reference information may also be information to identify the MPUs that make up the first content. In other words, the reference information is information for treating the first content (MPU) as an object included in the scene description.

[0174] Specifically, as shown in Figure 20, the presentation time of the first frame included in the scene description is determined by the absolute time information included in the scene description. The presentation times of the second and subsequent frames included in the scene description can be determined by the frame number included in the scene description and the frame rate of the second content. For example, considering the case where the frame rate is 30fps, the presentation time of the nth frame is determined by adding 1 / 30 × n to the time determined by the MPU timestamp descriptor. However, the frame number of the first frame included in the scene description is "0".

[0175] On the other hand, since the 2D video and audio are presented based on the MPU timestamp descriptor (simply "timestamp" in Figure 20), the 2D video and audio can be synchronized. Here, since the above-mentioned reference information is included in the scene description, the media processing device 200 can check whether or not there is a first content to be presented together with the second content based on the reference information included in the scene description.

[0176] In the second method, synchronization between 2D video and audio is achieved based on the MPU timestamp descriptor included in MMT-SI. However, in Modification Example 1, synchronization between 2D video and audio may also be achieved based on information elements (absolute time information and reference information) included in the scene description. In such cases, at least the MPU timestamp descriptor included in MMT-SI may be omitted. Furthermore, MMT-SI itself may be omitted.

[0177] Furthermore, if the reference time of the MPU timestamp descriptor included in MMT-SI (hereinafter referred to as the first reference time) and the reference time of the absolute time information included in the scene description (the second reference time) are different, at least one of the first control information (MMT-SI) and the second control information (scene description) may include conversion information between the first reference time and the second reference time. For example, MMT-SI may include an MPU timestamp descriptor expressed in the second reference time (e.g., a reference time other than UTC) in addition to an MPU timestamp descriptor expressed in the first reference time (e.g., UTC). The scene description may include absolute time information expressed in the first reference time (e.g., UTC) in addition to absolute time information expressed in the second reference time (e.g., a reference time other than UTC).

[0178] The MPU timestamp descriptor included in MMT-SI may also be referred to as the first absolute time information, and the absolute time information included in the scene description may also be referred to as the second absolute time information.

[0179] [Modification Example 2] Below, we will describe Modification Example 2 of the embodiment. Below, we will mainly describe the differences from the embodiment.

[0180] Specifically, in the embodiment described above (operation example 1), the media processing device 200 transmits the viewpoint information used to generate the specific content to the user terminal 300, associated with the sequence number attached to the specific content. In contrast, in modified example 2, the media processing device 200 transmits the viewpoint information used to generate the specific content to the user terminal 300 as an integral part of the specific content.

[0181] Here, the manner in which viewpoint information is integrated with specific content is sufficient if the user terminal 300 can extract viewpoint information used in the generation of the specific content in association with the specific content. The following methods can be considered to realize such a manner.

[0182] In the first method, as shown in Figure 21, the media processing device 200 (encoding processing unit 230) may link the viewpoint information used in generating the specific content to the specific content and transmit it to the user terminal 300.

[0183] For example, consider the case where the file format used for transmitting specific content is ISO Base Media File Format (ISOBMFF). In ISOBMFF, the video data (stream) that constitutes the specific content is contained in the movie data box (mdat). Metadata related to the video data (stream) is contained in the movie fragment box (moof). In ISOBMFF, viewpoint information used in generating the specific content may also be contained in the movie fragment box (moof).

[0184] Under the above-described premise, the media processing device 200 may link the movie fragment box (moof) with the movie data box (mdat) and transmit it to the user terminal 300. The linking may include a method of bundling the movie fragment box (moof) and the movie data box (mdat) via HTTP, or a method of including the movie fragment box (moof) and the movie data box (mdat) in the MPU of the MMT.

[0185] In the first method, the movie fragment box (moof) and movie data box (mdat) may be data structures defined in ISO / IEC 14496-12 “Information technology - Coding of audio-visual objects - Part 12: ISO base media file format”.

[0186] In the first method, transmission may be performed in units called segments, which link together movie fragment boxes (moof) and movie data boxes (mdat). The segments may be those defined in ISO / IEC 23009-1 “Information technology - Dynamic adaptive streaming over HTTP (DASH) Part 1: Media presentation description and segment formats”.

[0187] While not particularly limited, viewpoint information may be added on a Group of Picture (GOP) basis. For example, a GOP may consist of 30 frames, and one viewpoint information may be added to each of the 30 frames. In such cases, the viewpoint information may be included in a movie fragment box (moof) linked to the first frame among the frames that make up the GOP.

[0188] While not particularly limited, viewpoint information may be added on a frame-by-frame basis to constitute the GOP. For example, a GOP may consist of 30 frames, and one piece of viewpoint information may be added to each of the 30 target frames. The target frames may be all 30 frames or only some of the 30 frames. In such cases, the viewpoint information may be included in a movie fragment box (moof) linked to the target frame among the frames that constitute the GOP.

[0189] According to the first method, when the transmission method is MMTP or HTTP, the viewpoint information used to generate the specific content can be linked to the video data (stream) that constitutes the specific content.

[0190] In the second method, as shown in Figure 22, the media processing device 200 (encoding processing unit 230) may multiplex the viewpoint information used in generating the specific content onto the specific content and transmit it to the user terminal 300.

[0191] For example, consider a case where Supplemental Enhancement Information (SEI) or Video Usability Information (VUI) is assumed to be specific information multiplexed onto the video data (stream) that constitutes specific content. SEI and VUI may also be information defined in ISO / IEC 23002-7 “Information technology - MPEG video technologies Part 7: Versatile supplemental enhancement information messages for coded video bitstreams”. The viewpoint information used in generating the specific content may be included in the SEI or VUI.

[0192] According to the second method, when the compression encoding scheme is HEVC or VVC, viewpoint information used to generate specific content can be multiplexed onto the video data (stream) that constitutes the specific content.

[0193] While not particularly limited, SEI or VUI may be multiplexed on a frame-by-frame basis that constitutes the GOP. Whether or not viewpoint information is included in SEI or VUI may be identifiable in the NAL unit header. If viewpoint information is not included in the SEI or VUI multiplexed in a predetermined frame, viewpoint information may be applied to the predetermined frame in an SEI or VUI multiplexed in a frame prior to the predetermined frame. The prior frame may be the frame immediately preceding the predetermined frame, or it may be two or more frames prior to the predetermined frame.

[0194] While the second method primarily described the case where viewpoint information is multiplexed onto video data constituting specific content, viewpoint information may also be multiplexed onto audio data constituting specific content. Note that "audio" may be read as "sound." The audio data may be in a format compatible with MHAS (MPEG-H 3D Audio Stream). MHAS may be an audio compression method defined in ISO / IEC 23008-3 “Information technology—High efficiency coding and media delivery in heterogeneous environments Part 3: 3D audio.”

[0195] In modification example 2, the viewpoint information may have a similar structure to the viewpoint information included in the VP message described above (see Figure 5). For example, as shown in Figure 23, the viewpoint information may include fov, viewpoint_pos_x, viewpoint_pos_y, viewpoint_pos_z, viewpoint_yaw, viewpoint_pitch, viewpoint_roll, viewport_width, viewport_height, etc.

[0196] FOV is information that indicates the field of view.

[0197] `viewpoint_pos_x` is information indicating the x-coordinate of the viewpoint position. `viewpoint_pos_x` is an example of viewpoint information used in the generation of specific content.

[0198] `viewpoint_pos_y` is information indicating the y-coordinate of the viewpoint position. `viewpoint_pos_y` is an example of viewpoint information used in the generation of specific content.

[0199] `viewpoint_pos_z` is information indicating the z-coordinate of the viewpoint position. `viewpoint_pos_z` is an example of viewpoint information used in the generation of specific content.

[0200] `viewpoint_yaw` is information indicating the yaw of the viewpoint position. `viewpoint_yaw` is an example of viewpoint information used in the generation of specific content.

[0201] `viewpoint_pitch` is information indicating the pitch of the viewpoint position. `viewpoint_pitch` is an example of viewpoint information used in the generation of specific content.

[0202] `viewpoint_roll` is information indicating the view position roll. `viewpoint_roll` is an example of view information used in the generation of specific content.

[0203] `viewport_width` is information that indicates the width of the display area (specific content).

[0204] viewport_height is information that indicates the height of the display area (specific content).

[0205] (Operation and Effects) In Modification Example 2, the media processing device 200 transmits the viewpoint information used to generate the specific content to the user terminal 300 together with the specific content. With this configuration, even without using sequence numbers as in the above embodiment (Operation Example 1), the user terminal 300 can extract the viewpoint information used to generate the specific content in association with the specific content. Therefore, even if the viewpoint information used when generating the specific content in the media processing device 200 is different from the viewpoint information used when displaying the specific content in the user terminal 300, the specific content can be displayed appropriately.

[0206] [Modification Example 3] Below, Modification Example 3 of the embodiment will be described. Below, the differences from the embodiment will be mainly described. In Modification Example 3, the viewpoint information transmitted from the media processing device 200 to the user terminal 300 will be mainly described, so other detailed configurations, such as the viewpoint information transmitted from the user terminal 300 to the media processing device 200, will be omitted.

[0207] Specifically, in modification example 3, the media processing device 200 uses an API (Application Programming Interface) to multiplex the viewpoint information used to generate specific content. The user terminal 300 uses an API to separate the viewpoint information used to generate specific content. The API used by the media processing device 200 may be a common API for two or more media processing devices 200. The API used by the user terminal 300 may also be a common API for two or more user terminals 300. Viewpoint information is an example of metadata controlled by an API.

[0208] Firstly, as shown in Figure 24, the media processing device 200 includes a renderer 220 and an encoding processing unit 230, similar to the embodiment described above.

[0209] The renderer 220 outputs an API to the encoding processing unit 230, with the viewpoint information used in generating the specific content set as an argument. The API may also be called CtrlEncMux. The API with the viewpoint information set as an argument may also be called CtrlEncMux(viewpoint information).

[0210] Here, the renderer 220 outputs CtrlEncMux(viewpoint information) in association with specific content generated based on the viewpoint information. The association between the specific content and CtrlEncMux(viewpoint information) may be the same as in the embodiment described above. For example, the following options are possible.

[0211] In option 3-1, the renderer 220 may output the presentation time of the specific content along with the specific content to the encoding processing unit 230. The renderer 220 may also output CtrlEncMux (viewpoint information) to the encoding processing unit 230 in association with the presentation time of the specific content.

[0212] In option 3-2, the renderer 220 may output the sequence number of the specific content along with the specific content to the encoding processing unit 230. The renderer 220 may also output CtrlEncMux (viewpoint information) to the encoding processing unit 230 in association with the sequence number of the specific content.

[0213] The encoding processing unit 230 transmits viewpoint information used in generating specific content to the user terminal 300 using an API between the renderer 220 and the encoding processing unit 230. Specifically, the encoding processing unit 230 encodes the viewpoint information passed as an argument to CtrlEncMux. The viewpoint information passed as an argument to CtrlEncMux may be compressed and encoded using the same method as in the embodiment described above. The encoding processing unit 230 outputs (transmits) the viewpoint information passed as an argument to CtrlEncMux, associating it with the specific content generated based on the viewpoint information.

[0214] The encoding processing unit 230 may associate the viewpoint information passed as an argument to CtrlEncMux with specific content based on the presentation time, or it may associate the viewpoint information passed as an argument to CtrlEncMux with specific content based on the sequence number. The encoding processing unit 230 may also transmit the viewpoint information passed as an argument to CtrlEncMux together with the specific content, similar to the example in modification 2.

[0215] Secondly, as shown in Figure 24, the user terminal 300 has a decoding processing unit 320 and a renderer 330, similar to the embodiment described above.

[0216] The decoding processing unit 320 receives viewpoint information used in the generation of specific content. The decoding processing unit 320 decodes the viewpoint information. The viewpoint information may be decompressed and decoded in a manner similar to the compressed encoding described in the embodiment above.

[0217] The renderer 330 obtains viewpoint information used in the generation of specific content using an API between the decryption processing unit 320 and the renderer 330. Specifically, the renderer 330 outputs the API to the decryption processing unit 320. The API may also be called CtrlDecDeMux. The API requesting viewpoint information may have no arguments and may also be called CtrlDecDeMux. The renderer 330 obtains viewpoint information from the decryption processing unit 320 by outputting CtrlDecDeMux to the decryption processing unit 320.

[0218] As explained above, in Modification Example 3, the media processing unit 200 multiplexes viewpoint information into specific content (encoded signal) via an API between the renderer 220 and the encoding processing unit 230, and the user terminal 300 separates viewpoint information from the specific content (encoded signal) via an API between the decoding processing unit 320 and the renderer 330. The specific content (encoded signal) may be a video signal (frame) or an audio signal (frame).

[0219] In modification example 3, the renderer 220 may be configured as a renderer that generates specific content that includes at least content with degrees of freedom of viewpoint based on viewpoint information. The encoding processing unit 230 may be configured as a transmission unit that transmits the viewpoint information used in generating the specific content to the user terminal 300 using an API between the renderer 220 and the encoding processing unit 230 (for example, CtrlEncMux(viewpoint information)). Although omitted in Figure 24, the receiving unit 210 described in the embodiment may be configured as a receiving unit that receives viewpoint information from the user terminal 300.

[0220] In modification example 3, the decoding processing unit 320 may be configured as a receiving unit that receives specific content generated by the media processing unit 200 based on viewpoint information. The renderer 330 may be configured as a renderer that acquires viewpoint information used in the generation of specific content using an API (e.g., CtrlDecDeMux) between the decoding processing unit 320 and the renderer 330. Although omitted in Figure 24, the detection unit 310 described in the embodiment may be configured as a transmitting unit that transmits viewpoint information to the media processing unit 200.

[0221] (Function and Effects) In Modification Example 3, when viewpoint information used to generate specific content is transmitted from the media processing device 200 to the user terminal 300, the media processing device uses an application programming interface between the renderer and the transmission unit, and the user terminal uses an application programming interface between the reception unit and the renderer. With this configuration, even if a situation is anticipated where the configuration of viewpoint information corresponding to specific content may change, by handling viewpoint information separately from the specific content, it is possible to suppress modifications to the function that processes the specific content (for example, the encoding processing unit 230 or the decoding processing unit 320). Consequently, since the user terminal 300 can grasp the viewpoint information used to generate the specific content, even if a case is anticipated where the viewpoint information used when generating the specific content in the media processing device 200 is different from the viewpoint information used when displaying the specific content in the user terminal, the specific content can be displayed appropriately.

[0222] Specifically, in cases where two or more media processing units 200 are assumed, by adopting an API common to the two or more media processing units 200, even if the configuration of each renderer 220 of the two or more media processing units 200 is changed, modifications to the encoding processing units 230 of the two or more media processing units 200 can be suppressed. Similarly, in cases where two or more user terminals 300 are assumed, by adopting an API common to the two or more user terminals 300, even if the configuration of each renderer 330 of the two or more user terminals 300 is changed, modifications to the decoding processing units 320 of the two or more user terminals 300 can be suppressed.

[0223] Thus, in modification example 3, it is possible to increase the degree of freedom of viewpoint information while suppressing modifications to MMT-SI, scene description, etc., handled by the encoding processing unit 230 or the decoding processing unit 320.

[0224] [Modification Example 4] Below, Modification Example 4 of the embodiment will be described. Below, the differences from the embodiment will be mainly described. In Modification Example 4, the viewpoint information transmitted from the media processing device 200 to the user terminal 300 will be mainly described, so other detailed configurations, such as the viewpoint information transmitted from the user terminal 300 to the media processing device 200, will be omitted.

[0225] Specifically, in modification example 4, the media processing device 200 transmits the viewpoint information used in generating specific content to the user terminal 300 as general-purpose metadata. General-purpose metadata is metadata whose metadata configuration information is managed by the registry server.

[0226] Firstly, as shown in Figure 25, the transmission system 10 has a registry server 600. The registry server 600 manages configuration information for general-purpose metadata. The configuration information for general-purpose metadata includes at least configuration information for viewpoint information. The registry server 600 is accessible by the media processing device 200 and the user terminal 300. If there are two or more media processing devices 200, the registry server 600 may be accessible from each of the two or more media processing devices 200. If there are two or more user terminals 300, the registry server 600 may be accessible from each of the two or more user terminals 300.

[0227] Secondly, as shown in Figure 25, the media processing device 200 includes a renderer 220, an encoding processing unit 230, and a processing unit 290. The renderer 220 and the encoding processing unit 230 are the same as in the embodiments described above, so their details are omitted.

[0228] The processing unit 290 obtains viewpoint information from the renderer 220. The processing unit 290 queries the registry server 600 to determine whether the viewpoint information obtained from the renderer 220 is registered in the registry server 600 as general-purpose metadata.

[0229] If the viewpoint information obtained from the renderer 220 is registered as general-purpose metadata in the registry server 600, the processing unit 290 obtains configuration information of the general-purpose metadata corresponding to the viewpoint information obtained from the renderer 220 from the registry server 600. The configuration information of the general-purpose metadata includes information that identifies the general-purpose metadata and configuration information of the payload.

[0230] Here, generic metadata consists of information that identifies the generic metadata (e.g., metadata_type) and a payload that indicates the content of the generic metadata. The information that identifies the generic metadata may be represented by a URN (Uniform Resource Name). In other words, the generic metadata corresponding to viewpoint information consists of information that identifies the viewpoint information and a payload that indicates the content of the viewpoint information.

[0231] For example, generic metadata may include metadata_type and metadata_payload, as shown in Figure 26.

[0232] `metadata_type` is information indicating the type of metadata. `metadata_type` may include `metadata_urn`, which is a URN indicating the type of metadata. `metadata_urn` may also include information identifying viewpoint information (e.g., a string such as `urn:mpeg:g-meta:544` or `urn:mpeg:g-meta:545`).

[0233] For example, metadata_payload may include information such as fov, viewpoint_pos_x, viewpoint_pos_y, viewpoint_pos_z, viewpoint_yaw, viewpoint_pitch, viewpoint_roll, viewport_width, and viewport_height, as shown in Figure 27.

[0234] FOV is information that indicates the field of view.

[0235] `viewpoint_pos_x` is information indicating the x-coordinate of the viewpoint position. `viewpoint_pos_x` is an example of viewpoint information used in the generation of specific content.

[0236] `viewpoint_pos_y` is information indicating the y-coordinate of the viewpoint position. `viewpoint_pos_y` is an example of viewpoint information used in the generation of specific content.

[0237] `viewpoint_pos_z` is information indicating the z-coordinate of the viewpoint position. `viewpoint_pos_z` is an example of viewpoint information used in the generation of specific content.

[0238] `viewpoint_yaw` is information indicating the yaw of the viewpoint position. `viewpoint_yaw` is an example of viewpoint information used in the generation of specific content.

[0239] `viewpoint_pitch` is information indicating the pitch of the viewpoint position. `viewpoint_pitch` is an example of viewpoint information used in the generation of specific content.

[0240] `viewpoint_roll` is information indicating the view position roll. `viewpoint_roll` is an example of view information used in the generation of specific content.

[0241] `viewport_width` is information that indicates the width of the display area (specific content).

[0242] viewport_height is information that indicates the height of the display area (specific content).

[0243] The processing unit 290 converts general-purpose metadata (viewpoint information) according to the communication protocol between the media processing unit 200 and the user terminal 300, and transmits the converted general-purpose metadata (viewpoint information) to the user terminal 300.

[0244] The communication protocol may include a transport method. The transport method may include a method compliant with MMTP, or a method compliant with ISO / IEC 23009-1 (MPEG-DASH (Dynamic Adaptive Stream over HTTP)). For example, the processing unit 290 may convert general-purpose metadata (viewpoint information) into MMT-SI format, or convert general-purpose metadata (viewpoint information) into MPD (Media Presentation Description) format.

[0245] Alternatively, the processing unit 290 may convert general-purpose metadata (viewpoint information) along with the video signal in a format compliant with Supplemental Enhancement Information (SEI) or Video Usability Information (VUI). SEI or VUI may be in the format specified in ISO / IEC 23002-7 “Information technology - MPEG video technologies Part 7: Versatile supplemental enhancement information messages for coded video bitstreams”.

[0246] The processing unit 290 outputs (transmits) viewpoint information in association with specific content generated based on the viewpoint information.

[0247] The processing unit 290 may associate viewpoint information with specific content based on the presentation time, or it may associate viewpoint information with specific content based on the sequence number. The processing unit 290 may also transmit viewpoint information together with the specific content, similar to the example in modification 2.

[0248] Here, if the processing unit 290 already knows the configuration information of the general-purpose metadata (information that identifies the general-purpose metadata and the configuration information of the payload), it may omit the operation of querying the registry server 600.

[0249] Figure 25 illustrates a case where the processing unit 290 is separate from the renderer 220 or the encoding processing unit 230. However, the processing unit 290 may be part of the renderer 220 or part of the encoding processing unit 230.

[0250] Thirdly, as shown in Figure 25, the user terminal 300 includes a decoding processing unit 320, a renderer 330, and a processing unit 390. The decoding processing unit 320 and the renderer 330 are the same as those in the embodiments described above, so their details are omitted.

[0251] The processing unit 390 receives general-purpose metadata (viewpoint information) from the media processing unit 200. The processing unit 390 accesses the registry server 600 and notifies the registry server 600 of the metadata_type value (or metadata_urn value) of the general-purpose metadata (viewpoint information) received from the media processing unit 200.

[0252] The processing unit 390 obtains payload configuration information corresponding to the metadata_type value (or metadata_urn value) from the registry server 600. Based on the payload configuration information, the processing unit 390 analyzes general metadata (viewpoint information) and identifies the viewpoint information (for example, information such as fov, viewpoint_pos_x, viewpoint_pos_y, viewpoint_pos_z, viewpoint_yaw, viewpoint_pitch, viewpoint_roll, viewport_width, and viewport_height).

[0253] Here, if the processing unit 290 already knows the configuration information of the general-purpose metadata (the configuration information of the payload), it may omit the operation of notifying the registry server 600 of the value of metadata_type (or the value of metadata_urn).

[0254] Figure 25 illustrates a case where the processing unit 390 is separate from the decoding processing unit 320 or the renderer 330. However, the processing unit 290 may be part of the decoding processing unit 320 or part of the renderer 330.

[0255] In modification example 4, the renderer 220 may be configured as a renderer that generates specific content that includes at least content with degrees of freedom of viewpoint based on viewpoint information. The processing unit 290 may be configured as a transmission unit that transmits the viewpoint information used in generating the specific content to the user terminal 300 as general-purpose metadata. Although omitted in Figure 25, the receiving unit 210 described in the embodiment may be configured as a receiving unit that receives viewpoint information from the user terminal 300.

[0256] In modification example 4, the processing unit 390 may be configured as a receiving unit that receives viewpoint information used in the generation of specific content as general-purpose metadata from the media processing device 200. Although omitted in Figure 25, the detection unit 310 described in the embodiment may be configured as a transmitting unit that transmits viewpoint information to the media processing device 200.

[0257] (Function and Effects) In Modification Example 4, the viewpoint information used to generate the specific content is transmitted from the media processing device 200 to the user terminal 300 as general-purpose metadata. With this configuration, even if a situation arises where the configuration of viewpoint information corresponding to the specific content may change, by handling the viewpoint information separately from the specific content, it is possible to suppress modifications to the function that processes the specific content (for example, the encoding processing unit 230 or the decoding processing unit 320). Consequently, since the user terminal 300 can grasp the viewpoint information used to generate the specific content, even if a case is anticipated where the viewpoint information used when generating the specific content in the media processing device 200 differs from the viewpoint information used when displaying the specific content on the user terminal, the specific content can be displayed appropriately.

[0258] Specifically, this improves the flexibility of switching transport methods, replacing codecs (encoding processing unit 230, decoding processing unit 320), and converting to different file formats.

[0259] Thus, in modification example 4, it is possible to increase the degree of freedom of viewpoint information while suppressing modifications to MMT-SI, scene description, etc., handled by the encoding processing unit 230 or the decoding processing unit 320.

[0260] [Other Embodiments] Although this disclosure has been described in the above-mentioned disclosure, the statements and drawings that constitute part of this disclosure should not be understood as limiting this disclosure. Various alternative embodiments, examples and operational techniques will become apparent to those skilled in the art from this disclosure.

[0261] The disclosures described above illustrate cases where specific content includes both the first and second content, but the disclosures are not limited to these cases. Specific content only needs to include at least the second content.

[0262] Although not specifically mentioned in the disclosure above, terminology related to MMT may be interpreted based on the definitions provided in ISO / IEC 23008-1, ARIB STD-B60, ARIB TR-B39, etc.

[0263] In the disclosure described above, an MPU timestamp descriptor was given as an example of the first absolute time information included in the MMT-SI. However, the disclosure is not limited to this. The first absolute time information included in the MMT-SI may also be an MPU extended timestamp descriptor.

[0264] Although not specifically mentioned in the disclosure above, the media processing device 200 may, if necessary, request a portion of the second content from the transmitting device 100. With such a configuration, bandwidth associated with the transmission of the second content can be saved and the increase in processing load of the media processing device 200 can be suppressed.

[0265] In the disclosure described above, MMTP was given as an example of the transmission method for the first content. However, the disclosure is not limited to this. The transmission method for the first content may also be a method compliant with ISO / IEC 23009-1 (hereinafter referred to as MPEG-DASH (Dynamic Adaptive Stream over HTTP)). In such a case, the first control information may be an MPD (Media Presentation Description). That is, in the disclosure described above, MMT-SI may be read as MPD.

[0266] Although not specifically mentioned in the disclosure above, "acquisition" can be interpreted as "received."

[0267] While not particularly limited, Operation Example 2 may also be expressed as follows: The transmitting device 100 includes a transmitting unit that transmits a content configuration having freedom of viewpoint, and the transmitting unit transmits specific viewpoint information used to generate specific content that includes at least the content. The receiving device includes a receiving unit that receives a content configuration having freedom of viewpoint, and the receiving unit receives specific viewpoint information used to generate specific content that includes at least the content. In such a case, the receiving device may be a media processing device 200 or a user terminal 300.

[0268] While not particularly limited, Operation Example 3 may be expressed as follows: The transmitting device 100 includes a transmitting unit that transmits a content configuration having degrees of freedom of viewpoint, and the transmitting unit transmits quality information of a stream relating to a 3D object included in a specific content that includes at least the content, and the quality information includes quality information for each of two or more streams whose quality differs depending on the orientation of the 3D object. The receiving device includes a receiving unit that receives a content configuration having degrees of freedom of viewpoint, and the receiving unit receives quality information of a stream relating to a 3D object included in a specific content that includes at least the content, and the quality information includes quality information for each of two or more streams whose quality differs depending on the orientation of the 3D object. In such a case, the receiving device may be a media processing device 200 or a user terminal 300.

[0269] While not particularly limited, Operation Example 4 may be expressed as follows: The transmitting device 100 includes a transmitting unit that transmits the configuration of content having a degree of freedom of viewpoint, and the transmitting unit transmits importance information for each of two or more objects included in a specific content that includes at least the content. The receiving device includes a receiving unit that receives the configuration of content having a degree of freedom of viewpoint, and the receiving unit receives importance information for each of two or more objects included in a specific content that includes at least the content. In such a case, the receiving device may be a media processing device 200 or a user terminal 300.

[0270] While not particularly limited, Operation Example 4 may also be expressed as follows: The transmitting device 100 includes a transmitting unit that transmits a content configuration having degrees of freedom of viewpoint, and the transmitting unit transmits information elements that define the range of movement of the user's viewpoint position in a three-dimensional space composed of specific content that includes at least the content. The receiving device includes a receiving unit that receives a content configuration having degrees of freedom of viewpoint, and the receiving unit receives information elements that define the range of movement of the user's viewpoint position in a three-dimensional space composed of specific content that includes at least the content. In such a case, the receiving device may be a media processing device 200 or a user terminal 300.

[0271] While not particularly limited, the arguments set to the API described in Modification Example 3 may also be examples of general-purpose metadata described in Modification Example 4. In such cases, the arguments may be obtained from registry server 600.

[0272] Although not specifically mentioned in the disclosure above, a program may be provided that causes a computer to execute each of the processes performed by the transmitting device 100, the media processing device 200, and the user terminal 300. Furthermore, the program may be recorded on a computer-readable medium. Using a computer-readable medium, it is possible to install the program on a computer. Here, the computer-readable medium on which the program is recorded may be a non-transient recording medium. The non-transient recording medium is not particularly limited, but may include, for example, a CD-ROM and / or DVD-ROM.

[0273] Alternatively, a chip may be provided comprising a memory for storing programs for executing each of the processes performed by the transmitting device 100, the media processing device 200, and the user terminal 300, and a processor for executing the programs stored in the memory.

[0274] [Note] The features of the 1-1 are a media processing device comprising: a receiving unit that receives viewpoint information from a user terminal; a renderer that generates specific content including at least content having degrees of freedom of viewpoint based on the viewpoint information; and a transmitting unit that transmits the specific content generated by the renderer to the user terminal, wherein the transmitting unit transmits the viewpoint information used in generating the specific content to the user terminal using an application programming interface between the renderer and the transmitting unit.

[0275] The first and second features are that, in the first and second features, the interface information is associated with the specific content generated based on the viewpoint information, and the media processing device is such that

[0276] The first to third features include a transmitting unit that transmits viewpoint information to a media processing device, a receiving unit that receives specific content which includes at least content having degrees of freedom of viewpoint, and which is generated by the media processing device based on the viewpoint information, and a renderer that generates the specific content received by the receiving unit, wherein the renderer is a user terminal that acquires the viewpoint information used in generating the specific content using an application programming interface between the receiving unit and the renderer.

[0277] The first four features are a program that causes a computer to perform the following steps: step A, receiving viewpoint information from a user terminal; step B, generating specific content using a renderer that includes at least content having degrees of freedom of viewpoint based on the viewpoint information; and step C, transmitting the specific content generated in step B to the user terminal using a transmission unit, wherein step C includes transmitting the viewpoint information used in generating the specific content to the user terminal using an application programming interface between the renderer and the transmission unit.

[0278] The first five features are a program that causes a computer to perform the following steps: step A, which transmits viewpoint information to a media processing device; step B, which receives specific content that includes at least content having degrees of freedom of viewpoint, and which is generated by the media processing device based on the viewpoint information, via a receiving unit; and step C, which generates the specific content received by the receiving unit via a renderer, wherein step C includes the step of acquiring the viewpoint information used in generating the specific content using an application programming interface between the receiving unit and the renderer.

[0279] The second-instance of the media processing device is a media processing device comprising: a receiving unit that receives viewpoint information from a user terminal; a renderer that generates specific content including at least content having degrees of freedom of viewpoint based on the viewpoint information; and a transmitting unit that transmits the specific content generated by the renderer to the user terminal, wherein the transmitting unit transmits the viewpoint information used in generating the specific content to the user terminal as general-purpose metadata.

[0280] The second feature of the second feature is that, in the second feature of the second feature, the general-purpose metadata is a media processing device comprising information that identifies the viewpoint information and a payload that indicates the content of the viewpoint information.

[0281] The second-third feature is that, in the second-first or second-second feature, the transmission unit is a media processing device that acquires the configuration information of the general-purpose metadata from the media processing device and the registry server connected to the user terminal.

[0282] The second-fourth feature is that, in at least one of the features of the second-first to the second-third features, the transmitting unit is a media processing device that transmits the general-purpose metadata converted according to the communication protocol between the media processing device and the user terminal.

[0283] The second-fifth feature is a user terminal comprising: a transmitting unit that transmits viewpoint information to a media processing device; and a receiving unit that receives specific content, which includes at least content having a degree of freedom of viewpoint, and which is generated by the media processing device based on the viewpoint information, wherein the receiving unit receives the viewpoint information used in the generation of the specific content from the media processing device as general-purpose metadata.

[0284] The second-sixth feature is a program that causes a computer to perform the following steps: step A, receiving viewpoint information from a user terminal; step B, generating specific content that includes at least content having degrees of freedom of viewpoint based on the viewpoint information; and step C, transmitting the specific content generated by step B to the user terminal, wherein step C includes transmitting the viewpoint information used in generating the specific content to the user terminal as general-purpose metadata.

[0285] The second-seventh feature is a program that causes a computer to perform the following steps: step A, which transmits viewpoint information to a media processing device; and step B, which receives specific content that includes at least content having degrees of freedom of viewpoint, and which is generated by the media processing device based on the viewpoint information, wherein step B includes receiving the viewpoint information used in generating the specific content from the media processing device as general-purpose metadata.

Claims

1. A media processing device comprising: a receiving unit that receives viewpoint information from a user terminal; a renderer that generates specific content including at least content having degrees of freedom of viewpoint based on the viewpoint information; and a transmitting unit that transmits the specific content generated by the renderer to the user terminal, wherein the transmitting unit transmits the viewpoint information used in generating the specific content to the user terminal as general-purpose metadata.

2. The media processing apparatus according to claim 1, wherein the general-purpose metadata comprises information identifying the viewpoint information and a payload indicating the content of the viewpoint information.

3. The media processing apparatus according to claim 1, wherein the transmitting unit obtains the configuration information of the general-purpose metadata from the media processing apparatus and the registry server connected to the user terminal.

4. The media processing apparatus according to claim 1, wherein the transmitting unit transmits the general-purpose metadata converted in accordance with the communication protocol between the media processing apparatus and the user terminal.

5. A user terminal comprising: a transmitting unit that transmits viewpoint information to a media processing device; and a receiving unit that receives specific content, which includes at least content having degrees of freedom of viewpoint, and which is generated by the media processing device based on the viewpoint information, wherein the receiving unit receives the viewpoint information used in generating the specific content from the media processing device as general-purpose metadata.

6. A program that causes a computer to perform the following steps: Step A, receiving viewpoint information from a user terminal; Step B, generating specific content that includes at least content having degrees of freedom of viewpoint based on the viewpoint information; and Step C, transmitting the specific content generated by Step B to the user terminal, wherein Step C includes transmitting the viewpoint information used in generating the specific content to the user terminal as general-purpose metadata.

7. A program that causes a computer to perform the following steps: Step A, which transmits viewpoint information to a media processing device; and Step B, which receives specific content that includes at least content having degrees of freedom of viewpoint, and which is generated by the media processing device based on the viewpoint information, wherein Step B includes receiving the viewpoint information used in generating the specific content from the media processing device as general-purpose metadata.