Media processor and program

The media processing apparatus efficiently generates and outputs video and audio based on viewpoint information, addressing the lack of audio and tactile sensation in existing 6DoF content systems, thereby improving user experience in immersive media.

JP2025110772APending Publication Date: 2025-07-29NIPPON HOSO KYOKAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024004807
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing mechanisms for transmitting 6DoF content, such as 360° video and 3D objects, fail to efficiently generate and output audio and tactile sensations based on viewpoint information, which are essential components of the specific content.

Method used

A media processing apparatus with separate renderers for generating video and audio information based on viewpoint information, allowing for appropriate output of video and audio as part of specific content, including a receiving unit, a first renderer for video, a second renderer for audio, and a transmitting unit to send this content to a user terminal.

Benefits of technology

Enables the appropriate output of video and audio, along with potential tactile sensations, as part of specific content with a degree of freedom of viewpoint, enhancing user experience in immersive media environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025110772000001_ABST
    Figure 2025110772000001_ABST
Patent Text Reader

Abstract

To provide a media processor and a program which allow a proper output of a video constituting a specific content including contents with freedom of viewpoints and information other than the video (voice or sense of touch, for example).SOLUTION: The media processor includes: a reception unit for receiving viewpoint information from a user terminal; a first renderer for generating first information constituting a specific content including at least contents with freedom of viewpoints on the basis of the viewpoint information; a second renderer for generating second information constituting the specific content on the basis of the viewpoint information; and a transmission unit for transmitting the specific content including the first information and the second information to the user terminal.SELECTED DRAWING: Figure 21
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a media processing apparatus and a program.

Background Art

[0002] Conventionally, a mechanism for transmitting content such as 360° video and 3D objects has been proposed (for example, Non-Patent Document 1). As such a mechanism, 3DoF+ (Degree of Freedom) involving viewpoint movement within the range where the user moves their head while sitting, 6DoF involving viewpoint movement within the range where the user can move freely, etc. are known. In such a mechanism, the positional relationship between the 360° video and the 3D object is indicated by scene description.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Under the above-described background, in the case of 6DoF, in addition to the video constituting specific content including content having a degree of freedom of viewpoint, it is also necessary to generate the audio constituting the specific content according to viewpoint information (for example, information such as viewpoint position and line-of-sight direction).

[0005] As a result of intensive studies, the inventors have found that it is efficient to separately generate the video and information other than the video (for example, audio or tactile sensation) constituting the specific content when assuming the above-described case.

[0006] Therefore, the present invention has been made to solve the above-described problems, and an object thereof is to provide a media processing apparatus and a program that enable appropriate output of video and information other than video (for example, audio or tactile sensation) that constitute specific content including content having a degree of freedom of viewpoint.

Means for Solving the Problems

[0007] One aspect of the disclosure is a media processing apparatus including: a receiving unit that receives viewpoint information from a user terminal; a first renderer that generates first information constituting specific content including at least content having a degree of freedom of viewpoint based on the viewpoint information; a second renderer that generates second information constituting the specific content based on the viewpoint information; and a transmitting unit that transmits the specific content including the first information and the second information to the user terminal.

Effects of the Invention

[0008] According to the present invention, it is possible to provide a media processing apparatus and a program that enable appropriate output of video and information other than video (for example, audio or tactile sensation) that constitute specific content including content having a degree of freedom of viewpoint.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

[0010] Next, embodiments of the present invention will be described. In the following description of the drawings, the same or similar parts are denoted by the same or similar reference numerals. However, it should be noted that the drawings are schematic, and the ratios of the respective dimensions are different from the actual ones.

[0011] Therefore, specific dimensions and the like should be determined with reference to the following description. Of course, there are also portions where the relationships and ratios of the dimensions to each other are different among the drawings.

[0012] [Summary of Disclosure] The summary of the disclosure is a media processing apparatus including a receiving unit that receives viewpoint information from a user terminal, a first renderer that generates first information constituting specific content including at least content having a degree of freedom of viewpoint based on the viewpoint information, a second renderer that generates second information constituting the specific content based on the viewpoint information, and a transmitting unit that transmits the specific content including the first information and the second information to the user terminal.

[0013] In the summary of the disclosure, a first renderer that generates first information constituting specific content and a second renderer that generates second information constituting specific content are provided separately. According to such a configuration, the first information and the second information can be output appropriately.

[0014] Note that the first information may be video constituting specific content, and the second information may be audio constituting specific content. The video may be read as a video signal, video data, or video stream, and the audio may be read as an audio signal, audio data, or audio stream. The audio may be read as audio or sound.

[0015] Note that the specific content generated by the media processing apparatus is generated based on viewpoint information, and it should be noted that on the user terminal side, the video included in the specific content can be treated as a 2D video.

[0016] [Embodiment] (Transmission System) Hereinafter, the transmission system according to the embodiment will be described. FIG. 1 is a diagram showing a transmission system 10 according to the embodiment. As shown in FIG. 1, the digital wireless transmission system includes a transmission device 100, a media processing device 200, and a user terminal 300.

[0017] In the embodiment, the transmission device 100 transmits the first content without a degree of freedom of viewpoint and the second content with a degree of freedom of viewpoint to the media processing device 200. Further, the transmission device 100 transmits the first control information associated with the first content and the second control information associated with the second content to the media processing device 200.

[0018] The first content may include at least one of 2D video and audio. The first content and the first control information may be transmitted by a first method. The first method may be a method compliant with ISO / IEC 23008-1 (hereinafter, MMT (MPEG Media Transport)). Hereinafter, the case where the first method is MMTP (MMT Protocol) compliant with MMT will be exemplified. The first control information may be referred to as MMT-SI (Signaling Information).

[0019] The second content may include 360° video and 3D objects. The second content and the second control information may be transmitted by a second method. Alternatively, it may be transmitted by a protocol such as HTTP (Hyper Text Transfer Protocol). The second content may conform to 3DoF+ (Degree of Freedom) with a viewpoint movement within the range where the user moves their head while sitting, 6DoF with a viewpoint movement within the range where the user moves freely, etc. Since the second content has a degree of freedom of viewpoint, at the same time (frame), it may include two or more 360° videos and may include two or more 3D objects. The second control information may be referred to as a scene description.

[0020] Here, the second control information may be transmitted in the above-described first method. That is, the second control information may be transmitted in the same first method (e.g., MMTP) as the first control information. Alternatively, the second control information may be transmitted in a protocol such as HTTP.

[0021] The transmission from transmitting device 100 to media processing device 200 is not particularly limited, but may be transmission using satellite broadcasting, transmission using the Internet network, or transmission using a mobile communications network.

[0022] Although not particularly limited, the transmission system may be a digital wireless transmission system, and the digital wireless transmission system may be a system used for 4K and 8K satellite broadcasting.

[0023] Based on the viewpoint information received from user terminal 300, media processing device 200 generates specific content that includes at least the second content described above, and transmits the generated specific content to user terminal 300. Although not particularly limited, transmission of the specific content may be via the Internet network or a mobile communication network.

[0024] The user terminal 300 may be a user terminal such as a smartphone, a tablet terminal, or a head-mounted display. As shown in Fig. 1, two or more user terminals 300 may be provided as the user terminal 300. In other words, two or more user terminals 300 may request the media processing device 200 to generate specific content. Each user terminal 300 may transmit different viewpoint information to the media processing device 200.

[0025] (Media processing devices and user terminals) The media processing device and user terminal according to the embodiment are described below. Figure 2 is a block diagram showing a media processing device 200 and a user terminal 300 according to the embodiment.

[0026] As shown in FIG. 2, the media processing apparatus 200 includes a reception unit 210, a renderer 220, and an encoding processing unit 230.

[0027] The reception unit 210 receives viewpoint information. In an embodiment, the reception unit 210 constitutes a reception unit that receives viewpoint information from the user terminal 300. The viewpoint information includes an information element indicating the viewpoint position of the user of the user terminal 300 and an information element indicating the line-of-sight direction of the user of the user terminal 300.

[0028] The renderer 220 generates specific content including at least second content based on the viewpoint information. Since the specific content is generated based on the viewpoint information, at the same time (frame), it may include one 360° video or one 3D object. Hereinafter, a case where the specific content includes first content in addition to the second content will be exemplified.

[0029] First, the renderer 220 generates first content including a 2D video and audio as part of the specific content based on first control information (MMT-SI). Viewpoint information is not required for the generation of the first content.

[0030] Specifically, the renderer 220 acquires a 2D video, audio, and MMT-SI in the form of an MMTP packet in which the 2D video, audio, and MMT-SI are packetized.

[0031] For example, the MMTP packet is stored in an IP (Internet Protocol) packet. The IP packet may be transmitted using UDP (User Datagram Protocol) or TCP (Transmission Control Protocol).

[0032] Here, the first content is processed in units (hereinafter, MPU; Media Processing Unit) delimited by a certain time width. The MPU includes one or more access units. An access unit may be treated as an MFU (Media Fragment Unit). The MFU related to 2D video may be referred to as an NAL (Network Abstraction Layer) unit, and the MFU related to audio may be referred to as an MHAS (MPEG-H 3D Audio Stream) packet.

[0033] MMT-SI includes a PA (Package Access) message, and the PA message includes an MPT (MMT Package Table) indicating a list of the first content. Further, MMT-SI includes an MPU timestamp descriptor indicating the presentation time of the first content. The MPU timestamp descriptor may mean the presentation time of the MPU, that is, the time of the access unit presented first in the MPU.

[0034] The MPU timestamp descriptor may be generated with UTC (Coordinated Universal Time) as the reference time. As the reference time, TAI (International Atomic Time) may be used, or the time provided from GPS (Global Positioning System) may be used. The reference time may be the time provided from an NTP (Network Time Protocol) server, or the time provided from a PTP (Precision Time Protocol) server.

[0035] Second, the renderer 220 generates second content including 360° video and 3D objects as a part of specific content based on second control information (scene description). Viewpoint information is used in the generation of the second content.

[0036] Specifically, the renderer 220 may obtain the scene description in the form of an MMTP packet in which the scene description is packetized. The method for obtaining the 360° video and the 3D object is not particularly limited.

[0037] The 360° video may be converted into a 2D video by a projective transformation such as ERP (Equirectangular projection) or a cube map. Metadata indicating the type of projective transformation applied to the 360° video may be added. The 3D object may be encoded in a mesh format. As the encoding in the mesh format, ISO / IEC 14496-16 “Animation framework extension (AFX)” may be used. The 3D object may be encoded in a point cloud format. As the encoding in the point cloud format, ISO / IEC 23090-5 “Video-based Point Cloud Compression” may be used.

[0038] Here, the second content is grouped into one file in units delimited by a certain time width. The certain time width may be 500 ms. For example, when the frame rate is 60 fps (frame per second), one file contains 30 frames.

[0039] The scene description is generated for each file and includes information for specifying the 360° video and the 3D object for each frame. For example, the scene description includes an information element (object_name) indicating the name of the 3D object in the frame, an information element (frame_number) indicating the frame number, an information element (translation_object) indicating the position of the 3D object in the frame, an information element (rotation_object) indicating the rotation of the 3D object in the frame, an information element (scale_object) indicating the size of the 3D object in the frame, and the like.

[0040] Thirdly, the renderer 220 outputs specific content including the first content and the second content to the encoding processing unit 230. The renderer 220 may output the presentation time of the specific content to the encoding processing unit 230 together with the specific content.

[0041] Here, the presentation time of the specific content may be corrected based on the delay time between the media processing device 200 and the user terminal 300. Specifically, the renderer 220 may calculate the presentation time (T' = T + ΔT) of the specific content provided from the media processing device 200 to the user terminal 300 based on the presentation time (T) and the delay time (ΔT) of the specific content provided from the transmission device 100 to the media processing device 200. The delay time (ΔT) may be a value predetermined in the media processing device 200, or may be a different value for each user terminal 300.

[0042] Fourthly, the renderer 220 may constitute a transmission unit that transmits the viewpoint information used for generating the specific content to the user terminal 300. The viewpoint information used for generating the specific content may be transmitted from the encoding processing unit 230 to the user terminal 300.

[0043] For example, the transmission method of the viewpoint information and the specific content may be MMTP or HTTP. When MMTP is used as the transmission method of the specific content, the viewpoint information may be stored as metadata in the OMAF (Omnidirectional Media Format) defined in ISO / IEC 23090-2.

[0044] The encoding processing unit 230 encodes the specific content generated by the renderer 220. In an embodiment, the encoding processing unit 230 may be an example of a transmission unit that transmits the specific content to the user terminal 300.

[0045] Furthermore, the encoding processing unit 230 may encode the presentation time of the specific content. The encoding processing unit 230 may transmit an information element indicating the presentation time to the user terminal 300 together with the specific content.

[0046] Here, as the compression encoding method used by the encoding processing unit 230, any compression encoding method can be used. For example, the compression encoding method may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding).

[0047] As described above, since the second content included in the specific content is generated based on the viewpoint information, the video included in the specific content can be treated as a 2D video without the degree of freedom of the viewpoint.

[0048] For example, the transmission control method used at the start or end of viewing the specific content may include RTSP (Real Time Streaming Protocol). The transmission method may be MMTP or HTTP. When MMTP is used as the transmission method, the specific content may be stored in OMAF defined in ISO / IEC 23090-2.

[0049] As shown in FIG. 2, the user terminal 300 includes a detection unit 310, a decoding processing unit 320, and a renderer 330.

[0050] The detection unit 310 detects the viewpoint position and line-of-sight direction of the user. The detection unit 310 may include an acceleration sensor or a GPS (Global Positioning System) sensor. The detection unit 310 may include a user I / F (e.g., touch sensor, keyboard, mouse, controller, etc.) manually input by the user. The detection unit 310 may transmit the viewpoint information (viewpoint position and line-of-sight direction) to the media processing device 200. The detection unit 310 may output the viewpoint information (viewport) to the renderer 330.

[0051] The decoding processing unit 320 decodes the specific content received from the media processing device 200. The decoding processing unit 320 may decode the presentation time received from the media processing device 200. The decoding processing unit 320 may output the specific content to the renderer 330, or may output the presentation time to the renderer 330.

[0052] The renderer 330 outputs the specific content decoded by the decoding processing unit 320. The renderer 330 may output the specific content based on the presentation time decoded by the decoding processing unit 320. For example, the renderer 330 may output the video content included in the specific content to a display and output the audio content included in the specific content to a speaker.

[0053] Here, the renderer 330 may generate specific content with the viewpoint position and the line-of-sight direction corrected based on the difference between the viewpoint information received from the media processing device 200 and the viewpoint information input from the detection unit 310.

[0054] (Second Content) Hereinafter, the second content according to the embodiment will be described. Here, the second content at t = 0, t = 1, and t = 2 will be described. The time intervals of t = 0, t = 1, and t = 2 are not particularly limited.

[0055] For example, as shown in FIG. 3, at t = 0, a 360° video may be displayed without displaying the 3D object. The 360° video may be considered as the background video of the 3D object. At t = 1, the 3D object may be displayed in a form superimposed on the 360° video. Further, at t = 1, the position and rotation of the 3D object superimposed on the 360° video may be changed.

[0056] The above-described scene description includes information elements indicating the position, rotation, and size of the 3D object for each of t = 0, t = 1, and t = 2, and the 3D object can be appropriately superimposed on the 360° video.

[0057] (Video viewing method) Hereinafter, the video viewing method according to the embodiment will be described. Here, the viewing of specific content including the first content and the second content will be exemplified.

[0058] As shown in FIG. 4, in step S11, the user terminal 300 transmits RTSP SETUP to the media processing device. RTSP SETUP is a message indicating the start of viewing specific content.

[0059] Here, RTSP SETUP includes the IP address of the user terminal 300, the listening port number, the identification information (content ID) of the content, etc. RTSP SETUP may include the capability information of the user terminal 300 for viewing specific content. The capability information may include the frame rate, display resolution, etc. The display resolution may include the field of view (FoV). The capability information may include an information element indicating the encoding method and compression method supported by the user terminal 300.

[0060] Here, a case where the capability information of the user terminal 300 is directly notified to the media processing device 200 is exemplified, but the embodiment is not limited thereto. The capability information of the user terminal 300 may be notified to the transmission device 100 first, and then notified from the transmission device 100 to the media processing device 200.

[0061] In step S12, the media processing device 200 transmits a response to RTSP SETUP. Here, an ACK indicating that RTSP SETUP has been received is transmitted as the response.

[0062] In step S21, the user terminal 300 transmits initial viewpoint information to the media processing device 200. The initial viewpoint information may be transmitted in the form of MMT-SI.

[0063] In step S22, the media processing device 200 generates initial specific content based on the initial viewpoint information (rendering process). For example, the media processing device 200 generates the second content to be included in the initial specific content based on the initial viewpoint information and the scene description.

[0064] Here, the media processing device 200 may generate the initial specific content using a viewport that is wider than the display resolution of the user terminal 300. For example, the range wider than the display resolution may be a range of +20% in the horizontal direction and +20% in the vertical direction of the display resolution.

[0065] The media processing device 200 applies a compression encoding method to the initial specific content. Although not particularly limited, the compression encoding method may be HEVC or VVC.

[0066] In step S23, the media processing device 200 transmits the initial specific content corresponding to the initial viewpoint information to the user terminal 300. The media processing device 200 transmits the presentation time of the initial specific content to the user terminal 300. As described above, the presentation time (T’) provided to the user terminal 300 may be determined based on the delay time (ΔT).

[0067] When different values are used for each user terminal 300 as the delay time (ΔT), it is possible to specify on the side of the media processing device 200 by including the transmission time of the RTSP SETUP in the above-mentioned RTSP SETUP.

[0068] The user terminal 300 outputs the specific content based on the presentation time (T’). The user terminal 300 may generate specific content with the viewpoint position and the line-of-sight direction corrected based on the difference between the viewpoint information received from the media processing device 200 and the viewpoint information input from the detection unit 310.

[0069] In step S31, the user terminal 300 transmits viewpoint information to the media processing device 200. The viewpoint information may be transmitted in the form of MMT-SI. Here, the user terminal 300 may transmit the viewpoint information at a predetermined period (for example, 500 ms), or may transmit the viewpoint information in response to a change in at least one of the viewpoint position and the line-of-sight direction.

[0070] In step S32, the media processing device 200 generates specific content (rendering process) based on the viewpoint information received in step S31.

[0071] In step S33, the media processing device 200 transmits the specific content corresponding to the viewpoint information received in step S31 to the user terminal 300.

[0072] The processes of steps S31 to S33 are the same as the processes of steps S21 to S23, except that the viewpoint information received in step S31 is used instead of the initial viewpoint information. Therefore, the details of the processes of steps S31 to S33 are omitted. The processes of steps S31 to S33 may be repeated at a predetermined period, or may be repeated every time the user's viewpoint position or line-of-sight direction changes.

[0073] In step S41, the user terminal 300 transmits RTSP TEARDOWN to the media processing device. RTSP TEARDOWN is a message indicating the end of viewing specific content.

[0074] In step S42, the media processing device 200 transmits a response to RTSP TEARDOWN. Here, an ACK indicating that RTSP TEARDOWN has been received is transmitted as the response.

[0075] In FIG. 4, the case where steps S11 and S12 are executed based on RTSP is illustrated, but the embodiment is not limited thereto. Steps S11 and S12 may be executed based on MMTP or may be executed based on HTTP.

[0076] Similarly, the case where steps S41 and S42 are executed based on RTSP is illustrated, but the embodiment is not limited thereto. Steps S41 and S42 may be executed based on MMTP or may be executed based on HTTP.

[0077] In FIG. 4, the case where steps S31 to S33 are executed based on MMTP is illustrated, but the embodiment is not limited thereto. Steps S31 to S33 may be executed based on another method (for example, HTTP).

[0078] Similarly, the case where steps S41 to S43 are executed based on MMTP is illustrated, but the embodiment is not limited thereto. Steps S41 to S43 may be executed based on another method (for example, HTTP).

[0079] (Operation Example 1) The above-described embodiment may include Operation Example 1 shown below. In Operation Example 1, the media processing device 200 transmits the viewpoint information used in the generation of the specific content to the user terminal 300 in association with the sequence number added to the specific content.

[0080] Specifically, the media processing device 200 (renderer 220) generates second content including a 360° video and a 3D object as a part of the specific content based on the second control information (scene description) in the same manner as in the above-described embodiment. In the generation of the second content, the viewpoint information received from the user terminal 300 is used.

[0081] In operation example 1, the renderer 220 associates the viewpoint information used in the generation of specific content (here, the second content) with a sequence number. The renderer 220 transmits the viewpoint information used in the generation of the specific content to the user terminal 300 in association with the sequence number added to the specific content. The viewpoint information may be stored in a VP (View Port) message. The VP message may have the format of MMT-SI defined in ISO / IEC 23008-1, ARIB STD-B60, ARIB TR-B39, etc. The VP message may be transmitted for each frame. The VP message may be transmitted to the user terminal 300 as a message (MMT-SI) related to MMTP.

[0082] Although not particularly limited, the VP message may have the data structure shown in FIG. 5. As shown in FIG. 5, the VP message may include message_id, version, length, fov, viewpoint_pos_x, viewpoint_pos_y, viewpoint_pos_z, viewpoint_yaw, viewpoint_pitch, viewpoint_roll, viewport_width, viewport_height, mpu_sequence_number_flag, mpu_sequence_number, etc.

[0083] message_id is identification information indicating the VP message. message_id may be 0x0204.

[0084] version is information indicating the version of the MMTP protocol. version may be 0x00.

[0085] length is information indicating the length of the VP message.

[0086] fov is information indicating the field of view.

[0087] viewpoint_pos_x is information indicating the x-coordinate of the viewpoint position. viewpoint_pos_x is an example of the viewpoint information used in the generation of specific content.

[0088] viewpoint_pos_y is information indicating the y-coordinate of the viewpoint position. viewpoint_pos_y is an example of the viewpoint information used in the generation of specific content.

[0089] viewpoint_pos_z is information indicating the z-coordinate of the viewpoint position. viewpoint_pos_z is an example of the viewpoint information used in the generation of specific content.

[0090] viewpoint_yaw is information indicating the yaw of the viewpoint position. viewpoint_yaw is an example of the viewpoint information used in the generation of specific content.

[0091] viewpoint_pitch is information indicating the pitch of the viewpoint position. viewpoint_pitch is an example of the viewpoint information used in the generation of specific content.

[0092] viewpoint_roll is information indicating the roll of the viewpoint position. viewpoint_roll is an example of the viewpoint information used in the generation of specific content.

[0093] viewport_width is information indicating the width of the display area (specific content).

[0094] viewport_height is information indicating the height of the display area (specific content).

[0095] The mpu_sequence_number_flag is information indicating whether the field of mpu_sequence_number exists. For example, when the mpu_sequence_number_flag is 1, the field of mpu_sequence_number exists, and when the mpu_sequence_number_flag is 0, the field of mpu_sequence_number may not exist.

[0096] The mpu_sequence_number is the MPU sequence number of the video corresponding to the specific content indicated by the VP message. The mpu_sequence_number is an example of the sequence number added to the specific content.

[0097] Here, viewpoint_pos_x, viewpoint_pos_y, and viewpoint_pos_z are examples of information elements indicating the user's viewpoint position in the three-dimensional space constituted by the scene description. viewpoint_yaw, viewpoint_pitch, and viewpoint_roll are examples of information elements indicating the user's line-of-sight direction in the three-dimensional space constituted by the scene description. viewport_width and viewport_height are examples of information elements indicating the number of pixels of the video included in the specific content.

[0098] Under such a premise, the media processing apparatus 200 and the user terminal 300 may execute the operations shown below.

[0099] First, the media processing device 200 (renderer 220) may specify viewpoint_pos_x, viewpoint_pos_y, and viewpoint_pos_z based on the viewpoint position used for generating specific content. The renderer 220 may specify viewpoint_yaw, viewpoint_pitch, and viewpoint_roll based on the line-of-sight direction used for generating specific content. The renderer 220 may specify viewport_width based on the number of horizontal pixels of specific content, and may specify viewport_height based on the number of vertical pixels of specific content.

[0100] The media processing device 200 (encoding processing unit 230) may execute compression encoding of specific content generated by the renderer 220 and transmit the specific content. Here, the encoding processing unit 230 may transmit a VP message shown in FIG. 5 to the user terminal 300. That is, the encoding processing unit 230 may transmit the viewpoint information used for generating specific content to the user terminal 300 in association with the sequence number added to the specific content.

[0101] Second, the user terminal 300 (decoding processing unit 320) may configure a receiving unit that receives, from the media processing device 200, the viewpoint information used for generating specific content in association with the sequence number added to the specific content. That is, the decoding processing unit 320 may receive a VP message (MMT-SI) shown in FIG. 5 from the media processing device 200.

[0102] The user terminal 300 (renderer 330) may associate the video that constitutes the decoded specific content with the viewpoint information included in the VP message based on the sequence number (mpu_sequence_number). The renderer 330 may specify, from the display area defined by the information (viewport_width and viewport_height) included in the VP message, the video in the range specified by the viewpoint information detected by the detection unit 310 based on the difference between the viewpoint information detected by the detection unit 310 and the viewpoint information included in the VP message. The renderer 330 may display the specified video.

[0103] (Operation Example 2) The above-described embodiment may include the following Operation Example 2. In Operation Example 2, as shown in FIG. 6, a case where the first user terminal 400 that feeds back viewpoint information and the second user terminal 500 that does not feed back viewpoint information are mixed is assumed. The first user terminal 400 may be a terminal such as a head-mounted display. The first user terminal may have the same functions as the above-described user terminal 300. The second user terminal 500 may be a terminal such as a volumetric display.

[0104] Specifically, in Operation Example 2, as shown in FIG. 6, the media processing device 200 (renderer 220) may constitute a receiving unit that receives the content configuration and recommended viewport information from the transmitting device 100. The content configuration may be considered to include 2D video, audio, 360° video, and 3D objects. The content configuration may be considered to include MMT-SI and scene description. The recommended viewport information may be considered as an example of specific viewpoint information. The recommended viewport information may be information that defines (recommends) at which position, direction, and viewing angle to view the video in the three-dimensional space (the three-dimensional space constructed by the scene description) constituted by the specific content. The recommended viewport information may include at least one information element indicating an information element indicating the viewpoint position in the three-dimensional space constituted by the specific content and an information element indicating the line-of-sight direction in the three-dimensional space constituted by the specific content. The recommended viewport information may be considered to be mainly viewpoint information used in the second user terminal 500.

[0105] Hereinafter, unless explicitly stated otherwise, the viewpoint information used in the generation of specific content may include the specific viewpoint information (recommended viewport information) received from the transmitting device 100, or may include the viewpoint information received from the user terminal 300.

[0106] Although not particularly limited, the recommended viewport information may be included in the scene description in the manner shown in FIG. 7. As shown in FIG. 7, the recommended viewport information may include camera_orientation, frame_number, translation, and yfov.

[0107] camera_orientation is information indicating the direction to view the video in the three-dimensional space constructed by the scene description. camera_orientation may be considered synonymous with the line-of-sight direction.

[0108] The frame_number is information indicating the frame number of the video to which the camera_orientation, translation, and yfov are applied. The camera_orientation may be considered as an example of specific viewpoint information.

[0109] The translation is information indicating the position to view the video in the three-dimensional space constructed in the scene description. The translation may be considered synonymous with the viewpoint position. The translation may be considered as an example of specific viewpoint information. For example, in FIG. 7, when the frame number is 0, the viewpoint position is [0, 0, -50], and when the frame number is 2505, the case where the viewpoint position moves to [0, 0, -75] is illustrated.

[0110] The yfov is information indicating the viewing angle of the video in the three-dimensional space constructed in the scene description.

[0111] Although not particularly limited, the camera_orientation and translation may be given by the content producer. Alternatively, when assuming the case where the content is imaged by a camera, the camera_orientation and translation may be automatically given by the GPS and sensors provided in the camera.

[0112] Here, the translation is an example of an information element indicating the viewpoint position in the three-dimensional space constituted by specific content. The camera_orientation is an example of an information element indicating the line-of-sight direction in the three-dimensional space constituted by specific content.

[0113] Under such a premise, the media processing device 200 and the user terminal 300 may execute the operations shown below.

[0114] First, the media processing device 200 (renderer 220) may generate specific content based on specific viewpoint information (recommended viewport information). The media processing device 200 (encoding processing unit 230) may perform compression encoding of the specific content generated by the renderer 220 and transmit the specific content. Here, the encoding processing unit 230 may transmit the specific content generated based on the specific viewpoint information to the first user terminal 400, and may also transmit the specific content generated based on the specific viewpoint information to the second user terminal 500.

[0115] Second, when the media processing device 200 (reception unit 210) is the first user terminal 400 that feeds back viewpoint information, it may receive the viewpoint information from the first user terminal 400. The media processing device (renderer 220) may generate specific content based on the viewpoint information received from the first user terminal 400. In such a case, the media processing device 200 (reception unit 210) may receive a reset signal from the first user terminal 400. The media processing device (renderer 220) may generate specific content based on the specific viewpoint information (recommended viewport information) in response to the reset signal.

[0116] In such a case, the first user terminal 400 may have the same configuration as the user terminal 300. However, the first user terminal 400 (detection unit 310) may have a function of detecting a reset signal. The detection unit 310 may detect a user operation for inputting the reset signal. The detection unit 310 may transmit the reset signal to the media processing device 200.

[0117] Third, when the media processing device 200 (renderer 220) is the second user terminal 500 that does not feed back viewpoint information, it may generate specific content based on the specific viewpoint information (recommended viewport information).

[0118] In such a case, the second user terminal 500 may not have the detection unit 310 that detects viewpoint information. The second user terminal 500 may not have the renderer 330. The second user terminal 500 may have a configuration similar to that of the user terminal 300, except that it does not have the detection unit 310 and the renderer 330.

[0119] (Operation Example 3) The above-described embodiment may include the following Operation Example 3. In Operation Example 3, the media processing device 200 (for example, the selection unit 260 described later) may constitute a receiving unit that receives, from the transmission device 100, quality information regarding each of two or more streams whose quality varies depending on the orientation of the 3D object based on the viewpoint information, as the quality information of the stream regarding the 3D object included in the specific content.

[0120] Specifically, in Operation Example 3, as shown in FIG. 8, in addition to the configuration shown in FIG. 2, the media processing device 200 has a selection unit 260. The selection unit 260 receives the scene description and the 3D object from the transmission device 100. The selection unit 260 inputs the selected stream (3D object) among two or more streams to the renderer 220. The selection unit 260 may request the transmission device 100 to transmit the selected stream. When assuming a case where the media processing device 200 transmits specific content to a plurality of user terminals 300, the selection unit 260 may request the transmission device 100 to transmit the streams required by each of the plurality of user terminals 300, or may request the transmission device 100 to transmit all the streams.

[0121] Here, the selection unit 260 receives, from the transmission device 100, quality information regarding each of two or more streams whose quality varies depending on the orientation of the 3D object based on the viewpoint information, as the quality information of the stream regarding the 3D object.

[0122] The quality information may be information indicating the relative quality of each face constituting the bounding box for the 3D object. The bounding box may be represented by a three-dimensional rectangle that projects the 3D object. For example, as shown in FIG. 9, the bounding box may be defined by vertices A to H. In such a case, each face of the bounding box includes face #1 represented by vertices A, B, F, E, face #2 represented by vertices B, C, G, F, face #3 represented by vertices A, B, C, D, face #4 represented by vertices E, F, G, H, face #5 represented by vertices A, D, H, E, and face #6 represented by vertices D, C, G, H.

[0123] In such a case, assuming a case of viewing a 3D object in the three-dimensional space constructed by the scene description, it is assumed that three faces are mainly observed. In other words, it is assumed that the remaining three faces are not observed much.

[0124] In operation example 3, as streams related to the 3D object included in the specific content, two or more streams with different qualities depending on the orientation of the 3D object based on the viewpoint information are prepared.

[0125] Although not particularly limited, the quality information may be included in the scene description in the form shown in FIG. 10. In FIG. 10, six streams are illustrated as streams with different orientations of the 3D object. The quality information may be represented in the form of "quality" [#1, #2, #3, #4, #5, #6]. In [], #1 to #6 mean the quality indexes of face #1 to face #6. The quality index may take a value in the range of 1 to 9. A larger value of the quality index may mean higher quality. For example, in the stream identified by "id" = "1", the quality of #1, #2, #3 ("8") is high, and the quality of #4, #5, #6 ("3") is low. In the stream identified by "id" = "2", the quality of #1, #2, #3 ("3") is low, and the quality of #4, #5, #6 ("8") is high.

[0126] Under such a premise, the media processing apparatus 200 may execute the operations shown below. In the following, the selection of a stream (3D object) selected from two or more streams will be mainly described.

[0127] In mode 1, as shown in the upper part of FIG. 11, the media processing apparatus 200 (selection unit 260) identifies the vertex closest to the user's viewpoint position (for example, vertex B), and then may identify three faces having the closest vertex (for example, face #1 represented by vertices A, B, F, E; face #2 represented by vertices B, C, G, F; and face #3 represented by vertices A, B, C, D). The selection unit 260 may select a stream (in the example shown in FIG. 10, the stream identified by "id" = "1") for which the sum of the quality indexes of the three identified faces is the maximum.

[0128] Note that in mode 1, since the vertex closest to the user's viewpoint position is identified based on the viewpoint information, it may be considered that the selection unit 260 selects a stream to be transmitted to the user terminal 300 from two or more streams based on the viewpoint information and the quality information.

[0129] In Mode 2, it may be a mode applicable to cases where a reduced or enlarged display of a 3D object is executed. For example, as shown in the middle row of FIG. 11, the media processing device 200 (selection unit 260) identifies the vertex closest to the user's viewpoint position (for example, vertex B), and then may identify three faces having the closest vertex (for example, face #1 represented by vertices A, B, F, E, face #2 represented by vertices B, C, G, F, and face #3 represented by vertices A, B, C, D). In the case where a reduced display of the 3D object is executed, the pixels of the 3D object are decimated. Therefore, the selection unit 260 may select a stream with the minimum sum of the quality indices of the three identified faces (in the example shown in FIG. 10, the stream identified by "id" = "2"). On the other hand, in the case where an enlarged display of the 3D object is executed, the pixels of the 3D object are interpolated. Therefore, the selection unit 260 may select a stream with the maximum sum of the quality indices of the three identified faces (in the example shown in FIG. 10, the stream identified by "id" = "1").

[0130] Note that in Mode 2, since the vertex closest to the user's viewpoint position is identified based on the viewpoint information, it may be considered that the selection unit 260 selects a stream to be transmitted to the user terminal 300 from among two or more streams based on the viewpoint information and the quality information.

[0131] In mode 3, it may be a mode applied in a case where two 3D objects (3D object #1 and 3D object #2) overlap in the user's line of sight direction. Here, the selection of the stream related to 3D object #1 will be described. For example, as shown in the lower part of FIG. 11, the media processing device 200 (selection unit 260) identifies the vertex closest to the user's viewpoint position (for example, vertex B), and then may identify three faces having the closest vertex (for example, face #1 represented by vertices A, B, F, E, face #2 represented by vertices B, C, G, F, and face #3 represented by vertices A, B, C, D). Here, 3D object #2 overlaps on the line segment connecting the vertex closest to the user's viewpoint position (for example, vertex B) and the user's viewpoint position, and the three identified faces are blocked by 3D object #2. Therefore, the selection unit 260 may select a stream (in the example shown in FIG. 10, the stream identified by "id" = "2") with the minimum total quality index of the three identified faces.

[0132] Note that in mode 3, since the vertex closest to the user's viewpoint position is identified based on the viewpoint information, the selection unit 260 may be considered to select a stream to be transmitted to the user terminal 300 from among two or more streams based on the viewpoint information and the quality information. Furthermore, in mode 3, since the overlap of the two 3D objects is identified based on the viewpoint information and the arrangement information of the 3D objects, the selection unit 260 may be considered to select a stream to be transmitted to the user terminal 300 from among two or more streams based on the viewpoint information, the quality information, and the arrangement information. The arrangement information of the 3D objects (for example, "rotation_object", "scale_object", "translation_object" shown in FIG. 10) may be included in the scene description. The arrangement information of the 3D objects may be considered to be "link_area" shown in FIG. 10. That is, the selection unit 260 may receive the arrangement information of the 3D objects in the three-dimensional space from the transmission device 100.

[0133] In the quality information shown in FIG. 10, the sum of the quality indexes of the six surfaces is equal in each stream. However, the embodiment is not limited to this. The sum of the quality indexes of the six surfaces may be different between two or more streams.

[0134] (Operation Example 4) The above-described embodiment may include Operation Example 4 shown below. Here, Operation Example 4 includes the following operations in addition to Operation Example 3. In Operation Example 4, the media processing device 200 (for example, the selection unit 270 described later) may constitute a receiving unit that receives importance information regarding each of two or more objects included in specific content from the transmission device 100.

[0135] Specifically, in Operation Example 4, as shown in FIG. 12, in addition to the configuration shown in FIG. 2, the media processing device 200 has a selection unit 270. The selection unit 270 receives a scene description, 3D objects, and 360° video from the transmission device 100. The selection unit 270 inputs a selected stream (3D object) from among two or more streams to the renderer 220.

[0136] Here, the selection unit 270 receives importance information regarding each of two or more objects. The objects may include 3D objects and 360° video.

[0137] For example, the importance information may be information indicating the relative importance among two or more objects. For example, as shown in FIG. 13, consider a case where in a three-dimensional space constructed by a scene description, there are an object A (background), an object B (person), and an object C (dog). The object A (background) is an example of 360° video, and the objects B (person) and C (dog) are examples of 3D objects. In such a case, the importance information may be information indicating the relative importance among the object A (background), the object B (person), and the object C (dog).

[0138] Although not particularly limited, the importance information may be included in the scene description in the manner shown in FIG. 14. In FIG. 14, the importance information may be represented by weight. Weight may take a value in the range of 1 to 9. The larger the value of weight, the higher the importance may be meant. In FIG. 14, the weight ("9") of object A (background) identified by "object_id" = "0" is the highest, the weight ("3") of object B (person) identified by "object_id" = "1" is the lowest, and the case where the weight ("8") of object C (dog) identified by "object_id" = "2" is higher than the weight of object B (person) and lower than the weight of object A (background) is illustrated.

[0139] Under such a premise, the media processing apparatus 200 may execute the following operations. In the following, the selection of a stream (3D object) selected from among two or more streams will be mainly described.

[0140] First, the media processing apparatus 200 (selection unit 270) selects the stream with the highest quality for the 3D object with the highest importance. The method of selecting the stream may be the same as in operation example 3. For example, since the importance of object C (dog) is greater than the importance of object B (person), for object C (dog), the selection unit 270 selects the stream for which the sum of the quality indices of the three faces having the vertices closest to the user's viewpoint position is the largest.

[0141] Second, the media processing device 200 (selection unit 270) selects the stream with the lowest quality for 3D objects other than the 3D object with the highest importance. Subsequently, the selection unit 270 sequentially replaces the stream with the lowest quality with a stream with a high quality within the range where the specific conditions are satisfied, starting from the 3D objects with high importance. The specific conditions may include a first condition that the bandwidth of the line from the transmission device 100 to the media processing device 200 is equal to or less than a threshold value, and may include a second condition that the processing load of the media processing device 200 is equal to or less than a threshold value. The specific conditions may be defined by a combination of the first condition and the second condition. For example, since the importance of object B (person) is smaller than the importance of object C (dog), the selection unit 270 selects a stream with a high quality for object B (person) within the range where the specific conditions are satisfied.

[0142] As described above, it may be considered that the media processing device 200 (selection unit 270) selects a stream to be transmitted to the user terminal 300 from among two or more streams based on viewpoint information, quality information, and importance information. It may be considered that the media processing device 200 (selection unit 270) selects a stream to be transmitted to the user terminal 300 from among two or more streams based on viewpoint information, quality information, arrangement information, and importance information.

[0143] Note that in operation example 4, a case where there is one stream for 360° video has been illustrated. However, the embodiment is not limited to this. For 360° video, there may also be two or more streams with different qualities.

[0144] In operation example 4, a case where there are two or more streams with different qualities depending on the orientation of the 3D object based on the viewpoint information for the 3D object has been illustrated. However, the embodiment is not limited to this. For the 3D object, there may be two or more streams with different qualities regardless of the orientation of the 3D object.

[0145] (Operation Example 5) The above-described embodiment may include the following operation example 3. In operation example 3, the media processing device 200 (for example, the renderer 220) may include a receiving unit that receives from the transmission device 100 an information element that defines a movement range of the user's viewpoint position in a three-dimensional space (a three-dimensional space constructed by a scene description) constituted by specific content.

[0146] First, in operation example 5, the information element may include an information element (hereinafter, the first information element) that restricts the movement of the user's viewpoint position inside a 3D object included in specific content. For example, as shown in FIG. 15, in a case where a 3D object is arranged in a three-dimensional space constructed by a scene description, the movement of the viewpoint position inside the 3D object may be restricted. However, in a case where the 3D object is a building or in a case where another scene exists inside the 3D object, the movement of the viewpoint position inside the 3D object may be allowed.

[0147] Second, in operation example 5, the information element may include an information element (hereinafter, the second information element) that restricts the movement of the user's viewpoint position outside the three-dimensional space. For example, as shown in FIG. 16, the three-dimensional space may be defined by a combination of a rectangular parallelepiped and an ellipsoid of revolution. The number of rectangular parallelepipeds defining the three-dimensional space may be two or more, and the number of ellipsoids of revolution defining the three-dimensional space may be two or more. However, there may be a case where the movement of the user's viewpoint position outside the three-dimensional space is allowed.

[0148] Although not particularly limited, the first information element may be included in the scene description in the form shown in FIG. 17. In FIG. 17, the first information element may be represented by viewing_inside_object_flag. viewing_inside_object_flag may be set for each 3D object. For example, when viewing_inside_object_flag is "0", the movement of the viewpoint position into the 3D object may be restricted, and when viewing_inside_object_flag is "1", the movement of the viewpoint position into the 3D object may be permitted.

[0149] Although not particularly limited, the second information element may be included in the scene description in the manner shown in FIG. 18. In FIG. 18, the second information element may include information elements (cuboid_center_x, cuboid_center_y, cuboid_center_z, cuboid_size_x, cuboid_size_y, cuboid_size_z) that define a cuboid that defines a three-dimensional space. cuboid_center_x, cuboid_center_y, and cuboid_center_z are information elements indicating the center position of the cuboid, and cuboid_size_x, cuboid_size_y, and cuboid_size_z are information elements indicating the size of the cuboid. The second information element may include information elements (spheroid_center_x, spheroid_center_y, spheroid_center_z, spheroid_size_x, spheroid_size_y, spheroid_size_z) that define an ellipsoid of revolution that defines a three-dimensional space. spheroid_center_x, spheroid_center_y, and spheroid_center_z are information elements indicating the center position of the ellipsoid of revolution, and spheroid_size_x, spheroid_size_y, and spheroid_size_z are information elements indicating the size of the ellipsoid of revolution. Note that cuboid_enable is an information element indicating whether a three-dimensional space is defined by the cuboid, and spheroid_enable may be an information element indicating whether a three-dimensional space is defined by the ellipsoid of revolution. In FIG. 18, a case where a three-dimensional space is defined by two cuboids and two ellipsoids of revolution is illustrated.

[0150] Under such a premise, the media processing apparatus 200 (renderer 220) may execute the following operations.

[0151] First, when the user's viewpoint position moves outside the movement range, the renderer 220 may generate specific content with the intersection point of the trajectory of the user's viewpoint position and the boundary of the movement range as the viewpoint position. That is, the renderer 220 may fix the viewpoint position at the position (boundary position) when the viewpoint position is about to move outside the movement range.

[0152] Second, when the user's viewpoint position moves outside the movement range, the renderer 220 may notify the user that the movement of the viewpoint position is restricted. For example, the renderer 220 may display a message such as "You cannot move beyond this point."

[0153] (Operation and Effect) In the embodiment, the media processing device 200 generates specific content based on the viewpoint information and then transmits the specific content to the user terminal 300. According to such a configuration, it is not necessary for the user terminal 300 to generate specific content including the second content with a high degree of freedom of viewpoint. The user terminal 300 can present the specific content as long as it provides the viewpoint information to the media processing device 200. Therefore, although a delay occurs between the media processing device 200 and the user terminal 300, the processing load on the user terminal 300 can be reduced.

[0154] In Operation Example 1, the media processing device 200 transmits the viewpoint information used in generating the specific content to the user terminal 300 in association with the sequence number added to the specific content. According to such a configuration, the user terminal 300 can grasp the viewpoint information and the sequence number used in generating the specific content. Therefore, even in the case where it is assumed that the viewpoint information used when the media processing device 200 generates the specific content is different from the viewpoint information used when the user terminal 300 displays the specific content, the specific content can be appropriately displayed.

[0155] In Operation Example 2, the media processing device 200 receives the content configuration and specific viewpoint information (recommended viewport information) from the transmission device 100. According to such a configuration, the media processing device 200 can generate specific content based on the specific viewpoint information, and even in the case where the first user terminal 400 that feeds back the viewpoint information and the second user terminal 500 that does not feed back the viewpoint information are mixed, the specific content can be appropriately displayed.

[0156] In Operation Example 2, even when the user terminal is the first user terminal 400, the media processing device 200 generates specific content based on the specific viewpoint information in response to the reset signal. According to such a configuration, even in the case where the viewpoint position and the line-of-sight direction become unclear at the first user terminal 400 (the case of getting lost in the three-dimensional space) in the three-dimensional space constructed by the scene description, it is possible to return to the specific content based on the specific viewpoint information by the reset signal.

[0157] In Operation Example 3, the media processing device 200 receives, from the transmission device 100, quality information for each of two or more streams whose quality varies depending on the orientation of the 3D object based on the viewpoint information, as the quality information of the stream related to the 3D object included in the specific content. According to such a configuration, based on the new finding that the quality of each surface constituting the bounding box related to the 3D object does not have to be uniform, it is possible to appropriately display the 3D object while suppressing the transmission traffic.

[0158] In Operation Example 4, the media processing device 200 receives importance information for each of two or more objects included in the specific content from the transmission device 100. According to such a configuration, by introducing a mechanism for setting the importance for each object such as 360° video and 3D objects, it is possible to appropriately display each object included in the specific content while suppressing the transmission traffic.

[0159] In Operation Example 5, the media processing device 200 receives from the transmission device 100 an information element that defines the movement range of the user's viewpoint position in the three-dimensional space constituted by the specific content. According to such a configuration, it is possible to appropriately display specific content including content with a degree of freedom of viewpoint without causing a breakdown of the specific content displayed on the user terminal 300.

[0160] [Modification Example 1] Hereinafter, Modification Example 1 of the embodiment will be described. Hereinafter, the differences from the embodiment will be mainly described.

[0161] In Modification Example 1, when the specific content includes both the first content and the second content, a method for synchronizing the first content and the second content will be described.

[0162] Note that hereinafter, synchronization means that the presentation times of the first content (for example, MPU) and the second content (file) are appropriately aligned. Therefore, synchronization may include the alignment of the presentation times of a 2D video and a 3D object, and may also include the alignment of the presentation times of audio and a 3D object. Similarly, synchronization may include the alignment of the presentation times of a 2D video and a 360° video, and may also include the alignment of the presentation times of audio and a 360° video.

[0163] In the first method, a case where the media processing device 200 synchronizes the first content and the second content based on the first control information (MMT-SI) will be described. The media processing device 200 uses the MMT-SI as an entry point, checks the presence or absence of a scene description (second content), and when the scene description exists, diverts the MPU timestamp descriptor to specify the presentation time of the specific content including the first content and the second content.

[0164] Specifically, as shown in FIG. 19, since the 2D video and audio are presented based on the MPU timestamp descriptor (simply timestamp in FIG. 19), the 2D video and audio can be synchronized.

[0165] On the other hand, the presentation time of the first frame included in the scene description is specified by referring to the MPU timestamp descriptor included in the MMT-SI. The presentation times of the second and subsequent frames included in the scene description can be specified by the frame number included in the scene description and the frame rate of the second content. For example, considering the case where the frame rate is 30 fps, the presentation time of the nth frame is specified by adding 1 / 30×n to the time specified by the MPU timestamp descriptor. However, the frame number of the first frame included in the scene description is "0".

[0166] In the first method, the case where the presentation time of the first frame included in the scene description is not included in the scene description is exemplified, but the scene description may include the presentation time of the first frame included in the scene description.

[0167] In the second method, a case where the media processing device 200 synchronizes the first content and the second content based on the second control information (scene description) will be described. The media processing device 200 uses the scene description as an entry point, checks for the presence of the MMT-SI (first content), and if the MMT-SI exists, specifies the presentation times of the specific content including the first content and the second content based on the presentation times included in the scene description.

[0168] In such a case, the scene description includes absolute time information indicating the presentation time of the second content. The absolute time information may be the presentation time of the first frame included in the scene description.

[0169] For example, the absolute time information may be generated with UTC as the reference time. As the reference time, TAI may be used, or the time provided by GPS may be used. The reference time may be the time provided by an NTP server, or the time provided by a PTP server. Furthermore, the absolute time information may be generated based on the same reference time as the MPU time stamp descriptor.

[0170] Furthermore, the scene description includes reference information for identifying the first content. The reference information may be information for identifying the MPU that constitutes the first content. That is, the reference information is information for treating the first content (MPU) as an object included in the scene description.

[0171] Specifically, as shown in FIG. 20, the presentation time of the first frame included in the scene description is specified by the absolute time information included in the scene description. The presentation times of the second and subsequent frames included in the scene description can be specified by the frame number included in the scene description and the frame rate of the second content. For example, considering the case where the frame rate is 30 fps, the presentation time of the nth frame is specified by adding 1 / 30 × n to the time specified by the MPU time stamp descriptor. However, the frame number of the first frame included in the scene description is "0".

[0172] On the other hand, since the 2D video and audio are presented based on the MPU time stamp descriptor (simply "timestamp" in FIG. 20), the 2D video and audio can be synchronized. Here, since the above-described reference information is included in the scene description, the media processing device 200 can confirm the presence or absence of the first content to be presented together with the second content based on the reference information included in the scene description.

[0173] In the second method, the synchronization between the 2D video and the audio is performed based on the MPU time stamp descriptor included in the MMT-SI. However, in Modification Example 1, the synchronization between the 2D video and the audio may also be performed based on the information elements (absolute time information and reference information) included in the scene description. In such a case, at least the MPU time stamp descriptor included in the MMT-SI may be omitted. Furthermore, the MMT-SI itself may be omitted.

[0174] Note that when the reference time of the MPU time stamp descriptor included in the MMT-SI (hereinafter referred to as the first reference time) is different from the reference time of the absolute time information included in the scene description (the second reference time), at least one of the first control information (MMT-SI) and the second control information (scene description) may include conversion information between the first reference time and the second reference time. For example, the MMT-SI may include an MPU time stamp descriptor represented by the first reference time (e.g., UTC) in addition to an MPU time stamp descriptor represented by the second reference time (e.g., a reference time other than UTC). The scene description may include absolute time information represented by the first reference time (e.g., UTC) in addition to absolute time information represented by the second reference time (e.g., a reference time other than UTC).

[0175] Note that the MPU time stamp descriptor included in the MMT-SI may be referred to as first absolute time information, and the absolute time information included in the scene description may be referred to as second absolute time information.

[0176] [Modification Example 2] Hereinafter, Modification Example 2 of the embodiment will be described. Hereinafter, the differences from the embodiment will be mainly described.

[0177] In the above-described embodiment, the case where the information generated by the media processing device 200 using the viewpoint information is video has been described. In contrast, in Modification Example 2, the media processing device 200 generates, using the viewpoint information, not only video but also audio. For example, in Modification Example 2, a case where 6DoF is applied to audio is assumed.

[0178] (Media Processing Apparatus and User Terminal) Hereinafter, the media processing apparatus and the user terminal according to Modification Example 2 will be described. FIG. 21 is a block diagram showing a media processing apparatus 200 and a user terminal 300 according to Modification Example 2.

[0179] As shown in FIG. 21, the media processing apparatus 200 includes a reception unit 210, a video renderer 220A, an audio renderer 220B, a video encoding processing unit 230A, an audio encoding processing unit 230B, and a multiplexing processing unit 240.

[0180] The reception unit 210 receives viewpoint information in the same manner as in the embodiment. The viewpoint information includes an information element indicating the viewpoint position of the user of the user terminal 300 and an information element indicating the line-of-sight direction of the user of the user terminal 300. In Modification Example 2, the reception unit 210 constitutes a reception unit that receives viewpoint information from the user terminal 300. The reception unit 210 outputs the viewpoint information received from the user terminal 300 to the video renderer 220A and the audio renderer 220B. The timing of outputting the viewpoint information to the video renderer 220A and the audio renderer 220B may be the same.

[0181] The video renderer 220A generates a video constituting specific content based on the viewpoint information input from the reception unit 210. The video renderer 220A may perform the same processing as the renderer 220 described above for the video. The video renderer 220A outputs the video constituting the specific content to the video encoding processing unit 230A. The video renderer 220A may output the viewpoint information used for generating the specific content to the user terminal 300 (video renderer 330). The video may be read as a video signal or video data. In Modification Example 2, the video renderer 220A constitutes a first renderer that generates first information constituting specific content based on the viewpoint information.

[0182] The audio renderer 220B generates the audio that constitutes the specific content based on the viewpoint information input from the reception unit 210. The audio renderer 220B may execute the same processing as the renderer 220 described above for the audio. The audio renderer 220B outputs the audio that constitutes the specific content to the audio encoding processing unit 230B. The audio may be read as an audio signal or audio data. The audio may be read as audio or sound. In Modification 2, the audio renderer 220B constitutes a second renderer that generates second information that constitutes specific content based on the viewpoint information.

[0183] The video encoding processing unit 230A encodes the specific content (video) generated by the video renderer 220A. The video encoding processing unit 230A may execute the same processing as the encoding processing unit 230 described above for the video. The compression encoding method for the video may be HEVC or VVC.

[0184] The audio encoding processing unit 230B encodes the specific content (audio) generated by the audio renderer 220B. The audio encoding processing unit 230B may execute the same processing as the encoding processing unit 230 described above for the audio. The compression encoding method for the audio may be MPEG-4 Advanced Audio Coding or MPEG-H 3D Audio.

[0185] The multiplexing processing unit 240 multiplexes the output of the video encoding processing unit 230A and the output of the audio encoding processing unit 230B. In Modification 2, the multiplexing processing unit 240 may be an example of a transmission unit that transmits specific content including the first information and the second information to the user terminal 300. The multiplexing processing unit 240 may transmit an information element indicating the presentation time to the user terminal 300 together with the specific content.

[0186] As shown in FIG. 21, the user terminal 300 includes a detection unit 310, a video decoding processing unit 320A, an audio decoding processing unit 320B, a video renderer 330, and a separation processing unit 340.

[0187] The detection unit 310 detects the viewpoint position and the line-of-sight direction of the user, as in the embodiment.

[0188] The video decoding processing unit 320A performs the same processing as the decoding processing unit 320 described above for the video. The video decoding processing unit 320A outputs the decoded video to the video renderer 330.

[0189] The audio decoding processing unit 320B performs the same processing as the decoding processing unit 320 described above for the audio. The audio decoding processing unit 320B outputs the decoded specific content (audio).

[0190] The video renderer 330 may perform the same processing as the renderer 330 for the video. Similar to the embodiment, the video renderer 330 may generate specific content (video) with the viewpoint position and the line-of-sight direction corrected based on the difference between the viewpoint information received from the media processing device 200 and the viewpoint information input from the detection unit 310. The video renderer 330 outputs the generated specific content (video).

[0191] The separation processing unit 340 separates the video and the audio from the specific content. The separation processing unit 340 outputs the video constituting the specific content to the video decoding processing unit 320A and outputs the audio constituting the specific content to the audio decoding processing unit 320B. The video may be read as a video signal or video data. The audio may be read as an audio signal or audio data. The audio may be read as audio or sound.

[0192] In Modification 2, on the premise described above, a case where 6DoF is applied to the audio will be described. In such a case, a scene description is also introduced for the audio. The scene description of the audio is different from that of the video.

[0193] As shown in FIG. 22, the scene description of video includes information elements indicating the position, rotation, and size of 3D objects in a three-dimensional space. On the other hand, the scene description of audio includes information elements such as the position of the sound source in the three-dimensional space, the frequency of the sound output from the sound source, and the size of the sound output from the sound source. Further, as shown in FIG. 22, the scene description of audio includes information elements such as the position, size, and reflection coefficient of an object (e.g., a 3D object) that affects the propagation of sound in the three-dimensional space.

[0194] In a case where 6DoF is applied to audio, it is necessary to render the sound heard by the user based on the positional relationship between the sound source and the user. Therefore, in the rendering (generation) of audio, not only the audio signal itself but also the scene description of audio and the viewpoint information of the user are required.

[0195] Under such a background, the media processing device 200 may execute the following operations. In the following, different operations for the embodiments will be mainly described.

[0196] First, the reception unit 210 outputs the viewpoint information received from the user terminal 300 to the video renderer 220A and the audio renderer 220B. The timing of outputting the viewpoint information to the video renderer 220A and the audio renderer 220B may be the same. That is, the audio renderer 220B generates audio based on the viewpoint information used in the generation of video. In other words, the video renderer 220A and the audio renderer 220B share the viewpoint information received from the user terminal 300. Note that for audio, the line-of-sight direction (the direction in which the user is looking at an object) included in the viewpoint information may be read as the viewing direction (the direction in which the user hears the sound).

[0197] Second, the video renderer 220A acquires the scene description of video and the scene description of audio. The video renderer 220A outputs the scene description of audio to the audio renderer 220B.

[0198] Here, as shown in FIG. 23, the scene description of audio is <audioscene>Even if the video scene description and the audio scene description are written in one file, the video renderer 220A can separate the video scene description and the audio scene description from the file.

[0199] For example, as shown in Figure 23, the scene description is <audioscene>Under this node, elements such as Sources, Geometry, Transforms, Acoustics, Resources, Conditions, Updates, etc. These elements are known elements in 6DoF audio, so their details will be omitted.

[0200] Third, video renderer 220A controls audio renderer 220B based on predetermined control information. The predetermined control information may be information transmitted via an API (Application Programming Interface). The predetermined control information may be simply referred to as API. The API may include the information shown in FIG. 24. The predetermined control information may be in the form of OSC (Open Sound Control).

[0201] Init() is information that initializes audio renderer 220B. When audio renderer 220B receives Init(), it initializes its internal memory and internal clock and prepares to wait for the input of a new audio scene description. Alternatively, if an audio scene description has been received from video renderer 220A, audio renderer 220B may execute an operation to construct a three-dimensional space to be used for audio based on the audio scene description.

[0202] configure() is information that shares a timeline with audio renderer 220B. In other words, configure() is information that synchronizes the audio generated by audio renderer 220B with the video generated by video renderer 220A. The timeline may be represented by the elapsed time from the beginning of the content, by a clock of a specific frequency (for example, 27 MHz) from the beginning of the content, by absolute time (UTC), or by absolute time (UTC) and an offset.

[0203] Here, by sharing the timeline, the three-dimensional space is also shared. That is, configure() can be thought of as information that associates the three-dimensional space used by audio renderer 220B with the three-dimensional space used by video renderer 220A. By this association, audio renderer 220B performs association of 3D objects related to video and sound sources related to audio in the shared three-dimensional space.

[0204] In this way, because configure() allows the timeline and three-dimensional space to be shared, configure() can be thought of as information that realizes synchronized operation between video renderer 220A and audio renderer 220B.

[0205] It is expected that ultimately, video and audio will be synchronized on a frame-by-frame basis using timestamps or the like, but the rendering process increases the amount of information to be processed by the media processing device, and the processing times of video renderer 220A and audio renderer 220B differ. Therefore, aligning the timelines using configure() makes the operation of video renderer 220A and audio renderer 220B more efficient.

[0206] start() is information that instructs the audio renderer 220B to start operating. pause() is information that instructs the audio renderer 220B to temporarily stop operating. resume() is information that instructs the audio renderer 220B to resume operating (cancel the pause). stop() is information that instructs the audio renderer 220B to stop operating.

[0207] After video renderer 220A shares the timeline and three-dimensional space with audio renderer 220B using configure(), it outputs start(), pause(), resume(), and stop() to audio renderer 220B based on the scene description input to video renderer 220A. start() and resume() are examples of information that control the start of audio generation operation by audio renderer 220B. pause() and stop() are examples of information that control the stop of audio generation operation by audio renderer 220B.

[0208] Here, video renderer 220A may use viewpoint information received from user terminal 300 to instruct audio renderer 220B to render only the sounds generated by the sound sources required by the user.

[0209] In addition, with current video content, even if the video content is silent, an audio signal (silence) is present at all times, so if the synchronization operation of video renderer 220A and audio renderer 220B is first performed using configure(), synchronization will not be lost in subsequent operations.

[0210] On the other hand, with immersive content, there may be cases where the audio signal is not present at certain times, resulting in discontinuous audio signals. There may also be cases where an audio source outputs sound in response to events such as user actions or operations. In such cases, video renderer 220A can generate specific content while suppressing loss of synchronization by controlling audio renderer 220B using start(), pause(), resume(), and stop().

[0211] Fourth, the multiplexing processing unit 240 multiplexes the outputs of the video encoding processing unit 230A and the audio encoding processing unit 230B. The multiplexing processing unit 240 transmits specific content including video and audio corresponding to 6DoF to the user terminal 300. The transmission method of the specific content may be MPEG-2 Systems, MMT, or MPEG-DASH.

[0212] In Modification Example 2, the case where the second information is audio has been exemplified, but Modification Example 2 is not limited thereto. The second information may be information other than video and audio and may be information that requires viewpoint information. For example, the second information may be information related to touch (haptics). In such a case, audio may be read as touch or haptics. The signal output by the haptics renderer may be a touch signal such as vibration or force sensation.

[0213] (Operations and Effects) In Modification Example 2, the video renderer 220A that generates the video constituting the specific content and the audio renderer 220B that generates the audio constituting the specific content are provided separately. According to such a configuration, video and audio corresponding to 6DoF can be appropriately output.

[0214] In Modification Example 2, the video renderer 220A outputs the audio scene description and the predetermined control information to the audio renderer 220B. According to such a configuration, the mechanism by which the video renderer 220A and the audio renderer 220B cooperate is clarified, and video and audio corresponding to 6DoF can be appropriately output.

[0215] [Other Embodiments] Although the present invention has been described by the above disclosure, the discussions and drawings forming a part of this disclosure should not be understood as limiting the present invention. Various alternative embodiments, examples, and operation techniques will be apparent to those skilled in the art from this disclosure.

[0216] In the above disclosure, although a case where specific content includes both the first content and the second content has been exemplified, the above disclosure is not limited thereto. The specific content may include at least the second content.

[0217] Although not particularly mentioned in the above disclosure, terms related to MMT may be interpreted based on the content defined in ISO / IEC 23008-1, ARIB STD-B60, ARIB TR-B39, etc.

[0218] In the above disclosure, as the first absolute time information included in MMT-SI, an MPU time stamp descriptor has been exemplified. However, the above disclosure is not limited thereto. The first absolute time information included in MMT-SI may be an MPU extended time stamp descriptor.

[0219] Although not particularly mentioned in the above disclosure, the media processing device 200 may request a part of the second content from the transmission device 100 as necessary. According to such a configuration, it is possible to save the bandwidth associated with the transmission of the second content and suppress an increase in the processing load of the media processing device 200.

[0220] In the above disclosure, MMTP has been exemplified as the transmission method of the first content. However, the above disclosure is not limited thereto. The transmission method of the first content may be a method compliant with ISO / IEC 23009-1 (hereinafter, MPEG-DASH (Dynamic Adaptive Stream over HTTP)). In such a case, the first control information may be an MPD (Media Presentation Description). That is, in the above disclosure, MMT-SI may be read as an MPD.

[0221] Although not particularly mentioned in the above disclosure, "acquisition" may be read as "reception".

[0222] Although not particularly limited, Operation Example 2 may be expressed as follows. The transmission device 100 includes a transmission unit that transmits the configuration of content having a degree of freedom of viewpoint. The transmission unit transmits specific viewpoint information used for generating specific content including at least the content. The reception device includes a reception unit that receives the configuration of content having a degree of freedom of viewpoint. The reception unit receives specific viewpoint information used for generating specific content including at least the content. In such a case, the reception device may be the media processing device 200 or the user terminal 300.

[0223] Although not particularly limited, Operation Example 3 may be expressed as follows. The transmission device 100 includes a transmission unit that transmits the configuration of content having a degree of freedom of viewpoint. The transmission unit transmits quality information of a stream related to a three-dimensional object included in specific content including at least the content. The quality information includes quality information related to each of two or more streams whose quality varies depending on the orientation of the three-dimensional object. The reception device includes a reception unit that receives the configuration of content having a degree of freedom of viewpoint. The reception unit receives quality information of a stream related to a three-dimensional object included in specific content including at least the content. The quality information includes quality information related to each of two or more streams whose quality varies depending on the orientation of the three-dimensional object. In such a case, the reception device may be the media processing device 200 or the user terminal 300.

[0224] Although not particularly limited, Operation Example 4 may be expressed as follows. The transmission device 100 includes a transmission unit that transmits the configuration of content having a degree of freedom of viewpoint. The transmission unit transmits importance information related to each of two or more objects included in specific content including at least the content. The reception device includes a reception unit that receives the configuration of content having a degree of freedom of viewpoint. The reception unit receives importance information related to each of two or more objects included in specific content including at least the content. In such a case, the reception device may be the media processing device 200 or the user terminal 300.

[0225] Although not particularly limited, Operation Example 4 may be expressed as follows. The transmission device 100 includes a transmission unit that transmits the configuration of content having a degree of freedom of viewpoint. The transmission unit transmits an information element that defines a movement range of the user's viewpoint position in a three-dimensional space configured by specific content including at least the content. The reception device includes a reception unit that receives the configuration of content having a degree of freedom of viewpoint. The reception unit receives an information element that defines a movement range of the user's viewpoint position in a three-dimensional space configured by specific content including at least the content. In such a case, the reception device may be the media processing device 200 or the user terminal 300.

[0226] Although not particularly mentioned in the above disclosure, a program may be provided that causes a computer to execute each process performed by the transmission device 100, the media processing device 200, and the user terminal 300. Further, the program may be recorded on a computer-readable medium. By using a computer-readable medium, it is possible to install the program on a computer. Here, the computer-readable medium on which the program is recorded may be a non-transitory recording medium. The non-transitory recording medium is not particularly limited, and for example, it may be a recording medium such as a CD-ROM or a DVD-ROM.

[0227] Alternatively, a chip may be provided that includes a memory that stores a program for executing each process performed by the transmission device 100, the media processing device 200, and the user terminal 300, and a processor that executes the program stored in the memory.

[0228] [Supplementary Note] A first feature is a media processing apparatus including a receiving unit that receives viewpoint information from a user terminal, a first renderer that generates first information for configuring specific content including at least content having a degree of freedom of viewpoint based on the viewpoint information, a second renderer that generates second information for configuring the specific content based on the viewpoint information, and a transmitting unit that transmits the specific content including the first information and the second information to the user terminal.

[0229] A second feature is a media processing apparatus according to the first feature, in which the first information is video for configuring the specific content, and the second information is audio for configuring the specific content.

[0230] A third feature is a media processing apparatus according to the first or second feature, in which the second renderer generates the second information based on the viewpoint information used in the generation of the first information.

[0231] A fourth feature is a media processing apparatus according to at least any one of the first to third features, which controls the second renderer based on predetermined control information.

[0232] A fifth feature is a media processing apparatus according to at least any one of the first to fourth features, in which the first renderer controls an operation of sharing at least any one of a timeline and a three-dimensional space with the second renderer.

[0233] A sixth feature is a media processing apparatus according to at least any one of the first to fifth features, in which the first renderer controls at least any one of start and stop of an operation of generating the second information by the second renderer.

[0234] A seventh feature is a program that causes a computer to execute steps A of receiving viewpoint information from a user terminal, step B of generating first information for constructing specific content including at least content having a degree of freedom of viewpoint based on the viewpoint information, step C of generating second information for constructing the specific content based on the viewpoint information, and step D of transmitting the specific content including the first information and the second information to the user terminal.

Explanation of Signs

[0235] 10… Transmission system, 100… Transmission device, 200… Media processing device, 210… Reception unit, 220… Renderer, 230… Encoding processing unit, 260… Selection unit, 270… Selection unit, 300… User terminal, 310… Detection unit, 320… Decoding processing unit, 330… Renderer, 400… First user terminal, 500… Second user terminal< / audioscene> < / audioscene>

Claims

1. A receiving unit that receives viewpoint information from a user terminal; A first renderer that generates first information that constitutes specific content including at least content having a degree of freedom of viewpoint based on the viewpoint information; A second renderer that generates second information that constitutes the specific content based on the viewpoint information; A media processing apparatus comprising: a transmission unit that transmits the specific content including the first information and the second information to the user terminal.

2. The first information is video that constitutes the specific content, The second information is audio that constitutes the specific content, The media processing apparatus according to claim 1.

3. The second renderer generates the second information based on the viewpoint information used in the generation of the first information, The media processing apparatus according to claim 1.

4. The first renderer controls the second renderer based on predetermined control information, The media processing apparatus according to claim 1.

5. The first renderer controls an operation of sharing at least one of a timeline and a three-dimensional space with the second renderer, The media processing apparatus according to claim 1.

6. The first renderer controls at least one of start and stop of an operation of generating the second information by the second renderer, The media processing apparatus according to claim 1.

7. Step A of receiving viewpoint information from a user terminal; Step B of generating first information that constitutes specific content including at least content having a degree of freedom of viewpoint based on the viewpoint information; Step C of generating second information that constitutes the specific content based on the viewpoint information; A program that causes a computer to execute step D of transmitting the specific content including the first information and the second information to the user terminal.