Media processor, user terminal, and program
The described configuration ensures synchronized transmission of viewpoint information with content, addressing discrepancies in existing systems to achieve accurate rendering of 360° video and 3D objects on user terminals.
Patent Information
- Application Number
- JP2024004778
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-29
AI Technical Summary
Existing mechanisms for transmitting content with a degree of freedom of viewpoint, such as 360° video and 3D objects, face challenges due to discrepancies between the viewpoint information used in generation and display, leading to inappropriate rendering on user terminals.
A media processing apparatus and user terminal configuration that includes a receiving unit for viewpoint information, a renderer to generate specific content based on this information, and a transmission unit to send the viewpoint information along with the content, ensuring synchronized rendering.
Enables appropriate display of content with a degree of freedom of viewpoint by aligning the viewpoint information used in generation with display, thus improving rendering accuracy and consistency.
Smart Images

Figure 2025110755000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a media processing apparatus, a user terminal, and a program.
Background Art
[0002] Conventionally, a mechanism for transmitting contents such as 360° video and 3D objects has been proposed (for example, Non-Patent Document 1). As such mechanisms, 3DoF+ (Degree of Freedom) involving viewpoint movement within the range where the user moves their head while sitting, 6DoF involving viewpoint movement within the range where the user can move freely, etc. are known. In such mechanisms, the positional relationship between the 360° video and the 3D object is indicated by scene description.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Under the above-described background, a case is conceivable where, after a media processing apparatus generates specific content including content having a degree of freedom of viewpoint, the generated specific content is transmitted from the media processing apparatus to a user terminal.
[0005] As a result of intensive studies, the inventors have noted that, in the above-described case, it is assumed that the viewpoint information of the user terminal changes every moment, and the viewpoint information used when generating specific content in the media processing apparatus may be different from the viewpoint information used when displaying the specific content on the user terminal.
[0006] Therefore, the present invention has been made to solve the above-described problems, and an object thereof is to provide a media processing apparatus, a user terminal, and a program that enable appropriate display of specific content including content having a degree of freedom of viewpoint.
Means for Solving the Problems
[0007] An aspect of the disclosure includes a receiving unit that receives viewpoint information from a user terminal, a renderer that generates specific content including at least content having a degree of freedom of viewpoint based on the viewpoint information, and a transmission unit that transmits the specific content generated by the renderer to the user terminal. The transmission unit transmits the viewpoint information used in the generation of the specific content to the user terminal together with the specific content. It is a media processing apparatus.
[0008] An aspect of the disclosure includes a transmission unit that transmits viewpoint information to a media processing apparatus, and a reception unit that receives specific content including at least content having a degree of freedom of viewpoint, the specific content being generated by the media processing apparatus based on the viewpoint information. The reception unit receives the viewpoint information used in the generation of the specific content from the media processing apparatus together with the specific content. It is a user terminal.
Effects of the Invention
[0009] According to the present invention, it is possible to provide a media processing apparatus, a user terminal, and a program that enable appropriate display of specific content including content having a degree of freedom of viewpoint.
Brief Description of the Drawings
[0010] [Figure 1] FIG. 1 is a diagram showing a transmission system 10 according to an embodiment. [Diagram 2] FIG. 2 is a block diagram showing a media processing apparatus 200 and a user terminal 300 according to an embodiment. [Diagram 3]FIG. 3 is a diagram illustrating the second content according to the embodiment. [Figure 4] FIG. 4 is a diagram showing a method for viewing a specific content according to the embodiment. [Figure 5] FIG. 5 is a diagram for explaining the first operation example. [Figure 6] FIG. 6 is a diagram for explaining the second operation example. [Figure 7] FIG. 7 is a diagram for explaining the second operation example. [Figure 8] FIG. 8 is a diagram for explaining the third operation example. [Figure 9] FIG. 9 is a diagram for explaining the third operation example. [Figure 10] FIG. 10 is a diagram for explaining the third operation example. [Figure 11] FIG. 11 is a diagram for explaining the third operation example. [Figure 12] FIG. 12 is a diagram for explaining the fourth operation example. [Figure 13] FIG. 13 is a diagram for explaining the fourth operation example. [Figure 14] FIG. 14 is a diagram for explaining the fourth operation example. [Figure 15] FIG. 15 is a diagram for explaining the fifth operation example. [Figure 16] FIG. 16 is a diagram for explaining the fifth operation example. [Figure 17] FIG. 17 is a diagram for explaining the fifth operation example. [Figure 18] FIG. 18 is a diagram for explaining the fifth operation example. [Figure 19] FIG. 19 is a diagram for explaining the first method according to the first modification. [Figure 20] FIG. 20 is a diagram for explaining the second method according to the first modification. [Figure 21] FIG. 21 is a diagram for explaining the first method according to the second modification. [Figure 22] FIG. 22 is a diagram for explaining a second method according to the second modification. [Figure 23] FIG. 23 is a diagram illustrating viewpoint information according to the second modification. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Next, an embodiment of the present invention will be described. In the following description of the drawings, the same or similar parts are denoted by the same or similar reference numerals. However, it should be noted that the drawings are schematic and the ratios of the dimensions may differ from those of the actual parts.
[0012] Therefore, specific dimensions should be determined with reference to the following explanation. Of course, the dimensional relationships and ratios may differ between the drawings.
[0013] [Disclosure Summary] The outline of the disclosure is a media processing device comprising: a receiving unit that receives viewpoint information from a user terminal; a renderer that generates specific content including at least content with viewpoint freedom based on the viewpoint information; and a transmitting unit that transmits the specific content generated by the renderer to the user terminal, wherein the transmitting unit transmits the viewpoint information used in generating the specific content together with the specific content to the user terminal.
[0014] The outline of the disclosure is a user terminal that includes a transmitting unit that transmits viewpoint information to a media processing device, and a receiving unit that receives specific content that includes at least content with viewpoint freedom and that is generated by the media processing device based on the viewpoint information, and the receiving unit receives the viewpoint information used in generating the specific content from the media processing device together with the specific content.
[0015] In the summary of the disclosure, viewpoint information used in generating specific content is transmitted from a media processing device to a user terminal together with the specific content. According to such a configuration, since the user terminal can grasp the viewpoint information used in generating the specific content, even in a case where it is assumed that the viewpoint information used when generating the specific content by the media processing device is different from the viewpoint information used when displaying the specific content on the user terminal, the specific content can be appropriately displayed.
[0016] Note that the specific content generated by the media processing device is generated based on viewpoint information, and it should be noted that on the user terminal side, the video included in the specific content can be treated as a 2D video.
[0017] [Embodiment] (Transmission System) Hereinafter, the transmission system according to the embodiment will be described. FIG. 1 is a diagram showing a transmission system 10 according to the embodiment. As shown in FIG. 1, the digital wireless transmission system includes a transmission device 100, a media processing device 200, and a user terminal 300.
[0018] In the embodiment, the transmission device 100 transmits a first content having no degree of freedom of viewpoint and a second content having a degree of freedom of viewpoint to the media processing device 200. Further, the transmission device 100 transmits first control information associated with the first content and second control information associated with the second content to the media processing device 200.
[0019] The first content may include at least one of 2D video and audio. The first content and the first control information may be transmitted in a first method. The first method may be a method compliant with ISO / IEC 23008-1 (hereinafter, MMT (MPEG Media Transport)). In the following, a case where the first method is MMTP (MMT Protocol) compliant with MMT will be exemplified. The first control information may be referred to as MMT-SI (Signaling Information).
[0020] The second content may include 360° video and 3D objects. The second content and the second control information may be transmitted in a second method. Alternatively, it may be transmitted by a protocol such as HTTP (Hyper Text Transfer Protocol). The second content may comply with 3DoF+ (Degree of Freedom) with a viewpoint movement within the range where the user moves their head while sitting, 6DoF with a viewpoint movement within the range where the user moves freely, etc. Since the second content has a degree of freedom of viewpoint, at the same time (frame), it may include two or more 360° videos and may include two or more 3D objects. The second control information may be referred to as scene description.
[0021] Here, the second control information may be transmitted by the above-described first method. That is, the second control information may be transmitted by the same first method (for example, MMTP) as the first control information. Alternatively, it may be transmitted by a protocol such as HTTP.
[0022] The transmission from the transmission device 100 to the media processing device 200 is not particularly limited, and may be transmission using satellite broadcasting, may be transmission using the Internet network, or may be transmission using a mobile communication network.
[0023] Although not particularly limited, the transmission system may be a digital wireless transmission system. The digital wireless transmission system may be a system used in 4K and 8K satellite broadcasting.
[0024] Based on the viewpoint information received from the user terminal 300, the media processing device 200 generates specific content including at least the above-described second content, and transmits the generated specific content to the user terminal 300. Although not particularly limited, the transmission of the specific content may be transmission using the Internet network or may be transmission using a mobile communication network.
[0025] The user terminal 300 may be a user terminal such as a smartphone, a tablet terminal, or a head-mounted display. As shown in FIG. 1, two or more user terminals 300 may be provided as the user terminal 300. In other words, two or more user terminals 300 may request the media processing device 200 to generate specific content. Each user terminal 300 may transmit different viewpoint information to the media processing device 200.
[0026] (Media Processing Device and User Terminal) Hereinafter, the media processing device and the user terminal according to the embodiment will be described. FIG. 2 is a block diagram showing the media processing device 200 and the user terminal 300 according to the embodiment.
[0027] As shown in FIG. 2, the media processing device 200 includes a reception unit 210, a renderer 220, and an encoding processing unit 230.
[0028] The reception unit 210 receives viewpoint information. In the embodiment, the reception unit 210 constitutes a reception unit that receives viewpoint information from the user terminal 300. The viewpoint information includes an information element indicating the viewpoint position of the user of the user terminal 300 and an information element indicating the line-of-sight direction of the user of the user terminal 300.
[0029] The renderer 220 generates specific content including at least the second content based on the viewpoint information. Since the specific content is generated based on the viewpoint information, it may include one 360° video and one 3D object at the same time (frame). In the following, an example will be given in which the specific content includes the first content in addition to the second content.
[0030] First, the renderer 220 generates a first content including 2D video and audio as part of the specific content based on the first control information (MMT-SI). Viewpoint information is not required in generating the first content.
[0031] Specifically, the renderer 220 acquires the 2D video, audio, and MMT-SI in the form of MMTP packets in which the 2D video, audio, and MMT-SI are packetized.
[0032] For example, an MMTP packet is stored in an IP (Internet Protocol) packet, which may be transmitted using UDP (User Datagram Protocol) or TCP (Transmission Control Protocol).
[0033] Here, the first content is processed in units (hereinafter referred to as MPUs; Media Processing Units) that are separated by a fixed time width. An MPU includes one or more access units. An access unit may also be treated as an MFU (Media Fragment Unit). An MFU related to 2D video may be referred to as a NAL (Network Abstraction Layer) unit, and an MFU related to audio may be referred to as an MHAS (MPEG-H 3D Audio Stream) packet.
[0034] MMT-SI includes PA (Package Access) messages, and the PA messages include an MPT (MMT Package Table) indicating a list of first content. Further, MMT-SI includes an MPU timestamp descriptor indicating the presentation time of the first content. The MPU timestamp descriptor may mean the presentation time of the MPU, that is, the time of the access unit first presented in the MPU.
[0035] The MPU timestamp descriptor may be generated with UTC (Coordinated Universal Time) as the reference time. As the reference time, TAI (International Atomic Time) may be used, or the time provided from GPS (Global Positioning System) may be used. The reference time may be the time provided from an NTP (Network Time Protocol) server, or may be the time provided from a PTP (Precision Time Protocol) server.
[0036] Second, the renderer 220 generates second content including 360° video and 3D objects as part of specific content based on second control information (scene description). Viewpoint information is used in the generation of the second content.
[0037] Specifically, the renderer 220 may obtain the scene description in the form of an MMTP packet in which the scene description is packetized. The method for obtaining the 360° video and 3D objects is not particularly limited.
[0038] The 360° video may be converted into 2D video using a projective transformation such as ERP (Equirectangular Projection) or cube mapping. Metadata indicating the type of projective transformation applied to the 360° video may be added. The 3D object may be encoded in a mesh format. For mesh format encoding, ISO / IEC 14496-16 "Animation framework extension (AFX)" may be used. The 3D object may be encoded in a point cloud format. For point cloud format encoding, ISO / IEC 23090-5 "Video-based Point Cloud Compression" may be used.
[0039] Here, the second content is collected into a single file in units separated by a fixed time interval. The fixed time interval may be 500 ms. For example, if the frame rate is 60 fps (frames per second), one file contains 30 frames.
[0040] The scene description is generated for each file and includes information for specifying the 360-degree video and 3D objects for each frame. For example, the scene description includes an information element (object_name) indicating the name of the 3D object in the frame, an information element (frame_number) indicating the frame number, an information element (translation_object) indicating the position of the 3D object in the frame, an information element (rotation_object) indicating the rotation of the 3D object in the frame, and an information element (scale_object) indicating the size of the 3D object in the frame.
[0041] Third, the renderer 220 outputs the specific content including the first content and the second content to the encoding processing unit 230. The renderer 220 may output the presentation time of the specific content to the encoding processing unit 230 together with the specific content.
[0042] Here, the presentation time of the specific content may be corrected based on the delay time between the media processing device 200 and the user terminal 300. Specifically, the renderer 220 may calculate the presentation time (T') of the specific content provided from the media processing device 200 to the user terminal 300 based on the presentation time (T) and the delay time (ΔT) of the specific content provided from the transmission device 100 to the media processing device 200 (T' = T + ΔT). The delay time (ΔT) may be a value predetermined in the media processing device 200, or may be a different value for each user terminal 300.
[0043] Fourthly, the renderer 220 may constitute a transmission unit that transmits the viewpoint information used for generating the specific content to the user terminal 300. The viewpoint information used for generating the specific content may be transmitted from the encoding processing unit 230 to the user terminal 300.
[0044] For example, the transmission method of the viewpoint information and the specific content may be MMTP or HTTP. When MMTP is used as the transmission method of the specific content, the viewpoint information may be stored as metadata in the OMAF (Omnidirectional Media Format) defined in ISO / IEC 23090-2.
[0045] The encoding processing unit 230 encodes the specific content generated by the renderer 220. In an embodiment, the encoding processing unit 230 may be an example of a transmission unit that transmits the specific content to the user terminal 300.
[0046] Furthermore, the encoding processing unit 230 may encode the presentation time of the specific content. The encoding processing unit 230 may transmit an information element indicating the presentation time to the user terminal 300 together with the specific content.
[0047] Here, any compression encoding method can be used as the compression encoding method used by the encoding processing unit 230. For example, the compression encoding method may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding).
[0048] As described above, the second content included in the specific content is generated based on viewpoint information, and therefore the video included in the specific content can be treated as a 2D video that does not have freedom of viewpoint.
[0049] For example, the transmission control method used to start and end viewing of specific content may include RTSP (Real Time Streaming Protocol). The transmission method may be MMTP or HTTP. When MMTP is used as the transmission method, the specific content may be stored in OMAF defined in ISO / IEC 23090-2.
[0050] As shown in FIG. 2, the user terminal 300 includes a detection unit 310, a decoding unit 320, and a renderer 330.
[0051] The detector 310 detects the user's viewpoint position and line of sight direction. The detector 310 may include an acceleration sensor or a GPS (Global Positioning System) sensor. The detector 310 may also include a user I / F (for example, a touch sensor, keyboard, mouse, controller, etc.) that is manually input by the user. The detector 310 may transmit viewpoint information (viewpoint position and line of sight direction) to the media processing device 200. The detector 310 may output the viewpoint information (viewport) to the renderer 330.
[0052] The decoding processing unit 320 decodes the specific content received from the media processing device 200. The decoding processing unit 320 may decode the presentation time received from the media processing device 200. The decoding processing unit 320 may output the specific content to the renderer 330, or may output the presentation time to the renderer 330.
[0053] The renderer 330 outputs the specific content decoded by the decoding processing unit 320. The renderer 330 may output the specific content based on the presentation time decoded by the decoding processing unit 320. For example, the renderer 330 may output the video content included in the specific content to a display and output the audio content included in the specific content to a speaker.
[0054] Here, the renderer 330 may generate specific content with the viewpoint position and the line-of-sight direction corrected based on the difference between the viewpoint information received from the media processing device 200 and the viewpoint information input from the detection unit 310.
[0055] (Second Content) Hereinafter, the second content according to the embodiment will be described. Here, the second content at t = 0, t = 1, and t = 2 will be described. The time intervals of t = 0, t = 1, and t = 2 are not particularly limited.
[0056] For example, as shown in FIG. 3, at t = 0, a 360° video may be displayed without displaying a 3D object. The 360° video may be considered as the background video of the 3D object. At t = 1, a 3D object may be displayed in a form superimposed on the 360° video. Further, at t = 1, the position and rotation of the 3D object superimposed on the 360° video may be changed.
[0057] The above-described scene description includes information elements indicating the position, rotation, and size of the 3D object for each of t = 0, t = 1, and t = 2, and the 3D object can be appropriately superimposed on the 360° video.
[0058] (How to watch) A viewing method according to the embodiment will be described below, in which viewing of a specific content including a first content and a second content will be exemplified.
[0059] 4, in step S11, the user terminal 300 transmits an RTSP SETUP to the media processing device. The RTSP SETUP is a message indicating that viewing of specific content will begin.
[0060] Here, the RTSP SETUP includes the IP address of the user terminal 300, the listening port number, content identification information (content ID), etc. The RTSP SETUP may also include capability information of the user terminal 300 for viewing specific content. The capability information may include the frame rate, display resolution, etc. The display resolution may include the field of view (FoV). The capability information may also include information elements indicating the encoding method and compression method supported by the user terminal 300.
[0061] Here, a case is illustrated in which the capability information of the user terminal 300 is directly notified to the media processing device 200, but the embodiment is not limited to this. The capability information of the user terminal 300 may be notified to the transmitting device 100, and then the transmitting device 100 may notify the media processing device 200.
[0062] In step S12, the media processing device 200 transmits a response to the RTSP SETUP. Here, the response transmitted is an ACK indicating that the RTSP SETUP has been accepted.
[0063] In step S21, the user terminal 300 transmits initial viewpoint information to the media processing device 200. The initial viewpoint information may be transmitted in the MMT-SI format.
[0064] In step S22, the media processing device 200 generates initial specified content based on the initial viewpoint information (rendering process). For example, the media processing device 200 generates second content to be included in the initial specified content based on the initial viewpoint information and the scene description.
[0065] Here, media processing device 200 may generate initial specific content using a viewport that is wider than the display resolution of user terminal 300. For example, the range wider than the display resolution may be a range that is the display resolution + 20% in the horizontal direction and the display resolution + 20% in the vertical direction.
[0066] The media processing device 200 applies a compression encoding method to the initial specific content. Although not particularly limited, the compression encoding method may be HEVC or VVC.
[0067] In step S23, the media processing device 200 transmits the initial specified content corresponding to the initial viewpoint information to the user terminal 300. The media processing device 200 transmits the presentation time of the initial specified content to the user terminal 300. As described above, the presentation time (T') provided to the user terminal 300 may be determined based on the delay time (ΔT).
[0068] If a different value is used for the delay time (ΔT) for each user terminal 300, the media processing device 200 can specify this by including the transmission time of the RTSP SETUP in the RTSP SETUP described above.
[0069] The user terminal 300 outputs the specific content based on the presentation time (T'). The user terminal 300 may generate specific content with the viewpoint position and line of sight corrected based on the difference between the viewpoint information received from the media processing device 200 and the viewpoint information input from the detection unit 310.
[0070] In step S31, the user terminal 300 transmits viewpoint information to the media processing device 200. The viewpoint information may be transmitted in the form of MMT-SI. Here, the user terminal 300 may transmit the viewpoint information at a predetermined period (for example, 500 ms), or may transmit the viewpoint information in response to a change in at least one of the viewpoint position and the line-of-sight direction.
[0071] In step S32, the media processing device 200 generates specific content (rendering process) based on the viewpoint information received in step S31.
[0072] In step S33, the media processing device 200 transmits the specific content corresponding to the viewpoint information received in step S31 to the user terminal 300.
[0073] The processing of steps S31 to S33 is the same as the processing of steps S21 to S23, except that the viewpoint information received in step S31 is used instead of the initial viewpoint information. Therefore, the details of the processing of steps S31 to S33 are omitted. The processing of steps S31 to S33 may be repeated at a predetermined period, or may be repeated every time the user's viewpoint position or line-of-sight direction changes.
[0074] In step S41, the user terminal 300 transmits RTSP TEARDOWN to the media processing device. RTSP TEARDOWN is a message indicating the end of viewing specific content.
[0075] In step S42, the media processing device 200 transmits a response to RTSP TEARDOWN. Here, an ACK indicating that RTSP TEARDOWN has been received is transmitted as the response.
[0076] In FIG. 4, although the case where steps S11 and S12 are executed based on RTSP is illustrated, the embodiment is not limited thereto. Steps S11 and S12 may be executed based on MMTP or may be executed based on HTTP.
[0077] Similarly, although the case where steps S41 and S42 are executed based on RTSP is illustrated, the embodiment is not limited thereto. Steps S41 and S42 may be executed based on MMTP or may be executed based on HTTP.
[0078] In FIG. 4, although the case where steps S31 to S33 are executed based on MMTP is illustrated, the embodiment is not limited thereto. Steps S31 to S33 may be executed based on another method (for example, HTTP).
[0079] Similarly, although the case where steps S41 to S43 are executed based on MMTP is illustrated, the embodiment is not limited thereto. Steps S41 to S43 may be executed based on another method (for example, HTTP).
[0080] (Operation Example 1) The above-described embodiment may include Operation Example 1 shown below. In Operation Example 1, the media processing apparatus 200 transmits the viewpoint information used in the generation of the specific content to the user terminal 300 in association with the sequence number added to the specific content.
[0081] Specifically, the media processing apparatus 200 (renderer 220) generates second content including a 360° video and a 3D object as a part of the specific content based on the second control information (scene description) in the same manner as in the above-described embodiment. In the generation of the second content, the viewpoint information received from the user terminal 300 is used.
[0082] In operation example 1, the renderer 220 associates the viewpoint information used in the generation of specific content (here, the second content) with a sequence number. The renderer 220 transmits the viewpoint information used in the generation of the specific content to the user terminal 300 in association with the sequence number added to the specific content. The viewpoint information may be stored in a VP (View Port) message. The VP message may have the format of MMT-SI defined in ISO / IEC 23008-1, ARIB STD-B60, ARIB TR-B39, etc. The VP message may be transmitted for each frame. The VP message may be transmitted to the user terminal 300 as a message (MMT-SI) related to MMTP.
[0083] Although not particularly limited, the VP message may have the data structure shown in FIG. 5. As shown in FIG. 5, the VP message may include message_id, version, length, fov, viewpoint_pos_x, viewpoint_pos_y, viewpoint_pos_z, viewpoint_yaw, viewpoint_pitch, viewpoint_roll, viewport_width, viewport_height, mpu_sequence_number_flag, mpu_sequence_number, etc.
[0084] message_id is identification information indicating the VP message. message_id may be 0x0204.
[0085] version is information indicating the version of the MMTP protocol. version may be 0x00.
[0086] length is information indicating the length of the VP message.
[0087] fov is information indicating the field of view.
[0088] viewpoint_pos_x is information indicating the x coordinate of the viewpoint position. viewpoint_pos_x is an example of viewpoint information used in generating specific content.
[0089] viewpoint_pos_y is information indicating the y coordinate of the viewpoint position. viewpoint_pos_y is an example of viewpoint information used in generating specific content.
[0090] viewpoint_pos_z is information indicating the z coordinate of the viewpoint position. viewpoint_pos_z is an example of viewpoint information used in generating a specific content.
[0091] viewpoint_yaw is information indicating the yaw of the viewpoint position. viewpoint_yaw is an example of viewpoint information used in generating specific content.
[0092] viewpoint_pitch is information indicating the pitch of the viewpoint position. viewpoint_pitch is an example of viewpoint information used in generating a specific content.
[0093] viewpoint_roll is information indicating the role of the viewpoint position. viewpoint_roll is an example of viewpoint information used in generating specific content.
[0094] The viewport_width is information indicating the width of the display area (specific content).
[0095] The viewport_height is information indicating the height of the display area (specific content).
[0096] The mpu_sequence_number_flag is information indicating whether the mpu_sequence_number field exists. For example, if the mpu_sequence_number_flag is 1, the mpu_sequence_number field exists, and if the mpu_sequence_number_flag is 0, the mpu_sequence_number field does not necessarily exist.
[0097] The mpu_sequence_number is the MPU sequence number of the video corresponding to the specific content indicated by the VP message, and is an example of a sequence number added to the specific content.
[0098] Here, viewpoint_pos_x, viewpoint_pos_y, and viewpoint_pos_z are examples of information elements that indicate the user's viewpoint position in the three-dimensional space configured by the scene description. viewpoint_yaw, viewpoint_pitch, and viewpoint_roll are examples of information elements that indicate the user's line of sight in the three-dimensional space configured by the scene description. viewport_width and viewport_height are examples of information elements that indicate the number of pixels of the video included in specific content.
[0099] Under these conditions, the media processing device 200 and the user terminal 300 may perform the following operations.
[0100] First, the media processing device 200 (renderer 220) may specify viewpoint_pos_x, viewpoint_pos_y, and viewpoint_pos_z based on the viewpoint position used to generate the specific content. The renderer 220 may specify viewpoint_yaw, viewpoint_pitch, and viewpoint_roll based on the line-of-sight direction used to generate the specific content. The renderer 220 may specify viewport_width based on the number of pixels in the horizontal direction of the specific content, and may specify viewport_height based on the number of pixels in the vertical direction of the specific content.
[0101] The media processing device 200 (encoding processing unit 230) may perform compression encoding of the specific content generated by the renderer 220 and transmit the specific content. Here, the encoding processing unit 230 may transmit the VP message shown in FIG. 5 to the user terminal 300. That is, the encoding processing unit 230 may transmit the viewpoint information used in the generation of the specific content to the user terminal 300 in association with the sequence number added to the specific content.
[0102] Second, the user terminal 300 (decoding processing unit 320) may configure a receiving unit that receives, from the media processing device 200, the viewpoint information used in the generation of the specific content in association with the sequence number added to the specific content. That is, the decoding processing unit 320 may receive the VP message (MMT-SI) shown in FIG. 5 from the media processing device 200.
[0103] The user terminal 300 (renderer 330) may associate the video that constitutes the decoded specific content with the viewpoint information included in the VP message based on the sequence number (mpu_sequence_number). The renderer 330 may specify, from the display area defined by the information (viewport_width and viewport_height) included in the VP message, the video in the range specified by the viewpoint information detected by the detector 310 based on the difference between the viewpoint information detected by the detector 310 and the viewpoint information included in the VP message. The renderer 330 may display the specified video.
[0104] (Operation Example 2) The above-described embodiment may include the following Operation Example 2. In Operation Example 2, as shown in FIG. 6, a case where the first user terminal 400 that feeds back viewpoint information and the second user terminal 500 that does not feed back viewpoint information are mixed is assumed. The first user terminal 400 may be a terminal such as a head-mounted display. The first user terminal may have the same functions as the above-described user terminal 300. The second user terminal 500 may be a terminal such as a volumetric display.
[0105] Specifically, in operation example 2, as shown in FIG. 6 , the media processing device 200 (renderer 220) may constitute a receiving unit that receives content configuration and recommended viewport information from the transmitting device 100. The content configuration may be considered to include 2D video, audio, 360° video, and 3D objects. The content configuration may be considered to include MMT-SI and a scene description. The recommended viewport information may be considered to be an example of specific viewpoint information. The recommended viewport information may be information that defines (recommends) the position, direction, and angle of view at which to view video in a three-dimensional space formed by the specific content (a three-dimensional space constructed by the scene description). The recommended viewport information may include an information element that indicates at least one of an information element indicating a viewpoint position in the three-dimensional space formed by the specific content and an information element indicating a line-of-sight direction in the three-dimensional space formed by the specific content. The recommended viewport information may be considered to be viewpoint information primarily used by the second user terminal 500.
[0106] In the following, unless explicitly stated otherwise, the viewpoint information used in generating specific content may include specific viewpoint information (recommended viewport information) received from the transmitting device 100, or may include viewpoint information received from the user terminal 300.
[0107] Although not particularly limited, the recommended viewport information may be included in the scene description in the manner shown in Fig. 7. As shown in Fig. 7, the recommended viewport information may include camera_orientation, frame_number, translation, and yfov.
[0108] The camera_orientation is information that indicates the direction in which the image is viewed in the three-dimensional space constructed by the scene description. The camera_orientation may be considered synonymous with the line of sight.
[0109] The frame_number is information indicating the frame number of the video to which the camera_orientation, translation, and yfov are applied. The camera_orientation may be considered to be an example of specific viewpoint information.
[0110] Translation is information indicating the position at which an image is viewed in a three-dimensional space constructed by a scene description. Translation may be considered synonymous with viewpoint position. Translation may be considered an example of specific viewpoint information. For example, FIG. 7 illustrates a case where the viewpoint position is [0,0,-50] when the frame number is 0, and moves to [0,0,-75] when the frame number is 2505.
[0111] yfov is information that indicates the angle of view from which the video is viewed in the three-dimensional space constructed by the scene description.
[0112] Although not limited thereto, the camera_orientation and translation may be assigned by the content creator, or, assuming that the content is captured by a camera, the camera_orientation and translation may be assigned automatically by a GPS and a sensor provided on the camera.
[0113] Here, "translation" is an example of an information element that indicates the viewpoint position in the three-dimensional space configured by the specific content, and "camera_orientation" is an example of an information element that indicates the line of sight direction in the three-dimensional space configured by the specific content.
[0114] Under these conditions, the media processing device 200 and the user terminal 300 may perform the following operations.
[0115] First, the media processing device 200 (renderer 220) may generate specific content based on specific viewpoint information (recommended viewport information). The media processing device 200 (encoding processing unit 230) may perform compression encoding of the specific content generated by the renderer 220 and transmit the specific content. Here, the encoding processing unit 230 may transmit the specific content generated based on the specific viewpoint information to the first user terminal 400, and may also transmit the specific content generated based on the specific viewpoint information to the second user terminal 500.
[0116] Second, when the media processing device 200 (reception unit 210) is the first user terminal 400 that feeds back viewpoint information, it may receive the viewpoint information from the first user terminal 400. The media processing device (renderer 220) may generate specific content based on the viewpoint information received from the first user terminal 400. In such a case, the media processing device 200 (reception unit 210) may receive a reset signal from the first user terminal 400. The media processing device (renderer 220) may generate specific content based on the specific viewpoint information (recommended viewport information) in response to the reset signal.
[0117] In such a case, the first user terminal 400 may have the same configuration as the user terminal 300. However, the first user terminal 400 (detection unit 310) may have a function of detecting a reset signal. The detection unit 310 may detect a user operation for inputting the reset signal. The detection unit 310 may transmit the reset signal to the media processing device 200.
[0118] Third, when the media processing device 200 (renderer 220) is the second user terminal 500 that does not feed back viewpoint information, it may generate specific content based on the specific viewpoint information (recommended viewport information).
[0119] In such a case, the second user terminal 500 may not have the detection unit 310 that detects viewpoint information. The second user terminal 500 may not have the renderer 330. The second user terminal 500 may have a configuration similar to that of the user terminal 300, except that it does not have the detection unit 310 and the renderer 330.
[0120] (Operation Example 3) The above-described embodiment may include the following operation example 3. In operation example 3, the media processing device 200 (for example, the selection unit 260 described later) may constitute a receiving unit that receives, from the transmission device 100, quality information regarding each of two or more streams whose quality varies depending on the orientation of the 3D object based on the viewpoint information, as the quality information of the stream regarding the 3D object included in the specific content.
[0121] Specifically, in operation example 3, as shown in FIG. 8, in addition to the configuration shown in FIG. 2, the media processing device 200 has a selection unit 260. The selection unit 260 receives the scene description and the 3D object from the transmission device 100. The selection unit 260 inputs the selected stream (3D object) among two or more streams to the renderer 220. The selection unit 260 may request the transmission device 100 to transmit the selected stream. When assuming a case where the media processing device 200 transmits specific content to a plurality of user terminals 300, the selection unit 260 may request the transmission device 100 to transmit the streams required by each of the plurality of user terminals 300, or may request the transmission device 100 to transmit all the streams.
[0122] Here, the selection unit 260 receives, from the transmission device 100, quality information regarding each of two or more streams whose quality varies depending on the orientation of the 3D object based on the viewpoint information, as the quality information of the stream regarding the 3D object.
[0123] The quality information may be information indicating the relative quality of each face constituting a bounding box for a 3D object. The bounding box may be represented by a three-dimensional rectangle onto which the 3D object is projected. For example, as shown in FIG. 9, the bounding box may be defined by vertices A to H. In such a case, the faces of the bounding box include face #1 represented by vertices A, B, F, and E; face #2 represented by vertices B, C, G, and F; face #3 represented by vertices A, B, C, and D; face #4 represented by vertices E, F, G, and H; face #5 represented by vertices A, D, H, and E; and face #6 represented by vertices D, C, G, and H.
[0124] In such a case, when we imagine a 3D object being viewed in a 3D space constructed by a scene description, we assume that three sides are primarily observed, or in other words, that the remaining three sides are not observed very often.
[0125] In the third operational example, two or more streams with different qualities depending on the orientation of the 3D object based on viewpoint information are prepared as streams related to a 3D object included in specific content.
[0126] Although not particularly limited, the quality information may be included in the scene description in the manner shown in FIG. 10. FIG. 10 illustrates six streams with different orientations of 3D objects. The quality information may be expressed in the format of "quality" [#1, #2, #3, #4, #5, #6]. Note that within [ ], #1 to #6 represent the quality indexes of surfaces #1 to #6. The quality index may take on values ranging from 1 to 9. A larger value of the quality index may indicate higher quality. For example, in a stream identified by "id"="1", #1, #2, and #3 have high quality ("8"), while #4, #5, and #6 have low quality ("3"). In a stream identified by "id"="2", #1, #2, and #3 have low quality ("3"), while #4, #5, and #6 have high quality ("8")
[0127] Under these conditions, the media processing device 200 may perform the following operations: The following mainly describes the selection of a stream (3D object) selected from two or more streams.
[0128] In mode 1, as shown in the upper part of Fig. 11, the media processing device 200 (selector 260) may identify the vertex closest to the user's viewpoint (for example, vertex B), and then identify the three faces that have the closest vertex (for example, face #1 represented by vertices A, B, F, and E, face #2 represented by vertices B, C, G, and F, and face #3 represented by vertices A, B, C, and D). The selector 260 may select the stream that maximizes the sum of the quality indexes of the three identified faces (the stream identified by "id"="1" in the example shown in Fig. 10).
[0129] In mode 1, since the vertex closest to the user's viewpoint position is identified based on viewpoint information, it can be considered that the selection unit 260 selects the stream to be transmitted to the user terminal 300 from two or more streams based on the viewpoint information and quality information.
[0130] Mode 2 may be a mode applied when a 3D object is scaled down or enlarged. For example, as shown in the middle of FIG. 11 , the media processing device 200 (selector 260) may identify the vertex closest to the user's viewpoint (e.g., vertex B), and then identify the three faces that have the closest vertex (e.g., face #1 represented by vertices A, B, F, and E, face #2 represented by vertices B, C, G, and F, and face #3 represented by vertices A, B, C, and D). When a 3D object is scaled down, pixels of the 3D object are thinned out. Therefore, the selector 260 may select the stream with the smallest sum of the quality indexes of the three identified faces (the stream identified by "id"="2" in the example shown in FIG. 10). On the other hand, when a 3D object is scaled up, pixels of the 3D object are interpolated. Therefore, the selector 260 may select the stream with the largest sum of the quality indexes of the three identified faces (the stream identified by "id"="1" in the example shown in FIG. 10).
[0131] In mode 2, since the vertex closest to the user's viewpoint position is identified based on viewpoint information, it can be considered that the selection unit 260 selects the stream to be transmitted to the user terminal 300 from two or more streams based on the viewpoint information and quality information.
[0132] In mode 3, it may be a mode applied in a case where two 3D objects (3D object #1 and 3D object #2) overlap in the user's line of sight direction. Here, the selection of the stream related to 3D object #1 will be described. For example, as shown in the lower part of FIG. 11, the media processing device 200 (selection unit 260) identifies the vertex closest to the user's viewpoint position (for example, vertex B), and then may identify three faces having the closest vertex (for example, face #1 represented by vertices A, B, F, E; face #2 represented by vertices B, C, G, F; face #3 represented by vertices A, B, C, D). Here, 3D object #2 overlaps on the line segment connecting the vertex closest to the user's viewpoint position (for example, vertex B) and the user's viewpoint position, and the three identified faces are blocked by 3D object #2. Therefore, the selection unit 260 may select a stream with the minimum total quality index of the three identified faces (in the example shown in FIG. 10, the stream identified by "id" = "2").
[0133] Note that in mode 3, since the vertex closest to the user's viewpoint position is identified based on the viewpoint information, the selection unit 260 may be considered to select a stream to be transmitted to the user terminal 300 from among two or more streams based on the viewpoint information and the quality information. Further, in mode 3, since the overlap of the two 3D objects is identified based on the viewpoint information and the arrangement information of the 3D objects, the selection unit 260 may be considered to select a stream to be transmitted to the user terminal 300 from among two or more streams based on the viewpoint information, the quality information, and the arrangement information. The arrangement information of the 3D object (for example, "rotation_object", "scale_object", "translation_object" shown in FIG. 10) may be included in the scene description. The arrangement information of the 3D object may be considered to be "link_area" shown in FIG. 10. That is, the selection unit 260 may receive the arrangement information of the 3D object in the three-dimensional space from the transmission device 100.
[0134] In the quality information shown in Fig. 10, the sum of the quality indexes of the six aspects is the same for each stream. However, the embodiment is not limited to this. The sum of the quality indexes of the six aspects may be different between two or more streams.
[0135] (Example 4) The above-described embodiment may include the following Operation Example 4. Here, Operation Example 4 includes the following operations in addition to Operation Example 3. In Operation Example 4, media processing device 200 (for example, selection unit 270, described below) may constitute a receiving unit that receives importance information for each of two or more objects included in specific content from transmitting device 100.
[0136] Specifically, in operation example 4, as shown in Fig. 12, the media processing device 200 has a selection unit 270 in addition to the configuration shown in Fig. 2. The selection unit 270 receives a scene description, a 3D object, and a 360° video from the transmission device 100. The selection unit 270 inputs a stream (3D object) selected from two or more streams to the renderer 220.
[0137] Here, the selection unit 270 receives importance information regarding each of the two or more objects from the transmission device 100. The objects may include a 3D object and a 360-degree video.
[0138] For example, the importance information may be information indicating the relative importance between two or more objects. For example, as shown in FIG. 13, consider a case in which object A (background), object B (person), and object C (dog) exist in a three-dimensional space constructed by a scene description. Object A (background) is an example of a 360° image, and object B (person) and object C (dog) are examples of 3D objects. In such a case, the importance information may be information indicating the relative importance between each of object A (background), object B (person), and object C (dog).
[0139] Although not particularly limited, the importance information may be included in the scene description in the manner shown in FIG. 14. In FIG. 14, the importance information may be represented by weight. Weight may take a value in the range of 1 to 9. The larger the value of weight, the higher the importance may be meant. In FIG. 14, the weight ("9") of object A (background) identified by "object_id" = "0" is the highest, the weight ("3") of object B (person) identified by "object_id" = "1" is the lowest, and the case where the weight ("8") of object C (dog) identified by "object_id" = "2" is higher than the weight of object B (person) and lower than the weight of object A (background) is illustrated.
[0140] Under such a premise, the media processing apparatus 200 may execute the operations shown below. In the following, the selection of a stream (3D object) selected from among two or more streams will be mainly described.
[0141] First, the media processing apparatus 200 (selection unit 270) selects the stream with the highest quality for the 3D object with the highest importance. The method of selecting the stream may be the same as in operation example 3. For example, since the importance of object C (dog) is greater than the importance of object B (person), for object C (dog), the selection unit 270 selects the stream for which the sum of the quality indices of the three faces having the vertices closest to the user's viewpoint position is the largest.
[0142] Second, the media processing device 200 (selection unit 270) selects the stream with the lowest quality for 3D objects other than the 3D object with the highest importance. Subsequently, the selection unit 270 replaces the stream with the lowest quality with the stream with the highest quality within the range where the specific conditions are satisfied, in order from the 3D objects with high importance. The specific conditions may include a first condition that the bandwidth of the line from the transmission device 100 to the media processing device 200 is equal to or less than a threshold value, and may include a second condition that the processing load of the media processing device 200 is equal to or less than a threshold value. The specific conditions may be defined by a combination of the first condition and the second condition. For example, since the importance of object B (person) is smaller than the importance of object C (dog), the selection unit 270 selects the stream with the highest quality for object B (person) within the range where the specific conditions are satisfied.
[0143] As described above, it may be considered that the media processing device 200 (selection unit 270) selects the stream to be transmitted to the user terminal 300 from among two or more streams based on the viewpoint information, quality information, and importance information. It may be considered that the media processing device 200 (selection unit 270) selects the stream to be transmitted to the user terminal 300 from among two or more streams based on the viewpoint information, quality information, arrangement information, and importance information.
[0144] Note that in operation example 4, a case where there is one stream for 360° video has been illustrated. However, the embodiment is not limited to this. For 360° video as well, there may be two or more streams with different qualities.
[0145] In operation example 4, a case where there are two or more streams with different qualities depending on the orientation of the 3D object based on the viewpoint information for the 3D object has been illustrated. However, the embodiment is not limited to this. For the 3D object, there may be two or more streams with different qualities regardless of the orientation of the 3D object.
[0146] (Operation Example 5) The above-described embodiment may include the following Operation Example 3. In Operation Example 3, the media processing device 200 (e.g., the renderer 220) may constitute a receiving unit that receives, from the transmitting device 100, an information element that defines a range of movement of the user's viewpoint position in a three-dimensional space configured by specific content (a three-dimensional space constructed by a scene description).
[0147] First, in Operation Example 5, the information element may include an information element (hereinafter, referred to as a first information element) that restricts movement of the user's viewpoint position inside a 3D object included in specific content. For example, as shown in Fig. 15, in a case where a 3D object is placed in a three-dimensional space constructed by a scene description, movement of the viewpoint position inside the 3D object may be restricted. However, in a case where the 3D object is a building or another scene exists inside the 3D object, movement of the viewpoint position inside the 3D object may be permitted.
[0148] Second, in the fifth operational example, the information element may include an information element (hereinafter, referred to as a second information element) that restricts movement of the user's viewpoint outside the three-dimensional space. For example, as shown in FIG. 16, the three-dimensional space may be defined by a combination of a rectangular parallelepiped and a spheroid. The number of rectangular parallelepipeds that define the three-dimensional space may be two or more, and the number of spheroids that define the three-dimensional space may be two or more. However, there may be cases where movement of the user's viewpoint outside the three-dimensional space is permitted.
[0149] Although not particularly limited, the first information element may be included in the scene description in the form shown in FIG. 17. In FIG. 17, the first information element may be represented by viewing_inside_object_flag. viewing_inside_object_flag may be set for each 3D object. For example, when viewing_inside_object_flag is "0", the movement of the viewpoint position into the 3D object may be restricted, and when viewing_inside_object_flag is "1", the movement of the viewpoint position into the 3D object may be allowed.
[0150] Although not particularly limited, the second information element may be included in the scene description in the manner shown in Fig. 18. In Fig. 18, the second information element may include information elements (cuboid_center_x, cuboid_center_y, cuboid_center_z, cuboid_size_x, cuboid_size_y, cuboid_size_z) that define a rectangular parallelepiped that defines a three-dimensional space. The cuboid_center_x, cuboid_center_y, and cuboid_center_z are information elements that indicate the center position of the rectangular parallelepiped, and the cuboid_size_x, cuboid_size_y, and cuboid_size_z are information elements that indicate the size of the rectangular parallelepiped. The second information element may also include information elements (spheroid_center_x, spheroid_center_y, spheroid_center_z, spheroid_size_x, spheroid_size_y, spheroid_size_z) that define a spheroid that defines the three-dimensional space. spheroid_center_x, spheroid_center_y, and spheroid_center_z are information elements indicating the center position of a spheroid, and spheroid_size_x, spheroid_size_y, and spheroid_size_z are information elements indicating the size of the spheroid. Note that cuboid_enable is an information element indicating whether or not a three-dimensional space is defined by a rectangular parallelepiped, and spheroid_enable may be an information element indicating whether or not a three-dimensional space is defined by a spheroid. FIG. 18 illustrates a case in which a three-dimensional space is defined by two rectangular parallelepipeds and two spheroids.
[0151] Under these conditions, the media processing device 200 (renderer 220) may perform the following operations.
[0152] First, when the user's viewpoint position moves outside the movement range, the renderer 220 may generate specific content by using the intersection of the trajectory of the user's viewpoint position and the boundary of the movement range as the viewpoint position. In other words, the renderer 220 may fix the viewpoint position at the position (boundary position) at which the viewpoint position tries to move outside the movement range.
[0153] Second, when the user's viewpoint position moves outside the movement range, the renderer 220 may notify the user that movement of the viewpoint position is restricted. For example, the renderer 220 may display a message such as "You cannot move beyond this point."
[0154] (Action and effect) In the embodiment, media processing device 200 generates specific content based on viewpoint information and then transmits the specific content to user terminal 300. With this configuration, there is no need for user terminal 300 to generate specific content that includes second content with viewpoint flexibility; user terminal 300 can present the specific content simply by providing viewpoint information to media processing device 200. Therefore, although a delay occurs between media processing device 200 and user terminal 300, the processing load on user terminal 300 can be reduced.
[0155] In operation example 1, media processing device 200 associates the viewpoint information used in generating the specific content with the sequence number assigned to the specific content and transmits it to user terminal 300. This configuration allows user terminal 300 to grasp the viewpoint information and sequence number used in generating the specific content, and therefore allows the specific content to be displayed appropriately even in cases where the viewpoint information used when media processing device 200 generates the specific content differs from the viewpoint information used when user terminal 300 displays the specific content.
[0156] In operation example 2, the media processing device 200 receives the content configuration and the specific viewpoint information (recommended viewport information) from the transmission device 100. According to such a configuration, the media processing device 200 can generate specific content based on the specific viewpoint information, and even in the case where the first user terminal 400 that feeds back the viewpoint information and the second user terminal 500 that does not feed back the viewpoint information are mixed, the specific content can be appropriately displayed.
[0157] In operation example 2, even when the user terminal is the first user terminal 400, the media processing device 200 generates specific content based on the specific viewpoint information in response to the reset signal. According to such a configuration, even in the case where the viewpoint position and the line-of-sight direction become unclear at the first user terminal 400 in the three-dimensional space constructed by the scene description (the case of getting lost in the three-dimensional space), it is possible to return to the specific content based on the specific viewpoint information by the reset signal.
[0158] In operation example 3, the media processing device 200 receives, from the transmission device 100, quality information regarding each of two or more streams whose quality differs depending on the orientation of the 3D object based on the viewpoint information, as the quality information of the stream regarding the 3D object included in the specific content. According to such a configuration, based on the new finding that the quality of each surface constituting the bounding box regarding the 3D object does not have to be uniform, it is possible to appropriately display the 3D object while suppressing the transmission traffic.
[0159] In operation example 4, the media processing device 200 receives importance information regarding each of two or more objects included in the specific content from the transmission device 100. According to such a configuration, by introducing a mechanism for setting the importance for each object such as 360° video and 3D objects, it is possible to appropriately display each object included in the specific content while suppressing the transmission traffic.
[0160] In Operation Example 5, the media processing device 200 receives, from the transmission device 100, an information element that defines the movement range of the user's viewpoint position in the three-dimensional space configured by the specific content. According to such a configuration, it is possible to appropriately display specific content including content with a degree of freedom of viewpoint without causing a breakdown of the specific content displayed on the user terminal 300.
[0161] [Modification Example 1] Hereinafter, Modification Example 1 of the embodiment will be described. Hereinafter, the differences from the embodiment will be mainly described.
[0162] In Modification Example 1, when the specific content includes both the first content and the second content, a method for synchronizing the first content and the second content will be described.
[0163] Note that hereinafter, synchronization means that the presentation times of the first content (for example, MPU) and the second content (file) are appropriately aligned. Therefore, synchronization may include that the presentation times of a 2D video and a 3D object are aligned, or that the presentation times of audio and a 3D object are aligned. Similarly, synchronization may include that the presentation times of a 2D video and a 360° video are aligned, or that the presentation times of audio and a 360° video are aligned.
[0164] In the first method, a case where the media processing device 200 synchronizes the first content and the second content based on the first control information (MMT-SI) will be described. The media processing device 200 uses the MMT-SI as an entry point, checks the presence or absence of the scene description (second content), and when the scene description exists, diverts the MPU timestamp descriptor to specify the presentation time of the specific content including the first content and the second content.
[0165] Specifically, as shown in FIG. 19, since the 2D video and audio are presented based on the MPU timestamp descriptor (simply timestamp in FIG. 19), the 2D video and audio can be synchronized.
[0166] On the other hand, the presentation time of the first frame included in the scene description is specified by referring to the MPU timestamp descriptor included in the MMT-SI. The presentation times of the second and subsequent frames included in the scene description can be specified by the frame number included in the scene description and the frame rate of the second content. For example, considering the case where the frame rate is 30 fps, the presentation time of the nth frame is specified by adding 1 / 30 × n to the time specified by the MPU timestamp descriptor. However, the frame number of the first frame included in the scene description is "0".
[0167] In the first method, the case where the scene description does not include the presentation time of the first frame included in the scene description is exemplified, but the scene description may include the presentation time of the first frame included in the scene description.
[0168] In the second method, a case where the media processing device 200 synchronizes the first content and the second content based on the second control information (scene description) will be described. The media processing device 200 uses the scene description as an entry point, checks the presence or absence of the MMT-SI (first content), and when the MMT-SI exists, specifies the presentation times of the specific content including the first content and the second content based on the presentation times included in the scene description.
[0169] In such a case, the scene description includes absolute time information indicating the presentation time of the second content. The absolute time information may be the presentation time of the first frame included in the scene description.
[0170] For example, the absolute time information may be generated with UTC as the reference time. The reference time may be TAI, or the time provided by GPS. The reference time may also be the time provided by an NTP server or a PTP server. Furthermore, the absolute time information may be generated based on the same reference time as the MPU timestamp descriptor.
[0171] Furthermore, the scene description includes reference information for identifying the first content. The reference information may be information for identifying the MPU that constitutes the first content. That is, the reference information is information for treating the first content (MPU) as an object included in the scene description.
[0172] Specifically, as shown in FIG. 20, the presentation time of the first frame included in the scene description is specified by the absolute time information included in the scene description. The presentation times of the second and subsequent frames included in the scene description can be specified by the frame number included in the scene description and the frame rate of the second content. For example, considering the case where the frame rate is 30 fps, the presentation time of the nth frame is specified by adding 1 / 30 × n to the time specified by the MPU timestamp descriptor. However, the frame number of the first frame included in the scene description is "0".
[0173] On the other hand, since the 2D video and audio are presented based on the MPU timestamp descriptor (simply "timestamp" in FIG. 20), the 2D video and audio can be synchronized. Here, since the above-described reference information is included in the scene description, the media processing device 200 can confirm the presence or absence of the first content to be presented together with the second content based on the reference information included in the scene description.
[0174] In the second method, synchronization between 2D video and audio is achieved based on the MPU timestamp descriptor included in the MMT-SI, but in Modification 1, synchronization between 2D video and audio may also be achieved based on information elements included in the scene description (absolute time information and reference information). In such a case, at least the MPU timestamp descriptor included in the MMT-SI may be omitted. Furthermore, the MMT-SI itself may be omitted.
[0175] Note that, when the reference time of the MPU timestamp descriptor included in the MMT-SI (hereinafter referred to as the first reference time) differs from the reference time of the absolute time information included in the scene description (the second reference time), at least one of the first control information (MMT-SI) and the second control information (scene description) may include conversion information between the first reference time and the second reference time. For example, the MMT-SI may include an MPU timestamp descriptor expressed in the second reference time (e.g., a reference time other than UTC) in addition to an MPU timestamp descriptor expressed in the first reference time (e.g., UTC). The scene description may include absolute time information expressed in the first reference time (e.g., UTC) in addition to absolute time information expressed in the second reference time (e.g., a reference time other than UTC).
[0176] The MPU timestamp descriptor included in the MMT-SI may be referred to as first absolute time information, and the absolute time information included in the scene description may be referred to as second absolute time information.
[0177] [Change Example 2] Modification 2 of the embodiment will be described below, focusing mainly on the differences from the embodiment.
[0178] Specifically, in the above-described embodiment (operation example 1), the media processing device 200 transmits the viewpoint information used in the generation of the specific content to the user terminal 300 in association with the sequence number added to the specific content. In contrast, in modification example 2, the media processing device 200 transmits the viewpoint information used in the generation of the specific content to the user terminal 300 as an integral part of the specific content.
[0179] Here, the mode in which the viewpoint information is integral with the specific content may be any mode in which the viewpoint information used in the generation of the specific content can be extracted in association with the specific content at the user terminal 300. As a method for realizing such a mode, the following methods can be considered.
[0180] In the first method, as shown in FIG. 21, the media processing device 200 (encoding processing unit 230) may concatenate the viewpoint information used in the generation of the specific content with the specific content and transmit it to the user terminal 300.
[0181] For example, a case where the file format used for transmission of the specific content is ISO Base Media File Format (ISOBMFF) will be exemplified. In ISOBMFF, the video data (stream) constituting the specific content is included in the movie data box (mdat). Metadata and the like related to the video data (stream) are included in the movie fragment box (moof). In ISOBMFF, the viewpoint information used in the generation of the specific content may be included in the movie fragment box (moof).
[0182] On the premise described above, the media processing device 200 may concatenate the movie fragment box (moof) with the movie data box (mdat) and transmit it to the user terminal 300. The concatenation may include a mode of aggregating the movie fragment box (moof) and the movie data box (mdat) by HTTP, or may include a mode of including the movie fragment box (moof) and the movie data box (mdat) in the MPU of MMT.
[0183] In the first method, as the movie fragment box (moof) and the movie data box (mdat), data structures defined in ISO / IEC 14496-12 “Information technology - Coding of audio-visual objects - Part 12: ISO base media file format” may be used.
[0184] In the first method, transmission may be executed in units called segments that concatenate the movie fragment box (moof) and the movie data box (mdat). The segment may be a segment defined in ISO / IEC 23009-1 “Information technology - Dynamic adaptive streaming over HTTP (DASH) Part 1: Media presentation description and segment formats”.
[0185] Although not particularly limited, the viewpoint information may be added in units of GOP (Group Of Picture). For example, the GOP may be composed of 30 frames, and one piece of viewpoint information may be added to the 30 frames. In such a case, the viewpoint information may be included in the movie fragment box (moof) concatenated to the first frame among the frames constituting the GOP.
[0186] Although not particularly limited, the viewpoint information may be added in units of the frames constituting the GOP. For example, the GOP may be composed of 30 frames, and one piece of viewpoint information may be added to the target frames of the 30 frames. The target frames may be all of the 30 frames or a part of the 30 frames. In such a case, the viewpoint information may be included in the movie fragment box (moof) concatenated to the target frames among the frames constituting the GOP.
[0187] According to the first method, when the transmission method is MMTP or HTTP, the viewpoint information used to generate the specific content can be linked to the video data (stream) that constitutes the specific content.
[0188] In the second method, as shown in FIG. 22, the media processing device 200 (encoding processing unit 230) may multiplex the viewpoint information used in generating the specific content onto the specific content and transmit it to the user terminal 300.
[0189] For example, a case will be exemplified in which Supplemental Enhancement Information (SEI) or Video Usability Information (VUI) is assumed as specific information multiplexed onto video data (stream) constituting specific content. The SEI and VUI may be information specified in ISO / IEC 23002-7 "Information technology - MPEG video technologies Part 7: Versatile supplemental enhancement information messages for coded video bitstreams." Viewpoint information used in generating the specific content may be included in the SEI or VUI.
[0190] According to the second method, when the compression encoding method is HEVC or VVC, the viewpoint information used to generate the specific content can be multiplexed onto the video data (stream) that constitutes the specific content.
[0191] Although not particularly limited, the SEI or VUI may be multiplexed in units of frames constituting a GOP. Whether viewpoint information is included in the SEI or VUI may be identifiable by the header of the NAL unit. If viewpoint information is not included in the SEI or VUI multiplexed onto a predetermined frame, viewpoint information may be applied to the predetermined frame in the SEI or VUI multiplexed onto a frame prior to the predetermined frame. The previous frame may be the frame immediately preceding the predetermined frame, or may be two or more frames prior to the predetermined frame.
[0192] In the second method, the viewpoint information is mainly multiplexed onto video data constituting specific content. However, the viewpoint information may be multiplexed onto audio data constituting specific content. The term "audio" may be interpreted as "audio." The audio data may be in a format compatible with MHAS (MPEG-H 3D Audio Stream). MHAS may be an audio compression format defined in ISO / IEC 23008-3 "Information technology - High efficiency coding and media delivery in heterogeneous environments Part 3: 3D audio."
[0193] In Modification 2, the viewpoint information may have a configuration similar to that included in the above-mentioned VP message (see FIG. 5). For example, as shown in FIG. 23, the viewpoint information may include fov, viewpoint_pos_x, viewpoint_pos_y, viewpoint_pos_z, viewpoint_yaw, viewpoint_pitch, viewpoint_roll, viewport_width, viewport_height, etc.
[0194] fov is information indicating the field of view.
[0195] viewpoint_pos_x is information indicating the x - coordinate of the viewpoint position. viewpoint_pos_x is an example of the viewpoint information used in the generation of specific content.
[0196] viewpoint_pos_y is information indicating the y - coordinate of the viewpoint position. viewpoint_pos_y is an example of the viewpoint information used in the generation of specific content.
[0197] viewpoint_pos_z is information indicating the z - coordinate of the viewpoint position. viewpoint_pos_z is an example of the viewpoint information used in the generation of specific content.
[0198] viewpoint_yaw is information indicating the yaw of the viewpoint position. viewpoint_yaw is an example of the viewpoint information used in the generation of specific content.
[0199] viewpoint_pitch is information indicating the pitch of the viewpoint position. viewpoint_pitch is an example of the viewpoint information used in the generation of specific content.
[0200] viewpoint_roll is information indicating the roll of the viewpoint position. viewpoint_roll is an example of the viewpoint information used in the generation of specific content.
[0201] viewport_width is information indicating the width of the display area (specific content).
[0202] viewport_height is information indicating the height of the display area (specific content).
[0203] (Function and effect) In Modification 2, media processing device 200 transmits the viewpoint information used in generating the specific content together with the specific content to user terminal 300. With this configuration, it is possible to extract the viewpoint information used in generating the specific content in user terminal 300 in association with the specific content, without using sequence numbers as in the embodiment (Operation Example 1) described above. Therefore, even if the viewpoint information used when generating the specific content in media processing device 200 differs from the viewpoint information used when displaying the specific content on user terminal 300, the specific content can be displayed appropriately.
[0204] [Other embodiments] Although the present invention has been described by the above disclosure, the descriptions and drawings that form part of this disclosure should not be understood as limiting the present invention. From this disclosure, various alternative embodiments, examples, and operating techniques will become apparent to those skilled in the art.
[0205] In the above disclosure, a case where the specific content includes both the first content and the second content has been exemplified, but the above disclosure is not limited to this. The specific content may include at least the second content.
[0206] Although not specifically mentioned in the above disclosure, terms related to MMT may be interpreted based on the contents specified in ISO / IEC 23008-1, ARIB STD-B60, ARIB TR-B39, etc.
[0207] In the above disclosure, an MPU timestamp descriptor is exemplified as the first absolute time information included in the MMT-SI. However, the above disclosure is not limited to this. The first absolute time information included in the MMT-SI may also be an MPU extended timestamp descriptor.
[0208] Although not particularly mentioned in the above disclosure, the media processing device 200 may, if necessary, request a part of the second content from the transmission device 100. According to such a configuration, it is possible to save the bandwidth associated with the transmission of the second content and suppress an increase in the processing load of the media processing device 200.
[0209] In the above disclosure, MMTP was exemplified as the transmission method of the first content. However, the above disclosure is not limited to this. The transmission method of the first content may be a method compliant with ISO / IEC 23009-1 (hereinafter, MPEG-DASH (Dynamic Adaptive Stream over HTTP)). In such a case, the first control information may be an MPD (Media Presentation Description). That is, in the above disclosure, MMT-SI may be read as MPD.
[0210] Although not particularly mentioned in the above disclosure, "acquisition" may be read as "reception".
[0211] Although not particularly limited, the operation example 2 may be expressed as follows. The transmission device 100 includes a transmission unit that transmits the configuration of the content having a degree of freedom of viewpoint, and the transmission unit transmits specific viewpoint information used for generating specific content including at least the content. The reception device includes a reception unit that receives the configuration of the content having a degree of freedom of viewpoint, and the reception unit receives specific viewpoint information used for generating specific content including at least the content. In such a case, the reception device may be the media processing device 200 or the user terminal 300.
[0212] Although not particularly limited, Operation Example 3 may be expressed as follows. The transmission device 100 includes a transmission unit that transmits the configuration of the content having a degree of freedom of viewpoint. The transmission unit transmits the quality information of the stream related to the three-dimensional object included in the specific content including at least the content. The quality information includes the quality information related to each of two or more streams whose quality varies depending on the orientation of the three-dimensional object. The reception device includes a reception unit that receives the configuration of the content having a degree of freedom of viewpoint. The reception unit receives the quality information of the stream related to the three-dimensional object included in the specific content including at least the content. The quality information includes the quality information related to each of two or more streams whose quality varies depending on the orientation of the three-dimensional object. In such a case, the reception device may be the media processing device 200 or the user terminal 300.
[0213] Although not particularly limited, Operation Example 4 may be expressed as follows. The transmission device 100 includes a transmission unit that transmits the configuration of the content having a degree of freedom of viewpoint. The transmission unit transmits the importance information related to each of two or more objects included in the specific content including at least the content. The reception device includes a reception unit that receives the configuration of the content having a degree of freedom of viewpoint. The reception unit receives the importance information related to each of two or more objects included in the specific content including at least the content. In such a case, the reception device may be the media processing device 200 or the user terminal 300.
[0214] Although not particularly limited, Operation Example 4 may be expressed as follows. The transmission device 100 includes a transmission unit that transmits the configuration of content having a degree of freedom of viewpoint. The transmission unit transmits an information element that defines the movement range of the user's viewpoint position in a three-dimensional space configured by specific content including at least the content. The reception device includes a reception unit that receives the configuration of content having a degree of freedom of viewpoint. The reception unit receives an information element that defines the movement range of the user's viewpoint position in a three-dimensional space configured by specific content including at least the content. In such a case, the reception device may be the media processing device 200 or the user terminal 300.
[0215] Although not particularly mentioned in the above disclosure, a program may be provided to cause a computer to execute each process performed by the transmission device 100, the media processing device 200, and the user terminal 300. Further, the program may be recorded on a computer-readable medium. By using a computer-readable medium, it is possible to install the program on a computer. Here, the computer-readable medium on which the program is recorded may be a non-transitory recording medium. The non-transitory recording medium is not particularly limited, and for example, it may be a recording medium such as a CD-ROM or a DVD-ROM.
[0216] Alternatively, a chip may be provided that includes a memory that stores a program for executing each process performed by the transmission device 100, the media processing device 200, and the user terminal 300, and a processor that executes the program stored in the memory.
[0217] [Appendix] A first feature is a media processing device comprising: a receiving unit that receives viewpoint information from a user terminal; a renderer that generates specific content including at least content with viewpoint freedom based on the viewpoint information; and a transmitting unit that transmits the specific content generated by the renderer to the user terminal, wherein the transmitting unit transmits the viewpoint information used in generating the specific content together with the specific content to the user terminal.
[0218] A second feature is the media processing device according to the first feature, wherein the transmission unit links the viewpoint information used in generating the specific content to the specific content and transmits the linked viewpoint information to the user terminal.
[0219] A third feature is the media processing device according to the first feature, wherein the transmission unit multiplexes the viewpoint information used in generating the specific content onto the specific content and transmits the multiplexed viewpoint information to the user terminal.
[0220] A fourth feature is a user terminal comprising: a transmitter that transmits viewpoint information to a media processing device; and a receiver that receives specific content that includes at least content with viewpoint freedom and that is generated by the media processing device based on the viewpoint information, wherein the receiver receives the viewpoint information used in generating the specific content together with the specific content from the media processing device.
[0221] A fifth feature is a program that causes a computer to execute the steps of: receiving viewpoint information from a user terminal; generating specific content including at least content with viewpoint freedom based on the viewpoint information; and transmitting the specific content generated by the renderer to the user terminal, wherein step C includes a step of transmitting the viewpoint information used in generating the specific content together with the specific content to the user terminal.
[0222] The sixth feature is a program that causes a computer to execute: step A of transmitting viewpoint information to a media processing device; and step B of receiving specific content including at least content having a degree of freedom of viewpoint, the specific content being generated by the media processing device based on the viewpoint information, wherein step B includes the step of receiving, from the media processing device, the viewpoint information used in the generation of the specific content together with the specific content.
Explanation of Signs
[0223] 10…Transmission system, 100…Transmission device, 200…Media processing device, 210…Reception unit, 220…Renderer, 230…Encoding processing unit, 260…Selection unit, 270…Selection unit, 300…User terminal, 310…Detection unit, 320…Decoding processing unit, 330…Renderer, 400…First user terminal, 500…Second user terminal
Claims
1. A receiving unit that receives viewpoint information from a user terminal; A renderer that generates specific content including at least content having a degree of freedom of viewpoint based on the viewpoint information; A transmitting unit that transmits the specific content generated by the renderer to the user terminal, and The transmitting unit transmits the viewpoint information used in the generation of the specific content to the user terminal together with the specific content, a media processing device.
2. The transmitting unit connects the viewpoint information used in the generation of the specific content to the specific content and transmits it to the user terminal, the media processing device according to claim 1.
3. The transmitting unit multiplexes the viewpoint information used in the generation of the specific content with the specific content and transmits it to the user terminal, the media processing device according to claim 1.
4. A transmitting unit that transmits viewpoint information to a media processing device; A receiving unit that receives specific content including at least content having a degree of freedom of viewpoint, the specific content being generated by the media processing device based on the viewpoint information, and The receiving unit receives the viewpoint information used in the generation of the specific content from the media processing device together with the specific content, a user terminal.
5. Step A of receiving viewpoint information from a user terminal; Step B of generating specific content including at least content having a degree of freedom of viewpoint based on the viewpoint information; Step C of causing a computer to execute transmitting the specific content generated by the renderer to the user terminal, and Step C includes a step of transmitting the viewpoint information used in the generation of the specific content to the user terminal together with the specific content, a program.
6. Step A of transmitting viewpoint information to a media processing device; Step B of causing a computer to execute receiving specific content including at least content having a degree of freedom of viewpoint, the specific content being generated by the media processing device based on the viewpoint information, and Step B includes a step of receiving the viewpoint information used in the generation of the specific content from the media processing device together with the specific content, a program.