Media processing device, transmitting device, and receiving device
The media processing device addresses the issue of non-uniform quality in 3D object surfaces by generating and transmitting content with varying stream qualities, ensuring efficient and high-quality display of 3D objects with reduced transmission traffic.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-10
- Publication Date
- 2026-04-02
AI Technical Summary
Existing mechanisms for transmitting 3D objects and 360° video fail to ensure uniform quality of each surface constituting the bounding box, leading to inefficient display and increased transmission traffic.
A media processing device that receives viewpoint information from a user terminal, generates specific content with varying stream qualities based on orientation, and transmits these streams to the user terminal, allowing appropriate display of 3D objects while minimizing transmission traffic.
Enables appropriate display of 3D objects with reduced transmission traffic by adapting stream quality based on the orientation of the 3D object, optimizing resource usage and improving display quality.
Smart Images

Figure 0007839649000001 
Figure 0007839649000002 
Figure 0007839649000003
Abstract
Description
Technical Field
[0001] The present invention relates to a media processing device, a transmission device, and a reception device.
Background Art
[0002] Conventionally, a mechanism for transmitting contents such as 360° video and 3D objects has been proposed (for example, Non-Patent Document 1). As such mechanisms, 3DoF+ (Degree of Freedom) involving viewpoint movement within the range where the user moves their head while sitting, 6DoF involving viewpoint movement within the range where the user can move freely, etc. are known. In such mechanisms, the positional relationship between the 360° video and the 3D object is indicated by scene description.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Under the above-described background, when assuming specific contents including contents having a degree of freedom of viewpoint, a case where, after being generated by a media processing device, the generated specific contents are transmitted from the media processing device to a user terminal can be considered.
[0005] As a result of intensive studies, the inventors have focused on the fact that a part of the 3D object is not displayed according to the viewpoint information, and have found that the quality of each surface constituting the bounding box regarding the 3D object does not have to be uniform.
[0006] Therefore, the present invention has been made to solve the above-mentioned problems, and aims to provide a media processing device, a transmitting device, and a receiving device that enable the proper display of 3D objects while suppressing transmission traffic. [Means for solving the problem]
[0007] One aspect of the disclosure is a media processing device comprising: a receiving unit that receives viewpoint information from a user terminal; a renderer that generates specific content including at least content having degrees of freedom of viewpoint based on the viewpoint information; and a transmitting unit that transmits the specific content generated by the renderer to the user terminal, wherein the receiving unit receives quality information from the transmitting device for each of two or more streams whose quality differs depending on the orientation of the three-dimensional object based on the viewpoint information, as quality information for the stream relating to the three-dimensional object included in the specific content.
[0008] One aspect of the disclosure is a transmitting device comprising a transmitting unit that transmits a configuration of content having degrees of freedom of viewpoint, the transmitting unit transmits quality information of a stream relating to a three-dimensional object included in a specific content which includes at least the content, and the quality information includes quality information relating to each of two or more streams which have different quality depending on the orientation of the three-dimensional object.
[0009] One aspect of the disclosure is a receiving device comprising a receiving unit that receives a configuration of content having degrees of freedom of viewpoint, the receiving unit receiving quality information of a stream relating to a three-dimensional object contained in a specific content that includes at least the content, and the quality information includes quality information relating to each of two or more streams of different quality depending on the orientation of the three-dimensional object. [Effects of the Invention]
[0010] According to the present invention, it is possible to provide a media processing device, a transmitting device, and a receiving device that enable the appropriate display of 3D objects contained in specific content while suppressing transmission traffic. [Brief explanation of the drawing]
[0011] [Figure 1] Figure 1 is a diagram showing a transmission system 10 according to an embodiment. [Figure 2] Figure 2 is a block diagram showing a media processing device 200 and a user terminal 300 according to an embodiment. [Figure 3] Figure 3 is a diagram illustrating the second content according to the embodiment. [Figure 4] Figure 4 shows a method for viewing specific content according to an embodiment. [Figure 5] Figure 5 is a diagram illustrating Operation Example 1. [Figure 6] Figure 6 is a diagram illustrating example 2 of the operation. [Figure 7] Figure 7 is a diagram illustrating example 2 of the operation. [Figure 8] Figure 8 is a diagram illustrating example 3 of operation. [Figure 9] Figure 9 is a diagram illustrating example 3 of operation. [Figure 10] Figure 10 is a diagram illustrating example 3 of operation. [Figure 11] Figure 11 is a diagram illustrating example 3 of operation. [Figure 12] Figure 12 is a diagram illustrating example 4 of the operation. [Figure 13] Figure 13 is a diagram illustrating example 4 of the operation. [Figure 14] Figure 14 is a diagram illustrating example 4 of the operation. [Figure 15] Figure 15 is a diagram illustrating example 5 of the operation. [Figure 16] Figure 16 is a diagram illustrating example 5 of the operation. [Figure 17]FIG. 17 is a diagram for explaining Operation Example 5. [Figure 18] FIG. 18 is a diagram for explaining Operation Example 5. [Figure 19] FIG. 19 is a diagram for explaining the first method according to Modification Example 1. [Figure 20] FIG. 20 is a diagram for explaining the second method according to Modification Example 1. BEST MODE FOR CARRYING OUT THE INVENTION
[0012] Next, embodiments of the present invention will be described. In the following description of the drawings, the same or similar parts are denoted by the same or similar reference numerals. However, it should be noted that the drawings are schematic, and the ratios of the respective dimensions are different from the actual ones.
[0013] Therefore, specific dimensions and the like should be determined in consideration of the following description. Also, it is needless to say that there are portions where the dimensional relationships and ratios are different between the drawings.
[0014] [Summary of Disclosure] A media processing apparatus according to the summary of the disclosure includes a receiving unit that receives viewpoint information from a user terminal, a renderer that generates specific content including at least content having a degree of freedom of viewpoint based on the viewpoint information, and a transmitting unit that transmits the specific content generated by the renderer to the user terminal. The receiving unit receives quality information regarding each of two or more streams whose quality differs depending on the orientation of the three-dimensional object based on the viewpoint information as quality information of a stream regarding the three-dimensional object included in the specific content from the transmitting device.
[0015] The disclosure summary states that the media processing device receives quality information from the transmitting device for each of two or more streams, each with different quality depending on the orientation of the 3D object based on viewpoint information, as quality information for the stream containing the 3D object in the specific content. With this configuration, based on the new insight that the quality of each face constituting the bounding box of the 3D object does not need to be uniform, the 3D object can be displayed appropriately while suppressing transmission traffic.
[0016] It should be noted that specific content generated by the media processing device is generated based on viewpoint information, and the user terminal can treat the video contained in that content as 2D video.
[0017] [Embodiment] (Transmission system) The following describes a transmission system according to an embodiment. Figure 1 is a diagram showing a transmission system 10 according to an embodiment. As shown in Figure 1, the digital wireless transmission system comprises a transmitting device 100, a media processing device 200, and a user terminal 300.
[0018] In this embodiment, the transmitting device 100 transmits a first content that does not have freedom of viewpoint and a second content that does have freedom of viewpoint to the media processing device 200. Furthermore, the transmitting device 100 transmits first control information associated with the first content and second control information associated with the second content to the media processing device 200.
[0019] The first content may include at least one of 2D video and audio. The first content and the first control information may be transmitted using the first method. The first method may be a method compliant with ISO / IEC 23008-1 (hereinafter referred to as MMT (MPEG Media Transport)). In the following, an example will be given of the case where the first method is MMTP (MMT Protocol) compliant with MMT. The first control information may be referred to as MMT-SI (Signaling Information).
[0020] The second content may include 360° video and 3D objects. The second content and second control information may be transmitted using the second method, or they may be transmitted using a protocol such as HTTP (HyperText Transfer Protocol). The second content may conform to 3DoF+ (Degree of Freedom), which involves viewpoint movement within the range of head movement by the user while seated, or 6DoF, which involves viewpoint movement within the range of free movement by the user. Because the second content has degrees of freedom of viewpoint, it may include two or more 360° videos or two or more 3D objects at the same time (frame). The second control information may be called a scene description.
[0021] Here, the second control information may be transmitted using the first method described above. That is, the second control information may be transmitted using the same first method as the first control information (for example, MMTP). Alternatively, it may be transmitted using a protocol such as HTTP.
[0022] The transmission from the transmitting device 100 to the media processing device 200 is not particularly limited, but may be transmitted using satellite broadcasting, the Internet network, or a mobile communication network.
[0023] While not particularly limited, the transmission system may be a digital wireless transmission system. The digital wireless transmission system may be a system used for 4K or 8K satellite broadcasting.
[0024] The media processing device 200 generates specific content, which includes at least the second content described above, based on viewpoint information received from the user terminal 300, and transmits the generated specific content to the user terminal 300. Although not particularly limited, the transmission of the specific content may be via the Internet network or via a mobile communication network.
[0025] The user terminal 300 may be a smartphone, tablet, head-mounted display, or other user terminal. As shown in Figure 1, two or more user terminals 300 may be provided. In other words, two or more user terminals 300 may request the media processing device 200 to generate specific content. Each user terminal 300 may transmit different viewpoint information to the media processing device 200.
[0026] (Media processing equipment and user terminal) The media processing apparatus and user terminal according to the embodiment will be described below. Figure 2 is a block diagram showing the media processing apparatus 200 and user terminal 300 according to the embodiment.
[0027] Firstly, the media processing device 200 includes a receiving unit 210, a renderer 220, and an encoding processing unit 230.
[0028] The reception unit 210 receives viewpoint information. In this embodiment, the reception unit 210 constitutes a receiving unit that receives viewpoint information from the user terminal 300. The viewpoint information includes an information element indicating the viewpoint position of the user of the user terminal 300 and an information element indicating the direction of the user's line of sight of the user of the user terminal 300.
[0029] Renderer 220 generates specific content that includes at least the second content, based on viewpoint information. Since the specific content is generated based on viewpoint information, it may include one 360° video or one 3D object at the same time (frame). Below, we will illustrate the case where the specific content includes the first content in addition to the second content.
[0030] As shown in Figure 2, the renderer 220 generates a first content, including 2D video and audio, as part of a specific content, based on the first control information (MMT-SI). Viewpoint information is not required in the generation of the first content.
[0031] Specifically, renderer 220 acquires 2D video, audio, and MMT-SI in the form of MMTP packets, which are packets containing 2D video, audio, and MMT-SI.
[0032] For example, an MMTP packet is stored in an IP (Internet Protocol) packet. IP packets may be transmitted using UDP (User Datagram Protocol) or TCP (Transmission Control Protocol).
[0033] Here, the first content is processed in units divided into fixed time intervals (hereinafter referred to as MPU; Media Processing Unit). An MPU includes one or more access units. Access units are sometimes treated as MFUs (Media Fragment Units). An MFU related to 2D video may be called a NAL (Network Abstraction Layer) unit, and an MFU related to audio may be called an MHAS (MPEG-H 3D Audio Stream) packet.
[0034] The MMT-SI includes a PA (Package Access) message, which in turn includes an MPT (MMT Package Table) listing the first content. Furthermore, the MMT-SI includes an MPU timestamp descriptor indicating the presentation time of the first content. The MPU timestamp descriptor may represent the presentation time of the MPU, i.e., the time of the access unit that is first presented in the MPU.
[0035] The MPU timestamp descriptor may be generated using UTC (Coordinated Universal Time) as the base time. The base time may be TAI (International Atomic Time) or the time provided by GPS (Global Positioning System). The base time may be the time provided by an NTP (Network Time Protocol) server or the time provided by a PTP (Precision Time Protocol) server.
[0036] Secondly, the renderer 220 generates second content, including 360° video and 3D objects, as part of specific content, based on the second control information (scene description). Viewpoint information is used in the generation of the second content.
[0037] Specifically, renderer 220 may obtain the scene description in the form of MMTP packets in which the scene description is packetized. The method for obtaining 360° video and 3D objects is not particularly limited.
[0038] 360° video may be converted to 2D video by projection transformations such as ERP (Equirectangular projection) or cubemap. Metadata indicating the type of projection transformation applied to the 360° video may be added. 3D objects may be encoded in mesh format. ISO / IEC 14496-16 “Animation framework extension (AFX)” may be used for mesh format encoding. 3D objects may also be encoded in point cloud format. ISO / IEC 23090-5 “Video-based Point Cloud Compression” may be used for point cloud format encoding.
[0039] Here, the second content is compiled into a single file in units divided by a fixed time interval. This fixed time interval may be 500ms. For example, if the frame rate is 60fps (frames per second), one file will contain 30 frames.
[0040] A scene description is generated for each file and contains information for each frame that identifies the 360° video and 3D objects. For example, a scene description may include an information element indicating the name of the 3D object in the frame (object_name), an information element indicating the frame number (frame_number), an information element indicating the position of the 3D object in the frame (translation_object), an information element indicating the rotation of the 3D object in the frame (rotation_object), and an information element indicating the size of the 3D object in the frame (scale_object).
[0041] Thirdly, the renderer 220 outputs specific content, including the first content and the second content, to the encoding processing unit 230. The renderer 220 may also output the presentation time of the specific content to the encoding processing unit 230 along with the specific content.
[0042] Here, the presentation time of specific content may be modified based on the delay time between the media processing device 200 and the user terminal 300. Specifically, the renderer 220 may calculate the presentation time of specific content provided from the media processing device 200 to the user terminal 300 (T'=T+ΔT) based on the presentation time (T) and delay time (ΔT) of the specific content provided from the transmitting device 100 to the media processing device 200. The delay time (ΔT) may be a predetermined value in the media processing device 200, or it may be a different value for each user terminal 300.
[0043] Fourth, the renderer 220 may be configured as a transmission unit that transmits viewpoint information used to generate specific content to the user terminal 300. The viewpoint information used to generate specific content may be transmitted from the encoding processing unit 230 to the user terminal 300.
[0044] For example, the transmission method for viewpoint information and specific content may be MMTP or HTTP. If MMTP is used as the transmission method for specific content, viewpoint information may be stored as metadata in OMAF (Omnidirectional Media Format) as defined in ISO / IEC 23090-2.
[0045] The encoding processing unit 230 encodes specific content generated by the renderer 220. In this embodiment, the encoding processing unit 230 may be an example of a transmission unit that transmits the specific content to the user terminal 300.
[0046] Furthermore, the encoding processing unit 230 may encode the presentation time of the specific content. The encoding processing unit 230 may transmit an information element indicating the presentation time together with the specific content to the user terminal 300.
[0047] Here, any compression encoding scheme can be used by the encoding processing unit 230. For example, the compression encoding scheme may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding).
[0048] As mentioned above, since the second content included in a specific piece of content is generated based on viewpoint information, the video included in that specific piece of content can be treated as a 2D video that does not have freedom of viewpoint.
[0049] For example, the transmission control method used to start and end viewing of specific content may include RTSP (Real Time Streaming Protocol). The transmission method may be MMTP or HTTP. If MMTP is used as the transmission method, the specific content may be stored in OMAF as defined in ISO / IEC 23090-2.
[0050] As shown in Figure 2, the user terminal 300 includes a detection unit 310, a decoding processing unit 320, and a renderer 330.
[0051] The detection unit 310 detects the user's viewpoint position and gaze direction. The detection unit 310 may include an acceleration sensor and a GPS (Global Positioning System) sensor. The detection unit 310 may include a user interface (e.g., touch sensor, keyboard, mouse, controller, etc.) that is manually input by the user. The detection unit 310 may transmit viewpoint information (viewpoint position and gaze direction) to the media processing device 200. The detection unit 310 may output viewpoint information (viewport) to the renderer 330.
[0052] The decoding processing unit 320 decodes specific content received from the media processing unit 200. The decoding processing unit 320 may also decode the presentation time received from the media processing unit 200. The decoding processing unit 320 may output the specific content to the renderer 330, or it may output the presentation time to the renderer 330.
[0053] The renderer 330 outputs specific content decoded by the decoding processing unit 320. The renderer 330 may also output specific content based on the presentation time decoded by the decoding processing unit 320. For example, the renderer 330 may output the video content included in the specific content to a display and the audio content included in the specific content to a speaker.
[0054] Here, the renderer 330 may generate specific content in which the viewpoint position and line of sight direction have been corrected based on the difference between the viewpoint information received from the media processing device 200 and the viewpoint information input from the detection unit 310.
[0055] (Second Content) The second content according to the embodiment will be described below. Here, the second content at t=0, t=1, and t=2 will be described. The time intervals at t=0, t=1, and t=2 are not particularly limited.
[0056] For example, as shown in Figure 3, at t=0, a 360° video may be displayed instead of the 3D object. The 360° video can be considered a background video for the 3D object. At t=1, the 3D object may be displayed superimposed on the 360° video. Furthermore, at t=1, the position and rotation of the 3D object superimposed on the 360° video may be changed.
[0057] The scene description described above includes information elements indicating the position, rotation, and size of the 3D object for each of t=0, t=1, and t=2, allowing the 3D object to be appropriately superimposed onto the 360° image.
[0058] (How to watch) The following describes a viewing method according to an embodiment. Here, we will illustrate the viewing of specific content, including the first and second contents.
[0059] As shown in Figure 4, in step S11, the user terminal 300 sends an RTSP SETUP to the media processing unit. The RTSP SETUP is a message indicating that viewing of a specific content has begun.
[0060] Here, RTSP SETUP includes the IP address of the user terminal 300, the listening port number, and content identification information (content ID). RTSP SETUP may also include capability information for viewing specific content on the user terminal 300. Capability information may include frame rate, display resolution, etc. Display resolution may include field of view (FoV). Capability information may also include information elements indicating the encoding and compression methods supported by the user terminal 300.
[0061] Here, an example is given in which the capability information of the user terminal 300 is directly notified to the media processing device 200, but the embodiment is not limited to this. The capability information of the user terminal 300 may be notified to the transmitting device 100, and then notified from the transmitting device 100 to the media processing device 200.
[0062] In step S12, the media processing device 200 sends a response to the RTSP SETUP. Here, an ACK is sent as a response indicating that the RTSP SETUP has been received.
[0063] In step S21, the user terminal 300 transmits initial viewpoint information to the media processing device 200. The initial viewpoint information may be transmitted in MMT-SI format.
[0064] In step S22, the media processing device 200 generates initial specific content based on the initial viewpoint information (rendering process). For example, the media processing device 200 generates second content to be included in the initial specific content based on the initial viewpoint information and the scene description.
[0065] Here, the media processing device 200 may generate initial specific content using a viewport that is wider than the display resolution of the user terminal 300. For example, the wider range may be a range of display resolution + 20% in the horizontal direction and display resolution + 20% in the vertical direction.
[0066] The media processing device 200 applies a compression encoding scheme to the initial specified content. The compression encoding scheme may be HEVC or VVC, although it is not particularly limited.
[0067] In step S23, the media processing device 200 transmits initial specific content corresponding to the initial viewpoint information to the user terminal 300. The media processing device 200 transmits the presentation time of the initial specific content to the user terminal 300. As described above, the presentation time (T') provided to the user terminal 300 may be determined based on the delay time (ΔT).
[0068] Furthermore, if a different value is used for the delay time (ΔT) for each user terminal 300, the media processing device 200 can identify it by including the transmission time of the RTSP SETUP in the RTSP SETUP as described above.
[0069] The user terminal 300 outputs specific content based on the presentation time (T'). The user terminal 300 may also generate specific content with corrected viewpoint position and gaze direction based on the difference between viewpoint information received from the media processing device 200 and viewpoint information input from the detection unit 310.
[0070] In step S31, the user terminal 300 transmits viewpoint information to the media processing device 200. The viewpoint information may be transmitted in MMT-SI format. The user terminal 300 may transmit viewpoint information at predetermined intervals (e.g., 500ms), or it may transmit viewpoint information in response to a change in at least one of the viewpoint position and line of sight direction.
[0071] In step S32, the media processing device 200 generates specific content based on the viewpoint information received in step S31 (rendering process).
[0072] In step S33, the media processing device 200 transmits specific content corresponding to the viewpoint information received in step S31 to the user terminal 300.
[0073] The processing in steps S31 to S33 is the same as the processing in steps S21 to S23, except that the viewpoint information received in step S31 is used instead of the initial viewpoint information. Therefore, the details of the processing in steps S31 to S33 are omitted. The processing in steps S31 to S33 may be repeated at a predetermined interval, or it may be repeated each time the user's viewpoint position or line of sight changes.
[0074] In step S41, the user terminal 300 sends an RTSP TEARDOWN to the media processing unit. The RTSP TEARDOWN is a message indicating that viewing of a specific content has ended.
[0075] In step S42, the media processing device 200 sends a response to the RTSP TEARDOWN. Here, an ACK is sent as a response indicating that the RTSP TEARDOWN has been received.
[0076] Figure 4 illustrates a case where steps S11 and S12 are executed on an RTSP basis, but the embodiments are not limited to this. Steps S11 and S12 may also be executed on an MMTP basis or on an HTTP basis.
[0077] Similarly, while examples have been given of cases where steps S41 and S42 are performed on an RTSP basis, the embodiments are not limited thereto. Steps S41 and S42 may also be performed on an MMTP basis or on an HTTP basis.
[0078] Figure 4 illustrates a case where steps S31 to S33 are performed on an MMTP basis, but the embodiment is not limited to this. Steps S31 to S33 may also be performed on other methods (e.g., HTTP).
[0079] Similarly, while the examples illustrate the case where steps S41 to S43 are performed on an MMTP basis, the embodiments are not limited to this. Steps S41 to S43 may also be performed on other methods (e.g., HTTP).
[0080] (Example of operation 1) The embodiments described above may include the following Operation Example 1. In Operation Example 1, the media processing device 200 transmits viewpoint information used in generating the specific content to the user terminal 300, in association with the sequence number attached to the specific content.
[0081] Specifically, the media processing device 200 (renderer 220), similar to the embodiment described above, generates second content, including 360° video and 3D objects, as part of specific content, based on second control information (scene description). In generating the second content, viewpoint information received from the user terminal 300 is used.
[0082] In Operation Example 1, the renderer 220 associates the viewpoint information used in generating a specific content (in this case, the second content) with a sequence number. The renderer 220 associates the viewpoint information used in generating the specific content with the sequence number attached to the specific content and sends it to the user terminal 300. The viewpoint information may be stored in a VP (View Port) message. The VP message may have the MMT-SI format specified in ISO / IEC 23008-1, ARIB STD-B60, ARIB TR-B39, etc. The VP message may be sent frame by frame. The VP message may also be sent to the user terminal 300 as an MMTP-related message (MMT-SI).
[0083] While not particularly limited, a VP message may have the data structure shown in Figure 5. As shown in Figure 5, a VP message may include message_id, version, length, fov, viewpoint_pos_x, viewpoint_pos_y, viewpoint_pos_z, viewpoint_yaw, viewpoint_pitch, viewpoint_roll, viewport_width, viewport_height, mpu_sequence_number_flag, mpu_sequence_number, etc.
[0084] The message_id is an identifier that identifies a VP message. The message_id may also be 0x0204.
[0085] The `version` field indicates the version of the MMTP protocol. The `version` field may also be 0x00.
[0086] The `length` parameter indicates the length of the VP message.
[0087] FOV is information that indicates the field of view.
[0088] `viewpoint_pos_x` is information indicating the x-coordinate of the viewpoint position. `viewpoint_pos_x` is an example of viewpoint information used in the generation of specific content.
[0089] `viewpoint_pos_y` is information indicating the y-coordinate of the viewpoint position. `viewpoint_pos_y` is an example of viewpoint information used in the generation of specific content.
[0090] `viewpoint_pos_z` is information indicating the z-coordinate of the viewpoint position. `viewpoint_pos_z` is an example of viewpoint information used in the generation of specific content.
[0091] `viewpoint_yaw` is information indicating the yaw of the viewpoint position. `viewpoint_yaw` is an example of viewpoint information used in the generation of specific content.
[0092] `viewpoint_pitch` is information indicating the pitch of the viewpoint position. `viewpoint_pitch` is an example of viewpoint information used in the generation of specific content.
[0093] `viewpoint_roll` is information indicating the view position roll. `viewpoint_roll` is an example of view information used in the generation of specific content.
[0094] `viewport_width` is information that indicates the width of the display area (specific content).
[0095] `viewport_height` is information that indicates the height of the display area (specific content).
[0096] The `mpu_sequence_number_flag` indicates whether or not the `mpu_sequence_number` field exists. For example, if `mpu_sequence_number_flag` is 1, the `mpu_sequence_number` field exists, but if `mpu_sequence_number_flag` is 0, the `mpu_sequence_number` field does not necessarily have to exist.
[0097] `mpu_sequence_number` is the MPU sequence number of the video corresponding to the specific content indicated in the VP message. `mpu_sequence_number` is an example of a sequence number assigned to specific content.
[0098] Here, `viewpoint_pos_x`, `viewpoint_pos_y`, and `viewpoint_pos_z` are examples of information elements that indicate the user's viewpoint position in the 3D space constructed by the scene description. `viewpoint_yaw`, `viewpoint_pitch`, and `viewpoint_roll` are examples of information elements that indicate the user's line of sight direction in the 3D space constructed by the scene description. `viewport_width` and `viewport_height` are examples of information elements that indicate the number of pixels in the video contained in specific content.
[0099] Under these conditions, the media processing device 200 and the user terminal 300 may perform the following operations.
[0100] Firstly, the media processing device 200 (renderer 220) may determine viewpoint_pos_x, viewpoint_pos_y, and viewpoint_pos_z based on the viewpoint position used to generate the specific content. The renderer 220 may also determine viewpoint_yaw, viewpoint_pitch, and viewpoint_roll based on the viewing direction used to generate the specific content. The renderer 220 may determine viewport_width based on the number of pixels in the horizontal direction of the specific content, and determine viewport_height based on the number of pixels in the vertical direction of the specific content.
[0101] The media processing unit 200 (encoding processing unit 230) may perform compression encoding of specific content generated by the renderer 220 and transmit the specific content. Here, the encoding processing unit 230 may transmit the VP message shown in Figure 5 to the user terminal 300. That is, the encoding processing unit 230 may transmit the viewpoint information used in generating the specific content to the user terminal 300, in association with the sequence number attached to the specific content.
[0102] Secondly, the user terminal 300 (decoding processing unit 320) may be configured as a receiving unit that receives viewpoint information used in the generation of specific content from the media processing unit 200, in association with the sequence number attached to the specific content. That is, the decoding processing unit 320 may receive the VP message (MMT-SI) shown in Figure 5 from the media processing unit 200.
[0103] The user terminal 300 (renderer 330) may associate the video constituting the decoded specific content with the viewpoint information contained in the VP message based on the sequence number (mpu_sequence_number). Based on the difference between the viewpoint information detected by the detection unit 310 and the viewpoint information contained in the VP message, the renderer 330 may identify the video within the range specified by the viewpoint information detected by the detection unit 310 from the display area defined by the information contained in the VP message (viewport_width and viewport_height). The renderer 330 may display the identified video.
[0104] (Example of operation 2) The embodiments described above may include the following Operation Example 2. In Operation Example 2, as shown in Figure 6, a case is assumed in which a first user terminal 400 that provides viewpoint information feedback and a second user terminal 500 that does not provide viewpoint information feedback are mixed. The first user terminal 400 may be a terminal such as a head-mounted display. The first user terminal may have the same functions as the user terminal 300 described above. The second user terminal 500 may be a terminal such as a volumetric display.
[0105] Specifically, in Operation Example 2, as shown in Figure 6, the media processing device 200 (renderer 220) may be configured as a receiving unit that receives content configuration and recommended viewport information from the transmitting device 100. The content configuration may be considered to include 2D video, audio, 360° video, and 3D objects. The content configuration may also be considered to include MMT-SI and scene description. Recommended viewport information may be considered as an example of specific viewpoint information. Recommended viewport information may also be information that defines (recommends) the position, direction, and field of view at which to view the video in the three-dimensional space (three-dimensional space constructed by the scene description) configured by the specific content. Recommended viewport information may include information elements that indicate at least one of the following: an information element indicating the viewpoint position in the three-dimensional space configured by the specific content, and an information element indicating the direction of the line of sight in the three-dimensional space configured by the specific content. Recommended viewport information may be considered as viewpoint information mainly used by the second user terminal 500.
[0106] In the following, unless explicitly stated otherwise, the viewpoint information used in generating specific content may include specific viewpoint information (recommended viewport information) received from the transmitting device 100, and may also include viewpoint information received from the user terminal 300.
[0107] While not particularly limited, recommended viewport information may be included in the scene description in the manner shown in Figure 7. As shown in Figure 7, recommended viewport information may include camera_orientation, frame_number, translation, and yfov.
[0108] `camera_orientation` is information indicating the direction from which the image is viewed in the 3D space constructed by the scene description. `camera_orientation` can be considered synonymous with line of sight.
[0109] `frame_number` is information indicating the frame number of the video to which `camera_orientation`, `translation`, and `yfov` are applied. `camera_orientation` can be considered an example of specific viewpoint information.
[0110] Translation is information that indicates the viewing position of the image in the 3D space constructed by the scene description. Translation can be considered synonymous with viewpoint position. Translation can also be considered an example of specific viewpoint information. For example, Figure 7 illustrates the case where the viewpoint position is [0,0,-50] when the frame number is 0, and moves to [0,0,-75] when the frame number is 2505.
[0111] yfov is information that indicates the field of view in the 3D space constructed by the scene description.
[0112] While not strictly limited, camera_orientation and translation may be assigned by the content creator. Alternatively, if the content is captured by a camera, camera_orientation and translation may be automatically assigned by the GPS and sensors installed in the camera.
[0113] Here, `translation` is an example of an information element that indicates the viewpoint position in a 3D space composed of specific content. `camera_orientation` is an example of an information element that indicates the line of sight direction in a 3D space composed of specific content.
[0114] Under these conditions, the media processing device 200 and the user terminal 300 may perform the following operations.
[0115] Firstly, the media processing unit 200 (renderer 220) may generate specific content based on specific viewpoint information (recommended viewport information). The media processing unit 200 (encoding processing unit 230) may perform compression encoding of the specific content generated by the renderer 220 and transmit the specific content. Here, the encoding processing unit 230 may transmit the specific content generated based on the specific viewpoint information to the first user terminal 400, or it may transmit the specific content generated based on the specific viewpoint information to the second user terminal 500.
[0116] Secondly, the media processing device 200 (reception unit 210) may receive viewpoint information from the first user terminal 400 if the user terminal is the first user terminal 400 that provides viewpoint information as feedback. The media processing device (renderer 220) may generate specific content based on the viewpoint information received from the first user terminal 400. In such a case, the media processing device 200 (reception unit 210) may receive a reset signal from the first user terminal 400. The media processing device (renderer 220) may generate specific content based on specific viewpoint information (recommended viewport information) in response to the reset signal.
[0117] In such cases, the first user terminal 400 may have the same configuration as the user terminal 300. However, the first user terminal 400 (detection unit 310) may have a function to detect a reset signal. The detection unit 310 may detect a user operation that inputs a reset signal. The detection unit 310 may transmit the reset signal to the media processing device 200.
[0118] Thirdly, the media processing device 200 (renderer 220) may generate specific content based on specific viewpoint information (recommended viewport information) when the user terminal is a second user terminal 500 that does not provide viewpoint information feedback.
[0119] In such cases, the second user terminal 500 does not need to have a detection unit 310 for detecting viewpoint information. The second user terminal 500 does not need to have a renderer 330. The second user terminal 500 may have the same configuration as the user terminal 300, except that it does not have a detection unit 310 and a renderer 330.
[0120] (Example of operation 3) The embodiments described above may include the following Operation Example 3. In Operation Example 3, the media processing device 200 (for example, the selection unit 260 described later) may be configured as a receiving unit that receives quality information from the transmitting device 100 for each of two or more streams, each of which has a different quality depending on the orientation of the 3D object based on viewpoint information, as quality information for the stream relating to the 3D object contained in the specific content.
[0121] Specifically, in Operation Example 3, as shown in Figure 8, the media processing device 200 has a selection unit 260 in addition to the configuration shown in Figure 2. The selection unit 260 receives the scene description and 3D objects from the transmitting device 100. The selection unit 260 inputs the selected stream (3D object) from two or more streams to the renderer 220. The selection unit 260 may also request the transmitting device 100 to transmit the selected stream. In the case where the media processing device 200 transmits specific content to multiple user terminals 300, the selection unit 260 may request the transmitting device 100 to transmit the streams required by each of the multiple user terminals 300, or it may request the transmitting device 100 to transmit all streams.
[0122] Here, the selection unit 260 receives quality information for each of two or more streams, each with a different quality depending on the orientation of the 3D object based on viewpoint information, from the transmission device 100 as quality information for the 3D object stream.
[0123] Quality information may also be information indicating the relative quality of each face that makes up the bounding box for a 3D object. The bounding box may be represented by a three-dimensional rectangle onto which the 3D object is projected. For example, as shown in Figure 9, the bounding box may be defined by vertices A to H. In such a case, each face of the bounding box includes face #1 represented by vertices A, B, F, E; face #2 represented by vertices B, C, G, F; face #3 represented by vertices A, B, C, D; face #4 represented by vertices E, F, G, H; face #5 represented by vertices A, D, H, E; and face #6 represented by vertices D, C, G, H.
[0124] In such cases, assuming we view a 3D object in a three-dimensional space constructed by the scene description, it is assumed that three faces will be primarily observed. In other words, the remaining three faces are assumed to be observed less.
[0125] In example 3, two or more streams are prepared for 3D objects included in specific content, each with a different quality depending on the orientation of the 3D object based on viewpoint information.
[0126] While not particularly limited, quality information may be included in the scene description in the manner shown in Figure 10. In Figure 10, six streams are exemplified as streams with different orientations of 3D objects. Quality information may be represented in the format "quality" [#1,#2,#3,#4,#5,#6]. Within the brackets [ ], #1 to #6 represent the quality index of faces #1 to #6. The quality index may take values in the range of 1 to 9. A higher quality index value may indicate higher quality. For example, in the stream identified by "id"="1", the quality of #1, #2, and #3 is high ("8"), and the quality of #4, #5, and #6 is low ("3"). In the stream identified by "id"="2", the quality of #1, #2, and #3 is low ("3"), and the quality of #4, #5, and #6 is high ("8").
[0127] Under these conditions, the media processing device 200 may perform the operations shown below. The following description will mainly focus on the selection of a selected stream (3D object) from among two or more streams.
[0128] In Mode 1, as shown in the upper part of Figure 11, the media processing device 200 (selection unit 260) may first identify the vertex closest to the user's viewpoint (for example, vertex B), and then identify three faces with the closest vertex (for example, face #1 represented by vertices A, B, F, E; face #2 represented by vertices B, C, G, F; and face #3 represented by vertices A, B, C, D). The selection unit 260 may also select the stream that has the maximum sum of the quality indices of the three identified faces (in the example shown in Figure 10, the stream identified by "id" = "1").
[0129] In Mode 1, since the vertex closest to the user's viewpoint is identified based on viewpoint information, it can be considered that the selection unit 260 selects the stream to be transmitted to the user terminal 300 from among two or more streams based on viewpoint information and quality information.
[0130] Mode 2 may be the mode applied when a 3D object is to be displayed in a reduced or enlarged state. For example, as shown in the middle of Figure 11, the media processing device 200 (selection unit 260) may identify the vertex closest to the user's viewpoint (e.g., vertex B) and then identify three faces with the closest vertex (e.g., face #1 represented by vertices A, B, F, E; face #2 represented by vertices B, C, G, F; and face #3 represented by vertices A, B, C, D). When a 3D object is to be displayed in a reduced state, pixels of the 3D object are downsampled. Therefore, the selection unit 260 may select the stream that minimizes the sum of the quality indices of the three identified faces (in the example shown in Figure 10, the stream identified by "id"="2"). On the other hand, when a 3D object is to be displayed in an enlarged state, pixels of the 3D object are interpolated. Therefore, the selection unit 260 may select the stream that maximizes the sum of the quality indices of the three identified faces (in the example shown in Figure 10, the stream identified by "id"="1").
[0131] In Mode 2, since the vertex closest to the user's viewpoint is identified based on viewpoint information, it can be considered that the selection unit 260 selects the stream to be transmitted to the user terminal 300 from among two or more streams based on viewpoint information and quality information.
[0132] Mode 3 may be applied in cases where two 3D objects (3D object #1 and 3D object #2) overlap in the user's line of sight. Here, the selection of a stream related to 3D object #1 is described. For example, as shown in the lower part of Figure 11, the media processing device 200 (selection unit 260) may identify the vertex closest to the user's viewpoint (e.g., vertex B) and then identify three faces with the closest vertices (e.g., face #1 represented by vertices A, B, F, E; face #2 represented by vertices B, C, G, F; and face #3 represented by vertices A, B, C, D). Here, 3D object #2 overlaps with the line segment connecting the vertex closest to the user's viewpoint (e.g., vertex B) and the user's viewpoint, and the three identified faces are obscured by 3D object #2. Therefore, the selection unit 260 may select the stream that minimizes the sum of the quality indices of the three identified faces (in the example shown in Figure 10, the stream identified by "id" = "2").
[0133] In Mode 3, since the vertex closest to the user's viewpoint is identified based on viewpoint information, the selection unit 260 can be thought of as selecting a stream to send to the user terminal 300 from among two or more streams based on viewpoint information and quality information. Furthermore, in Mode 3, since the overlap of two 3D objects is identified based on viewpoint information and 3D object placement information, the selection unit 260 can be thought of as selecting a stream to send to the user terminal 300 from among two or more streams based on viewpoint information, quality information, and placement information. The 3D object placement information (for example, “rotation_object”, “scale_object”, and “translation_object” shown in Figure 10) may be included in the scene description. The 3D object placement information may be thought of as “link_area” shown in Figure 10. That is, the selection unit 260 may receive the 3D object placement information in three-dimensional space from the transmission device 100.
[0134] In the quality information shown in Figure 10, the sum of the quality indices for the six faces is equal in each stream. However, the embodiments are not limited to this. The sum of the quality indices for the six faces may differ between two or more streams.
[0135] (Example of operation 4) The embodiments described above may also include the following Operation Example 4. Here, Operation Example 4 includes the following operations in addition to Operation Example 3. In Operation Example 4, the media processing device 200 (for example, the selection unit 270 described later) may be configured as a receiving unit that receives importance information for each of two or more objects included in a specific content from the transmitting device 100.
[0136] Specifically, in Operation Example 4, as shown in Figure 12, the media processing device 200 has a selection unit 270 in addition to the configuration shown in Figure 2. The selection unit 270 receives the scene description, 3D objects, and 360° video from the transmission device 100. The selection unit 270 inputs the selected stream (3D object) from two or more streams to the renderer 220.
[0137] Here, the selection unit 270 receives importance information for each of two or more objects from the transmitting device 100. The objects may include 3D objects and 360° images.
[0138] For example, importance information may indicate the relative importance between two or more objects. For instance, consider the case shown in Figure 13, where object A (background), object B (person), and object C (dog) exist in a three-dimensional space constructed by a scene description. Object A (background) is an example of a 360° video, and object B (person) and object C (dog) are examples of 3D objects. In such a case, importance information may indicate the relative importance between each of object A (background), object B (person), and object C (dog).
[0139] While not particularly limited, importance information may be included in the scene description in the manner shown in Figure 14. In Figure 14, importance information may be represented by weight. Weight may take values in the range of 1 to 9. A higher weight value may indicate higher importance. In Figure 14, an example is shown where object A (background) identified by "object_id"="0" has the highest weight ("9"), object B (person) identified by "object_id"="1" has the lowest weight ("3"), and object C (dog) identified by "object_id"="2" has a weight ("8") that is higher than object B (person) but lower than object A (background).
[0140] Under these conditions, the media processing device 200 may perform the operations shown below. The following description will mainly focus on the selection of a selected stream (3D object) from among two or more streams.
[0141] Firstly, the media processing unit 200 (selection unit 270) selects the stream with the highest quality for the 3D object with the highest importance. The method of selecting the stream may be the same as in operation example 3. For example, since the importance of object C (dog) is greater than that of object B (person), the selection unit 270 selects the stream for object C (dog) that maximizes the sum of the quality indices of the three faces having the vertices closest to the user's viewpoint.
[0142] Secondly, the media processing unit 200 (selection unit 270) selects the stream with the lowest quality for all 3D objects except the one with the highest importance. Subsequently, the selection unit 270 replaces the stream with the lowest quality with a stream with a higher quality for each 3D object, starting with the most important ones, within the range where specific conditions are met. The specific conditions may include a first condition that the bandwidth of the line from the transmitter 100 to the media processing unit 200 is below a threshold, and a second condition that the processing load of the media processing unit 200 is below a threshold. The specific conditions may also be defined by a combination of the first and second conditions. For example, since the importance of object B (person) is less than that of object C (dog), the selection unit 270 selects a stream with a higher quality for object B (person) within the range where the specific conditions are met.
[0143] As described above, the media processing device 200 (selection unit 270) may be thought to select a stream to be transmitted to the user terminal 300 from among two or more streams based on viewpoint information, quality information, and importance information. The media processing device 200 (selection unit 270) may be thought to select a stream to be transmitted to the user terminal 300 from among two or more streams based on viewpoint information, quality information, placement information, and importance information.
[0144] In Operation Example 4, we illustrated a case where there is one stream of 360° video. However, the embodiments are not limited to this. There may be two or more streams of 360° video with different qualities.
[0145] In Operation Example 4, we illustrated a case where, for a 3D object, there are two or more streams with different qualities depending on the orientation of the 3D object based on viewpoint information. However, the embodiments are not limited to this. For a 3D object, there may be two or more streams with different qualities regardless of the orientation of the 3D object.
[0146] (Example of operation 5) The embodiments described above may include the following Operation Example 3. In Operation Example 3, the media processing device 200 (for example, the renderer 220) may be configured as a receiving unit that receives information elements from the transmitting device 100 that define the range of movement of the user's viewpoint position in a three-dimensional space composed of specific content (a three-dimensional space constructed by a scene description).
[0147] Firstly, in Operation Example 5, the information element may include an information element (hereinafter referred to as the first information element) that restricts the movement of the user's viewpoint position inside a 3D object contained in specific content. For example, as shown in Figure 15, in the case where a 3D object is placed in a three-dimensional space constructed by a scene description, the movement of the viewpoint position inside the 3D object may be restricted. However, in cases such as when the 3D object is a building or when another scene exists inside the 3D object, the movement of the viewpoint position inside the 3D object may be permitted.
[0148] Secondly, in Operation Example 5, the information element may include an information element (hereinafter referred to as the second information element) that restricts the movement of the user's viewpoint outside the three-dimensional space. For example, as shown in Figure 16, the three-dimensional space may be defined by a combination of a cuboid and a spheroid. The number of cuboids defining the three-dimensional space may be two or more, and the number of spheroids defining the three-dimensional space may be two or more. However, there may be cases in which the movement of the user's viewpoint outside the three-dimensional space is permitted.
[0149] While not particularly limited, the first information element may be included in the scene description in the manner shown in Figure 17. In Figure 17, the first information element may be represented by viewing_inside_object_flag. viewing_inside_object_flag may be set for each 3D object. For example, if viewing_inside_object_flag is "0", movement of the viewpoint position within the 3D object may be restricted, and if viewing_inside_object_flag is "1", movement of the viewpoint position within the 3D object may be permitted.
[0150] While not particularly limited, the second information element may be included in the scene description in the manner shown in Figure 18. In Figure 18, the second information element may include information elements that define a cuboid that defines a three-dimensional space (cuboid_center_x, cuboid_center_y, cuboid_center_z, cuboid_size_x, cuboid_size_y, cuboid_size_z). cuboid_center_x, cuboid_center_y, and cuboid_center_z are information elements that indicate the center position of the cuboid, and cuboid_size_x, cuboid_size_y, and cuboid_size_z are information elements that indicate the size of the cuboid. The second information element may also include information elements that define a spheroid that defines a three-dimensional space (spheroid_center_x, spheroid_center_y, spheroid_center_z, spheroid_size_x, spheroid_size_y, spheroid_size_z). spheroid_center_x, spheroid_center_y, and spheroid_center_z are information elements indicating the center position of the ellipsoid, while spheroid_size_x, spheroid_size_y, and spheroid_size_z are information elements indicating the size of the ellipsoid. Note that cuboid_enable may be an information element indicating whether or not a cuboid defines a 3D space, and spheroid_enable may be an information element indicating whether or not a 3D space defines a ellipsoid. Figure 18 illustrates a case in which a 3D space is defined by two cuboids and two ellipsoids.
[0151] Under these conditions, the media processing device 200 (renderer 220) may perform the following operations.
[0152] Firstly, when the user's viewpoint moves outside the movement range, the renderer 220 may generate specific content using the point of intersection between the user's viewpoint trajectory and the boundary of the movement range as the viewpoint position. In other words, the renderer 220 may fix the viewpoint position at the position (boundary position) at the time the viewpoint position was about to move outside the movement range.
[0153] Secondly, the renderer 220 may notify the user that the movement of the viewpoint is restricted when the user's viewpoint moves outside the range of movement. For example, the renderer 220 may display a message such as "You cannot move beyond this point."
[0154] (Mechanism of Action and Effects) In this embodiment, the media processing device 200 generates specific content based on viewpoint information and then transmits the specific content to the user terminal 300. With this configuration, there is no need for the user terminal 300 to generate specific content, including second content with degrees of freedom of viewpoint, and the user terminal 300 can present the specific content by providing viewpoint information to the media processing device 200. Therefore, although a delay occurs between the media processing device 200 and the user terminal 300, the processing load on the user terminal 300 can be reduced.
[0155] In Operation Example 1, the media processing device 200 transmits viewpoint information used to generate the specific content to the user terminal 300, associated with the sequence number attached to the specific content. With this configuration, the user terminal 300 can understand the viewpoint information and sequence number used to generate the specific content. Therefore, even if the viewpoint information used by the media processing device 200 to generate the specific content differs from the viewpoint information used by the user terminal 300 to display the specific content, the specific content can be displayed appropriately.
[0156] In Operation Example 2, the media processing device 200 receives content configuration and specific viewpoint information (recommended viewport information) from the transmitting device 100. With this configuration, the media processing device 200 can generate specific content based on the specific viewpoint information, and can display the specific content appropriately even in cases where there is a mix of first user terminals 400 that provide viewpoint information feedback and second user terminals 500 that do not.
[0157] In Operation Example 2, the media processing device 200 generates specific content based on specific viewpoint information in response to a reset signal, even when the user terminal is the first user terminal 400. With this configuration, even in cases where the viewpoint position and line of sight direction become unknown to the first user terminal 400 in the three-dimensional space constructed by the scene description (cases where the user gets lost in the three-dimensional space), the reset signal allows the system to return to specific content based on specific viewpoint information.
[0158] In Operation Example 3, the media processing device 200 receives quality information from the transmitting device 100 for each of two or more streams, each with different quality depending on the orientation of the 3D object based on viewpoint information, as quality information for the streams relating to the 3D object contained in the specific content. With this configuration, based on the new insight that the quality of each face constituting the bounding box of the 3D object does not need to be uniform, it is possible to display the 3D object appropriately while suppressing transmission traffic.
[0159] In Operation Example 4, the media processing device 200 receives importance information for each of two or more objects included in a specific content from the transmitting device 100. With this configuration, by introducing a mechanism to set the importance of each object, such as 360° video and 3D objects, it is possible to appropriately display each object included in a specific content while suppressing transmission traffic.
[0160] In Operation Example 5, the media processing device 200 receives information elements from the transmitting device 100 that define the range of movement of the user's viewpoint position in a three-dimensional space composed of specific content. With this configuration, specific content, including content with freedom of viewpoint, can be displayed appropriately without causing any distortion of the specific content displayed on the user terminal 300.
[0161] [Example of change 1] The following describes Example 1 of the modified embodiment. The following mainly describes the differences from the embodiment.
[0162] Change Example 1 describes how to synchronize the first and second content when a specific piece of content includes both the first and second content.
[0163] In the following, synchronization means that the presentation times of the first content (e.g., MPU) and the second content (file) are properly aligned. Therefore, synchronization may include the presentation times of 2D video and 3D objects being aligned, or the presentation times of audio and 3D objects being aligned. Similarly, synchronization may include the presentation times of 2D video and 360° video being aligned, or the presentation times of audio and 360° video being aligned.
[0164] The first method describes a case in which the media processing device 200 synchronizes the first content and the second content based on the first control information (MMT-SI). The media processing device 200 uses the MMT-SI as an entry point to check for the presence or absence of a scene description (second content). If a scene description exists, it uses the MPU timestamp descriptor to determine the presentation time of a specific content including the first and second content.
[0165] Specifically, as shown in Figure 19, the 2D video and audio are presented based on the MPU timestamp descriptor (simply called "timestamp" in Figure 19), thus enabling synchronization of the 2D video and audio.
[0166] On the other hand, the presentation time of the first frame included in the scene description is determined by referring to the MPU timestamp descriptor included in MMT-SI. The presentation times of the second and subsequent frames included in the scene description can be determined by the frame number included in the scene description and the frame rate of the second content. For example, in the case where the frame rate is 30fps, the presentation time of the nth frame is determined by adding 1 / 30 × n to the time determined by the MPU timestamp descriptor. However, the frame number of the first frame included in the scene description is "0".
[0167] In the first method, we exemplified a case where the scene description does not include the presentation time of the first frame included in the scene description, but the scene description may include the presentation time of the first frame included in the scene description.
[0168] The second method describes a case in which the media processing device 200 synchronizes the first content and the second content based on the second control information (scene description). The media processing device 200 uses the scene description as an entry point to check for the presence or absence of MMT-SI (first content), and if MMT-SI is present, it identifies the presentation time of a specific content including the first content and the second content based on the presentation time included in the scene description.
[0169] In such cases, the scene description includes absolute time information indicating the presentation time of the second content. The absolute time information may also be the presentation time of the first frame included in the scene description.
[0170] For example, absolute time information may be generated using UTC as the reference time. The reference time may be TAI, or the time provided by GPS. The reference time may be the time provided by an NTP server, or the time provided by a PTP server. Furthermore, absolute time information may be generated based on the same reference time as the MPU timestamp descriptor.
[0171] Furthermore, the scene description includes reference information to identify the first content. This reference information may also be information to identify the MPUs that make up the first content. In other words, the reference information is information for treating the first content (MPU) as an object included in the scene description.
[0172] Specifically, as shown in Figure 20, the presentation time of the first frame included in the scene description is determined by the absolute time information included in the scene description. The presentation times of the second and subsequent frames included in the scene description can be determined by the frame number included in the scene description and the frame rate of the second content. For example, considering the case where the frame rate is 30fps, the presentation time of the nth frame is determined by adding 1 / 30 × n to the time determined by the MPU timestamp descriptor. However, the frame number of the first frame included in the scene description is "0".
[0173] On the other hand, since the 2D video and audio are presented based on the MPU timestamp descriptor (simply "timestamp" in Figure 20), the 2D video and audio can be synchronized. Here, since the above-mentioned reference information is included in the scene description, the media processing device 200 can check whether or not there is a first content to be presented together with the second content based on the reference information included in the scene description.
[0174] In the second method, the synchronization of 2D video and audio is based on the MPU timestamp descriptor included in MMT-SI. However, in Modification Example 1, the synchronization of 2D video and audio may also be based on information elements (absolute time information and reference information) included in the scene description. In such cases, at least the MPU timestamp descriptor included in MMT-SI may be omitted. Furthermore, MMT-SI itself may be omitted.
[0175] Furthermore, if the reference time of the MPU timestamp descriptor included in MMT-SI (hereinafter referred to as the first reference time) and the reference time of the absolute time information included in the scene description (the second reference time) are different, at least one of the first control information (MMT-SI) and the second control information (scene description) may include conversion information between the first reference time and the second reference time. For example, MMT-SI may include an MPU timestamp descriptor expressed in the second reference time (e.g., a reference time other than UTC) in addition to an MPU timestamp descriptor expressed in the first reference time (e.g., UTC). The scene description may include absolute time information expressed in the first reference time (e.g., UTC) in addition to absolute time information expressed in the second reference time (e.g., a reference time other than UTC).
[0176] The MPU timestamp descriptor included in MMT-SI may also be referred to as the first absolute time information, and the absolute time information included in the scene description may also be referred to as the second absolute time information.
[0177] [Other embodiments] Although the present invention has been described by the above disclosure, the discussion and drawings that constitute part of this disclosure should not be understood as limiting the invention. Various alternative embodiments, examples, and operational techniques will become apparent to those skilled in the art from this disclosure.
[0178] The disclosures described above illustrate cases where specific content includes both the first and second content, but the disclosures are not limited to these cases. Specific content only needs to include at least the second content.
[0179] Although not specifically mentioned in the disclosure above, terminology related to MMT may be interpreted based on the definitions provided in ISO / IEC 23008-1, ARIB STD-B60, ARIB TR-B39, etc.
[0180] In the disclosure described above, an MPU timestamp descriptor was given as an example of the first absolute time information included in the MMT-SI. However, the disclosure is not limited to this. The first absolute time information included in the MMT-SI may also be an MPU extended timestamp descriptor.
[0181] Although not specifically mentioned in the disclosure above, the media processing device 200 may, if necessary, request a portion of the second content from the transmitting device 100. With such a configuration, bandwidth associated with the transmission of the second content can be saved and the increase in processing load of the media processing device 200 can be suppressed.
[0182] In the disclosure described above, MMTP was given as an example of the transmission method for the first content. However, the disclosure is not limited to this. The transmission method for the first content may also be a method compliant with ISO / IEC 23009-1 (hereinafter, MPEG-DASH (Dynamic Adaptive Stream over HTTP)). In such a case, the first control information may be an MPD (Media Presentation Description). That is, in the disclosure described above, MMT-SI may be read as MPD.
[0183] Although not specifically mentioned in the disclosure above, "acquisition" can be interpreted as "received."
[0184] While not particularly limited, Operation Example 2 may also be expressed as follows: The transmitting device 100 includes a transmitting unit that transmits a content configuration having freedom of viewpoint, and the transmitting unit transmits specific viewpoint information used to generate specific content that includes at least the content. The receiving device includes a receiving unit that receives a content configuration having freedom of viewpoint, and the receiving unit receives specific viewpoint information used to generate specific content that includes at least the content. In such a case, the receiving device may be a media processing device 200 or a user terminal 300.
[0185] While not particularly limited, Operation Example 3 may be expressed as follows: The transmitting device 100 includes a transmitting unit that transmits a content configuration having degrees of freedom of viewpoint, and the transmitting unit transmits quality information of a stream relating to a 3D object included in a specific content that includes at least the content, and the quality information includes quality information for each of two or more streams whose quality differs depending on the orientation of the 3D object. The receiving device includes a receiving unit that receives a content configuration having degrees of freedom of viewpoint, and the receiving unit receives quality information of a stream relating to a 3D object included in a specific content that includes at least the content, and the quality information includes quality information for each of two or more streams whose quality differs depending on the orientation of the 3D object. In such a case, the receiving device may be a media processing device 200 or a user terminal 300.
[0186] While not particularly limited, Operation Example 4 may be expressed as follows: The transmitting device 100 includes a transmitting unit that transmits the configuration of content having a degree of freedom of viewpoint, and the transmitting unit transmits importance information for each of two or more objects included in a specific content that includes at least the content. The receiving device includes a receiving unit that receives the configuration of content having a degree of freedom of viewpoint, and the receiving unit receives importance information for each of two or more objects included in a specific content that includes at least the content. In such a case, the receiving device may be a media processing device 200 or a user terminal 300.
[0187] While not particularly limited, Operation Example 4 may also be expressed as follows: The transmitting device 100 includes a transmitting unit that transmits a content configuration having degrees of freedom of viewpoint, and the transmitting unit transmits information elements that define the range of movement of the user's viewpoint position in a three-dimensional space composed of specific content that includes at least the content. The receiving device includes a receiving unit that receives a content configuration having degrees of freedom of viewpoint, and the receiving unit receives information elements that define the range of movement of the user's viewpoint position in a three-dimensional space composed of specific content that includes at least the content. In such a case, the receiving device may be a media processing device 200 or a user terminal 300.
[0188] Although not specifically mentioned in the disclosure above, a program may be provided that causes a computer to execute each of the processes performed by the transmitting device 100, the media processing device 200, and the user terminal 300. Furthermore, the program may be recorded on a computer-readable medium. Using a computer-readable medium, it is possible to install the program on a computer. Here, the computer-readable medium on which the program is recorded may be a non-transient recording medium. The non-transient recording medium is not particularly limited, but may include, for example, a CD-ROM or DVD-ROM.
[0189] Alternatively, a chip may be provided comprising a memory for storing programs for executing each of the processes performed by the transmitting device 100, the media processing device 200, and the user terminal 300, and a processor for executing the programs stored in the memory. [Explanation of Symbols]
[0190] 10...Transmission system, 100...Transmitting device, 200...Media processing device, 210...Reception unit, 220...Renderer, 230...Encoding processing unit, 260...Selection unit, 270...Selection unit, 300...User terminal, 310...Detection unit, 320...Decoding processing unit, 330...Renderer, 400...First user terminal, 500...Second user terminal
Claims
1. A receiving unit that receives viewpoint information from the user terminal, A renderer that generates specific content including at least content having degrees of freedom of viewpoint based on the aforementioned viewpoint information, The system includes a transmission unit that transmits the specific content generated by the renderer to the user terminal, The receiving unit receives quality information from the transmitting device for each of two or more streams, each of which has a different quality depending on the orientation of the 3D object based on the viewpoint information, as quality information for the stream relating to the 3D object included in the specific content. The renderer is a media processing device that, based on the quality information, selects a stream from the two or more streams to be transmitted to the user terminal.
2. The media processing apparatus according to claim 1, wherein the renderer selects a stream to be transmitted to the user terminal from among the two or more streams based on the viewpoint information and the quality information.
3. The media processing apparatus according to claim 1 or claim 2, wherein the quality information is information indicating the relative quality of each face constituting the bounding box of the three-dimensional object.
4. The media processing apparatus according to any one of claims 1 to 3, wherein the receiving unit receives arrangement information of three-dimensional objects in three-dimensional space from the transmitting device.
5. The media processing apparatus according to claim 4, wherein the renderer selects a stream to be transmitted to the user terminal from among the two or more streams based on the viewpoint information, the quality information, and the placement information.
6. The media processing apparatus according to claim 1, wherein the quality information is transmitted by a second method different from the first method used for transmitting first control information associated with content that does not have a degree of freedom of viewpoint, and is included in the second control information associated with content that has a degree of freedom of viewpoint.
7. A generation unit that generates content configurations with a degree of freedom of viewpoint, The system includes a transmission unit that transmits the configuration of the aforementioned content, The transmission unit transmits quality information of a stream relating to a 3D object included in a specific content which includes at least the content, The quality information includes quality information for each of two or more streams, each with different quality depending on the orientation of the three-dimensional object, and is used to select a stream from the two or more streams to be displayed by the receiving device; this is a transmitting device.
8. A receiving unit that receives the configuration of content with a degree of freedom of viewpoint, The system comprises a display unit for playing the aforementioned content, The receiving unit receives quality information of a stream relating to a 3D object included in a specific content which includes at least the content, The quality information includes quality information for each of two or more streams, each with different quality depending on the orientation of the three-dimensional object, and is used to select a stream to be displayed on the display unit from among the two or more streams, in the receiving device.
Citation Information
Patent Citations
Method, device and computer program for enhancing streaming of virtual reality media content
JP2019524004A
Image processing device and file generating device
WO2019054202A1