Transmission method and transmission apparatus
By generating and transmitting a combined video and subtitle stream with embedded abstract information, the processing load for displaying TTML-formatted subtitles is reduced, allowing efficient control of display timing and state without full file scanning.
Patent Information
- Application Number
- JP2025192732
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2015-07-16
- Filing Date
- 2025-11-12
- Publication Date
- 2026-01-23
AI Technical Summary
Existing subtitle display systems require extensive file scanning to obtain important parameters, leading to increased processing load on the receiving side when handling TTML-formatted text information.
A video encoding unit generates a video stream with encoded data, a subtitle encoding unit creates a subtitle stream including display timing and abstract information, and a transmitting unit sends a container format combining both streams, allowing the receiving side to control subtitle display timing and state without full file scanning.
This approach reduces the processing load for displaying subtitles by enabling the receiving side to utilize abstract information for timing and state control, thereby optimizing subtitle processing efficiency.
Smart Images

Figure 2026012503000001_ABST
Abstract
Description
[Technical Field]
[0001] The present technology relates to a transmission method and a transmission device. [Background technology]
[0002] In the past, for example, in DVB (Digital Video Broadcasting) broadcasts, subtitle information was transmitted as bitmap data. Recently, it has been proposed to transmit subtitle information in text character code, i.e., text-based. In this case, the font is expanded according to the resolution on the receiving side.
[0003] Furthermore, when transmitting subtitle information in a text-based format, it has been proposed to include timing information in the text information. For example, the World Wide Web Consortium (W3C) has proposed TTML (Timed Text Markup Language) as this text information (see Patent Document 1). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-169885 Summary of the Invention [Problem to be solved by the invention]
[0005] The text information of subtitles expressed in TTML is handled as a file in the form of a markup language. In this case, there are no restrictions on the order in which parameters can be transmitted, so the receiver must scan the entire file to obtain important parameters.
[0006] An object of this technology is to reduce the processing load for displaying subtitles on the receiving side. [Means for solving the problem]
[0007] The concept of this technology is: a video encoding unit for generating a video stream including encoded video data; a subtitle encoding unit that generates a subtitle stream including subtitle text information having display timing information and abstract information having information corresponding to a part of the plurality of pieces of information indicated by the text information; a transmitting unit for transmitting a container in a predetermined format including the video stream and the subtitle stream; Located in the transmitting device.
[0008] In this technology, a video encoding unit generates a video stream including encoded video data. A subtitle encoder generates a subtitle stream including subtitle text information having display timing information and abstract information having information corresponding to a portion of the multiple pieces of information indicated by the text information. For example, the subtitle text information may be in TTML or a format derived from TTML. Then, a transmitting unit transmits a container in a predetermined format including the video stream and the subtitle stream.
[0009] For example, the abstract information may include subtitle display timing information, which allows the receiving side to control the subtitle display timing based on the subtitle display timing information included in the abstract information without scanning the subtitle text information.
[0010] In this case, for example, the subtitle display timing information may include information on the display start timing and display period. In this case, the subtitle stream may be configured with PES packets each consisting of a PES header and a PES payload, the subtitle text information and abstract information may be placed in the PES payload, and the display start timing may be indicated by a display offset from a PTS (Presentation Time Stamp) inserted in the PES header.
[0011] Furthermore, for example, the abstract information may include display control information for controlling the display state of the subtitles, so that the receiving side can control the display state of the subtitles based on the display control information included in the abstract information without scanning the text information of the subtitles.
[0012] In this case, for example, the display control information may include at least information on any one of the subtitle display position, color gamut, and dynamic range, and may further include information on the target video.
[0013] Furthermore, for example, the abstract information may include notification information notifying that there is a change in an element of the subtitle text information, which allows the receiving side to easily recognize that there is a change in the element of the subtitle text information and to efficiently scan the element of the subtitle text information.
[0014] Alternatively, for example, the subtitle encoding unit may segment the subtitle text information and abstract information to generate a subtitle stream having a predetermined number of segments, in which case the receiving side can easily obtain the abstract information by extracting the segment containing the abstract information from the subtitle stream.
[0015] In this case, for example, the subtitle stream may be arranged so that the abstract information segment is arranged first, followed by the subtitle text information segment, which allows the receiving side to easily and efficiently extract the abstract information segment from the subtitle stream.
[0016] In this way, in this technology, the subtitle stream contains not only the subtitle text information but also abstract information corresponding to the text information, so that the receiving side can perform processing for displaying the subtitles using the abstract information, thereby reducing the processing load.
[0017] Another concept of the present technology is a receiving unit for receiving a container in a predetermined format including a video stream and a subtitle stream; the video stream comprises encoded video data; the subtitle stream includes subtitle text information having display timing information and abstract information having information corresponding to a part of the plurality of pieces of information indicated by the text information, The video decoding device further includes a control unit for controlling a video decoding process for decoding the video stream to obtain video data, a subtitle decoding process for decoding the subtitle stream to obtain subtitle bitmap data and extracting the abstract information, a video overlay process for overlaying the subtitle bitmap data on the video data to obtain video data for display, and a bitmap data process for processing the subtitle bitmap data overlaid on the video data based on the abstract information. It is in the receiving device.
[0018] In this technology, a container in a predetermined format including a video stream and a subtitle stream is received. The video stream includes encoded video data. The subtitle stream includes subtitle text information having display timing information and abstract information having information corresponding to a portion of the multiple pieces of information indicated by the text information.
[0019] The control unit controls video decoding, subtitle decoding, video overlay, and bitmap data processing. In the video decoding process, the video stream is decoded to obtain video data. In the subtitle decoding process, the subtitle stream is decoded to obtain subtitle bitmap data and extract abstract information.
[0020] In the video overlay process, subtitle bitmap data is overlaid on video data to obtain video data for display. In the bitmap data process, the subtitle bitmap data to be overlaid on video data is processed based on abstract information.
[0021] For example, the abstract information may include subtitle display timing information, and in the bitmap data processing, the timing of superimposing the subtitle bitmap data onto the video data may be controlled based on the subtitle display timing information.
[0022] Also, for example, the abstract information may include display control information that controls the display state of the subtitle, and in the bitmap data processing, the state of the bitmap of the subtitle superimposed on the video data may be controlled based on the display control information.
[0023] In this way, the present technology processes the bitmap data of the subtitles superimposed on the video data based on the abstract information extracted from the subtitle stream, thereby reducing the processing load for displaying the subtitles.
[0024] Another concept of the present technology is a video encoding unit for generating a video stream including encoded video data; a subtitle encoding unit that generates one or more segments in which elements of subtitle text information having display timing information are arranged, and generates a subtitle stream including the one or more segments; a transmitting unit for transmitting a container in a predetermined format including the video stream and the subtitle stream; Located in the transmitting device.
[0025] In this technology, a video encoding unit generates a video stream including encoded video data. A subtitle encoding unit generates one or more segments in which elements of subtitle text information having display timing information are arranged, and a subtitle stream including the one or more segments is generated. For example, the subtitle text information may be in TTML or a format derived from TTML. A transmitting unit transmits a container in a predetermined format including the video stream and the subtitle stream.
[0026] In this way, with this technology, subtitle text information with display timing information is segmented and transmitted in a subtitle stream, which allows the receiving side to receive each element of the subtitle text information satisfactorily.
[0027] In the present technology, for example, when the subtitle encoding unit generates one segment in which all elements of subtitle text information are arranged, the subtitle encoding unit may insert information regarding the transmission order and / or whether or not the subtitle text information has been updated into the layer of the segment or the layer of the elements arranged therein. By inserting information regarding the transmission order of the subtitle text information, the receiving side can recognize the transmission order of the subtitle text information, thereby enabling efficient decoding processing. Furthermore, by inserting information regarding whether or not the subtitle text information has been updated, the receiving side can easily know whether or not the subtitle text information has been updated. [Effects of the Invention]
[0028] According to the present technology, it is possible to reduce the processing load for displaying subtitles on the receiving side. Note that the effects described in this specification are merely examples and are not limiting, and additional effects may also be provided. [Brief explanation of the drawings]
[0029] [Figure 1] 1 is a block diagram showing an example of the configuration of a transmission / reception system according to an embodiment; [Figure 2] FIG. 2 is a block diagram illustrating an example of the configuration of a transmitting device. [Figure 3] FIG. 10 is a diagram illustrating an example of photoelectric conversion characteristics. [Figure 4] FIG. 10 is a diagram illustrating an example of the structure of a dynamic range SEI message and the contents of the main information in the example structure. [Figure 5] FIG. 1 is a diagram illustrating a TTML structure. [Figure 6] FIG. 1 is a diagram illustrating a TTML structure. [Figure 7] An example of the structure of metadata (TTM: TTML Metadata) present in the header (head) of the TTML structure is shown. [Figure 8]FIG. 10 is a diagram showing an example of the structure of styling (TTS: TTML Styling) present in the header (head) of the TTML structure. [Figure 9] FIG. 1 is a diagram showing an example of the structure of a styling extension (TTSE: TTML Styling Extension) present in the header (head) of the TTML structure. [Figure 10] FIG. 10 is a diagram showing an example of the structure of a layout (TTL: TTML layout) present in the header (head) of the TTML structure. [Figure 11] FIG. 10 is a diagram showing an example of the structure of a body of a TTML structure. [Figure 12] FIG. 10 is a diagram illustrating an example of the configuration of a PES packet. [Figure 13] This figure shows the segment interfaces inside the PES. [Figure 14] A diagram showing an example of the structure of "TimedTextSubtitling_segments()" placed in the PES data payload. [Figure 15] A figure showing another example structure of ``TimedTextSubtitling_segments()'' placed in the PES data payload. [Figure 16] FIG. 10 is a diagram illustrating an example of the structure of a THMS (text_header_metadata_segment) in which metadata (TTM) is placed. [Figure 17] FIG. 10 is a diagram showing an example of the structure of a THSS (text_header_styling_segment) in which styling (TTS) is placed. [Figure 18] This is a diagram showing an example of the structure of THSES (text_header_styling_extension_segment) in which styling extensions (TTML) are placed. [Figure 19] FIG. 10 is a diagram illustrating an example of the structure of a THLS (text_header_layout_segment) in which a layout (TTL) is placed. [Figure 20]FIG. 10 is a diagram showing an example of the structure of a TBS (text_body_segment) in which a body of a TTML structure is placed. [Figure 21] FIG. 10 is a diagram showing an example structure of THAS (text_header_all_segment) in which the header (head) of the TTML structure is placed. [Figure 22] FIG. 10 is a diagram showing an example of the structure of a TWS (text whole segment) in which the entire TTML structure is placed. [Figure 23] FIG. 10 is a diagram showing another example structure of a TWS (text whole segment) in which the entire TTML structure is placed. [Figure 24] FIG. 10 is a diagram (1 / 2) showing an example of the structure of an APTS (abstract_parameter_TimedText_segment) in which abstract information is placed. [Figure 25] FIG. 2 is a diagram (2 / 2) showing an example of the structure of an APTS (abstract_parameter_TimedText_segment) in which abstract information is placed. [Figure 26] FIG. 1 is a diagram (1 / 2) showing the contents of the main information in an example of the APTS structure. [Figure 27] FIG. 2 is a diagram (2 / 2) showing the contents of the main information in an example of the APTS structure. [Figure 28] FIG. 10 is a diagram for explaining the settings of "PTS," "start_time_offset," and "end_time_offset" when converting TTML into segments. [Figure 29] FIG. 2 is a block diagram illustrating an example of the configuration of a receiving device. [Figure 30] FIG. 10 is a block diagram showing an example of the configuration of a subtitle decoder. [Figure 31] FIG. 2 is a diagram illustrating an example of the configuration of a color gamut / luminance conversion unit. [Figure 32] 10 is a diagram showing an example of the configuration of a component related to a luminance signal Y included in a luminance conversion unit. FIG. [Figure 33] FIG. 4 is a diagram illustrating the operation of a luminance conversion unit. [Figure 34] 10A and 10B are diagrams for explaining position conversion in a position / size conversion unit. [Figure 35] 10A and 10B are diagrams for explaining size conversion in a position / size conversion unit. [Figure 36] FIG. 10 is a diagram illustrating an example of display control of subtitles in chronological order. [Figure 37] FIG. 10 is a diagram illustrating an example of display control of subtitles in chronological order. DETAILED DESCRIPTION OF THE INVENTION
[0030] Hereinafter, modes for carrying out the invention (hereinafter referred to as "embodiments") will be described. The description will be made in the following order. 1. Embodiment 2. Variations
[0031] <1. Embodiment> [Example of transmission and reception system configuration] 1 shows an example of the configuration of a transmission / reception system 10 according to an embodiment. The transmission / reception system 10 includes a transmission device 100 and a reception device 200.
[0032] The transmitting device 100 generates an MPEG2 transport stream TS as a container and transmits this transport stream TS via broadcast waves or network packets. This transport stream TS includes a video stream having coded video data.
[0033] This transport stream TS also includes a subtitle stream. This subtitle stream includes text information of subtitles (captions) having display timing information, and abstract information having information corresponding to some of the multiple pieces of information indicated by this text information. In this embodiment, the text information is, for example, TTML (Timed Text Markup Language) proposed by the W3C (World Wide Web Consortium).
[0034] In this embodiment, the abstract information includes subtitle display timing information. This display timing information has information about the display start timing and display period. Here, the subtitle stream is composed of PES packets each consisting of a PES header and a PES payload, and the subtitle text information and display timing information are placed in the PES payload. For example, the display start timing is indicated by a display offset from the PTS inserted in the PES header.
[0035] In this embodiment, the abstract information also includes display control information for controlling the display state of the subtitles. In this embodiment, the display control information also includes information on the display position, color gamut, and dynamic range of the subtitles. In this embodiment, the abstract information also includes information on the target video.
[0036] The receiving device 200 receives the transport stream TS transmitted by broadcast waves from the transmitting device 100. As described above, this transport stream TS includes a video stream including coded video data and a subtitle stream including subtitle text information and abstract information.
[0037] The receiving device 200 obtains video data from the video stream, obtains subtitle bitmap data from the subtitle stream, and extracts abstract information. The receiving device 200 superimposes the subtitle bitmap data on the video data to obtain video data for display. The television receiver 200 processes the subtitle bitmap data superimposed on the video data based on the abstract information.
[0038] In this embodiment, the abstract information includes subtitle display timing information, and the receiving device 200 controls the timing of superimposing the subtitle bitmap data on the video data based on the display timing information. Also, in this embodiment, the abstract information includes display control information for controlling the subtitle display state (display position, color gamut, dynamic range, etc.), and the receiving device 200 controls the subtitle bitmap state based on the display control information.
[0039] "Example of transmitter configuration" 2 shows an example of the configuration of the transmitting device 100. This transmitting device 100 has a control unit 101, a camera 102, a video photoelectric conversion unit 103, an RGB / YCbCr conversion unit 104, a video encoder 105, a subtitle generation unit 106, a text format conversion unit 107, a subtitle encoder 108, a system encoder 109, and a transmitting unit 110.
[0040] The control unit 101 is configured with a CPU (Central Processing Unit) and controls the operation of each unit of the transmission device 100 based on a control program. The camera 102 captures an image of a subject and outputs HDR (High Dynamic Range) or SDR (Standard Dynamic Range) video data (image data). An HDR image has a contrast ratio of 0 to 100%*N (N is a number greater than 1), for example, 0 to 1000%, which exceeds the white peak brightness of an SDR image. Here, a level of 100% corresponds to a white luminance value of 100 cd / m2, for example.
[0041] The video photoelectric conversion unit 103 performs photoelectric conversion on the video data obtained by the camera 102 to obtain transmission video data V1. In this case, if the video data is SDR video data, the video photoelectric conversion unit 103 applies SDR photoelectric conversion characteristics to perform photoelectric conversion, thereby obtaining SDR transmission video data (transmission video data having SDR photoelectric conversion characteristics). On the other hand, if the video data is HDR video data, the video photoelectric conversion unit 103 applies HDR photoelectric conversion characteristics to perform photoelectric conversion, thereby obtaining HDR transmission video data (transmission video data having HDR photoelectric conversion characteristics).
[0042] The RGB / YCbCr conversion unit 104 converts the transmission video data V1 from the RGB domain to the YCbCr (luminance / chrominance) domain. The video encoder 105 encodes the transmission video data V1 converted to the YCbCr domain using, for example, MPEG4-AVC or HEVC, and generates a video stream (PES stream) VS including the encoded video data.
[0043] At this time, the video encoder 105 inserts meta information such as information indicating the electro-optical conversion characteristics (transfer function) corresponding to the electro-optical conversion characteristics of the transmitted video data V1, information indicating the color gamut of the transmitted video data V1, and information indicating the reference level into the VUI (video usability information) area of the SPS NAL unit of the access unit (AU).
[0044] In addition, the video encoder 105 inserts a newly defined dynamic range SEI message (Dynamic_range SEI message) into the "SEIs" portion of the access unit (AU), which contains meta information such as information (transfer function) indicating the electrical-optical conversion characteristics corresponding to the electrical-optical conversion characteristics of the transmitted video data V1, and reference level information.
[0045] Here, the reason why the dynamic range SEI message contains information indicating the electro-optical conversion characteristics is that even if the transmission video data V1 is HDR transmission video data, if the HDR electro-optical conversion characteristics are compatible with the SDR electro-optical conversion characteristics, information indicating the electro-optical conversion characteristics (gamma characteristics) corresponding to the SDR electro-optical conversion characteristics is inserted into the VUI of the SPS NAL unit, and therefore information indicating the electro-optical conversion characteristics corresponding to the HDR electro-optical conversion characteristics is required in a location other than the VUI.
[0046] FIG. 3 shows an example of photoelectric conversion characteristics. In this figure, the horizontal axis represents the input luminance level, and the vertical axis represents the transmission code value. Curve a shows an example of SDR photoelectric conversion characteristics. Curve b1 shows an example of HDR photoelectric conversion characteristics (not compatible with SDR photoelectric conversion characteristics). Curve b2 shows an example of HDR photoelectric conversion characteristics (compatible with SDR photoelectric conversion characteristics). In this example, the input luminance level matches the SDR photoelectric conversion characteristics up to the compatibility limit value. When the input luminance level is at the compatibility limit value, the transmission code value is at a compatible level.
[0047] The dynamic range SEI message contains reference level information because, when the transmission video data V1 is SDR transmission video data, information indicating the electro-optical conversion characteristics (gamma characteristics) corresponding to the SDR electro-optical conversion characteristics is inserted into the VUI of the SPS NAL unit, but the standard does not specify the insertion of a reference level.
[0048] Figure 4(a) shows an example of the structure (Syntax) of the Dynamic Range SEI message. Figure 4(b) shows the main information content (Semantics) of this example structure. The 1-bit flag information "Dynamic_range_cancel_flag" indicates whether to refresh the "Dynamic_range" message. "0" indicates that the message is refreshed, and "1" indicates that the message is not refreshed, i.e., the previous message is maintained as is.
[0049] When "Dynamic_range_cancel_flag" is "0", the following fields exist. The 8-bit field "coded_data_bit_depth" indicates the number of coded pixel bits. The 8-bit field "reference_level" indicates the reference luminance level value as the reference level. The 1-bit field "modify_tf_flag" indicates whether to modify the Transfer Function (TF) indicated in the VUI (video usability information). "0" indicates that the TF indicated in the VUI is the target, and "1" indicates that the TF of the VUI is modified by the TF specified in "transfer_function" of this SEI. The 8-bit field "transfer_function" indicates the electrical-to-optical conversion characteristics corresponding to the optical-to-electrical conversion characteristics of the transmitted video data V1.
[0050] 2, a subtitle generation unit 106 generates text data (character code) DT as subtitle information. A text format conversion unit 107 receives the text data DT and obtains subtitle text information in a predetermined format, which in this embodiment is TTML (Timed Text Markup Language).
[0051] Figure 5 shows an example of a TTML (Timed Text Markup Language) structure. TTML is written in XML format. Figure 6(a) also shows an example of a TTML structure. As in this example, it is also possible to specify a subtitle area at the root container position using "tts:extent". Figure 6(b) shows a subtitle area of 1920 pixels horizontally and 1080 pixels vertically, specified by "tts:extent="1920px 1080px"".
[0052] TTML consists of a header and a body. The header contains elements such as metadata, styling, styling extensions, and layout. Figure 7 shows an example of the structure of metadata (TTM: TTML Metadata). This metadata includes information on the metadata title and copyright information.
[0053] Figure 8(a) shows an example of the structure of styling (TTS: TTML Styling). In addition to an identifier (id), this styling includes information such as the region position, size, color, font (fontFamily), font size (fontSize), and text alignment (textAlign).
[0054] "tts:origin" specifies the start position of the region, which is the display area for subtitles, in pixels. In this example, it is "tts:origin“480px 600px”", which indicates that the start position (see arrow P) is (480,600), as shown in Figure 8(b). "tts:extent" specifies the end position of the region in terms of the number of horizontal and vertical pixel offsets from the start position. In this example, it is "tts:extent“560px 350px”", which indicates that the end position (see arrow Q) is (480 + 560,600 + 350), as shown in Figure 8(b). Here, the number of offset pixels corresponds to the horizontal and vertical size of the region.
[0055] "tts:opacity="1.0"" indicates the mixing ratio of subtitles (subtitles) and background video. For example, "1.0" indicates that the subtitles are 100% and the background video is 0%, while "0.1" indicates that the subtitles (subtitles) are 0% and the background video is 100%. In the example shown, it is set to "1.0".
[0056] Figure 9 shows an example of the structure of a styling extension (TTML Styling Extension). In addition to an identifier (id), this styling extension contains information on color space and dynamic range. The color gamut information specifies the color gamut expected for the subtitle. In the example shown, it is "ITUR2020". The dynamic range information specifies whether the dynamic range expected for the subtitle is SDR or HDR. In the example shown, it is SDR.
[0057] Figure 10 shows an example of the layout (region: TTML layout) structure. This layout includes information such as the identifier (id) of the region where the subtitles are placed, as well as the offset (padding), background color (backgroundColor), and alignment (displayAlign).
[0058] Fig. 11 shows an example of the structure of a body. In the example shown, information on three subtitles, subtitle 1, subtitle 2, and subtitle 3, is included. For each subtitle, the display start timing and display end timing are recorded, along with text data. For example, for subtitle 1, the display start timing is "T1," the display end timing is "T3," and the text data is "ABC."
[0059] Returning to FIG. 2, the subtitle encoder 108 converts the TTML obtained by the text format conversion unit 107 into various segments, and generates a subtitle stream SS made up of PES packets in which these segments are arranged in the payload.
[0060] Figure 12 shows an example of the structure of a PES packet. The PES header contains a Presentation Time Stamp (PTS). The PES data payload contains the following segments: abstract_parameter_TimedText_segment (APTS), text_header_metadata_segment (THMS), text header styling segment (THSS), text_header_styling_extension_segment (THSES), text_header_layout_segment (THLS), and text_body_segment (TBS).
[0061] The PES data payload may contain the following segments: APTS (abstract_parameter_TimedText_segment), THAS (text_header_all_segment), and TBS (text_body_segment). The PES data payload may also contain the following segments: APTS (abstract_parameter_TimedText_segment) and TWS (text_whole_segment).
[0062] Figure 13 shows the segment interface within PES. The "PES_data_field" indicates the container portion of the PES data payload of a PES packet. The 8-bit field "data_identifier" indicates an ID that identifies the type of data transmitted in the container portion. Conventional subtitles (in the case of bitmaps) are indicated by "0x20," but text can also be identified by a new value, for example, "0x21."
[0063] The 8-bit field "subtitle_stream_id" indicates an ID that identifies the type of subtitle stream. A subtitle stream that transmits text information can be identified by a new value, such as "0x01," to distinguish it from the conventional subtitle stream "0x00" that transmits bitmap information.
[0064] Segments are placed in the "TimedTextSubtitling_segments()" field. Figure 14 shows an example of the structure of "TimedTextSubtitling_segments()" when the following segments are placed in the PES data payload: APTS (abstract_parameter_TimedText_segment), THMS (text_header_metadata_segment), THSS (text header styling segment), THSES (text_header_styling_extension_segment), THLS (text_header_layout_segment), and TBS (text_body_segment).
[0065] Figure 15(a) shows an example of the structure of "TimedTextSubtitling_segments()" when the segments APTS (abstract_parameter_TimedText_segment), THAS (text_header_all_segment), and TBS (text_body_segment) are placed in the PES data payload. Figure 15(b) shows an example of the structure of "TimedTextSubtitling_segments()" when the segments APTS (abstract_parameter_TimedText_segment) and TWS (text_whole_segment) are placed in the PES data payload.
[0066] The insertion of each segment into the subtitle stream is flexible. For example, if there are no changes other than the displayed subtitles, the stream will consist of only two segments: an APTS (abstract_parameter_TimedText_segment) and a TBS (text_body_segment). In either case, the PES data payload will contain the APTS segment containing abstract information first, followed by the other segments. This arrangement allows the receiving side to easily and efficiently extract the abstract information segment from the subtitle stream.
[0067] FIG. 16(a) shows an example structure (syntax) of THMS (text_header_metadata_segment). This structure includes the following information: "sync_byte," "segment_type," "page_id," "segment_length," "thm_version_number," and "segment_payload()." "segment_type" is 8-bit data indicating the segment type, and in this case is set to, for example, "0x20" indicating THMS. "segment_length" is 8-bit data indicating the length (size) of the segment. Metadata such as that shown in FIG. 16(b) is placed as XML information within "segment_payload()." This metadata is the same as the metadata elements present in the TTML header (see FIG. 7).
[0068] FIG. 17(a) shows an example structure (syntax) of THSS (text_header_styling_segment). This structure includes the following information: "sync_byte," "segment_type," "page_id," "segment_length," "ths_version_number," and "segment_payload()." "segment_type" is 8-bit data indicating the segment type, and in this case is set to, for example, "0x21" indicating THSS. "segment_length" is 8-bit data indicating the length (size) of the segment. Metadata such as that shown in FIG. 17(b) is placed as XML information within "segment_payload()." This metadata is the same as the styling element present in the TTML header (see FIG. 8(a)).
[0069] Figure 18(a) shows an example structure (syntax) of THSES (text_header_styling_extension_segment). This structure includes the following information: "sync_byte," "segment_type," "page_id," "segment_length," "thse_version_number," and "segment_payload()." "segment_type" is 8-bit data indicating the segment type, and in this case is set to "0x22," for example, indicating THSES. "segment_length" is 8-bit data indicating the length (size) of the segment. Metadata such as that shown in Figure 18(b) is placed as XML information within "segment_payload()." This metadata is the same as the styling extension (styling_extension) element in the TTML header (see Figure 9(a)).
[0070] FIG. 19(a) shows an example structure (syntax) of THLS (text_header_layout_segment). This structure includes the following information: "sync_byte," "segment_type," "page_id," "segment_length," "thl_version_number," and "segment_payload()." "segment_type" is 8-bit data indicating the segment type, and in this case is set to, for example, "0x23," indicating THLS. "segment_length" is 8-bit data indicating the length (size) of the segment. Metadata such as that shown in FIG. 19(b) is placed as XML information within "segment_payload()." This metadata is the same as the layout element present in the TTML header (see FIG. 10).
[0071] FIG. 20(a) shows an example structure (syntax) of a TBS (text_body_segment). This structure includes the following information: "sync_byte," "segment_type," "page_id," "segment_length," "tb_version_number," and "segment_payload()." "segment_type" is 8-bit data indicating the segment type, and in this case is set to, for example, "0x24," indicating TBS. Metadata such as that shown in FIG. 20(b) is placed as XML information within "segment_payload()." This metadata is the same as the TTML body (see FIG. 11).
[0072] FIG. 21(a) shows an example structure (syntax) of THAS (text_header_all_segment). This structure includes the following information: "sync_byte", "segment_type", "page_id", "segment_length", "tha_version_number", and "segment_payload()". "segment_type" is 8-bit data indicating the segment type, and in this case is set to, for example, "0x25" indicating THAS. "segment_length" is 8-bit data indicating the length (size) of the segment. Metadata such as that shown in FIG. 21(b) is placed in "segment_payload()" as XML information. This metadata is the entire header (head).
[0073] FIG. 22(a) shows an example structure (syntax) of a TWS (text whole segment). This structure includes the following information: "sync_byte," "segment_type," "page_id," "segment_length," "tw_version_number," and "segment_payload()." "segment_type" is 8-bit data indicating the segment type, and in this case is set to, for example, "0x26," indicating TWS. "segment_length" is 8-bit data indicating the length (size) of the segment. Metadata such as that shown in FIG. 22(b) is placed as XML information within "segment_payload()." This metadata is the entire TTML (see FIG. 5). This structure is intended to maintain compatibility across the entire TTML, and places the entire TTML in a single segment.
[0074] When all TTML elements are placed in one segment and sent in this way, two new elements, "ttnew:sequentialinorder" and "ttnew:partialupdate", are inserted into the element layer as shown in Figure 22(b). Note that these do not have to be inserted at the same time.
[0075] "ttnew:sequentialinorder" configures information about the transmission order of TTML. This "ttnew:sequentialinorder" is placed before . "ttnew:sequentialinorder=true (=1)" indicates that there is a restriction on the transmission order. In that case, the inside of is <metadata> 、 <styling>、<styling extension>、 <layout>are arranged in the order of ", followed by the contents of " 、 text " indicates that a comma follows.<styling extension> If does not exist, <metadata> 、 <styling> 、 <layout>On the other hand, "ttnew:sequentialinorder=false (=0)" indicates that there is no such constraint.
[0076] By inserting the "ttnew:sequentialinorder" element in this way, the receiving side can recognize the transmission order of the TTML, so even if all TTML elements are sent together, it can recognize that the TTML transmission order is in accordance with a specified order, simplifying the process up to decoding and enabling the decoding process to be carried out efficiently.
[0077] Additionally, "ttnew:partialupdate" constitutes information regarding whether or not TTML has been updated. This "ttnew:partialupdate" is placed before . "ttnew:partialupdate=true (=1)" indicates that an update to one of the elements will occur before . On the other hand, "ttnew:partialupdate=false (=0)" indicates that there will be no such update. By inserting the "ttnew:sequentialinorder" element in this way, the receiving side can easily determine whether or not TTML has been updated.
[0078] In the above example, two new elements, "ttnew:sequentialinorder" and "ttnew:partialupdate", are inserted into the element layer. However, as shown in Figure 23(a), these new elements may also be inserted into the segment layer. Figure 23(b) shows the metadata (XML information) placed in "segment_payload()" in this case.
[0079] "APTS(abstract_parameter_TimedText_segment) segment" Here, we will explain the APTS (abstract_parameter_TimedText_segment) segment. This APTS segment contains abstract information. This abstract information contains information that corresponds to some of the multiple pieces of information indicated in TTML.
[0080] Figures 24 and 25 show an example structure (syntax) of an APTS (abstract_parameter_TimedText_segment). Figures 26 and 27 show the contents (semantics) of the main information in this example structure. Like other segments, this structure includes the information "sync_byte," "segment_type," "page_id," and "segment_length." "segment_type" is 8-bit data indicating the segment type, and in this case is set to, for example, "0x19," indicating APTS. "segment_length" is 8-bit data indicating the length (size) of the segment.
[0081] The 4-bit field "APT_version_number" indicates whether there is a change from the content previously sent in the APTS (abstract_parameter_TimedText_segment) element, and if there is a change, the value is increased by 1. The 4-bit field "TTM_version_number" indicates whether there is a change from the content previously sent in the THMS (text_header_metadata_segment) element, and if there is a change, the value is increased by 1. The 4-bit field "TTS_version_number" indicates whether there is a change from the content previously sent in the THSS (text_header_styling_segment) element, and if there is a change, the value is increased by 1.
[0082] The 4-bit field "TTSE_version_number" indicates whether there is a change from the content previously sent in the THSES (text_header_styling_extension_segment) element, and if there is a change, the value is increased by 1. The 4-bit field "TTL_version_number" indicates whether there is a change from the content previously sent in the THLS (text_header_layout_segment) element, and if there is a change, the value is increased by 1.
[0083] The 4-bit field "TTHA_version_number" indicates whether there is a change from the content previously sent in the THAS (text_header_all_segment) element, and if there is a change, the value is increased by 1. The 4-bit field "TW_version_number" indicates whether there is a change from the content previously sent in the TWS (text whole segment) element, and if there is a change, the value is increased by 1.
[0084] The 4-bit field of "subtitle_display_area" specifies the subtitle display area. For example, "0x1" specifies 640h*480v, "0x2" specifies 720h*480v, "0x3" specifies 720h*576v, "0x4" specifies 1280h*720v, "0x5" specifies 1920h*1080v, "0x6" specifies 3840h*2160v, and "0x7" specifies 7680h*4320v.
[0085] The 4-bit field "subtitle_color_gamut_info" specifies the color gamut assumed for the subtitle. The 4-bit field "subtitle_dynamic_range_info" specifies the dynamic range assumed for the subtitle. For example, "0x1" indicates SDR, and "0x2" indicates HDR. Specifying HDR for the subtitle indicates that the brightness of the subtitle is expected to be kept below the standard white level of the video.
[0086] The 4-bit field "target_video_resolution" specifies the expected video resolution. For example, "0x1" specifies 640h*480v, "0x2" specifies 720h*480v, "0x3" specifies 720h*576v, "0x4" specifies 1280h*720v, "0x5" specifies 1920h*1080v, "0x6" specifies 3840h*2160v, and "0x7" specifies 7680h*4320v.
[0087] The 4-bit field "target_video_color_gamut_info" specifies the expected video color gamut. For example, "0x1" indicates "BT.709", and "0x2" indicates "BT.2020". The 4-bit field "target_video_dynamic_range_info" specifies the expected video dynamic range. For example, "0x1" indicates "BT.709", "0x2" indicates "BT.202x", and "0x3" indicates "Smpte 2084".
[0088] The 4-bit field "number_of_regions" specifies the number of regions. The following fields are repeated for each region. The 16-bit field "region_id" indicates the region ID.
[0089] The 8-bit field "start_time_offset" indicates the subtitle display start time as an offset value from the PTS. The offset value of this "start_time_offset" is signed, and a negative value indicates that the display will start earlier than the PTS. When the offset value of this "start_time_offset" is 0, it means that the display will start at the timing of the PTS. The precision of the value in 8-bit representation is up to one decimal place, obtained by dividing the sign value by 10.
[0090] The 8-bit field "end_time_offset" indicates the end time of the subtitle display as an offset value from "start_time_offset". This offset value indicates the display period. When the offset value of "start_time_offset" mentioned above is 0, it indicates that the display will end at the timing of the value obtained by adding the offset value of "end_time_offset" to the PTS. In the case of 8-bit representation, the precision of the value will be up to one decimal place, obtained by dividing the sign value by 10.
[0091] It is also possible to transmit "start_time_offset" and "end_time_offset" with the same 90 kHz precision as the PTS. In that case, 32 bits of space are reserved for each of the "start_time_offset" and "end_time_offset" fields.
[0092] As shown in Fig. 28, when converting TTML into segments, the subtitle encoder 108 sets the "PTS," "start_time_offset," and "end_time_offset" of each subtitle based on the description of the display start timing (begin) and display end timing (end) of each subtitle included in the TTML body, and by referencing the system time information (PCR, video / audio synchronization time). At this time, the subtitle encoder 108 may use a decoder buffer model to set the "PTS," "start_time_offset," and "end_time_offset" while verifying that the receiving side is operating correctly.
[0093] The 16-bit field "region_start_horizontal" indicates the horizontal pixel position of the upper left corner point of the region in the subtitle display area specified by the above-mentioned "subtitle_display_area" (see point P in Figure 8(b)). The 16-bit field "region_start_vertical" indicates the vertical pixel position of the upper left corner point of the region in the subtitle display area. The 16-bit field "region_end_horizontal" indicates the horizontal pixel position of the lower right corner point of the region in the subtitle display area (see point Q in Figure 8(b)). The 16-bit field "region_end_vertical" indicates the vertical pixel position of the lower right corner point of the region in the subtitle display area.
[0094] Returning to Fig. 2, the system encoder 109 generates a transport stream TS including the video stream VS generated by the video encoder 105 and the subtitle stream SS generated by the subtitle encoder 108. The transmitter 110 transmits this transport stream TS to the receiving device 200 via broadcast waves or network packets.
[0095] The operation of the transmitting device 100 shown in Fig. 2 will be briefly described. Video data (image data) captured by the camera 102 is supplied to the video photoelectric conversion unit 103. The video photoelectric conversion unit 103 performs photoelectric conversion on the video data captured by the camera 102 to obtain transmission video data V1.
[0096] In this case, if the video data is SDR video data, the SDR photoelectric conversion characteristics are applied to perform photoelectric conversion, resulting in SDR transmission video data (transmission video data with SDR photoelectric conversion characteristics).On the other hand, if the video data is HDR video data, the HDR photoelectric conversion characteristics are applied to perform photoelectric conversion, resulting in HDR transmission video data (transmission video data with HDR photoelectric conversion characteristics).
[0097] The transmission video data V1 obtained by the video photoelectric conversion unit 103 is converted from the RGB domain to the YCbCr (luminance and chrominance) domain by the RGB / YCbCr conversion unit 104, and then supplied to the video encoder 105. The video encoder 105 encodes the transmission video data V1 using, for example, MPEG4-AVC or HEVC, and generates a video stream (PES stream) VS including the encoded video data.
[0098] In addition, in the video encoder 105, meta information such as information indicating the electrical-to-optical conversion characteristics corresponding to the electrical-to-electrical conversion characteristics of the transmitted video data V1 (transfer function), information indicating the color gamut of the transmitted video data V1, and information indicating the reference level is inserted into the VUI area of the SPS NAL unit of the access unit (AU).
[0099] In addition, the video encoder 105 inserts a newly defined dynamic range SEI message (see Figure 4) into the "SEIs" portion of the access unit (AU), which contains meta information such as information (transfer function) indicating the electrical-to-optical conversion characteristics corresponding to the electrical-to-electrical conversion characteristics of the transmitted video data V1, and reference level information.
[0100] The subtitle generation unit 106 generates text data (character code) DT as subtitle information. This text data DT is supplied to a text format conversion unit 107. The text format conversion unit 107 converts the text data DT into subtitle text information having display timing information, i.e., TTML (see FIGS. 3 and 4). This TTML is supplied to a subtitle encoder 108.
[0101] The subtitle encoder 108 converts the TTML obtained by the text format conversion unit 107 into various segments, and generates a subtitle stream SS made up of PES packets with these segments arranged in their payloads. In this case, the payload of the PES packets begins with an APTS segment carrying abstract information (see Figures 24 to 27), followed by a segment carrying subtitle text information (see Figure 12).
[0102] The video stream VS generated by the video encoder 105 is supplied to the system encoder 109. The subtitle stream SS generated by the subtitle encoder 108 is supplied to the system encoder 109. The system encoder 109 generates a transport stream TS including the video stream VS and the subtitle stream SS. This transport stream TS is transmitted to the receiving device 200 via broadcast waves or network packets by the transmitting unit 110.
[0103] "Example of receiving device configuration" 29 shows an example configuration of a receiving device 200. This receiving device 200 has a control unit 201, a user operation unit 202, a receiving unit 203, a system decoder 204, a video decoder 205, a subtitle decoder 206, a color gamut / brightness conversion unit 207, and a position / size conversion unit 208. The receiving device 200 also has a video superimposition unit 209, a YCbCr / RGB conversion unit 210, an electro-optical conversion unit 211, a display mapping unit 212, and a CE monitor 213.
[0104] The control unit 201 is configured with a CPU (Central Processing Unit) and controls the operation of each unit of the receiving device 200 based on a control program. The user operation unit 202 is a switch, a touch panel, a remote control transmission unit, etc. that allows a user such as a viewer to perform various operations. The receiving unit 203 receives a transport stream TS transmitted from the transmitting device 100 via broadcast waves or network packets.
[0105] The system decoder 204 extracts the video stream VS and the subtitle stream SS from the transport stream TS. The system decoder 204 also extracts various pieces of information inserted in the transport stream TS (container) and sends them to the control unit 201.
[0106] The video decoder 205 performs a decoding process on the video stream VS extracted by the system decoder 204 and outputs transmission video data V1. The video decoder 205 also extracts parameter sets and SEI messages inserted in each access unit constituting the video stream VS and sends them to the control unit 201.
[0107] The VUI area of the SPS NAL unit contains information (transfer function) indicating electrical-optical conversion characteristics corresponding to the electrical-optical conversion characteristics of the transmitted video data V1, information indicating the color gamut of the transmitted video data V1, information indicating a reference level, etc. This SEI message also includes a dynamic range SEI message (see FIG. 4) that contains information (transfer function) indicating electrical-optical conversion characteristics corresponding to the electrical-optical conversion characteristics of the transmitted video data V1, information on the reference level, etc.
[0108] The subtitle decoder 206 processes the segment data of each region included in the subtitle stream SS and outputs bitmap data of each region to be superimposed on the video data. The subtitle decoder 206 also extracts abstract information included in the APTS segments and sends it to the control unit 201.
[0109] This abstract information includes subtitle display timing information, subtitle display control information (subtitle display position, color gamut, and dynamic range information), and target video information (resolution, color gamut, and dynamic range information).
[0110] Here, subtitle display timing information and display control information are included in the XML information placed in "segment_payload()" of segments other than APTS, so it is possible to obtain them by scanning that XML information, but they can also be obtained easily by simply extracting abstract information from APTS segments.Incidentally, target video information (resolution, color gamut, dynamic range information) can be obtained from the video stream VS system, but it can also be obtained easily by simply extracting abstract information from APTS segments.
[0111] 30 shows an example of the configuration of the subtitle decoder 206. The subtitle decoder 206 has a coded buffer 261, a subtitle segment decoder 262, a font development unit 263, and a bitmap buffer 264.
[0112] The coded buffer 261 temporarily stores the subtitle stream SS. The subtitle segment decoder 262 decodes the segment data of each region stored in the coded buffer 261 at a predetermined timing to obtain the text data and control codes of each region.
[0113] The font expansion unit 263 acquires subtitle bitmap data for each region by expanding the font based on the text data and control codes of each region obtained by the subtitle segment decoder 262. In this case, the font expansion unit 263 uses, for example, position information included in the abstract information ("region_start_horizontal," "region_start_vertical," "region_end_horizontal," "region_end_vertical") as position information for each region.
[0114] The subtitle bitmap data is obtained in the RGB domain. The color gamut of the subtitle bitmap data matches the color gamut indicated by the subtitle color gamut information included in the abstract information. The dynamic range of the subtitle bitmap data matches the dynamic range indicated by the subtitle dynamic range information included in the abstract information.
[0115] For example, if the dynamic range information is "SDR," the subtitle bitmap data has an SDR dynamic range and has undergone photoelectric conversion using SDR photoelectric conversion characteristics. Also, if the dynamic range information is "HDR," the subtitle bitmap data has an HDR dynamic range and has undergone photoelectric conversion using HDR photoelectric conversion characteristics. In this case, the brightness range is limited to the HDR reference level, assuming superimposition on HDR video.
[0116] The bitmap buffer 264 temporarily stores the bitmap data of each region obtained by the font development unit 263. The bitmap data of each region stored in the bitmap buffer 264 is read out from the display start timing and superimposed on the image data, and this continues for the display period.
[0117] Here, the subtitle segment decoder 262 extracts the PTS from the PES header of the PES packet. The subtitle segment decoder 262 also extracts abstract information from the APTS segment. This information is sent to the control unit 201. The control unit 201 controls the timing of reading the bitmap data of each region from the bitmap buffer 264 based on the PTS and the information "start_time_offset" and "end_time_offset" included in the abstract information.
[0118] 29, under the control of the control unit 201, the color gamut / luminance conversion unit 207 adjusts the color gamut of the subtitle bitmap data to the color gamut of the video data based on the color gamut information of the subtitle bitmap data ("subtitle_color_gamut_info") and the color gamut information of the video data ("target_video_color_gamut_info"). Also, under the control of the control unit 201, the color gamut / luminance conversion unit 207 adjusts the maximum luminance level of the subtitle bitmap data to be equal to or lower than the reference luminance level of the video data based on the dynamic range information of the subtitle bitmap data ("subtitle_dynamic_range_info") and the dynamic range information of the video data ("target_video_dynamic_range_info").
[0119] 31 shows an example of the configuration of the color gamut / luminance conversion unit 207. This color gamut / luminance conversion unit 210 has an electric-to-optical conversion unit 221, a color gamut conversion unit 222, an optical-to-electrical conversion unit 223, an RGB / YCbCr conversion unit 224, and a luminance conversion unit 225.
[0120] The electro-optical converter 221 performs photoelectric conversion on the input subtitle bitmap data. Here, when the dynamic range of the subtitle bitmap data is SDR, the electro-optical converter 221 applies SDR electro-optical conversion characteristics to perform electro-optical conversion to create a linear state. On the other hand, when the dynamic range of the subtitle bitmap data is HDR, the electro-optical converter 221 applies HDR electro-optical conversion characteristics to perform electro-optical conversion to create a linear state. Note that it is also possible that the input subtitle bitmap data is in a linear state without having been subjected to photoelectric conversion. In that case, the electro-optical converter 221 is unnecessary.
[0121] The color gamut conversion unit 222 matches the color gamut of the subtitle bitmap data output from the electro-optical conversion unit 221 to the color gamut of the video data. For example, when the color gamut of the subtitle bitmap data is "BT.709" and the color gamut of the video data is "BT.2020", the color gamut of the subtitle bitmap data is converted from "BT.709" to "BT.2020". Note that when the color gamut of the subtitle bitmap data is the same as the color gamut of the video data, the color gamut conversion unit 222 does nothing substantially and outputs the input subtitle bitmap data as is.
[0122] The photoelectric conversion unit 223 performs photoelectric conversion on the subtitle bitmap data output from the color gamut conversion unit 222 by applying the same photoelectric conversion characteristics as those applied to the video data. The RGB / YCbCr conversion unit 224 converts the subtitle bitmap data output from the photoelectric conversion unit 223 from the RGB domain to the YCbCr (luminance / chrominance) domain.
[0123] The luminance conversion unit 225 obtains output bitmap data by adjusting the subtitle bitmap data output from the RGB / YCbCr conversion unit 224 so that the maximum luminance level of the subtitle bitmap data is equal to or lower than the reference level of luminance of the video data, or equal to the reference white level. In this case, if the luminance adjustment of the subtitle bitmap data has already been performed in consideration of rendering to HDR video, and the video data is HDR, the input subtitle bitmap data is output as is without actually doing anything.
[0124] 32 shows an example of the configuration of a configuration unit 225Y related to the luminance signal Y included in the luminance conversion unit 225. This configuration unit 225Y has an encoding pixel bit number adjustment unit 231 and a level adjustment unit 232.
[0125] The coding pixel bit rate adjustment unit 231 adjusts the coding pixel bit rate of the luminance signal Ys of the subtitle bitmap data to match the coding pixel bit rate of the video data. For example, when the coding pixel bit rate of the luminance signal Ys is 8 bits and the coding pixel bit rate of the video data is 10 bits, the coding pixel bit rate of the luminance signal Ys is converted from 8 bits to 10 bits. The level adjustment unit 232 adjusts the maximum level of the luminance signal Ys, whose coding pixel bit rate has been adjusted, so that it is equal to or lower than the luminance reference level of the video data or the reference white level, and outputs the adjusted level as the output luminance signal Ys'.
[0126] Fig. 33 schematically shows the operation of component 225Y shown in Fig. 32. The example shown in the figure shows a case where the video data is HDR. The reference level corresponds to the boundary between the non-shiny part and the shiny part.
[0127] The reference level exists between the maximum level (sc_high) and minimum level (sc_low) of the luminance signal Ys after the encoding pixel bit rate has been adjusted. In this case, the maximum level (sc_high) is adjusted to be equal to or lower than the reference level. In this case, using the clipping method would result in a whiteout effect, so a linear scaling down method, for example, is used.
[0128] By adjusting the level of the luminance signal Ys in this way, when the subtitle bitmap data is superimposed on the video data, the subtitle is prevented from being displayed shining brightly against the background video, making it possible to maintain high picture quality.
[0129] The above description concerns the component 225Y (see FIG. 32) related to the luminance signal Ys included in the luminance converter 225. In the luminance converter 225, only the process of adjusting the number of encoded pixel bits for the color difference signals Cb and Cr to the number of encoded pixel bits for the video data is performed. For example, in the conversion from an 8-bit space to a 10-bit space, the entire range expressed by the bit width is set to 100%, and the median value is used as the reference value, and the conversion is performed so that the fluctuation range is 50% in the positive direction and 50% in the negative direction from the reference value.
[0130] 29, the position / size conversion unit 208 performs position conversion processing on the subtitle bitmap data obtained by the color gamut / brightness conversion unit 207 under the control of the control unit 201. When the supported resolution of the subtitle (indicated by the information in "subtitle_display_area") differs from the resolution of the video ("target_video_resolution"), the position / size conversion unit 208 converts the position of the subtitle so that the subtitle is displayed at an appropriate position in the background video.
[0131] For example, consider a case where the subtitles are HD resolution compatible and the video is UHD resolution, where UHD resolution is higher than HD resolution and includes 4K resolution or 8K resolution.
[0132] Figure 34(a) shows an example where the video is UHD resolution and the subtitles are HD resolution compatible. The subtitle display area is represented by the "subtitle area" in the figure. The positional relationship between the "subtitle area" and the video is represented by sharing the reference position of both, i.e., the top left (left-top). The pixel position of the start point of the region is (a, b), and the pixel position of its end point is (c, d). In this case, because the resolution of the background video is higher than the resolution supported by the subtitles, the display position of the subtitles on the background video is biased to the upper right, rather than the position intended by the production side.
[0133] Figure 34(b) shows an example of a case where position conversion processing has been performed. The pixel position of the start point of the region, which is the subtitle display area, is (a', b'), and the pixel position of its end point is (c', d'). In this case, the position coordinates of the region before position conversion are coordinates of the HD display area, so they are converted to coordinates of the UHD display area in accordance with their relationship with the video picture frame, and therefore based on the ratio of UHD resolution to HD resolution. Note that in this example, subtitle size conversion processing is also performed simultaneously with the position conversion.
[0134] In addition, the position / size conversion unit 208 performs size conversion processing on the subtitle bitmap data obtained by the color gamut / brightness conversion unit 207 under the control of the control unit 201, for example, in response to an operation by a user such as a viewer, or automatically based on the relationship between the video resolution and the corresponding resolution of the subtitle.
[0135] As shown in Figure 35(a), the distance from the center position of the display area (dc: display center) to the center position of the region, i.e., the point dividing the region in half horizontally and vertically (region center position: rc), is determined in proportion to the video resolution. For example, if the video resolution is assumed to be HD and the center position rc of the region is defined from the center position dc of the subtitle display area, then when the video resolution is 4K (=3840x2160), the position is controlled so that the distance from dc to rc is doubled in terms of the number of pixels.
[0136] As shown in Figure 35(b), when the size of a region is changed from r_org (Region 00) to r_mod (Region 01), the start position (rsx1, rsy1) and end position (rex1, rey1) are modified to the start position (rsx2, rsy2) and end position (rex2, rey2), respectively, so as to satisfy Ratio = (r_mod / r_org).
[0137] In other words, the ratio of the distance from rc to (rsx2,rsy2) to the distance from rc to (rsx1,rsy1), and the ratio of the distance from rc to (rex2,rey2) to the distance from rc to (rex1,rey1) are made to match Ratio. By doing this, the center position rc of the region remains the same even when size conversion is performed, and it is possible to convert the size of the subtitle (region) while maintaining a constant relative positional relationship across the entire display area.
[0138] 29 , the video superimposing unit 209 superimposes the subtitle bitmap data output from the position / size conversion unit 208 onto the transmission video data V1 output from the video decoder 205. In this case, the video superimposing unit 209 mixes the subtitle bitmap data at a mixing ratio indicated by the mixing ratio information (Mixing data) obtained by the subtitle decoder 206.
[0139] The YCbCr / RGB converter 210 converts the transmission video data V1' on which the subtitle bitmap data is superimposed from the YCbCr (luminance / chrominance) domain to the RGB domain. In this case, the YCbCr / RGB converter 210 performs the conversion using a conversion formula corresponding to the color gamut based on the color gamut information.
[0140] The electro-optical conversion unit 211 performs electro-optical conversion on the transmission video data V1' converted into the RGB domain by applying electro-optical conversion characteristics corresponding to the applied photoelectric conversion characteristics to obtain display video data for displaying an image. The display mapping unit 212 adjusts the display brightness of the display video data in accordance with the maximum brightness display capability of the CE monitor 213, etc. The CE monitor 213 displays an image based on the display video data for which the display brightness has been adjusted. The CE monitor 213 is configured, for example, with an LCD (Liquid Crystal Display), an organic EL display (organic electroluminescence display), etc.
[0141] The operation of the receiving device 200 shown in Fig. 29 will be briefly described. The receiving unit 203 receives the transport stream TS transmitted from the transmitting device 100 via broadcast waves or network packets. This transport stream TS is supplied to the system decoder 204. The system decoder 204 extracts a video stream VS and a subtitle stream SS from this transport stream TS. The system decoder 204 also extracts various pieces of information inserted in the transport stream TS (container) and sends the extracted information to the control unit 201.
[0142] The video stream VS extracted by the system decoder 204 is supplied to the video decoder 205. The video decoder 205 performs a decoding process on the video stream VS to obtain transmission video data V1. The video decoder 205 also extracts parameter sets and SEI messages inserted in each access unit constituting the video stream VS and sends them to the control unit 201.
[0143] The VUI area of the SPS NAL unit contains information (transfer function) indicating electrical-optical conversion characteristics corresponding to the electrical-optical conversion characteristics of the transmitted video data V1, information indicating the color gamut of the transmitted video data V1, information indicating a reference level, etc. This SEI message also includes a dynamic range SEI message (see FIG. 4) that contains information (transfer function) indicating electrical-optical conversion characteristics corresponding to the electrical-optical conversion characteristics of the transmitted video data V1, information on the reference level, etc.
[0144] The subtitle stream SS extracted by the system decoder 204 is supplied to the subtitle decoder 206. The subtitle decoder 206 decodes the segment data of each region included in the subtitle stream SS to obtain bitmap data of the subtitle of each region to be superimposed on the video data.
[0145] The subtitle decoder 206 also extracts abstract information contained in the APTS segments (see FIGS. 24 and 25) and sends it to the control unit 201. This abstract information includes subtitle display timing information, subtitle display control information (subtitle display position, color gamut, and dynamic range information), and target video information (resolution, color gamut, and dynamic range information).
[0146] Under the control of the control unit 201, the subtitle decoder 206 controls the output timing of the subtitle bitmap data for each region based on, for example, the subtitle display timing information ("start_time_offset", "end_time_offset") included in the abstract information.
[0147] The subtitle bitmap data for each region obtained by the subtitle decoder 206 is supplied to a color gamut and brightness conversion unit 207. Under the control of the control unit 201, the color gamut and brightness conversion unit 207 adjusts the color gamut of the subtitle bitmap data to match the color gamut of the video data, for example, based on color gamut information ("subtitle_color_gamut_info", "target_video_color_gamut_info") included in the abstract information.
[0148] In addition, under the control of the control unit 201, the color gamut / brightness conversion unit 207 adjusts the maximum brightness level of the subtitle bitmap data to be below the reference brightness level of the video data, for example, based on the dynamic range information ("subtitle_dynamic_range_info", "target_video_dynamic_range_info") included in the abstract information.
[0149] The subtitle bitmap data for each region obtained by the color gamut / brightness conversion unit 207 is supplied to the position / size conversion unit 208. Under the control of the control unit 201, the position / size conversion unit 208 performs position conversion processing on the subtitle bitmap data for each region based on, for example, resolution information ("subtitle_display_area", "target_video_resolution") included in the abstract information.
[0150] Furthermore, the position and size conversion unit 208 performs size conversion processing on the subtitle bitmap data obtained by the color gamut and brightness conversion unit 207 under the control of the control unit 201, for example, in response to an operation by a user such as a viewer, or automatically based on the relationship between the video resolution and the corresponding resolution of the subtitle.
[0151] The transmission video data V1 obtained by the video decoder 204 is supplied to a video superimposing unit 209. In addition, the subtitle bitmap data of each region obtained by the position / size conversion unit 208 is supplied to the video superimposing unit 209. The video superimposing unit 209 superimposes the subtitle bitmap data of each region on the transmission video data V1. In this case, the subtitle bitmap data is mixed at a mixing ratio indicated by mixing ratio information (Mixing data).
[0152] The transmission video data V1' obtained by the video superimposing unit 209 and onto which the subtitle bitmap data for each region is superimposed is converted from the YCbCr (luminance / chrominance) domain to the RGB domain by a YCbCr / RGB conversion unit 210 and supplied to an electric-to-optical conversion unit 211. The electric-to-optical conversion unit 211 applies electric-to-optical conversion characteristics corresponding to the applied optical-to-electrical conversion characteristics to the transmission video data V1', thereby performing electric-to-optical conversion, and obtains display video data for displaying an image.
[0153] The display video data is supplied to a display mapping unit 212. In this display mapping unit 212, the display brightness of the display video data is adjusted in accordance with the maximum brightness display capability of the CE monitor 213. The display video data that has undergone this display brightness adjustment is supplied to the CE monitor 213. An image is displayed on the CE monitor 213 based on this display video data.
[0154] As explained above, in the transmission / reception system 10 shown in Fig. 1, the subtitle stream contains not only the subtitle text information but also abstract information corresponding to that text information. Therefore, the receiving side can perform processing for displaying the subtitles using the abstract information, thereby reducing the processing load.
[0155] In this case, the processing load on the receiving side is reduced, and it becomes easy to handle chronological display control in which the subtitle display changes relatively quickly. For example, consider the case in which the subtitle display changes as shown in Figure 36(a)-(f).
[0156] In this case, first, as shown in FIG. 37(a), for example, a subtitle stream SS is transmitted, which includes PES packets in which APTS (abstract_parameter_TimedText_segment) and TBS (text body segment) segments are arranged in the PES data payload. On the receiving side, subtitle bitmap data for displaying the wording "ABC" at the position of the region "region r1" is generated based on the TBS segment data and the region position information (Region_position) included in the APTS segment.
[0157] The receiving side then outputs this bitmap data from display start timing T1 to display end timing T3 based on PTS1 and the display timing information (STS1, ETS1) included in the APTS segments. As a result, the receiving side displays the wording "ABC" continuously on the screen from T1 to T3, as shown in Figure 36.
[0158] Next, as shown in FIG. 37(b), for example, a subtitle stream SS is transmitted, which includes PES packets in which APTS (abstract_parameter_TimedText_segment) and TBS (text body segment) segments are arranged in the PES data payload. The receiving side generates subtitle bitmap data for displaying the wording "DEF" at the position of the region "region r2" based on the TBS segment data and the region position information (Region_position) included in the APTS segment.
[0159] The receiving side then outputs this bitmap data from display start timing T2 to display end timing T5 based on PTS2 and the display timing information (STS2, ETS2) included in the APTS segments. As a result, the receiving side displays the word "DEF" continuously on the screen from T2 to T5, as shown in Figure 36.
[0160] Next, as shown in FIG. 37(c), for example, a subtitle stream SS is transmitted, which includes PES packets in which APTS (abstract_parameter_TimedText_segment) and TBS (text body segment) segments are arranged in the PES data payload. The receiving side generates subtitle bitmap data for displaying the wording "GHI" at the position of the region "region r3" based on the TBS segment data and the region position information (Region_position) included in the APTS segment.
[0161] The receiving side then outputs this bitmap data from display start timing T4 to display end timing T6 based on PTS3 and the display timing information (STS3, ETS3) included in the APTS segments. As a result, the receiving side displays the message "GHI" continuously on the screen from T4 to T6, as shown in Figure 36.
[0162] <2. Modifications> In the above-described embodiment, an example has been shown in which TTML is used as text information for subtitles in a predetermined format having display timing information. However, the present technology is not limited to this, and other timed text information having information equivalent to TTML may also be used. For example, a format derived from TTML may be used.
[0163] In addition, in the above-described embodiment, an example has been shown in which the container is a transport stream (MPEG-2 TS). However, the present technology is not limited to MPEG-2 TS containers, and can be similarly implemented with containers of other formats, such as MMT or ISOBMFF.
[0164] In the above-described embodiment, an example has been shown in which TTML and abstract information are placed in segments and then placed in the PES data payload of a PES packet. However, the present technology also contemplates placing TTML and abstract information directly in the PES data payload.
[0165] Furthermore, in the above-described embodiment, the transmission / reception system 10 including the transmission device 100 and the reception device 200 has been described, but the configuration of the transmission / reception system to which the present technology can be applied is not limited to this. For example, the reception device 200 may be configured as a set-top box and a monitor connected via a digital interface such as HDMI (High-Definition Multimedia Interface). Note that "HDMI" is a registered trademark.
[0166] The present technology can also be configured as follows. (1) a video encoding unit that generates a video stream including encoded video data; a subtitle encoding unit that generates a subtitle stream including subtitle text information having display timing information and abstract information having information corresponding to a part of the plurality of pieces of information indicated by the text information; a transmitter for transmitting a container including the video stream and the subtitle stream; Transmitting device. (2) The above abstract information includes subtitle display timing information. The transmitting device according to (1) above. (3) The display timing information of the subtitles includes information on the display start timing and display period. The transmitting device according to (2) above. (4) The subtitle stream is composed of PES packets each consisting of a PES header and a PES payload, The subtitle text information and the abstract information are placed in the PES payload, The display start timing is indicated by a display offset from the PTS inserted in the PES header. The transmitting device according to (3) above. (5) The abstract information includes display control information for controlling the display state of the subtitles. The transmitting device according to any one of (1) to (4). (6) The display control information includes at least one of the following information: a subtitle display position, a color gamut, and a dynamic range. The transmitting device according to (5) above. (7) The display control information further includes information about the target video. The transmitting device according to (6) above. (8) The abstract information includes notification information that notifies that there is a change in the elements of the text information of the subtitle. The transmitting device according to any one of (1) to (7). (9) The subtitle encoding unit is Segmenting the subtitle text information and the abstract information to generate the subtitle stream having a predetermined number of segments. The transmitting device according to any one of (1) to (8). (10) The above subtitle stream includes: The abstract information segment is placed first, followed by the subtitle text information segment. The transmitting device according to (9) above. (11) The text information of the subtitles is in TTML or a format derived from TTML. The transmitting device according to any one of (1) to (10). (12) a video encoding step of generating a video stream including encoded video data; a subtitle encoding step of generating a subtitle stream including subtitle text information having display timing information and abstract information having information corresponding to a part of the plurality of pieces of information indicated by the text information; a transmitting step of transmitting a container including the video stream and the subtitle stream by a transmitting unit. Sending method. (13) A receiving unit for receiving a container in a predetermined format including a video stream and a subtitle stream, the video stream comprises encoded video data; the subtitle stream includes subtitle text information having display timing information and abstract information having information corresponding to a part of the plurality of pieces of information indicated by the text information, a video decoding unit that decodes the video stream to obtain video data; a subtitle decoding unit that decodes the subtitle stream to obtain subtitle bitmap data and extracts the abstract information; a video superimposing unit that superimposes the subtitle bitmap data on the video data to obtain video data for display; a control unit that controls bitmap data of a subtitle superimposed on the video data based on the abstract information. Receiving device. (14) The abstract information includes subtitle display timing information, The control unit The timing of superimposing the subtitle bitmap data on the video data is controlled based on the display timing information of the subtitle. The receiving device according to (13) above. (15) The abstract information includes display control information for controlling the display state of the subtitles, The control unit A state of the bitmap of the subtitle superimposed on the video data is controlled based on the display control information. The receiving device according to (13) or (14). (16) A receiving step of receiving, by a receiving unit, a container in a predetermined format including a video stream and a subtitle stream, the video stream comprises encoded video data; the subtitle stream includes subtitle text information having display timing information and abstract information having information corresponding to a part of the plurality of pieces of information indicated by the text information, a video decoding step of decoding the video stream to obtain video data; a subtitle decoding step of decoding the subtitle stream to obtain subtitle bitmap data and extracting the abstract information; a video superimposing step of superimposing the subtitle bitmap data on the video data to obtain video data for display; The method further includes a control step of controlling bitmap data of a subtitle superimposed on the video data based on the abstract information. Receiving method. (17) a video encoding unit for generating a video stream including encoded video data; a subtitle encoding unit that generates one or more segments in which elements of subtitle text information having display timing information are arranged, and generates a subtitle stream including the one or more segments; a transmitting unit for transmitting a container in a predetermined format including the video stream and the subtitle stream; Transmitting device. (18) The subtitle encoding unit is To generate one segment containing all the elements of the subtitle text information, Inserting information about the transmission order and / or update status of the subtitle text information into the segment layer or the element layer. The transmitting device according to (17) above. (19) The text information of the subtitles is in TTML or a format derived from TTML. The transmitting device according to (17) or (18). (20) a video encoding step of generating a video stream including encoded video data; a subtitle encoding step of generating one or more segments in which elements of subtitle text information having display timing information are arranged, and generating a subtitle stream including the one or more segments; a transmitting step of transmitting a container in a predetermined format including the video stream and the subtitle stream by a transmitting unit; Sending method.
[0167] The main feature of this technology is that it reduces the processing load for displaying subtitles on the receiving side by including abstract information corresponding to the text information of the subtitles in the subtitle stream (see Figure 12). [Explanation of symbols]
[0168] 10. Transmitting and receiving system 100 Transmitting device 101 Control unit 102···Camera 103 Video photoelectric conversion unit 104 RGB / YCbCr conversion section 105...Video Encoder 106 Subtitle generator 107 Text format conversion section 108···Subtitle Encoder 109 System Encoder 110 Transmitter 200 Receiving device 201 Control unit 202 User operation unit 203 Receiving unit 204 System Decoder 205...Video decoder 206 Subtitle Decoder 207 Color gamut and brightness conversion unit 208 Position and size conversion section 209 Video Overlay Unit 210 YCbCr / RGB conversion section 211 Electrical-optical converter 212 Display mapping section 213···CE Monitor 221 Electrical-optical converter 222 Color gamut conversion unit 223... Photoelectric conversion unit 224···RGB / YCbCr conversion section 225...Luminance conversion unit 225Y...Component 231 Encoding pixel bit number adjustment unit 232 Level adjustment section 261 Coded Buffer 262 Subtitle Segment Decoder 263 Font expansion section 264 Bitmap Buffer< / layout> < / styling> < / metadata> < / layout> < / styling> < / metadata>
Claims
1. a transport stream generating step of generating a transport stream including encoded video data, subtitle text information, and control information related to the text information; a transmitting step of transmitting the transmission stream; the control information regarding the text information includes information regarding subtitle display timing and information regarding subtitle display state control, The subtitle text information includes information in TTML or a TTML-derived format. Sending method.
2. a transport stream generating unit for generating a transport stream including coded video data, subtitle text information, and control information related to the text information; a transmitter for transmitting the transmission stream, the control information regarding the text information includes information regarding subtitle display timing and information regarding subtitle display state control, The subtitle text information includes information in TTML or a TTML-derived format. Transmitting device.
Citation Information
Patent Citations
Display control method, recording medium and display control unit
JP2012169885A
Receiving device and data processing method
JP5713142B1