Transmitting device and receiving device
By employing multi-layer coding with a base layer for primary video and an enhancement layer for secondary information, the system achieves low-latency video distribution, ensuring timely delivery of emergency alerts in digital broadcasting.
Patent Information
- Application Number
- JP2021195343
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-02
- Filing Date
- 2021-12-01
- Publication Date
- 2025-09-11
- Estimated Expiration
- 2041-12-01
AI Technical Summary
Existing video distribution systems experience significant transmission delays due to encoding processes, which hinder the timely delivery of urgent information such as disaster alerts, especially in digital broadcasting where latency can exceed several seconds, undermining safety and efficiency.
A transmitting device and receiving device utilize multi-layer coding technology, encoding the primary video as a base layer with coding efficiency priority and the secondary video as an enhancement layer with low-delay processing, allowing for synchronized transmission and decoding of urgent information with reduced latency.
This approach enables low-delay video distribution, enabling immediate display of superimposed emergency alerts by encoding the primary video as a base layer and secondary information like captions as an enhancement layer, reducing latency to less than conventional methods.
Smart Images

Figure 0007737880000001 
Figure 0007737880000002 
Figure 0007737880000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a transmitting device and a receiving device for a video transmission system. [Background technology]
[0002] Broadcasting services often superimpose captions onto images to widely inform disaster areas and related people about natural disasters such as heavy rain and earthquakes, as well as serious incidents and accidents. In recent years, in order to mitigate damage, captions are sometimes superimposed onto images to warn of upcoming disasters that may cause damage, such as emergency earthquake alerts, several seconds in advance.
[0003] However, in digital broadcasting, video signals are encoded, and while this encoding process makes it possible to transmit high-quality video with a small amount of information, it also involves various storage processes, which causes transmission delays, leaving little time between receiving the information and taking action.
[0004] Specifically, with the aim of making effective use of limited radio wave resources, compression technology has become more sophisticated, resulting in increased latency. Video coding technology accumulates and rearranges multiple frames of data to improve coding efficiency, and when coding is performed in a coding efficiency priority mode, delays of several hundred milliseconds occur, and the total delay time, including delays on the decoding side and delays associated with transmission, can reach several seconds.
[0005] In coding settings aimed at low latency, it is possible to achieve latency of within a few tens of milliseconds by using parallel processing without reordering, but since the coding efficiency for general natural images is extremely low, it is extremely rare to transmit video using such settings. In systems where video is distributed continuously, such as broadcasting, frequent changes in latency for each program would undermine the stability of the system, so video distribution is generally continued with settings that prioritize coding efficiency.
[0006] Disaster alerts such as earthquake early warnings are intended to inform people of disasters that may occur in the next few seconds, but for the reasons mentioned above, there is always a transmission delay of about one or two seconds. As a result, one or two seconds are wasted in transmitting notifications of disasters that may occur in the next few seconds, which is a major issue in ensuring the safety of life and property.
[0007] In order to deliver urgent information with as little delay as possible, methods have been developed such as using asynchronous subtitles, which have relatively little broadcast delay, using event messages in data broadcasting, and using AC (Auxiliary Channel).Current digital broadcasting uses these methods to present information more than one second earlier than the video (see non-patent document 1).
[0008] On the other hand, broadcast video transmission basically uses a mode that can only transmit one video per channel. Furthermore, one coding technology for video transmission is multi-layer coding (also called hierarchical coding or scalable coding). Multi-layer coding is a technology that transmits coded streams with multiple layers. While high-quality video can be reproduced by decoding all coded streams, it is also a technology that can reproduce low-quality video by decoding only a part of the coded stream (specifically, the base layer). [Prior art documents] [Non-patent literature]
[0009] [Non-Patent Document 1] Hiroyuki Hamazumi, "Trends in Terrestrial Digital Broadcasting", [online], [Retrieved October 1, 2020], Internet<URL:http: / / tokkyo.shinsakijun.com / information / newtech.html> Summary of the Invention [Problem to be solved by the invention]
[0010] The technology described in Non-Patent Document 1 uses digital broadcasting functions to present information by means other than video, and while network-based video distribution services are expected to expand in the future, these functions may not be available when using a network, making low-latency video distribution technology an issue. Note that while mobile phones such as smartphones can instantly report an incident using a mass call function, this is also a method that does not involve video.
[0011] On the other hand, although the above-mentioned multi-layer coding technology can efficiently realize video services with different resolutions and qualities, and is used in video distribution services that involve a mix of line qualities, such as wireless and wired transmission, there is a problem in that it is not being fully utilized.
[0012] Therefore, an object of the present invention is to provide a transmitting device and a receiving device that utilize multi-layer coding technology to achieve low-delay video distribution in a video transmission system. [Means for solving the problem]
[0013] A transmitting device according to a first aspect is a transmitting device for a video transmission system, and includes: a coding unit that encodes a video signal of a primary video service as a base layer and a secondary video signal of a secondary video service as an enhancement layer using a multi-layer coding technique; and a transmitting unit that transmits each of the coded streams of the base layer and the enhancement layer output by the coding unit. The coding unit codes the video signal in a first mode when coding the base layer, and codes the secondary video signal in a second mode when coding the enhancement layer, which has a smaller coding processing delay than the first mode.
[0014] A receiving device according to a second aspect is a receiving device for a video transmission system using a multi-layer coding technique, and includes a receiving unit that receives coded streams of a base layer and an enhancement layer, and a decoding unit that decodes a video signal of a primary video service from the coded stream of the base layer and a secondary video signal of a secondary video service from the coded stream of the enhancement layer. The decoding unit decodes the video signal in a first mode when decoding the base layer, and decodes the secondary video signal in a second mode when decoding the enhancement layer, which has a smaller decoding processing delay than the first mode. [Effects of the Invention]
[0015] According to the present invention, it is possible to provide a transmitting device and a receiving device that utilize multi-layer coding technology to achieve low-delay video distribution in a video transmission system. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 is a diagram illustrating a configuration of a transmission device according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of an encoding unit according to the embodiment. [Figure 3] FIG. 10 is a diagram illustrating a specific example of the operation of the transmitting device according to the embodiment. [Figure 4] FIG. 1 is a diagram illustrating a frame reference relationship in conventional multi-layer coding. [Figure 5] FIG. 1 is a diagram illustrating a frame reference relationship in multi-layer coding according to an embodiment. [Figure 6] 10 is a diagram showing the encoding timing of each frame, the output timing of the superimposed video (overlaid video) according to this embodiment, and the timing of the superimposed video according to a conventional method for comparison. FIG. [Figure 7] FIG. 1 is a diagram illustrating a configuration of a receiving device according to an embodiment. [Figure 8] FIG. 10 is a diagram illustrating a modification of the transmitting device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0017] A transmitting device and a receiving device for a video transmission system according to an embodiment will be described with reference to the drawings. In the following description of the drawings, the same or similar parts are denoted by the same or similar reference numerals.
[0018] (Transmitting device) First, a transmitting device according to this embodiment will be described. In this embodiment, an example of realizing an emergency news service as an example of a sub-video service in a video transmission system will be described.
[0019] 1 is a diagram showing the configuration of a transmission device 10 according to this embodiment. As shown in FIG. 1, the transmission device 10 includes a superimposing unit 11, an encoding unit 12, and a transmission unit 13.
[0020] The superimposing unit 11 superimposes the sub-video signal of the sub-video service on the video signal of the main video service. In this embodiment, the sub-video is assumed to be subtitles of an emergency news service (for example, a disaster telop). The subtitles consist of characters including numbers, figures, symbols, or a combination of these. The superimposing unit 11 superimposes the subtitles on a partial area of the video frame of the main video (for example, the upper area or the lower area of the frame). The superimposing unit 11 then outputs the subtitled video signal to the encoding unit 12 as a video signal with sub-video.
[0021] The encoding unit 12 uses multi-layer encoding technology to encode the video signal of the main video service as a base layer, and encodes the video signal with secondary video output by the superimposition unit 11 as an enhancement layer. The multi-layer encoding method can use the Scalable Main 10 profile of an encoding method called HEVC (High efficiency video coding). HEVC is an encoding method currently used in 4K8K satellite broadcasting. In 4K8K satellite broadcasting, video is efficiently compressed and transmitted using a group of encoding tools called the Main 10 profile. The Scalable Main 10 profile is an encoding mechanism that can transmit video equivalent to the Main 10 profile as a base layer, while adding an enhancement layer and superimposing it on the base layer video.
[0022] Alternatively, the multilayer encoding method may be the Multilayer Main 10 profile, which is a latest encoding method called VVC (Versatile Video Coding). The Multilayer Main 10 profile, like the Scalable Main 10 profile described above, is a profile that realizes the transmission of multiple layers compared to the Main 10 profile, and is capable of transmitting video of multiple layers.
[0023] Coding technology with such multi-layer functionality makes it possible to superimpose video from an enhancement layer, which is a higher layer, on the base layer, which is the main video service. Note that although the HEVC Scalable Main 10 profile and the VVC Multilayer Main 10 profile have been given as examples of multi-layer coding, other coding methods may also be used, and are not limited to HEVC and VVC.
[0024] In this embodiment, the encoding unit 12 encodes the video signal of the primary video service in a first mode in the base layer. The encoding unit 12 encodes the video signal with secondary video in a second mode in the enhancement layer, which has a smaller encoding processing delay than the first mode. This makes it possible to utilize multi-layer encoding technology to achieve low-delay video distribution in a video transmission system.
[0025] The first mode applied to the base layer is a coding efficiency priority mode in which video frames are rearranged when encoding a video signal. Rearrangement of video frames refers to a process of rearranging video frames into an order different from the display order. The encoding unit 12 prioritizes coding efficiency for the video signal of the primary video service, accumulates and rearranges data of multiple frames, and then encodes the data.
[0026] On the other hand, the second mode applied to the enhancement layer is a low-delay mode in which the video frames are not rearranged when encoding the video signal with sub-picture. For example, the encoding unit 12 does not rearrange the video frames, but performs encoding using parallel processing with a delay of within several tens of milliseconds.
[0027] The transmitter 13 transmits, via a transmission path, the coded stream of the base layer and the coded stream of the enhancement layer output by the encoder 12. Regarding the transmission path, for example, the coded stream of the base layer and all the coded streams of the enhancement layers may be transmitted via the same path using broadcast waves or a network, or may be transmitted via different paths, such as transmitting some of the coded streams via broadcast waves and transmitting the remaining coded streams via a network.
[0028] The transmitter 13 includes a multiplexer 13a that multiplexes the base layer encoded stream and the enhancement layer encoded stream. The multiplexer 13a assigns a service identification identifier to each of the base layer encoded stream and the enhancement layer encoded stream, and transmits the multiple encoded streams in synchronization using a media transmission function such as MPEG (Moving Picture Experts Group)-2 / 4 Transport stream (TS), MPEG Media Transport (MMT), or MPEG Dynamic Adaptive Streaming over HTTP (DASH). Note that although MPEG-2 / 4 TS, MMT, and DASH are given as examples, the present invention is not limited to MPEG-2 / 4 TS, MMT, and DASH, and other media transmission functions may also be used.
[0029] Fig. 2 is a diagram showing an example of the configuration of the encoding unit 12 according to this embodiment. As shown in Fig. 2, the encoding unit 12 includes a base layer encoding unit 121 that performs base layer encoding, and an enhancement layer encoding unit 122 that performs encoding corresponding to an upper layer.
[0030] The base layer encoding unit 121 includes a frame rearrangement unit 1211, an encoding processing unit 121a, an entropy encoding unit 121b, a decoding processing unit 121c, and a DPB (Decoded Picture Buffer) 121d.
[0031] The frame rearrangement unit 1211 accumulates and rearranges the video frames included in the video signal of the main video service, and then outputs the accumulated video frames to the encoding processing unit 121a. Specifically, the frame rearrangement unit 1211 rearranges the video frames so that the encoding processing unit 121a can efficiently perform inter-frame prediction.
[0032] The encoding processing unit 121a performs prediction processing on a video frame (original image) of the video signal of the main video service by referring to the DPB 121d, transforms and quantizes a prediction residual, which is the difference between the predicted image and the original image, and outputs the quantized transform coefficient. In addition, the encoding processing unit 121a outputs information related to the prediction processing (for example, motion vector information) and the like to the enhancement layer encoding unit 122 as encoding control information.
[0033] The entropy coding unit 121b performs entropy coding processing on the coding information such as the quantized transform coefficients and motion vector information output by the coding processing unit 121a, and outputs a coded stream (bit stream) of the base layer.
[0034] The decoding processing unit 121c dequantizes and inversely transforms the quantized transformation coefficients output by the encoding processing unit 121a to restore the prediction residual, synthesizes the restored prediction image and the prediction residual using encoding control information such as motion vector information to restore (decode) the original image, and outputs the decoded image to the DPB 121d and the encoding processing unit 121a.
[0035] The DPB 121d stores the decoded image output by the decoding processing unit 121c. The decoded image output by the DPB 121d is used as a predictive reference image for prediction processing in the encoding processing unit 121a. In addition, the DPB 121d outputs the decoded image to the enhancement layer encoding unit 122 as one of the predictive reference signals used in the encoding processing unit 122a.
[0036] Similar to the base layer encoding unit 121, the enhancement layer encoding unit 122 includes an encoding processing unit 122a, an entropy encoding unit 122b, a decoding processing unit 122c, and a DPB 122d.
[0037] The DPB 122d stores the decoded image output by the DPB 121d of the base layer encoding unit 121 and the decoded image output by the decoding processing unit 122c as predicted reference images for the enhancement layer.
[0038] The encoding processing unit 122a performs prediction processing on the video frame of the sub-video video signal by referring to the DPB 122d, transforms and quantizes the prediction residual, which is the difference between the predicted image and the original image, and outputs the quantized transform coefficients. Here, the encoding processing unit 122a uses encoding control information (e.g., motion vector information, decoded image, etc.) output by the base layer encoding unit 121 for the prediction processing.
[0039] The entropy coding unit 122b performs entropy coding processing on the coding information such as the quantized transform coefficients and motion vector information output by the coding processing unit 122a, and outputs a coded stream (bit stream) of the enhancement layer.
[0040] The decoding processing unit 122c dequantizes and inversely transforms the quantized transformation coefficients output by the encoding processing unit 122a to restore the prediction residual, combines the restored prediction image and the prediction residual using encoding control information such as motion vector information to restore (decode) the original image, and outputs the decoded image to the DPB 122d.
[0041] In this way, the enhancement layer encoding unit 122 also uses the decoded image of the base layer in the prediction process, and performs processing to encode the difference (prediction residual) between this decoded image and the video signal with sub-picture. Therefore, the processing is to encode the difference signal between the image obtained by decoded the encoded image of the base layer and the original image in the enhancement layer, and generally encodes a signal that compensates for encoding degradation. In this embodiment, by intentionally using a signal in which subtitles different from the image of the base layer are added in the enhancement layer, the subtitle portion corresponds to the difference, and therefore the subtitle portion is encoded only in the enhancement layer.
[0042] Note that multi-layer coding, in which video of the same resolution is coded in the base layer and enhancement layer as in this embodiment, is sometimes called SNR (Signal to Noise Ratio) scalable coding, which transmits signals of different quality because information compensating for coding degradation occurring in the base layer is generally coded as the enhancement layer and the coding quality differs between the layers. With an SNR scalable configuration, video in which an enhancement layer of the same resolution is superimposed on the base layer can be provided to the receiving side. However, this embodiment differs from general SNR scalable coding in that a video signal with a sub-video different from the video signal coded by the base layer is coded in the enhancement layer.
[0043] In this embodiment, the base layer encoding unit 121 performs frame rearrangement processing, but the enhancement layer encoding unit 122 does not. As a result, the correspondence between video frames in the base layer and the enhancement layer differs. Therefore, the enhancement layer encoding unit 122 imposes a constraint on the encoding process, controlling the encoding process so that the base layer is copied (no difference is sent) in encoding the enhancement layer except for the subtitle portion, which is the superimposed video portion. In other words, the encoding unit 12 controls the encoding process so that, in encoding the enhancement layer, an area other than the sub-video area included in the video frame of the video signal with sub-video is the area to copy the base layer at the same position.
[0044] FIG. 3 is a diagram showing a specific example of the operation of the transmission device 10 according to this embodiment.
[0045] As shown in FIG. 3(a), video frames included in a video signal of a primary video service are input to the superimposing unit 11 and the encoding unit 12. As shown in FIG. 3(b), the superimposing unit 11 superimposes subtitles (disaster captions) of an emergency news service as a secondary video service onto the input video frames. FIG. 3(b) shows an example in which the emergency news service is an emergency earthquake alert. Such a subtitle area is called a subtitle area. The superimposing unit 11 outputs the subtitled video frames shown in FIG. 3(b) to the encoding unit 12.
[0046] In the encoding unit 12, the base layer encoding unit 121 encodes the video frame shown in Fig. 3(a). The enhancement layer encoding unit 122 encodes the subtitled video frame shown in Fig. 3(b). Here, as shown in Fig. 3(c), the enhancement layer encoding unit 122 controls the encoding process in the enhancement layer so that an area other than the subtitle area (sub-video area) included in the subtitled video frame is an area where the base layer is copied (an area where the base layer at the same position is referenced and no difference is sent).
[0047] However, when encoding in multiple layers, rearrangement processing is usually performed on each of the base layer and the enhancement layer to ensure that the correspondence between video frames is the same in the base layer and the enhancement layer.However, in this embodiment, the enhancement layer is encoded in low-delay mode, and the base layer is encoded in encoding efficiency priority mode.
[0048] Fig. 4 is a diagram showing frame reference relationships in conventional multi-layer coding. Fig. 5 is a diagram showing frame reference relationships in multi-layer coding according to this embodiment. Figs. 4 and 5 show an example in which the interval "M" between low-layer frames in the base layer is 4. Arrows also indicate reference relationships, and show an example in which the third frame in the base layer references the second and fourth frames, for example.
[0049] As shown in Figure 4, in the base layer, video frames are rearranged to encode the video frame to be encoded by referring to video frames temporally preceding and succeeding the video frame to be encoded. Specifically, after the 0th frame in display order, the 4th frame is encoded, followed by the 2nd frame, the 1st frame, and the 3rd frame, in that order.
[0050] In the enhancement layer, a video frame to be coded is coded after a video frame corresponding to the video frame to be coded in the enhancement layer has been coded in the base layer. Here, the reference frame of the video frame to be coded in the enhancement layer is the video frame with the same number (same time) in the base layer. By coding the difference signal between the decoded video of the base layer and the video before coding for the video frame with the same time, a signal that compensates for degradation of the base layer is coded as the enhancement layer.
[0051] In contrast to this, in this embodiment, in encoding in the enhancement layer, the encoding process is controlled so that an area other than the subtitle area included in the video frame of the subtitled video signal is set as an area to copy the base layer, so that the reference frame of the video frame to be encoded in the enhancement layer can be set to a video frame with a different number (different time) in the base layer.As a result, the video frame to be encoded in the enhancement layer can be encoded without waiting for the video frame corresponding to the video frame to be encoded in the enhancement layer to be encoded in the base layer, so low-delay encoding is possible without performing rearrangement.
[0052] As shown in Fig. 5, in the base layer according to this embodiment, first, the 0th frame is coded after the -4th frame in display order, followed by coding the -2nd frame, the -1st frame, and the -3rd frame in that order. Second, the 4th frame is coded after the 0th frame in display order, followed by coding the 2nd frame, the 1st frame, and the 3rd frame in that order.
[0053] In the enhancement layer according to this embodiment, video frames are coded in display order without rearrangement as in the base layer. Therefore, the reference frame of a video frame to be coded in the enhancement layer may be a video frame with a different number (different time) in the base layer. Specifically, the first frame of the enhancement layer references the -2nd frame of the base layer, the second frame of the enhancement layer references the -3rd frame of the base layer, and the third frame of the enhancement layer references the -1st frame of the base layer.
[0054] FIG. 6 is a diagram showing the encoding timing of each frame, the output timing of the superimposed video (overlaid video) according to this embodiment, and, for comparison, the timing of the superimposed video according to a conventional method.
[0055] As shown in Figure 6, when superimposing on conventional base layer video, a delay occurs due to the rearrangement of M+1 frames, but in this embodiment, superimposing can be achieved with a delay of 2 frames. For example, in the case of a 60 fps video signal, one frame time is 1 / 60 seconds ≈ 17 msec, so if we ignore the encoding processing time, in general broadcasting, a coding efficiency priority mode uses M = 32, which causes a delay of 550 msec, but superimposing can be achieved with a delay of 33 msec.
[0056] Furthermore, since HEVC encoding using existing low-delay technologies can encode in a few tens of milliseconds, it is expected that by combining these low-delay encoding techniques with enhancement layer encoding, it will be possible to encode with a delay of 100 milliseconds or less.
[0057] (receiving device) Next, a receiving device according to this embodiment will be described. Fig. 7 is a diagram showing the configuration of a receiving device 20 according to this embodiment.
[0058] As shown in FIG. 7, the receiving device 20 includes a receiving unit 21, a selecting unit 22, an extracting unit 23, and a decoding unit 24.
[0059] The receiving unit 21 receives the coded stream of the base layer and the coded stream of the enhancement layer from the transmitting device 10 via a transmission path, and outputs each of the received coded streams to the extracting unit 23 .
[0060] The selection unit 22 selects (tunes in to) an encoded stream to be played back based on an operation from a viewer, and outputs an identifier indicating the selected encoded stream to the extraction unit 23 as control information.
[0061] The extraction unit 23 extracts a desired encoded stream from the base layer encoded stream and the enhancement layer encoded stream based on the identifier indicated by the control information output by the selection unit 22, and outputs the extracted encoded stream to the decoding unit 24.
[0062] When video playback of a sub-video service is not selected by the selection unit 22, the extraction unit 23 extracts only the encoded stream of the base layer and outputs only the encoded stream of the base layer to the decoding unit 24. On the other hand, when video playback of a sub-video service is selected by the selection unit 22, the extraction unit 23 extracts the encoded stream of the base layer and the encoded stream of the enhancement layer and outputs the encoded stream of the base layer and the encoded stream of the enhancement layer to the decoding unit 24.
[0063] When the video playback of the secondary video service is not selected by the selection unit 22, the decoding unit 24 decodes (plays) the video signal of the primary video service from the base layer encoded stream output by the extraction unit 23, and outputs the decoded video signal.
[0064] On the other hand, when video playback of a secondary video service is selected by the selection unit 22, the decoding unit 24 decodes the video signal of the primary video service from the coded stream of the base layer, and decodes the video signal with secondary video on which the secondary video signal of the secondary video service is superimposed from this video signal and the coded stream of the enhancement layer. Here, the decoding unit 24 decodes the video signal in the base layer in a coding efficiency priority mode (first mode), and decodes the video signal with secondary video in the enhancement layer in a low delay mode (second mode) in which the decoding processing delay is smaller than in the coding efficiency priority mode (first mode).
[0065] As described above, the coding efficiency priority mode (first mode) is a mode in which video frames are rearranged when decoding a video signal of the base layer, and the low latency mode (second mode) is a mode in which video frames are not rearranged when decoding a video signal with sub-pictures of the enhancement layer.
[0066] The decoding unit 24 controls the decoding process so that, in decoding in the enhancement layer, an area other than the sub-video area (for example, the above-mentioned subtitle area) included in the video frame of the video signal with sub-video is set as an area to which the base layer is copied.
[0067] (Actions and Effects) According to the video transmission system of this embodiment, it is possible to quickly present highly urgent superimposed information, such as an emergency earthquake alert, in digital broadcasting.
[0068] In conventional broadcasting, captions (subtitles) are superimposed on video services as a reliable means of conveying information, since they can be reliably displayed as expected on any properly functioning receiver. Specifically, the sender adds text and graphics to a portion of the video being served using video processing, and the video is transmitted without changing the broadcast encoding mode (in this case, the profile or prediction structure). Because of this advanced compression and transmission process, the video arrives at the receiver with some delay. If the encoding mode were to be forcibly changed mid-encoding to a low-latency processing mode and then transmitted, the receiver would have to perform initialization processing to change the mode, which would result in a delay greater than the encoding delay associated with transmission. Therefore, this method is not used. As a result, it is not possible to reduce the delay, resulting in a delay of one or two seconds before the caption is displayed on the receiving end.
[0069] In contrast, in this embodiment, the video of the main video service is transmitted as a base layer, while the video on which captions (subtitles) are superimposed is transmitted as an enhancement layer, and a low-latency mode is used as the encoding mode for the enhancement layer, thereby utilizing multi-layer encoding technology to achieve video distribution with lower latency than conventional methods. In this case, although a mismatch may occur in the correspondence between video frames between the base layer and the enhancement layer, the base layer is used as a copy area for areas other than the caption (subtitle area), making it possible to decode and play back the video on which captions are superimposed.
[0070] (Other embodiments) In the above-described embodiment, the transmitting device 10 includes the superimposing unit 11, but the transmitting device 10 may be configured without the superimposing unit 11. FIG. 8 is a diagram showing a modified example of the transmitting device 10.
[0071] The transmitting device 10 shown in FIG. 8 has a configuration in which the superimposing unit 11 of FIG. 1 is omitted, and a sub-video signal such as a subtitle video is input to an encoding unit 12. The encoding unit 12 can detect the subtitle area because the area other than the subtitle area in the sub-video signal is zero (or a predetermined value). The encoding unit 12 encodes the area other than the subtitle area as an area that references the base layer, and encodes only the subtitle area to encode the enhancement layer in low delay mode (same as in the above-mentioned embodiment). This makes it possible to obtain the same effect as in the above-mentioned embodiment without using the superimposing unit 11.
[0072] In the above embodiment, an example has been described in which the sub-video service is an emergency news service. However, the sub-video service may be any video service that requires low latency, such as a time signal service.
[0073] A program may be provided that causes a computer to execute each process performed by the transmitting device 10. Similarly, a program may be provided that causes a computer to execute each process performed by the receiving device 20. The program may be recorded on a computer-readable medium. Using the computer-readable medium, the program can be installed on a computer. Here, the computer-readable medium on which the program is recorded may be a non-transitory recording medium. The non-transitory recording medium is not particularly limited, and may be, for example, a recording medium such as a CD-ROM or a DVD-ROM.
[0074] The circuits that execute the processes performed by the transmitting device 10 may be integrated, and the transmitting device 10 may be configured as a semiconductor integrated circuit (chip set, SoC). Similarly, the circuits that execute the processes performed by the receiving device 20 may be integrated, and the receiving device 20 may be configured as a semiconductor integrated circuit (chip set, SoC).
[0075] The above describes the embodiments in detail with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes can be made within the scope that does not deviate from the gist of the invention. [Explanation of symbols]
[0076] 10: Transmitting device 11: Superimposed part 12: Encoding section 13: Transmitter 13a: Multiplexing section 20: Receiving device 21: Receiving unit 22: Selection section 23:Extraction part 24: Decryption unit 121: Base layer coding unit 121a: Encoding processing unit 121b: Entropy encoding section 121c: Decryption processing unit 121d :DPB 122: Enhancement layer coding unit 122a: Encoding processing unit 122b: Entropy coding unit 122c: Decryption processing unit 122d :DPB 1211: Frame reordering unit
Claims
1. A transmitting device for a video transmission system, comprising: an encoding unit that encodes a video signal of a primary video service as a base layer and encodes a secondary video signal of a secondary video service as an enhancement layer using a multi-layer encoding technique; a transmitting unit that transmits the encoded streams of the base layer and the enhancement layer output by the encoding unit, The encoding unit In the encoding in the base layer, the video signal is encoded in a first mode; A transmitting device, characterized in that, in encoding in the enhancement layer, the secondary video signal is encoded in a second mode in which encoding processing delay is smaller than that in the first mode.
2. a superimposing unit that superimposes the sub-video signal of the sub-video service on the video signal of the main video service, The transmitting device according to claim 1, wherein the encoding unit encodes the video signal with sub-picture output by the superimposing unit as the sub-picture signal in the second mode in encoding in the enhancement layer.
3. the first mode is a coding efficiency priority mode in which video frames are rearranged when the video signal is coded; 3. The transmitting device according to claim 1, wherein the second mode is a low-delay mode in which video frames are not rearranged when the sub-video signal is encoded.
4. The transmitting device according to any one of claims 1 to 3, characterized in that the encoding unit controls the encoding process so that, in encoding in the enhancement layer, an area other than the sub-video area included in the video frame of the sub-video signal is used as an area to copy the base layer.
5. 5. The transmitting device according to claim 1, wherein the sub-video service is an emergency news service or a time signal service.
6. A receiving device for a video transmission system using a multi-layer coding technique, comprising: a receiving unit that receives encoded streams of the base layer and the enhancement layer; a decoding unit that decodes a video signal of a primary video service from the coded stream of the base layer and decodes a secondary video signal of a secondary video service from the coded stream of the enhancement layer, The decoding unit In the decoding at the base layer, the video signal is decoded in a first mode; A receiving device characterized in that, in decoding in the enhancement layer, the secondary video signal is decoded in a second mode in which a decoding processing delay is smaller than that in the first mode.
7. The receiving device according to claim 6, wherein the decoding unit decodes the video signal and the encoded stream of the enhancement layer into a video signal with a sub-video signal in which the sub-video signal is superimposed on the video signal.
8. the first mode is a coding efficiency priority mode in which video frames are rearranged when the video signal is decoded; 8. The receiving device according to claim 6, wherein the second mode is a low-delay mode in which video frames are not rearranged when the sub-video signal is decoded.
9. The receiving device according to any one of claims 6 to 8, characterized in that the decoding unit controls the decoding process so that, in decoding the enhancement layer, an area other than the sub-video area included in the video frame of the sub-video signal is used as an area to copy the base layer.
10. 10. The receiving device according to claim 6, wherein the sub-video service is an emergency news service or a time signal service.
Citation Information
Patent Citations
Method and device for multiplex transmission, and program and recording medium thereof
JP2006041869A
Transmitter and transmission method for transmitting payload data and emergency information
WO2014195303A1
Receiving device and data processing method
WO2018016295A1