Video encoding device and video decoding device

By dynamically switching frame sizes between 8K and 4K based on scene complexity and using reference picture scaling, the video encoding and decoding devices maintain high video quality and reduce data volume, addressing the challenge of limited transmission capacity in ultra-high definition broadcasting.

JP2025178389APending Publication Date: 2025-12-05NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025162731
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

The challenge of maintaining high video quality for ultra-high definition videos, such as 8K video, in next-generation terrestrial broadcasting with limited transmission capacity is significant, especially in scenes with complex images or movement, as existing technologies like HEVC struggle to reduce bit rates sufficiently.

Method used

A video encoding device and method that dynamically switches the image size of frames between 8K and 4K based on scene complexity, using reference picture scaling to maintain high image quality and reduce data volume, while a video decoding device demultiplexes and scales the image sizes accordingly for smooth playback.

Benefits of technology

The solution effectively maintains high video quality by dynamically adjusting frame sizes between 8K and 4K, reducing data volume in complex scenes and preventing noticeable image degradation, ensuring smooth playback without requiring re-loading of video bitstreams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025178389000001_ABST
    Figure 2025178389000001_ABST
Patent Text Reader

Abstract

To provide a video encoding device and a video decoding device that can maintain high image quality for ultra-high definition video.SOLUTION: A video encoding device includes multiplexing means for multiplexing the maximum image width and maximum image height of the luminance samples of all frames into a bitstream, and means for periodically switching the image width and image height of the luminance samples of the frames between a first image width and image height and a second image width and image height in a predetermined AU, and the multiplexing means is controlled to multiplex the determined image width and image height of the luminance samples into the bitstream.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video encoding device and a video decoding device that utilizes reference picture scaling. [Background technology]

[0002] Non-Patent Document 1 discloses the specifications of the VVC (Versatile Video Coding) method, which can reduce the bit rate to approximately half that of the HEVC (High Efficiency Video Coding) method while maintaining the same image quality.

[0003] Non-Patent Document 2 defines video signal compression based on the HEVC standard for digital broadcasting and introduces the concept of SOP (Set of Pictures). SOP is a unit that describes the coding order and reference relationship of each AU (Access Unit) when performing time-direction hierarchical coding. SOP structures include L0 structure, L1 structure, L2 structure, L3 structure, and L4 structure.

[0004] For the VVC system, by defining an SOP structure for the VVC system, digital broadcasting similar to that for the HEVC system can be operated. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] "Versatile Video Coding (Draft 8)", JVET-Q2001, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 17th Meeting: Brussels, BE, 7-17 January 2020. [Non-patent document 2] ARIB (Association of Radio Industries and Businesses) Standard STD-B32 Version 3.11 July 26, 2018 Association of Radio Industries and Businesses [Non-patent document 3] "Supplemental enhancement information for coded video bitstreams (Draft 3)", JVET-Q2007, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 17th Meeting: Brussels, BE, 7-17 January 2020. Summary of the Invention [Problem to be solved by the invention]

[0006] In Japan, the new 4K / 8K satellite broadcasting that began in December 2018 has a transmission capacity of approximately 100Mbps, and one 8K video is transmitted using the HEVC format. Therefore, even if the video bit rate can be halved by adopting the VVC format, it will be difficult to maintain the quality of 8K video at the service quality level in scenes with complex images or movement with the transmission capacity of approximately 40Mbps for next-generation terrestrial broadcasting.

[0007] An object of the present invention is to provide a video encoding device and a video decoding device that can maintain high video quality for ultra-high definition video. [Means for solving the problem]

[0008] A video encoding device according to the present invention includes multiplexing means for multiplexing the maximum image width and maximum image height of luma samples of all frames into a bitstream, and means for periodically switching the image width and image height of luma samples of the frames between a first image width and image height and a second image width and image height in a predetermined AU, the multiplexing means being controlled to multiplex the determined image width and image height of luma samples into the bitstream.A video encoding method according to the present invention multiplexes the maximum image width and maximum image height of luma samples of all frames into a bitstream, and periodically switching the image width and image height of luma samples of the frames between a first image width and image height and a second image width and image height in a predetermined AU, the multiplexing process being controlled to multiplex the determined image width and image height of luma samples into the bitstream.

[0009] A video decoding device according to the present invention includes demultiplexing means for demultiplexing a maximum image width and a maximum image height of luma samples of all frames from a bitstream and for demultiplexing the image width and image height of luma samples from the bitstream for each frame, wherein the demultiplexing means obtains image widths and image heights of luma samples that are periodically switched between a first image width and image height and a second image width and image height in a predetermined AU. A video decoding method according to the present invention demultiplexes a maximum image width and a maximum image height of luma samples of all frames from the bitstream and for demultiplexing the image width and image height of luma samples from the bitstream for each frame, wherein the demultiplexing process obtains image widths and image heights of luma samples that are periodically switched between a first image width and image height and a second image width and image height in the predetermined AU. [Effects of the Invention]

[0010] According to the present invention, it is possible to maintain high image quality of ultra-high definition images. [Brief explanation of the drawings]

[0011] [Figure 1]FIG. 10 is an explanatory diagram showing an example of 65 types of angular intra prediction. [Figure 2] FIG. 1 is an explanatory diagram showing an example of inter-frame prediction. [Figure 3] 10 is an explanatory diagram showing an example of CTU division of frame t and an example of CU division of CTU8 of frame t. FIG. [Figure 4] 1 is a block diagram illustrating an example of the configuration of a video encoding device according to a first embodiment. [Figure 5] 10 is a flowchart illustrating the operation of an encoding controller. [Figure 6] 10 is a flowchart showing the operation of the video encoding device. [Figure 7] FIG. 1 is an explanatory diagram showing the L2 structure of an SOP. [Figure 8] FIG. 1 is an explanatory diagram showing the L3 structure of an SOP. [Figure 9] FIG. 1 is an explanatory diagram showing the L4 structure of an SOP. [Figure 10] FIG. 10 is an explanatory diagram illustrating a method for switching image sizes depending on the difficulty of video encoding of a scene. [Figure 11] FIG. 1 is a block diagram illustrating an example of the configuration of a video decoding device. [Figure 12] FIG. 1 is a block diagram illustrating an example of the configuration of a video system. [Figure 13] FIG. 1 is a block diagram showing an example configuration of an information processing system capable of implementing the functions of a video encoding device and a video decoding device. [Figure 14] 1 is a block diagram showing the main parts of a video encoding device; [Figure 15] FIG. 2 is a block diagram showing the main parts of a video decoding device. DETAILED DESCRIPTION OF THE INVENTION

[0012] To facilitate understanding of the following description, intra prediction, inter-frame prediction, coding tree units (CTUs) and coding units (CUs) will be first described.

[0013] Each frame of digitized video is divided into CTUs, and each CTU is coded in raster scan order.

[0014] Each CTU is divided into CUs in a quad-tree (QT) or multi-tree (MT) structure and then coded.

[0015] Each CU is subjected to predictive coding. Predictive coding includes intra-frame prediction and inter-frame prediction. The prediction error of each CU is then transformed based on frequency transformation.

[0016] Intra prediction is a prediction that generates a predicted image from a reconstructed image that has the same display time as the frame to be coded. Non-Patent Document 1 defines 65 types of angular intra prediction as shown in FIG. 1. Angular intra prediction generates an intra prediction signal by extrapolating reconstructed pixels around the block to be coded in one of 65 directions. In addition to angular intra prediction, Non-Patent Document 1 also defines DC intra prediction, which averages reconstructed pixels around the block to be coded, and planar intra prediction, which linearly interpolates reconstructed pixels around the block to be coded. Hereinafter, a CU coded based on intra prediction is called an intra CU.

[0017] Inter-frame prediction is a method of generating a predicted image from a reconstructed image (reference picture) that has a different display time from the current frame to be coded. Hereinafter, inter-frame prediction is also referred to as inter-prediction.

[0018] FIG. 2 is an explanatory diagram showing an example of inter-frame prediction. x , mv y ) indicates the translational movement amount of the reconstructed image block of the reference picture relative to the block to be coded. Inter prediction generates an inter prediction signal based on the reconstructed image block of the reference picture (using pixel interpolation if necessary). Hereinafter, a CU coded based on inter-frame prediction will be referred to as an inter CU.

[0019] A frame coded using only intra CUs is called an I-frame (or I-picture). A frame coded using not only intra CUs but also inter CUs is called a P-frame (or P-picture). A frame coded using not only one reference picture but also two reference pictures simultaneously for block inter prediction is called a B-frame (or B-picture).

[0020] Note that inter prediction using one reference picture is called unidirectional prediction, and inter prediction using two reference pictures at the same time is called bidirectional prediction.

[0021] FIG. 3 is an explanatory diagram showing an example of CTU division of frame t when the number of pixels of the frame is CIF (Common Intermediate Format) and the CTU size is 64, and an example of division of the eighth CTU (CTU8) included in frame t.

[0022] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0023] Embodiment 1. 4 is a block diagram showing an example of the configuration of a video encoding device according to Embodiment 1. The video encoding device 100 of this embodiment includes a transformer / quantizer 101, an entropy encoder 102, an inverse transformer / inverse quantizer 103, a buffer 104, a predictor 105, a multiplexer 106, a pixel number converter 107, and an encoding controller 108.

[0024] The encoding controller controls the pixel number converter 107 and other components. The pixel number converter 107 has a function of converting the image size of the input video into a pixel size determined by the encoding controller .

[0025] An ultra-high definition video frame (image signal) is input to the pixel number converter 107. The transformer / quantizer 101 frequency-transforms a prediction error image obtained by subtracting a prediction signal from the image signal supplied from the pixel number converter 107, to obtain frequency transform coefficients. Furthermore, the transformer / quantizer 101 quantizes the frequency-transformed prediction error image (frequency transform coefficients) using a predetermined quantization step width. Hereinafter, the quantized frequency transform coefficients are referred to as transformed quantization values.

[0026] The entropy encoder 102 entropy encodes the cu_split_flag, syntax value, pred_mode_flag, syntax value, intra prediction direction, motion vector difference information, and transformed / quantized value determined by the predictor 105.

[0027] The inverse transform / inverse quantizer 103 inversely quantizes the transformed and quantized values ​​using a predetermined quantization step width. Furthermore, the inverse transform / inverse quantizer 103 inversely frequency transforms the inversely quantized frequency transform coefficients. A prediction signal is added to the reconstructed prediction error image obtained by the inverse frequency transform, and the reconstructed prediction error image is supplied to a buffer 104. The buffer 104 stores the supplied reconstructed image.

[0028] The multiplexer 106 multiplexes the output data of the entropy encoder 102 and outputs the multiplexed data.

[0029] Next, the operation of the encoding controller 108 in the video encoding device 100 will be described with reference to the flowchart in Fig. 5. Note that an example will be taken in which the input video, which is an ultra-high definition video input to the pixel number converter 107, is an 8K video (7680 pixels horizontally and 4320 pixels vertically).

[0030] The encoding controller 108 determines the image size of the image frame to be processed (processing target frame) (step S101). How this is determined will be described later.

[0031] The encoding controller 108 controls the operation of the pixel number converter 107 for the frame to be processed based on the determined image size (step S102).

[0032] When the frame to be processed is to be processed as 8K video, the encoding controller 108 controls the pixel number converter 107 so that the image size of the frame output remains 8K (7,680 pixels horizontally, 4,320 pixels vertically). That is, the encoding controller 108 provides a command to the pixel number converter 107 to do so. Otherwise (when processing as 4K video), the encoding controller 108 provides a command to the pixel number converter 107 to do so so that the image size of the output frame from the pixel number converter 107 becomes 4K (3,840 pixels horizontally, 2,160 pixels vertically). That is, the encoding controller 108 provides a command to the pixel number converter 107 to do so. The pixel number converter 107 reduces the number of pixels of the frame in accordance with the command.

[0033] Next, the encoding controller 108 controls the multiplexer 106 based on the determined image size (step S103). The encoding controller 108 controls the multiplexer 106, for example, as follows.

[0034] The encoding controller 108 controls the multiplexer 106 to set the values ​​of the pic_width_max_in_luma_samples syntax (corresponding to the maximum image width of the luminance samples) and the pic_height_max_in_luma_samples syntax (corresponding to the maximum image height of the luminance samples) in the sequence parameter set output by the multiplexer 106 to 7680 and 4320, respectively. That is, the encoding controller 108 provides the multiplexer 106 with a command to do so.

[0035] Furthermore, when processing a frame to be processed as 8K video, the encoding controller 108 controls the multiplexer 106 to output a picture parameter set for the frame to be processed such that the values ​​of the pic_width_in_luma_samples syntax (corresponding to the image width of the luminance samples) and the pic_height_in_luma_samples syntax (corresponding to the image height of the luminance samples) become 7680 and 4320, respectively. That is, the encoding controller 108 provides the multiplexer 106 with a command to do so.

[0036] Otherwise (when processing as 4K video), the encoding controller 108 controls the multiplexer 106 to set the values ​​of the pic_width_in_luma_samples syntax (corresponding to the image width of the luminance samples) and the pic_height_in_luma_samples syntax (corresponding to the image height of the luminance samples) in the picture parameter set of the frame to be processed to 3840 and 2160, respectively. In other words, the encoding controller 108 provides the multiplexer 106 with a command to do so.

[0037] The multiplexer 106 multiplexes the pic_width_max_in_luma_samples syntax value and the pic_height_max_in_luma_samples syntax value for all frames into a bitstream under the control of the encoding controller 108. The multiplexer 106 also multiplexes the pic_width_in_luma_samples syntax value and the pic_height_in_luma_samples syntax value for each frame into a bitstream.

[0038] Furthermore, in order to scale the image size of the frame to be processed to the image size of a previously processed frame, the encoding controller 108 derives a reference picture scale ratio RefPicScale for each previously processed frame and supplies it to the predictor 105 (step S104).

[0039] RefPicScale is expressed by the following equation, which is described in 8.3.2 Decoding process for reference picture lists construction in Non-Patent Document 1. RefPicScale[ i ][ j ]

[0000] = ( ( fRefWidth << 14 ) + ( PicOutputWidthL >> 1 ) ) / PicOutputWidthL RefPicScale[ i ][ j ]

[0001] = ( ( fRefHeight << 14 ) + ( PicOutputHeightL >> 1 ) ) / PicOutputHeightL ···(1)

[0040] where PicOutputWidthL=pic_width_in_luma_samples, PicOutputHeightL=pic_height_in_luma_samples, and fRefWidth and fRefHeight are the values ​​of the pic_width_in_luma_samples syntax and the pic_height_in_luma_samples syntax, respectively, that were set for the target previously processed frame.

[0041] As can be seen from equation (1), the reference picture scale ratio is the ratio between the image size of a previously processed frame and the image size of the frame to be processed.

[0042] Next, the overall operation of the video encoding device 100 will be described with reference to the flowchart of FIG.

[0043] The predictor 105 performs predictive coding. That is, the predictor 105 first determines, for each CTU, a cu_split_flag syntax value that determines a CU partition shape that minimizes the coding cost (step S201). Next, the predictor 105 determines, for each CU, coding parameters that minimize the coding cost (such as a pred_mode_flag syntax value that determines intra prediction / inter prediction, an intra prediction direction, and motion vector difference information) (step S202).

[0044] Furthermore, the predictor 105 generates a prediction signal for the input image signal of each CU based on the determined cu_split_flag syntax value, pred_mode_flag syntax value, intra prediction direction, motion vector, reference picture scale ratio, etc. (Step S203). The prediction signal is generated based on intra prediction or inter-frame prediction.

[0045] As described above, the pixel number converter 107 scales the frame to be processed so that it has the image size determined by the encoding controller 108 .

[0046] The transformer / quantizer 101 frequency-transforms a prediction error image obtained by subtracting the prediction signal from the image signal supplied from the pixel number converter 107 (step S204). Furthermore, the transformer / quantizer 101 quantizes the frequency-transformed prediction error image (frequency transform coefficients) (step S205).

[0047] The entropy encoder 102 entropy encodes the cu_split_flag syntax value, pred_mode_flag syntax value, intra prediction direction, motion vector difference information, and quantized frequency transform coefficients (transformed and quantized values) determined by the predictor 105 (step S206).

[0048] The multiplexer 106 multiplexes the entropy-encoded data supplied from the entropy encoder 102 and outputs the multiplexed data as a bit stream (step S207).

[0049] The inverse transform / inverse quantizer 103 inversely quantizes the transformed and quantized values. Furthermore, the inverse transform / inverse quantizer 103 inversely frequency transforms the inversely quantized frequency transform coefficients. The inversely frequency transformed reconstructed prediction error image is added with a prediction signal and supplied to a buffer 104. The buffer 104 stores the reconstructed image.

[0050] Through the above-described operations, the video encoding device of this embodiment generates a bitstream.

[0051] <Example of how to determine image size> As an example of how to determine the image size, a method of switching the image size of the frame to be processed between 8K and 4K according to the Temporal ID of the SOP structure will be described. Note that the Temporal ID of an AU is the value obtained by subtracting 1 from nuh_temporal_id_plus1 in the NALU (Network Abstraction Layer Unit) header within the AU.

[0052] Fig. 7 is an explanatory diagram showing the L2 structure of an SOP, Fig. 8 is an explanatory diagram showing the L3 structure of an SOP, and Fig. 9 is an explanatory diagram showing the L4 structure of an SOP.

[0053] 7 to 9 show an example in which frames included in an AU whose Temporal ID value is equal to or greater than a predetermined threshold are set to a small image size (4K), and frames of other AUs are set to the same image size (8K). However, in the example shown in FIGS. 7 to 9, the predetermined threshold is 2.

[0054] As described above, when a video encoding device is configured to switch between 8K and 4K, the high-resolution 8K image is periodically displayed, resulting in an afterimage effect. In other words, the high definition of 8K video can be perceived. Furthermore, since the data volume is reduced in 4K frames, degradation due to video encoding is prevented even in scenes with complex patterns or movement. In other words, high video quality can be maintained. Furthermore, since there is no need for reloading the video bitstream on the receiving terminal side, such as a video decoding device, video can be played back smoothly on the receiving terminal side even when the image size is switched.

[0055] Note that the threshold value of 2 for the Temporal ID value used to determine the AU to be processed at the small image size described above is just an example, and other values ​​may be used.

[0056] Furthermore, in cases where video encoding is easy, the encoding controller 108 may also keep the image size of frames included in AUs whose Temporal ID values ​​are equal to or greater than a predetermined threshold. That is, the encoding controller 108 may keep the image size or a smaller image size of frames included in AUs whose Temporal ID values ​​are equal to or greater than a predetermined threshold, and always keep the image size of frames in other AUs.

[0057] Furthermore, in order to obtain a desirable afterimage effect, it is desirable to process frames included in an AU including an I-picture whose Temporal ID value is less than a predetermined threshold value at a larger image size than other frames. On the other hand, in order to maximize the effect of reducing the amount of data, it is desirable to make the image size of frames included in an AU whose Temporal ID value is equal to or greater than a predetermined threshold value larger than the image size of frames included in an AU whose Temporal ID value is less than the predetermined threshold value.

[0058] <Another example of how to determine image size> As another example of how to determine the image size, the encoding controller 108 may switch the image size of the frame to be processed between 8K and 4K depending on the difficulty (degree of difficulty) of video encoding of the scene, as illustrated in FIG. 10.

[0059] The difficulty of video encoding can be determined based on the results of monitoring the characteristics of the input video (such as the complexity of the image and movement) and the output characteristics of the entropy encoder 102 (such as the coarseness of quantization).

[0060] To absorb the difference in image quality at the transition between 4K and 8K, it is desirable to use the frame before the transition as a reference picture for the leading picture of the first I-picture after the transition. This is because a smoothing effect can be achieved by generating a predicted image for the leading picture using bidirectional prediction that combines 4K and 8K images.

[0061] Furthermore, for the leading picture after switching to 8K, it is desirable to process it in 4K to reduce the amount of data, but on the other hand, it is desirable to process it in 8K to maximize the smoothing effect.

[0062] Next, the configuration and operation of a video decoding device will be described. Fig. 11 is a block diagram showing an example configuration of a video decoding device according to this embodiment. The video decoding device 200 shown in Fig. 11 is capable of receiving a bitstream from the video encoding device 100 shown in Fig. 4 and performing video decoding processing. However, the source of the bitstream is not limited to the video encoding device 100 shown in Fig. 4.

[0063] The video decoding device shown in FIG. 11 includes a demultiplexer 201, an entropy decoder 202, an inverse transformer / inverse quantizer 203, a predictor 204, a buffer 205, a pixel number converter 206, and a decoding control unit 208.

[0064] The demultiplexer 201 demultiplexes the input bitstream to extract the entropy-encoded data.

[0065] The entropy decoder 202 entropy-decodes the entropy-encoded data. The entropy decoder 202 supplies the entropy-decoded transformed and quantized values ​​to the inverse transformer / inverse quantizer 203, and further supplies the cu_split_flag, pred_mode_flag, intra prediction direction, and motion vector to the predictor 204.

[0066] In this embodiment, data representing the maximum image width and maximum image height of luma samples of all frames (e.g., the pic_width_max_in_luma_samples syntax value and the pic_height_max_in_luma_samples syntax value) are multiplexed into the bitstream. Data representing the image width and image height of luma samples for each frame (e.g., the pic_width_in_luma_samples syntax value and the pic_height_in_luma_samples syntax value) are also multiplexed into the bitstream. The entropy decoder 202 supplies the entropy-decoded data to the decoding controller 208.

[0067] The decoding controller 208 derives the reference picture scale ratio RefPicScale for each frame from the pic_width_in_luma_samples syntax value and the pic_height_in_luma_samples syntax value based on, for example, equation (1). The decoding controller 208 supplies the reference picture scale ratio RefPicScale for each frame to the predictor 204. The decoding controller 208 also supplies the pic_width_max_in_luma_samples syntax value, the pic_height_max_in_luma_samples syntax value, the pic_width_in_luma_samples syntax value, and the pic_height_in_luma_samples syntax value to the pixel number converter 206.

[0068] The inverse transform / inverse quantizer 203 inversely quantizes the transformed and quantized values ​​using a predetermined quantization step width, and further inversely frequency transforms the inversely quantized frequency transform coefficients.

[0069] The predictor 204 generates a prediction signal based on the cu_split_flag, the pred_mode_flag, the intra prediction direction, the motion vector, and the reference picture scale ratio RefPicScale. The prediction signal is generated based on intra prediction or inter-frame prediction.

[0070] The reconstructed prediction error image that has been inverse frequency transformed by the inverse transformer / inverse quantizer 203 is added with a prediction signal supplied from the predictor 204, and is supplied as a reconstructed image to the buffer 205. Then, the reconstructed picture stored in the buffer 205 is output as a decoded video.

[0071] Through the above-described operations, the video decoding device of this embodiment generates decoded video.

[0072] The decoded video data is supplied to a display device or storage device as display video data, and the pixel number converter 206 scales each decoded video to a predetermined image width and image height so that all display video data has the same image size. For example, the maximum image width and maximum image height can be used as the predetermined image width and image height. In this case, the pixel number converter 206 can derive the ratio for the scaling using the pic_width_in_luma_samples syntax value, the pic_height_max_in_luma_samples syntax value, and the pic_width_max_in_luma_samples syntax value and the pic_height_in_luma_samples syntax value.

[0073] In this embodiment, the image size of a reconstructed image frame may differ from frame to frame. Therefore, in this embodiment, the video decoding device 200 is configured so that the pixel number converter 206 performs size conversion so that the image size of the reconstructed image frame becomes the size indicated by the values ​​of the pic_width_max_in_luma_samples syntax and the pic_height_max_in_luma_samples syntax included in the sequence parameter set, in order to align the displayed image sizes. This allows for smooth video playback even when the image size is changed.

[0074] As described above, in this embodiment, the video encoding device performs video encoding by switching image sizes so as to maintain video quality at a service quality level even in scenes with complex patterns or movement. The video encoding device also utilizes reference picture scaling in the video encoding so that a receiving terminal, such as a video decoder, does not need to reload a video bitstream due to the image size switch. Furthermore, the video encoding device can also control the video encoding so that the image size switch is visually less noticeable.

[0075] Therefore, video quality can be maintained at the service quality level even in scenes with complex patterns or movement. Furthermore, there is no need for the receiving terminal to reload the video bitstream, allowing for smooth video playback even when the image size changes. Furthermore, the change in image size becomes less noticeable, allowing video quality to be maintained at the service quality level at the moment the image size changes.

[0076] In the above embodiment, when the input video is an 8K video, the video is switched between 8K video (7680 pixels horizontally, 4320 pixels vertically) and 4K video (3840 pixels horizontally, 2160 pixels vertically) while maintaining the same aspect ratio, but in another embodiment, it is also possible to switch the aspect ratio.

[0077] For example, it may be possible to switch between 8K video with an aspect ratio of 16:9 (7680 pixels horizontally, 4320 pixels vertically) and 8K video with an aspect ratio of 4:3 (5760 pixels horizontally, 4320 pixels vertically). In this case, however, the VUI (Video Usability Information) and Sample aspect ratio information SEI (Supplemental Enhancement Information) messages will be as follows:

[0078] [VUI] The value of vui_aspect_ratio_constant_flag included in the VUI is 0.

[0079] [Sample aspect ratio information SEI message] Each AU contains a Sample aspect ratio information SEI message. To ensure that each playback video of AUs coded with different aspect ratios is displayed at the same size, the pixel aspect expressed by sari_aspect_ratio_idc, sari_sar_width, and sari_sar_height in the SEI message of an AU coded with an image size of one aspect ratio is different from the sari_aspect_ratio_idc, sari_sar_width, and sari_sar_height in the SEI message of an AU coded with an image size of the other aspect ratio.

[0080] In the above example, when vui_aspect_ratio_idc is 1, sari_aspect_ratio_idc of the SEI message of the AU coded with 8K video having an aspect ratio of 16:9 is 1, and sari_aspect_ratio_idc of the SEI message of the AU coded with 8K video having an aspect ratio of 4:3 is 14.

[0081] Embodiment 2. Fig. 12 is a block diagram showing an example of the configuration of a video system, in which the video encoding device 100 and the video decoding device 200 are connected via a wireless transmission path or a wired transmission path 300.

[0082] In the video system 300, the video encoding device 100 can generate the bitstream as described above. Also, in the video system 300, the video decoding device 200 can decode the bitstream as described above.

[0083] Furthermore, each of the above embodiments can be configured by hardware, but can also be realized by a computer program.

[0084] 13 includes a processor 1001 including a CPU, a program memory 1002, a storage medium 1003 for storing video data, and a storage medium 1004 for storing a bitstream. The storage medium 1003 and the storage medium 1004 may be separate storage media or may be storage areas of the same storage medium. A magnetic storage medium such as a hard disk can be used as the storage medium.

[0085] In the information processing system shown in Fig. 13, a program memory 1002 stores a program (a video encoding program or a video decoding program) for implementing the functions of each block (excluding the buffer block) shown in Fig. 4 and Fig. 11. The processor 1001 executes processing in accordance with the program stored in the program memory 1002, thereby implementing the functions of the video encoding device or the video decoding device shown in Fig. 4 and Fig. 11, respectively.

[0086] Note that some of the functions of the video encoding device or video decoding device shown in FIG. 4 or FIG. 11 may be implemented by a semiconductor integrated circuit, and the remaining functions may be implemented by the processor 1000 or the like.

[0087] The program memory 1002 is, for example, a non-transitory computer-readable medium. The non-transitory computer-readable medium includes various types of tangible storage media. Specific examples of the non-transitory computer-readable medium include semiconductor memory, magnetic storage media (e.g., hard disks), and magneto-optical storage media (e.g., magneto-optical disks).

[0088] The program may also be stored in various types of transitory computer-readable media, and the program may be supplied to the transitory computer-readable media (e.g., flash ROM) via a wired or wireless communication path, i.e., via an electrical signal, an optical signal, or an electromagnetic wave.

[0089] Fig. 14 is a block diagram showing the main components of a video encoding device. The video encoding device 10 shown in Fig. 14 includes a multiplexing unit (multiplexing means) 11 (implemented by a multiplexer 106 in the embodiment) that multiplexes the maximum image width (specifically, data representing the maximum image width, for example, pic_width_max_in_luma_samples syntax) and the maximum image height (specifically, data representing the maximum image width, for example, pic_height_max_in_luma_samples syntax) of luminance samples of all frames into a bitstream, and a multiplexing unit (multiplexing means) 12 (implemented by a multiplexer 106 in the embodiment) that multiplexes the image width (specifically, data representing the image width, for example, pic_width_in_luma_samples syntax) of luminance samples that are equal to or less than the maximum image width and the maximum image height for each frame. The multiplexing unit 11 includes a determination unit (determination means) 12 (implemented by the encoding control unit 108 in the embodiment) that determines the image width and image height of the luminance samples (specifically, data representing the height of the image; for example, the pic_height_in_luma_samples syntax), and a derivation unit (deriving means) 13 (implemented by the encoding control unit 108 in the embodiment) that multiplexes the determined image width and image height of the luminance samples into a bitstream and derives a reference picture scale ratio for scaling the image width and image height of the luminance samples of the frame to be processed to the image width and image height of the luminance samples of a previously processed frame.

[0090] Fig. 15 is a block diagram showing the main components of a video decoding device. The video decoding device 20 shown in Fig. 15 includes a demultiplexing unit (demultiplexing means) 21 (implemented by a demultiplexer 201 in the embodiment) that demultiplexes the maximum image width and maximum image height of luminance samples of all frames from the bitstream and demultiplexes the image width and image height of luminance samples from the bitstream for each frame, a derivation unit (derivation means) 22 (implemented by a decoding controller 208 in the embodiment) that derives a reference picture scale ratio for scaling the image width and image height of luminance samples of a frame to be processed to the image width and image height of luminance samples of a previously processed frame, and a scaling unit (scaling means) 23 (implemented by a pixel number converter 206 in the embodiment) that scales the image size of a frame to be output for display to the maximum image width and maximum image height based on information related to the reference picture scale ratio (e.g., the reference picture scale ratio RefPicScale itself or a syntax value for deriving the reference picture scale ratio RefPicScale).

[0091] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes.

[0092] (Appendix 1) A computer-readable recording medium on which a video encoding program is recorded, The video encoding program is installed on a computer. multiplexing the maximum image width and maximum image height of all frame luma samples into the bitstream; determining an image width and an image height of luminance samples for each frame that are less than or equal to the maximum image width and the maximum image height; multiplexing the determined image width and image height of the luminance samples into a bitstream; deriving a reference picture scale ratio for scaling the image width and image height of the luminance samples of the frame to be processed to the image width and image height of the luminance samples of the previously processed frame; Execute the following.

[0093] (Appendix 2) A computer-readable recording medium on which a video decoding program is recorded, The video decoding program is installed on a computer. demultiplexing the maximum image width and maximum image height of all frame luma samples from the bitstream; demultiplexing the image width and image height of luma samples from the bitstream for each frame; deriving a reference picture scale ratio for scaling the width and height of the luminance samples of the current frame to the width and height of the luminance samples of a previously processed frame; a process of scaling the image size of the frame to be output for display to the maximum image width and the maximum image height; Execute the following.

[0094] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. [Explanation of symbols]

[0095] 10,100 Video Encoding Device 11 Multiplexer 12 Decision Section 13 Derivation part 20,200 video decoder 21 Demultiplexer 22 Derivation part 23 Scaling section 101 Transform / Quantizer 102 Entropy Encoder 103 Inverse Transform / Inverse Quantizer 104 buffers 105 Predictor 106 Multiplexer 107 Pixel Converter 108 Coding controller 201 Demultiplexer 202 Entropy Decoder 203 Inverse Transform / Inverse Quantizer 204 Predictor 205 buffers 206 Pixel Converter 208 Decoding control unit 300 Video System 1001 processor 1002 program memory 1003,1004 Storage medium

Claims

1. multiplexing means for multiplexing a maximum image width and a maximum image height of luminance samples of all frames into a bitstream; means for periodically switching an image width and an image height of a luminance sample of a frame between a first image width and an image height and a second image width and an image height in a predetermined AU (Access Unit); The multiplexing means is controlled to multiplex the determined image width and image height of the luminance samples into a bitstream. Video encoding device.

2. demultiplexing means for demultiplexing a maximum image width and a maximum image height of luminance samples of all frames from the bitstream, and for demultiplexing an image width and an image height of luminance samples for each frame from the bitstream; The demultiplexing means acquires an image width and an image height of luminance samples that are periodically switched between a first image width and an image height and a second image width and an image height in a predetermined AU (Access Unit). Video decoder.

3. multiplexing the maximum image width and maximum image height of all frame luma samples into the bitstream; periodically switching an image width and an image height of the luminance samples of the frame between a first image width and an image height and a second image width and an image height in a predetermined access unit (AU); The multiplexing process is controlled to multiplex the determined image width and image height of the luminance samples into a bitstream. Video coding method.

4. demultiplexing the maximum image width and maximum image height of luma samples of all frames from the bitstream, and demultiplexing the image width and image height of luma samples for each frame from the bitstream; In the demultiplexing process, the image width and image height of the luminance samples are periodically switched between the first image width and image height and the second image width and image height in a predetermined AU (Access Unit). Video decoding method.

Citation Information

Patent Citations

  • Stereoscopic image data generation device and stereoscopic image data generation method

    JP2012044537A