High-level signaling method and apparatus for weighted prediction
By reordering and restricting the coding of syntax elements in video coding, especially for weighted prediction parameters, the method improves compression efficiency and maintains picture quality during lighting transitions, addressing inefficiencies in existing video coding technologies.
Patent Information
- Application Number
- JP2025045581
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-06
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2040-09-07
AI Technical Summary
Existing video coding technologies face challenges in efficiently compressing video data while maintaining high quality, particularly in handling lighting changes such as fades, due to the inefficiencies in signaling and coding of weighted prediction parameters.
The method involves reordering syntax elements to include high-level syntax (HLS) weighted prediction parameters before coding the reference picture list structure, and restricting binarization based on these parameters, especially when delta POC values are zero, to reduce the amount of weighted prediction parameters coded.
This approach enhances compression efficiency by reducing the amount of weighted prediction parameters, thereby improving coding efficiency and maintaining picture quality during lighting transitions.
Smart Images

Figure 2025100557000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This patent application claims priority to International Patent Application No. PCT / RU2019 / 000625, filed on September 6, 2019. The disclosure of the above - mentioned patent application is hereby incorporated by reference in its entirety.
[0002] Embodiments of the present disclosure generally relate to the field of picture processing, and more particularly, to shape - adaptive resampling of residual blocks for still image and video coding.
Background Art
[0003] Video coding (video encoding and video decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the Internet and mobile networks, real - time conversational applications such as video chat, video conferencing, DVD and Blu - ray (registered trademark) disks, video content collection and editing systems, and camcorders for security applications.
[0004] The amount of video data required to depict even relatively short videos can be significant, which can pose difficulties when the data is to be streamed or otherwise transmitted across a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being transmitted across modern telecommunications networks. The size of the video can also be an issue when the video is stored in a memory device since memory resources can be limited. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. With limited network resources and an increasing demand for even higher video quality, improved compression and decompression techniques that improve the compression ratio with little to no sacrifice in picture quality are desirable.
[0005] Weighted Prediction (WP) is a particularly useful tool for coding fades. Weighted Prediction can compensate for lighting changes such as fades in, fades out, or cross fades. The Weighted Prediction (WP) tool is adopted in the main and extended profiles of the H.264 video coding standard to improve coding efficiency by applying multiplicative weight factors and additive offsets to motion compensated prediction to form weighted prediction. In the explicit mode, the weight factors and offsets may be coded within the slice header for each acceptable reference picture index. In the implicit mode, the weight factors are not coded but are derived based on the relative picture order count (POC) distance of two reference pictures.
[0006] The relationship of pictures with respect to ordering and distance when used for prediction is represented by the POC. The POC value is an index number that defines the output position of the current picture within the coded video sequence. The POC value is used to identify pictures within the decoded picture buffer. For identification, the value of the POC increases strictly along with the output order of the coded pictures. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0007] According to a first aspect of the present disclosure, an encoding method is provided. The method includes determining a syntax element to be coded, where the syntax element includes a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter; coding at least one HLS weighted prediction parameter; and coding the reference picture list structure following the coding of at least one HLS weighted prediction parameter. The syntax element is reordered so that coding of the reference picture list structure can be based on the value of at least one HLS weighted prediction parameter.
[0008] In a first implementation form of the first aspect itself, the reference picture list derived from the reference picture list structure includes reference pictures having the same picture order count (POC) parameter.
[0009] In any preceding implementation form of the first aspect or in a second implementation form of the first aspect itself, at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted single prediction.
[0010] In any preceding implementation form of the first aspect or the third implementation form of the first aspect itself, at least one HLS weighted prediction parameter comprises a sequence parameter set flag for weighted bi-prediction.
[0011] In any preceding implementation form of the first aspect or the fourth implementation form of the method according to the first aspect itself, the coding of the reference picture list structure comprises a restriction on the binarization of at least a part of the reference picture list structure.
[0012] In the fifth implementation form of the fourth implementation form of the first aspect, the restriction on the binarization of at least a part of the reference picture list structure comprises coding a modified delta POC value for an element of the reference picture list when the sequence parameter set flag for weighted uni-prediction is set to 0, and the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process.
[0013] In the sixth implementation form of the fourth implementation form of the first aspect, the restriction on binarization is that (i) when at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted bi-prediction is set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, where the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, the said coding, or (ii) when at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted bi-prediction and a sequence parameter set flag for weighted uni-prediction, and at least one of the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction is set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, where the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, the said coding, (iii) when at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted bi-prediction and a sequence parameter set flag for weighted uni-prediction, and both the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction are set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, where the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, the said coding.
[0014] In the seventh implementation form of the fifth or sixth implementation form of the first aspect, the modified delta POC value is 1 less than the delta POC value used in the coding process.
[0015] According to a second aspect of the present disclosure, there is provided a method of determining a syntax element to be coded, the syntax element including a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter, the reference picture list derived from the reference picture list structure comprising reference pictures having the same picture order count (POC) parameter, and coding the syntax element determined in coding order with a restriction on the binarization of the syntax element having a later position in the coding order. When at least one HLS weighted prediction parameter is coded after the reference picture list structure in coding order, the restriction on the binarization of the syntax element comprises coding at least one HLS weighted prediction parameter only when the reference picture list has at least one element having a delta POC value equal to 0. This has the advantage that the amount of weighted prediction parameters to be coded can be reduced.
[0016] According to a third aspect of the present disclosure, there is provided a decoding method by a decoder, comprising receiving a bitstream, entropy decoding the bitstream to obtain a syntax element, the syntax element comprising a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter, wherein among the syntax elements, at least one HLS weighted prediction parameter is entropy decoded before the reference picture list structure, performing a prediction based on the obtained syntax element to obtain a prediction block, reconstructing a reconstructed block based on the prediction block, and obtaining a decoded picture based on the reconstructed block.
[0017] In a first implementation form of the third aspect itself, at least one HLS weighted prediction parameter includes at least one of a sequence parameter set flag for weighted single prediction and a sequence parameter set flag for weighted bi-prediction.
[0018] According to a fourth aspect of the present disclosure, there is provided a decoding method by a decoder, including steps of receiving a bitstream, entropy-decoding the bitstream to obtain syntax elements, where the syntax elements include a reference picture list structure and a preset flag, and a value of the preset flag indicates whether the syntax elements include at least one high-level syntax (HLS) weighted prediction parameter; performing a prediction based on the obtained syntax elements to obtain a predicted block; reconstructing a reconstructed block based on the predicted block; and obtaining a decoded picture based on the reconstructed block.
[0019] In a first implementation form of the fourth aspect itself, the value of the preset flag corresponds to whether a reference picture list derived from the reference picture list structure has at least one element having a delta POC value equal to 0.
[0020] In a second implementation form of the first implementation form of the fourth aspect, when the value of the preset flag corresponding to the reference picture list has at least one element having a delta POC value equal to 0 and the syntax elements include at least one HLS weighted prediction parameter, or when the value of the preset flag corresponding to the reference picture list has no element having a delta POC value equal to 0, the syntax elements do not include at least one HLS weighted prediction parameter.
[0021] In a third implementation form of the method according to any preceding implementation form of the fourth aspect or the fourth aspect itself, at least one HLS weighted prediction parameter includes at least one of a sequence parameter set flag for weighted single prediction and a sequence parameter set flag for weighted bi-prediction.
[0022] In a fourth implementation form of the method according to any preceding implementation form of the fourth aspect or the fourth aspect itself, the preset flag is RestrictWPFlag which is set to true in the coding process when a delta POC value (AbsDeltaPocSt) with a value of 0 appears during the inspection of each element of the reference picture list.
[0023] According to a fifth aspect of the present disclosure, an encoder is provided that includes a processing circuit for performing the method according to the first aspect, any one of the first to seventh implementation forms of the first aspect, or the second aspect.
[0024] According to a sixth aspect of the present disclosure, a decoder is provided that includes a processing circuit for performing the method according to the third aspect, the first implementation form of the third aspect, the fourth aspect, or any one of the first to fourth implementation forms of the fourth aspect.
[0025] According to a seventh aspect of the present disclosure, a computer program product is provided that includes program code for performing the method according to the first aspect, any one of the first to seventh implementation forms of the first aspect, the second aspect, the third aspect, the first implementation form of the third aspect, the fourth aspect, or any one of the first to fourth implementation forms of the fourth aspect.
[0026] According to an eighth aspect of the present disclosure, there is provided a decoder including one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming, when executed by the processors, configuring the decoder to execute a method according to any one of the third aspect, the first implementation form of the third aspect, the fourth aspect, or the first to fourth implementation forms of the fourth aspect.
[0027] According to a ninth aspect of the present disclosure, there is provided a decoder including receiving means for receiving a bitstream, entropy decoding means for entropy decoding the bitstream to obtain syntax elements, the syntax elements including a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter, and at least one of the HLS weighted prediction parameters being entropy decoded before the reference picture list structure, prediction means for performing a prediction based on the obtained syntax elements to obtain a predicted block, reconstruction means for reconstructing a reconstructed block based on the predicted block, and obtaining means for obtaining a decoded picture based on the reconstructed block.
[0028] In a first implementation form of the ninth aspect itself, the value of a preset flag corresponds to whether a reference picture list derived from a reference list structure has at least one element having a delta POC value equal to 0.
[0029] In the second implementation form of the first implementation form of the ninth aspect, the value of the preset flag corresponding to the reference picture list has at least one element having a delta POC value equal to 0, the syntax element includes at least one HLS weighted prediction parameter, or when the value of the preset flag corresponding to the reference picture list does not have any element having a delta POC value equal to 0, the syntax element does not include at least one HLS weighted prediction parameter.
[0030] In the third implementation form of any preceding implementation form of the ninth aspect or the ninth aspect itself, at least one HLS weighted prediction parameter includes at least one of the sequence parameter set flag for weighted single prediction and the sequence parameter set flag for weighted bi-prediction.
[0031] In the fourth implementation form of any preceding implementation form of the ninth aspect or the ninth aspect itself, the preset flag is RestrictWPFlag which is set to true in the coding process when a delta POC value (AbsDeltaPocSt) of 0 appears during the inspection of each element of the reference picture list.
[0032] According to the tenth aspect of the present disclosure, there is provided an encoder including one or more processors and a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming causing the processor to configure the encoder to execute the method according to the first aspect, any one of the first to seventh implementation forms of the first aspect, or the second aspect when executed.
[0033] According to an eleventh aspect of the present disclosure, there is provided a determination means for determining a syntax element to be coded, the syntax element including a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter, and a coding means for coding at least one HLS weighted prediction parameter and for coding a reference picture list structure following the coding of at least one HLS weighted prediction parameter.
[0034] In a first implementation form of the eleventh aspect itself, the reference picture list derived from the reference picture list structure comprises reference pictures having the same picture order count (POC) parameter.
[0035] In a second implementation form of the eleventh aspect itself or the first implementation form of the eleventh aspect, at least one HLS weighted prediction parameter comprises a sequence parameter set flag for weighted single prediction.
[0036] In a third implementation form of any one of the eleventh aspect itself or the preceding implementation forms of the eleventh aspect, at least one HLS weighted prediction parameter comprises a sequence parameter set flag for weighted bi-prediction.
[0037] In a fourth implementation form of any one of the eleventh aspect itself or the preceding implementation forms of the eleventh aspect, the coding of the reference picture list structure comprises a restriction on the binarization of at least a part of the reference picture list structure.
[0038] In the fifth implementation form of the fourth implementation form of the eleventh aspect, the restriction on the binarization of at least a part of the reference picture list structure comprises signaling a modified delta POC value for an element of the reference picture list when the sequence parameter set flag for weighted single prediction is set to 0, and the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process.
[0039] In the sixth implementation form of the fourth implementation form of the eleventh aspect, the restriction on binarization is as follows: (i) when at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted bi-prediction is set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, where the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process; or (ii) when at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted bi-prediction and a sequence parameter set flag for weighted uni-prediction, and at least one of the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction is set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, where the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process; or (iii) when at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted bi-prediction and a sequence parameter set flag for weighted uni-prediction, and both the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction are set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, where the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process.
[0040] In the seventh implementation form of the fifth or sixth implementation form of the eleventh aspect, the modified delta POC value is 1 less than the delta POC value used in the coding process.
[0041] According to the eleventh aspect of the present disclosure, when executed by a computer device, a non - transitory computer - readable medium carrying program code for causing the computer device to execute a method according to the first aspect, any one of the first to seventh implementation forms of the first aspect, the second aspect, the third aspect, the first implementation form of the third aspect, the fourth aspect, or any one of the first to fourth implementation forms of the fourth aspect is provided.
[0042] Embodiments provide a method for encoding and decoding a video sequence using joint signaling of high - level syntax weighted prediction parameters and a reference picture list.
[0043] Embodiments provide an efficient encoding and / or decoding that uses signaling - related information only in the slice header for slices that allow or enable bidirectional inter - prediction, for example, within a bidirectional (B) prediction slice, also called a B slice.
[0044] The above and other objects are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description, and the figures.
[0045] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, objects, and advantages will become apparent from the description, the drawings, and the claims.
[0046] Hereinafter, embodiments of the invention will be described in more detail with reference to the accompanying figures and drawings.
Brief Description of the Drawings
[0047]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Mode for Carrying Out the Invention
[0048] Hereinafter, the same reference numerals refer to the same or at least functionally equivalent features, unless explicitly specified otherwise.
[0049] In the following description, reference is made to the accompanying drawings, which form a part of the disclosure and illustrate specific aspects of the embodiments of the invention or specific aspects in which embodiments of the invention may be used. It is understood that embodiments of the invention may be used in other aspects and may include structural or logical changes not depicted in the figures. Accordingly, the following detailed description should not be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0050] It is understood that the disclosure regarding the methods described, for example, may equally apply to corresponding devices or systems configured to perform the methods and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units for performing the one or more method steps described, even if such one or more units are not explicitly described or illustrated in the figures, for example, functional units (e.g., one unit for performing one or more steps, or multiple units each performing one or more of the multiple steps). On the other hand, for example, if a specific device is described based on one or more units, for example, functional units, the corresponding method may include one step for performing the functions of the one or more units (e.g., one step for performing the functions of one or more units, or multiple steps each performing one or more of the functions of the multiple units), even if such one or more steps are not explicitly described or illustrated in the figures. Furthermore, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, particularly if not otherwise noted.
[0051] Video coding typically refers to the processing of a sequence of pictures that form a video or video sequence. In the field of video coding, the terms "frame" or "image" may be used synonymously instead of the term "picture". Video coding (or generally coding) comprises two parts, video encoding and video decoding. Video encoding is performed on the source side and typically involves processing the original video picture (e.g., by compression) in order to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves the reverse process compared to the encoder for reconstructing the video picture. Embodiments that refer to the "coding" of a video picture (or generally a picture) are to be understood as relating to the "encoding" or "decoding" of the video picture or respective video sequence. The combination of the encoding part and the decoding part is also called a CODEC (Coding and Decoding).
[0052] In the case of lossless video coding, the original video picture can be reconstructed, i.e., assuming no transmission loss or other data loss during storage or transmission, the reconstructed video picture has the same quality as the original video picture. In the case of lossy video coding, further compression is performed, e.g., by quantization, in order to reduce the amount of data representing the video picture, and the video picture cannot be fully reconstructed at the decoder, i.e., the quality of the reconstructed video picture is lower or worse compared to the quality of the original video picture.
[0053] Some video coding standards belong to the group of "lossy hybrid video codecs" (i.e., they combine spatial and temporal prediction in the sample domain with 2D transform coding for applying quantization in the transform domain). Each picture of a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in the encoder, the video is typically processed, i.e., encoded, at the block (video block) level by, for example, generating prediction blocks using spatial (intra-picture) prediction and / or temporal (inter-picture) prediction, subtracting the prediction blocks from the current block (the block being currently processed / to be processed) to obtain a residual block, transforming the residual block, and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compression), while in the decoder, the reverse process compared to the encoder is applied to the encoded or compressed blocks to reconstruct the current block for presentation. Further, the encoder duplicates the decoder processing loop so that both generate the same prediction (e.g., intra and inter prediction) and / or reconstruction for processing, i.e., coding, subsequent blocks.
[0054] Embodiments of a video coding system 10, a video encoder 20, and a video decoder 30 are described below with reference to FIGS. 1 through 3.
[0055] FIG. 1A is a schematic block diagram illustrating an exemplary coding system 10 that may utilize the techniques of this application, e.g., a video coding system 10 (or simply coding system 10). The video encoder 20 (or simply encoder 20) and the video decoder 30 (or simply decoder 30) of the video coding system 10 represent examples of devices that may be configured to perform techniques according to various examples described in this application.
[0056] As shown in FIG. 1A, the coding system 10 includes a source device 12 configured to provide encoded picture data 21 to a destination device 14, for example, to decode the encoded picture data 13.
[0057] The source device 12 includes an encoder 20 and, in addition, or optionally, may include a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22.
[0058] The picture source 16 may include any kind of picture capture device, for example, a camera for capturing real-world pictures, and / or any kind of picture generation device, for example, a computer graphics processor for generating computer-animated pictures, or a real-world picture, a computer-generated picture (e.g., screen content, virtual reality (VR) picture), and / or any other device for acquiring and / or providing any combination thereof (e.g., augmented reality (AR) picture), or it may be those. The picture source may be any kind of memory or storage device that stores any of the above-described pictures.
[0059] Distinguished from the processing performed by the preprocessor 18 and the preprocessing unit 18, the picture or picture data 17 may also be referred to as unprocessed picture or unprocessed picture data 17.
[0060] The preprocessor 18 is configured to receive the (unprocessed) picture data 17 and perform preprocessing on the picture data 17 to obtain the preprocessed picture 19 or preprocessed picture data 19. The preprocessing performed by the preprocessor 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It is understood that the preprocessing unit 18 may be an optional component.
[0061] The video encoder 20 is configured to receive the preprocessed picture data 19 and provide the encoded picture data 21 (further details will be described below, for example, based on Figure 2).
[0062] The communication interface 22 of the source device 12 is configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) on the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.
[0063] The destination device 14 includes a decoder 30 (e.g., a video decoder 30), and in addition, i.e., optionally, may include a communication interface or communication unit 28, a postprocessor 32 (or postprocessing unit 32), and a display device 34.
[0064] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further processed version thereof), for example, directly from the source device 12 or from any other source, such as a storage device, e.g., an encoded picture data storage device, and provide the encoded picture data 21 to the decoder 30.
[0065] Communication interfaces 22 and 28 may be configured to transmit or receive encoded picture data 21 or encoded data 13 via a direct communication link between source device 12 and destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired or wireless network or any combination thereof, or any type of private and public network, or any combination of any type thereof.
[0066] Communication interface 22 may be configured to, for example, package the encoded picture data 21 in a suitable format, such as in packets, and / or process the encoded picture data using any type of transmission encoding or processing for transmission over a communication link or communication network.
[0067] Communication interface 28, which forms the other side of communication interface 22, may be configured to, for example, receive the transmitted data and process the transmitted data using any type of corresponding transmission decoding or processing and / or packet removal to obtain the encoded picture data 21.
[0068] Both communication interface 22 and communication interface 28 may be configured as a unidirectional communication interface, as indicated by the arrow for communication channel 13 in FIG. 1A, which indicates from source device 12 to destination device 14, or as a bidirectional communication interface, and may be configured to, for example, send and receive messages to affirmatively respond and exchange any other information regarding the communication link and / or data transmission, such as encoded picture data transmission, for example, to set up a connection.
[0069] Decoder 30 is configured to receive the encoded picture data 21 and provide the decoded picture data 31 or the decoded picture (further details will be described below, for example, based on FIG. 3 or FIG. 5).
[0070] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as the reconstructed picture data), for example, the decoded picture 31, to obtain the post-processed picture data 33, for example, the post-processed picture 33. The post-processing executed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, or resampling, or any other processing for preparing the decoded picture data 31 for display by the display device 34, for example.
[0071] The display device 34 of the destination device 14 is configured to receive, for example, the post-processed picture data 33 for displaying a picture to a user or viewer. The display device 34 may be or include any type of display for representing the reconstructed picture, for example, an integrated or external display or monitor. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0072] FIG. 1A depicts the source device 12 and the destination device 14 as separate devices, but embodiments of the device may also include both or either the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality. In such embodiments, the source device 12 or corresponding functionality, and the destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, or by separate hardware and / or software or any combination thereof.
[0073] As will become apparent to those skilled in the art based on the description, the presence and (exact) partitioning of the functionality of the different units or of the functionality within the source device 12 and / or destination device 14 as represented in FIG. 1A may vary depending on the actual device and application.
[0074] The encoder 20 (e.g., video encoder 20) or decoder 30 (e.g., video decoder 30), or both the encoder 20 and decoder 30, can be implemented via a processing circuit as shown in FIG. 1B, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof. The encoder 20 can be implemented via processing circuit 46 to embody various modules as discussed with respect to encoder 20 of FIG. 2 and / or any other encoder system or subsystem described herein. The decoder 30 can be implemented via processing circuit 46 to embody various modules as discussed with respect to decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuit can be configured to perform various operations as discussed later. As shown in FIG. 5, if the technique is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Either the video encoder 20 or the video decoder 30 can be integrated within a single device, for example, as part of a combined encoder / decoder (CODEC) as shown in FIG. 1B.
[0075] The source device 12 and the destination device 14 may comprise any kind of handheld or stationary device, such as a notebook or laptop computer, mobile phone, smartphone, tablet or tablet computer, camera, desktop computer, set-top box, television, display device, digital media player, video game console, a video streaming device (such as a content service server or a content delivery server), a broadcast receiver device, a broadcast transmitter device, or the like, may not use an operating system, or may use any kind of operating system. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Accordingly, the source device 12 and the destination device 14 may be wireless communication devices.
[0076] In some cases, the video coding system 10 illustrated in FIG. 1A is merely an example, and the techniques of the present application do not necessarily include any data communication between the encoding and decoding devices and may be applied to video coding settings (e.g., video encoding or video decoding). In other examples, the data may be retrieved from local memory, streamed over a network, or the like. The video encoding device may encode the data and store it in memory and / or the video decoding device may retrieve the data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other but merely encode data into memory and / or retrieve and decode data from memory.
[0077] For the sake of convenience of explanation, embodiments of the invention are described herein by reference to, for example, the reference software of Versatile Video Coding (VVC), which is the next-generation video coding standard developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG) for High-Efficiency Video Coding (HEVC). Those skilled in the art of this technology will understand that the embodiments of the invention are not limited to HEVC or VVC.
[0078] Encoder and encoding method FIG. 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the techniques of the present application. In the example of FIG. 2, the video encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output 272 (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 as represented in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder by a hybrid video codec.
[0079] The residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, and the mode selection unit 260 may be referred to as forming the forward signal path of the encoder 20. On the other hand, the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 may be referred to as forming the reverse signal path of the video encoder 20. The reverse signal path of the video encoder 20 corresponds to the signal path of the decoder (see the video decoder 30 in FIG. 3). The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also referred to as forming the "built-in decoder" of the video encoder 20.
[0080] Picture and picture segment (picture and block) The encoder 20 may be configured to receive a picture 17 (or picture data 17), for example, a picture of a sequence of pictures forming a video or video sequence, via the input 201. The received picture or picture data may also be the pre-processed picture 19 (or pre-processed picture data 19). For the sake of brevity, the following description refers to the picture 17. The picture 17 may also be referred to as the current picture or the picture to be coded (especially in video coding) to distinguish it from other pictures of the same video sequence, i.e., also pictures of the video sequence that includes the current picture, for example, previously encoded and / or decoded pictures.
[0081] (Digital) pictures are, or can be considered to be, two-dimensional arrays or matrices of samples having intensity values. Samples within an array may also be referred to as pixels (short for picture elements) or pels. The number of samples in the horizontal and vertical directions (or axes) of an array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are employed, i.e., a picture may represent or include three sample arrays. In the RGB format or color space, a picture comprises corresponding red, green, and blue sample arrays. However, in video coding, each pixel is typically represented in a luminance and chrominance format or color space, e.g., YCbCr, which comprises a luminance component denoted by Y (sometimes L is also used instead), and two chrominance components denoted by Cb and Cr. The luminance (or short for luma) component Y represents luminance or gray-level intensity (such as in a grayscale picture), while the two chrominance (or short for chroma) components Cb and Cr represent chrominance or color information components. Thus, a picture in the YCbCr format comprises a luminance sample array of luminance sample values (Y), and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format may be converted or transformed to the YCbCr format, and vice versa, and the process is also known as color conversion or transformation. If a picture is monochrome, the picture may comprise only a luminance sample array. Thus, a picture may be, for example, an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0082] An embodiment of the video encoder 20 may comprise a picture partitioning unit (not depicted in FIG. 2) configured to partition picture 17 into a plurality of (typically non-overlapping) picture blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB) or coding tree units (CTU) (H.265 / HEVC and VVC). The picture partitioning unit may use the same block size and corresponding grid defining the block size for all pictures of the video sequence, or may vary the block size between pictures or subsets or groups of pictures and be configured to partition each picture into corresponding blocks.
[0083] In a further embodiment, the video encoder may be configured to directly receive blocks 203 of picture 17, for example, one, some, or all of the blocks forming picture 17. Picture block 203 may also be referred to as the current picture block or the picture block to be coded.
[0084] Similar to picture 17, picture block 203 can again be or be regarded as a two-dimensional array or matrix of samples having intensity values (sample values), but of a smaller dimension than picture 17. In other words, block 203 may comprise, for example, one sample array (e.g., a luminance array in the case of a monochrome picture 17, or a luminance or chroma array in the case of a color picture), or three sample arrays (e.g., a luminance and two chroma arrays in the case of a color picture 17), or any other number and / or kind of arrays depending on the color format applied. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Thus, the block may be, for example, an M×N (M columns × N rows) array of samples, or an M×N array of transform coefficients.
[0085] An embodiment of the video encoder 20 as shown in FIG. 2 may be configured to encode picture 17 block by block. For example, encoding and prediction may be performed for each block 203.
[0086] An embodiment of the video encoder 20 as shown in FIG. 2 may be further configured to partition and / or encode a picture by using slices (also called video slices). The picture may be partitioned into one or more slices (typically non-overlapping), or encoded using the slices. Each slice may include one or more blocks (e.g., CTUs).
[0087] An embodiment of the video encoder 20 as shown in FIG. 2 may be further configured to partition and / or encode a picture by using tile groups (also called video tile groups) and / or tiles (also called video tiles). The picture may be partitioned into one or more tile groups (typically non-overlapping), or encoded using the tile groups. Each tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles. Each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTUs), e.g., complete or fragmented blocks.
[0088] Residual calculation The residual calculation unit 204 may be configured to calculate a residual block 205 (also called residual 205) based on the picture block 203 and the prediction block 265 (further details about the prediction block 265 will be provided later) by subtracting the sample values of the prediction block 265 from the sample values of the picture block 203 for each sample (pixel by pixel), so as to obtain the residual block 205 in the sample area.
[0089] Transformation The conversion processing unit 206 may be configured to apply a conversion, for example, a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain the conversion coefficients 207 in the conversion domain. The conversion coefficients 207, also referred to as conversion residual coefficients, may represent the residual block 205 in the conversion domain.
[0090] The conversion processing unit 206 may be configured to apply an integer approximation of DCT / DST, such as the conversion specified for H.265 / HEVC. Compared with the orthogonal DCT transform, such an integer approximation is typically scaled by a certain coefficient. To maintain the norm of the residual block processed by the forward and inverse transforms, an additional scaling coefficient is applied as part of the conversion process. The scaling coefficient is typically selected based on certain constraints, such as the scaling coefficient being a power of 2 for shift operations, the bit depth of the conversion coefficients, and the trade-off between accuracy and implementation cost. For example, specific scaling coefficients are specified for the inverse transform (and, for example, the corresponding inverse transform by the inverse transform processing unit 312 in the video decoder 30), and the corresponding scaling coefficients for the forward transform, for example, by the conversion processing unit 206 in the encoder 20, may be specified accordingly.
[0091] Embodiments of the video encoder 20 (each, the conversion processing unit 206) may be configured to output conversion parameters, for example, one or more types of conversions, encoded or compressed, for example, directly or via the entropy encoding unit 270, whereby, for example, the video decoder 30 may receive and use the conversion parameters for decoding.
[0092] Quantization The quantization unit 208 can be configured to obtain the quantized coefficient 209 by quantizing the transform coefficient 207, for example, by applying scalar quantization or vector quantization. The quantized coefficient 209 can also be referred to as the quantized transform coefficient 209 or the quantized residual coefficient 209.
[0093] The quantization process may reduce the bit depth associated with some or all of the conversion coefficients 207. For example, an n-bit conversion coefficient may be truncated to an m-bit conversion coefficient during quantization, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (Quantization Parameter (QP)). For example, for scalar quantization, different scalings may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to coarser quantization. The applicable quantization step size may be indicated by the quantization parameter (QP). The quantization parameter may be, for example, an index to a default set of applicable quantization step sizes. For example, a small quantization parameter may correspond to fine quantization (small quantization step size), a large quantization parameter may correspond to coarse quantization (large quantization step size), or vice versa. Quantization may include division by a quantization step size. For example, the corresponding and / or inverse dequantization by the inverse quantization unit 210 may include multiplication by the quantization step size. Some standards, such as embodiments according to HEVC, may be configured to determine the quantization step size using the quantization parameter. In general, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation that includes division. Additional scaling factors may be introduced for quantization and dequantization to restore the norm of the residual block, which may be modified due to the scaling used in the fixed-point approximation of the equations for the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of inverse transformation and dequantization may be combined. Alternatively, a customized quantization table may be used and signaled, for example, in the bitstream, from the encoder to the decoder. Quantization is a lossy operation, and the loss increases with increasing quantization step size.
[0094] Embodiments of the video encoder 20 (each, quantization unit 208) may be configured to output quantization parameters (QPs) encoded, for example, directly or via the entropy encoding unit 270, whereby, for example, the video decoder 30 may receive and apply the quantization parameters for decoding.
[0095] Inverse quantization The inverse quantization unit 210 is configured to obtain dequantized coefficients 211 by applying inverse quantization of the quantization unit 208 to the quantized coefficients, for example, based on or using the same quantization step size as the quantization unit 208, or by applying the inverse of the quantization method applied by the quantization unit 208. The dequantized coefficients 211, also referred to as dequantized residual coefficients 211, are typically not identical to the transform coefficients due to losses caused by quantization, but may correspond to the transform coefficients 207.
[0096] Inverse transform The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or inverse discrete sine transform (DST), or other inverse transform, to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as the transform block 213.
[0097] Reconstruction The reconstruction unit 214 (e.g., adder or accumulator 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 by adding, for example, the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265 on a sample-by-sample basis to obtain a reconstructed block 215 in the sample domain.
[0098] Filter processing The loop filter unit 220 (or, abbreviated as "loop filter" 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or generally, to filter the reconstructed samples to obtain filtered samples. The loop filter unit is configured to, for example, smooth pixel transitions or otherwise improve video quality. The loop filter unit 220 may comprise a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening, a smoothing filter, or a collaborative filter, or any combination thereof. Although the loop filter unit 220 is shown in FIG. 2 as an in-loop filter, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstructed block 221.
[0099] Embodiments of the video encoder 20 (each, the loop filter unit 220) may be configured to output loop filter parameters (such as sample-adaptive offset information), for example, directly or via the entropy encoding unit 270, encoded, such that, for example, the decoder 30 may receive and apply the same loop filter parameters or respective loop filters for decoding.
[0100] Decoded picture buffer The decoded picture buffer (DPB) 230 may be a memory for storing reference pictures or generally reference picture data for encoding video data by the video encoder 20. The DPB 230 may be formed by any of various memory devices, such as a dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM (registered trademark)), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may be further configured to store other previously filtered blocks, e.g., previously reconstructed and filtered blocks 221 of the same current picture or of a different picture, e.g., previously reconstructed pictures, for example, for inter prediction, to provide previously reconstructed, i.e., decoded, complete pictures (and corresponding reference blocks and samples), and / or partially reconstructed current pictures (and corresponding reference blocks and samples). For example, if the reconstructed block 215 is not filtered by the loop filter unit 220 or any other further processed version of the reconstructed block or sample, the decoded picture buffer (DPB) 230 may also be configured to store one or more unfiltered reconstructed blocks 215, or generally, unfiltered reconstructed samples.
[0101] Mode Selection (Partitioning and Prediction) The mode selection unit 260 includes a classification unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original picture data, for example, the original block 203 (the current block 203 of the current picture 17), and reconstructed picture data from the same (current) picture and / or from one or more previously decoded pictures, for example, from the decoded picture buffer 230 or other buffers (e.g., an unrepresented line buffer), such as filtered and / or unfiltered reconstructed samples or blocks. The reconstructed picture data is used as reference picture data for prediction, for example, inter prediction or intra prediction, to obtain the prediction block 265 or predictor 265.
[0102] The mode selection unit 260 may be configured to classify and determine or select a prediction mode (e.g., an intra or inter prediction mode) for the current block prediction mode (excluding classification), and generate a corresponding prediction block 265 that is used for the calculation of the residual block 205 and for the reconstruction of the reconstructed block 215.
[0103] Embodiments of the mode selection unit 260 may be configured to select segmentation and prediction modes (e.g., from those supported by or available to the mode selection unit 260) that provide the best match, or in other words the least residual (the least residual implies better compression for transmission or storage), or the least signaling overhead (the least signaling overhead implies better compression for transmission or storage), or both, or a trade-off, or a balance. The mode selection unit 260 may be configured to determine the segmentation and prediction modes based on Rate Distortion Optimization (RDO), i.e., to select the prediction mode that provides the least rate distortion. Terms such as "best", "least", "optimal", etc. in this context do not necessarily refer to an overall "best", "least", "optimal", etc., but may refer to the fulfillment of a criterion for termination or selection, such as a value above or below a threshold or other constraint, potentially leading to a "quasi-optimal selection" while reducing complexity and processing time.
[0104] In other words, the segmentation unit 262 may be configured to repeatedly use, for example, quad-tree-partitioning (QT), binary partitioning (BT), or triple-tree-partitioning (TT), or any combination thereof, to partition the block 203 into smaller block segments or sub-blocks (which again form blocks), and, for example, to perform prediction for each of the block segments or sub-blocks. Mode selection may include the selection of the tree structure of the block 203 to be segmented, and the prediction mode is applied to each of the block segments or sub-blocks.
[0105] The segmentation and prediction processing (e.g., by the segmentation unit 260) performed by the exemplary video encoder 20 will be described in more detail below.
[0106] Partition The partitioning unit 262 can partition the current block 203 into smaller partitions, for example, smaller blocks of square or rectangular size. These smaller blocks (which may also be called sub-blocks) can be further partitioned into even smaller partitions. This is also called a tree partition or a hierarchical tree partition. For example, a root block at the root tree level 0 (hierarchical level 0, depth 0) can be recursively partitioned, for example, into two or more blocks at the next lower tree level, for example, nodes at tree level 1 (hierarchical level 1, depth 1), and these blocks can again be partitioned, for example, into two or more blocks at the next lower level, for example, tree level 2 (hierarchical level 2, depth 2), etc., until the termination criterion is satisfied, for example, until the maximum tree depth or the minimum block size is reached and the partitioning is terminated. Blocks that are not further partitioned are also called leaf blocks or leaf nodes of the tree. A tree using partitioning into two partitions is called a binary tree (BT), a tree using partitioning into three partitions is called a ternary tree (TT), and a tree using partitioning into four partitions is called a quadtree (QT).
[0107] As described above, the term "block" as used herein may be a portion of a picture, particularly a square or rectangular portion. For example, referring to HEVC and VVC, a block may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB), or may correspond thereto.
[0108] For example, a coding tree unit (CTU) may be a CTB of luma samples of a picture having three sample arrays, two corresponding CTBs of chroma samples, or a CTB of samples of a picture coded using a monochrome picture or three separate color planes, and a syntax structure used to code the samples, or may comprise them. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some values of N such that the division of components into CTBs is in segments. A coding unit (CU) may be a coding block of luma samples of a picture having three sample arrays, two corresponding coding blocks of chroma samples, or a coding block of samples of a picture coded using a monochrome picture or three separate color planes, and a syntax structure used to code the samples, or may comprise them. Correspondingly, a coding block (CB) may be an M×N block of samples for some values of M and N such that the division of CTBs into coding blocks is in segments.
[0109] For example, in an embodiment according to HEVC, a coding tree unit (CTU) can be divided into coding units (CUs) by using a quadtree structure represented as a coding tree. The decision as to whether to code a picture area using (temporal) inter-picture prediction or (spatial) intra-picture prediction is made at the CU level. Each CU can be further divided into one, two, or four prediction units (PUs) according to the PU division type. Inside one PU, the same prediction process is applied, and the relevant information is transmitted to the decoder for each PU. After obtaining a residual block by applying a prediction process based on the PU division type, the CU can be divided into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU.
[0110] For example, in an embodiment according to the latest video coding standard under development, called Versatile Video Coding (VVC), a combined quadtree and binary tree (Quad-Tree and Binary Tree (QTBT)) partitioning is used, for example, to partition coding blocks. In the QTBT block structure, a CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quadtree structure. A quadtree leaf node is further partitioned by a binary tree or a ternary tree (or triple tree) structure. The partitioning tree leaf node is called a coding unit (CU), and its segmentation is used for prediction and transform processing without further partitioning. This means that the CU, PU, and TU have the same block size in the QTBT coding block structure. In parallel, multiple partitions, for example, ternary tree partitions, can be used together with the QTBT block structure.
[0111] In one example, the mode selection unit 260 of the video encoder 20 can be configured to perform any combination of the partitioning techniques described herein.
[0112] As described above, the video encoder 20 is configured to determine or select the best or optimal prediction mode from a set of prediction modes (e.g., pre-determined). The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0113] Intra prediction The set of intra prediction modes may include, for example, 35 different intra prediction modes as defined in HEVC, such as non-directional modes like DC (or average) mode and planar mode, or directional modes, or, for example, 67 different intra prediction modes as defined for VVC, such as non-directional modes like DC (or average) mode and planar mode, or directional modes.
[0114] The intra prediction unit 254 is configured to use the reconstructed samples of adjacent blocks of the same current picture to generate an intra prediction block 265 according to the intra prediction mode of the set of intra prediction modes.
[0115] The intra prediction unit 254 (or generally the mode selection unit 260) is further configured to output to the entropy encoding unit 270 the intra prediction parameters (or generally the information indicating the selected intra prediction mode for a block) in the form of a syntax element 266 for inclusion in the encoded picture data 21, whereby, for example, the video decoder 30 may receive and use the prediction parameters for decoding.
[0116] Inter prediction The set of inter prediction modes (or possible inter prediction modes) depends on the available reference pictures (i.e., previously decoded pictures that are at least partially stored in, for example, DBP 230), and other inter prediction parameters, for example, whether the entire reference picture is used to search for the best matching reference block, or only a part of the reference picture, for example, the search window area around the area of the current block, and / or, for example, whether pixel interpolation, for example, half / semi pixel and / or quarter pixel interpolation, is applied or not.
[0117] In addition to the above prediction modes, a skip mode and / or a direct mode may be applied.
[0118] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both not shown in FIG. 2). The motion estimation unit may be configured to receive or obtain, for motion estimation, a picture block 203 (the current picture block 203 of the current picture 17) and a decoded picture 231, or at least one or a plurality of previously reconstructed blocks, for example, reconstructed blocks of one or a plurality of other / different previously decoded pictures 231. For example, the video sequence may comprise the current picture and the previously decoded picture 231, or in other words, the current picture and the previously decoded picture 231 may be part of a sequence of pictures forming the video sequence, or may form them.
[0119] The encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different pictures among a plurality of other pictures, and provide a reference picture (or a reference picture index), and / or an offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block to the motion estimation unit as an inter-prediction parameter. This offset is also called a motion vector (MV).
[0120] The motion compensation unit is configured to obtain, for example, receive, an inter-prediction parameter and perform inter-prediction based on or using the inter-prediction parameter to obtain an inter-prediction block 265. The motion compensation performed by the motion compensation unit may involve fetching or generating a prediction block based on the motion / block vector determined by motion estimation, and perhaps performing interpolation to sub-pixel accuracy. The interpolation filtering process may generate additional pixel samples from known pixel samples, thus potentially increasing the number of candidate prediction blocks that can be used to code a picture block. Upon receiving a motion vector for the PU of the current picture block, the motion compensation unit may locate the prediction block indicated by the motion vector within one of the reference picture lists.
[0121] The motion compensation unit may also generate syntax elements associated with the block and the video slice for use by the video decoder 30 when decoding the picture block of the video slice. In addition to or as an alternative to the slice and its respective syntax elements, tile groups and / or tiles and their respective syntax elements may be generated or used.
[0122] Entropy coding The entropy encoding unit 270 applies, in the quantized coefficients 209, inter prediction parameters, intra prediction parameters, loop filter parameters, and / or other syntax elements, for example, an entropy encoding algorithm or method (e.g., variable length coding (VLC) method, context adaptive VLC scheme (CAVLC), arithmetic coding method, binarization, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy encoding methodology or technique), or bypass (no compression), to obtain encoded picture data 21 that can be output via output 272, for example, in the form of an encoded bitstream 21, whereby, for example, the video decoder 30 can receive and use the parameters for decoding. The encoded bitstream 21 can be transmitted to the video decoder 30 or stored in memory for later transmission or retrieval by the video decoder 30.
[0123] Other structural variations of the video encoder 20 can be used to encode the video stream. For example, a non-transform-based encoder 20 can directly quantize the residual signal for a block or frame without the transform processing unit 206. In another implementation, the encoder 20 can have a quantization unit 208 and an inverse quantization unit 210 combined in a single unit.
[0124] Decoder and Decoding Method FIG. 3 shows an example of a video decoder 30 configured to implement the technique of this present application. The video decoder 30 is configured to receive, for example, encoded picture data 21 (e.g., an encoded bitstream 21) encoded by an encoder 20 and obtain a decoded picture 331. The encoded picture data or bitstream includes information for decoding the encoded picture data, for example, data representing picture blocks of an encoded video slice (and / or a tile group or tile), and associated syntax elements.
[0125] In the example of FIG. 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or include a motion compensation unit. The video decoder 30 may, in some examples, execute a decoding path that is generally complementary to the encoding path described with respect to the video encoder 100 from FIG. 2.
[0126] As described with respect to the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 344, and the intra prediction unit 354 are also referred to as forming the "built-in decoder" of the video encoder 20. Thus, the inverse quantization unit 310 may be identical in function to the inverse quantization unit 110, the inverse transform processing unit 312 may be identical in function to the inverse transform processing unit 212, the reconstruction unit 314 may be identical in function to the reconstruction unit 214, the loop filter 320 may be identical in function to the loop filter 220, and the decoded picture buffer 330 may be identical in function to the decoded picture buffer 230. Accordingly, the description provided for each unit and function of the video 20 encoder is correspondingly applicable to each unit and function of the video decoder 30.
[0127] Entropy decoding The entropy decoding unit 304 syntax-analyzes the bitstream 21 (or generally the encoded picture data 21), and for example, performs entropy decoding on the encoded picture data 21 to obtain, for example, the quantized coefficients 309 and / or the decoded coding parameters (not shown in FIG. 3), such as inter prediction parameters (e.g., reference picture index and motion vector), intra prediction parameters (e.g., intra prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or any or all of other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or method corresponding to the encoding method as described for the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to the mode application unit 360 and other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level. In addition to or instead of slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be received and / or used.
[0128] Inverse quantization The inverse quantization unit 310 receives the quantized parameters (quantization parameter (QP)) (or generally information regarding inverse quantization) and the quantized coefficients from the encoded picture data 21 (e.g., by the entropy decoding unit 304, e.g., by syntax analysis and / or decoding), and is configured to apply inverse quantization to the decoded quantized coefficients 309 based on the quantization parameters to obtain the dequantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include the use of quantization parameters determined by the video encoder 20 for each video block in a video slice (or tile or tile group) to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied.
[0129] Inverse transformation The inverse transformation processing unit 312 receives the dequantized coefficients 311, which may also be referred to as transform coefficients 311, and is configured to apply a transformation to the dequantized coefficients 311 to obtain the reconstructed residual block 213 in the sample region. The reconstructed residual block 213 may also be referred to as the transform block 313. The transformation may be an inverse transformation, e.g., inverse DCT, inverse DST, inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 may be further configured to receive transformation parameters or corresponding information from the encoded picture data 21 (e.g., by the entropy decoding unit 304, e.g., by syntax analysis and / or decoding) to determine the transformation to be applied to the dequantized coefficients 311.
[0130] Reconstruction The reconstruction unit 314 (e.g., adder or summer 314) is configured to add the reconstructed residual block 313 to the prediction block 365, e.g., by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365, to obtain the reconstructed block 315 in the sample region.
[0131] Filter processing (Either within or after the coding loop) The loop filter unit 320, for example, is configured to filter the reconstructed block 315 to obtain a filtered block 321 in order to smooth pixel transitions or otherwise improve video quality. The loop filter unit 320 may comprise a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), sharpening, a smoothing filter, or a collaborative filter, or any combination thereof. Although the loop filter unit 320 is depicted in FIG. 3 as an in-loop filter, in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.
[0132] Decoded picture buffer The decoded video block 321 of the picture is then stored in a decoded picture buffer 330 that stores the decoded picture 331 as a reference picture for subsequent motion compensation for other pictures and / or for each output display to be output.
[0133] The decoder 30 is configured to output the decoded picture 311, for example, via output 312, for presentation or viewing by the user.
[0134] Prediction The inter prediction unit 344 may be the same as the inter prediction unit 244 (especially the motion compensation unit), and the intra prediction unit 354 may be the same as the inter prediction unit 254 in function. Based on each piece of information received from the segmentation or classification determination and prediction, or the encoded picture data 21 (e.g., by the entropy decoding unit 304, e.g., by syntax analysis and / or decoding), the segmentation or classification determination and prediction are performed. The mode application unit 360 may be configured to perform prediction (intra or inter prediction) for each block based on the (filtered or unfiltered) reconstructed picture, block, or each sample to obtain the prediction block 365.
[0135] When a video slice is coded as an intra-coded (I) slice, the intra prediction unit 354 of the mode application unit 360 is configured to generate a prediction block 365 for a picture block of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the mode application unit 360 is configured to create a prediction block 365 for a video block of the current video slice based on a motion vector and other syntax elements received from the entropy decoding unit 304. For inter prediction, the prediction block may be created from one of the reference pictures in one of the reference picture lists. The video decoder 30 may configure the reference frame lists, list 0 and list 1, using default construction techniques based on the reference pictures stored in the DPB 330. In addition to or instead of slices (e.g., video slices), the same or similar may apply to embodiments using tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles), e.g., the video may be coded using I, P, or B tile groups and / or tiles.
[0136] The mode application unit 360 is configured to determine prediction information for video blocks of a current video slice by parsing motion vectors or related information and other syntax elements, and use the prediction information to create a prediction block for the currently decoded video block. For example, the mode application unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra or inter prediction) used to code the video blocks of the video slice, the inter prediction slice type (e.g., B slice, P slice, or GPB slice), the configuration information for one or more of the reference picture lists for the slice, the motion vector for each inter-encoded video block of the slice, the inter prediction status for each inter-coded video block of the slice, and other information for decoding the video blocks within the current video slice. In addition to or instead of a slice (e.g., a video slice), the same or similar may apply to embodiments that use a tile group (e.g., a video tile group) and / or a tile (e.g., a video tile), e.g., the video may be coded using I, P, or B tile groups and / or tiles.
[0137] An embodiment of the video decoder 30 as represented in FIG. 3 may be configured to partition and / or decode a picture by using slices (also referred to as video slices), where the picture may be partitioned into or decoded using one or more (typically non-overlapping) slices, and each slice may comprise one or more blocks (e.g., CTUs).
[0138] An embodiment of video decoder 30 as shown in FIG. 3 may be configured to partition and / or decode a picture by using tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles), where the picture may be partitioned into one or more (typically non-overlapping) tile groups or decoded using them, and each tile group may comprise, for example, one or more blocks (e.g., CTUs) or one or more tiles, and each tile may be, for example, rectangular in shape and may comprise one or more blocks (e.g., CTUs), e.g., complete or fragmentary blocks.
[0139] Other variations of video decoder 30 can be used to decode the encoded picture data 21. For example, decoder 30 can produce an output video stream without loop filter processing unit 320. For example, a non-transform-based decoder 30 can directly inverse quantize the residual signal for a block or frame without inverse transform processing unit 312. In another implementation, video decoder 30 can have an inverse quantization unit 310 and an inverse transform processing unit 312 combined in a single unit.
[0140] It should be understood that in encoder 20 and decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filter processing, motion vector derivation, or loop filter processing, further operations such as clip or shift may be performed on the processing result of interpolation filter processing, motion vector derivation, or loop filter processing.
[0141] It should be noted that further operations can be applied to the derived motion vectors of the current block (including, but not limited to, the control point motion vectors in affine mode, affine, planar, sub-block motion vectors in ATMVP mode, temporal motion vectors, etc.). For example, the value of the motion vector is constrained to a predefined range according to its representation bits. If the representation bits of the motion vector are bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" means exponentiation. For example, if bitDepth is set equal to 16, the range is -32768 to 32767, and if bitDepth is set equal to 18, the range is -131072 to 131071. For example, the value of the derived motion vector (e.g., the MV of four 4×4 sub-blocks within one 8×8 block) is constrained such that the maximum difference between the integer parts of the four 4×4 sub-block MVs is not greater than N pixels, where N is not greater than 1 pixel. Here, two methods are provided for constraining the motion vector according to bitDepth.
[0142] Method 1: Remove the overflow MSB (most significant bit) by a flow operation. ux = ( mvx + 2 bitDepth ) % 2 bitDepth (1) mvx = ( ux >= 2 bitDepth-1 )? ( ux - 2 bitDepth ) : ux (2) uy = ( mvy + 2 bitDepth ) % 2 bitDepth (3) mvy = ( uy >= 2 bitDepth-1 )? ( uy - 2 bitDepth ) : uy (4) Here, mvx is the horizontal component of the motion vector of the image block or sub-block, mvy is the vertical component of the motion vector of the image block or sub-block, and ux and uy represent intermediate values.
[0143] For example, if the value of mvx is -32769, after applying equations (1) and (2), the resulting value is 32767. In a computer system, decimal numbers are stored as two's complements. The two's complement of -32769 is 1,0111,1111,1111,1111 (17 bits), and then the MSB is discarded, so the resulting two's complement is the same as the output by applying equations (1) and (2), which is 0111,1111,1111,1111 (decimal 32767). ux = (mvpx + mvdx + 2 bitDepth ) % 2 bitDepth (5) mvx = (ux >= 2 bitDepth-1 )? (ux - 2 bitDepth ) : ux (6) uy = (mvpy + mvdy + 2 bitDepth ) % 2 bitDepth (7) mvy = (uy >= 2 bitDepth-1 )? (uy - 2 bitDepth ) : uy (8)
[0144] As represented by equations (5) to (8), the operation can be applied between the sum of mvp and mvd.
[0145] Method 2: Remove the overflow MSB by clipping the value. vx = Clip3(-2 bitDepth-1 , 2 bitDepth-1 -1, vx) vy = Clip3(-2 bitDepth-1 , 2 bitDepth-1 -1, vy) Here, vx is the horizontal component of the motion vector of an image block or sub-block, vy is the vertical component of the motion vector of an image block or sub-block, x, y, and z respectively correspond to the three input values of the MV clipping process, and the definition of the function Clip3 is as follows.
[0146] [Number]
[0147] Figure 4 is a schematic diagram of a video coding device 400 according to an embodiment of the disclosure. The video coding device 400 is suitable for implementing the disclosed embodiments as described herein. In one embodiment, the video coding device 400 may be a decoder such as the video decoder 30 of FIG. 1A, or an encoder such as the video encoder 20 of FIG. 1A.
[0148] The video coding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an outlet port 450 (or output port 450) for transmitting data, and a memory 460 for storing data. The video coding device 400 may also include optoelectrical (optical-to-electrical (OE)) components and electro-optical (electrical-to-optical (EO)) components coupled to the inlet port 410, the receiver unit 420, the transmitter unit 440, and the outlet port 450 for an outlet or inlet of optical or electrical signals.
[0149] Processor 430 is implemented by hardware and software. Processor 430 can be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. Processor 430 communicates with an input port 410, a receiver unit 420, a transmitter unit 440, an output port 450, and a memory 460. Processor 430 includes a coding module 470. Coding module 470 implements the disclosed embodiments described above. For example, coding module 470 implements, processes, prepares, or provides various coding operations. Thus, the inclusion of coding module 470 provides a significant improvement in the functionality of video coding device 400 and results in the conversion of video coding device 400 to different states. Alternatively, coding module 470 is implemented as instructions stored in memory 460 and executed by processor 430.
[0150] Memory 460 may comprise one or more disks, tape drives, and solid state drives and may be used as an overflow data storage device to store programs when such a program is selected for execution and to store instructions and data read during program execution. Memory 460 may be, for example, volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0151] FIG. 5 is a simplified block diagram of an apparatus 500 that can be used as either or both of the source device 12 and the destination device 14 from FIG. 1 according to an exemplary embodiment.
[0152] The processor 502 within the device 500 can be a central processing unit. Alternatively, the processor 502 can be any other type of device or devices capable of manipulating or processing information, existing or to be developed in the future. The disclosed implementation can be carried out using a single processor, e.g., the processor 502, as represented, but advantages in terms of speed and efficiency can be achieved using more than one processor.
[0153] The memory 504 within the device 500 can be, in one implementation, a read-only memory (ROM) device or a random access memory (RAM) device in one implementation. Any other suitable type of storage device can be used as the memory 504. The memory 504 can include code and data 506 accessed by the processor 502 using the bus 512. The memory 504 can further include an operating system 508 and application programs 510, and the application programs 510 include at least one program that enables the processor 502 to execute the methods described herein. For example, the application programs 510 can include applications 1 through N, and applications 1 through N further include video coding applications that execute the methods described herein.
[0154] The device 500 can also include one or more output devices such as a display 518. The display 518 can be, in one example, a touch-sensitive display that combines a display with a touch-sensitive element operable to sense touch input. The display 518 can be coupled to the processor 502 via the bus 512.
[0155] Although depicted here as a single bus, the bus 512 of the apparatus 500 can consist of multiple buses. Further, the secondary storage device 514 can be directly coupled to other components of the apparatus 500 or can be accessed via a network and can comprise a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, the apparatus 500 can be implemented in a wide variety of configurations.
[0156] As described in the document J.M. Boyce, "Weighted prediction in the H.264 / MPEG AVC video coding standard", IEEE International Symposium on Circuits and Systems, May 2004, Canada, pages 789 - 792, weighted prediction (WP) is a particularly useful tool for coding fades. A weighted prediction (WP) tool is adopted in the main and extended profiles of the H.264 video coding standard to improve coding efficiency by applying multiplicative weight factors and additive offsets to motion - compensated prediction to form weighted prediction. In the explicit mode, the weight factors and offsets may be coded in the slice header for each acceptable reference picture index. In the implicit mode, the weight factors are not coded but are derived based on the relative picture order count (POC) distance of two reference pictures. Experimental results are provided that measure the improvement in coding efficiency using WP. When coding a fading sequence, a bit - rate reduction of up to 67% was achieved.
[0157] When applied to a single prediction as in the P picture, WP is similar to the leak prediction previously proposed for error resilience. The leak prediction is a special case of WP with a scaling coefficient limited to the range 0 ≦ α ≦ 1. H.264 WP allows negative scaling coefficients and scaling coefficients greater than 1. For efficient compression of the covered and uncovered regions, the weighting coefficients are applied pixel by pixel using the coded label field. An important difference of the H.264 WP tool from previous proposals with weighted prediction for compression efficiency is the association of the reference picture index with the weighting coefficient parameter, which enables efficient signaling of these parameters in a multiple reference picture environment. As written in the document R. Zhang and G. Cote, "accurate parameter estimation and efficient fade detection for weighted prediction in H.264 video compression", the 15th IEEE International Conference on Image Processing, October 2008, San Diego, California, USA, pages 2836 - 2839, the procedure for applying WP in a real - time encoding system can be formalized as a sequence of steps represented in Figure 6. First, some statistical values 611 are generated through video analysis 610. The statistical values 611 within a small window from several previous pictures to the current picture are then used to detect fades. Each picture is assigned a state value 631 indicating whether the picture is in the NORMAL state or the FADE state. Such state values are saved for each picture. When encoding a picture, if there is a FADE state in either the current picture or one of its reference pictures, WP is used for this current reference pair, and the statistical values of the current picture and the corresponding reference picture are processed in step 650 to estimate the WP parameters. These parameters are then passed to the encoding engine 660.Otherwise, normal encoding is performed.
[0158] As described in document A. Leontaris and A. M. Tourapis, "Weighted prediction methods for improved motion compensation", 16th IEEE International Conference on Image Processing (ICIP), November 2009, Cairo, Egypt, pages 1029 - 1032, macroblocks in H.264 are divided into macroblock partitions. For each macroblock partition, a reference is selected from each of the available reference lists (often denoted as RefPicList in the specification), list 0 for P - or B - coded slices, or reference list 1 for B - coded slices. The references used may differ for each partition. Using these references, prediction blocks are generated for each list, i.e., P for single - list prediction and P and P1 for dual - prediction O and P1 are generated using motion information with optional sub - pixel accuracy. The prediction blocks may be further processed depending on the availability of weighted prediction for the current slice. For P - slices, the WP parameter is transmitted in the slice header. For B - slices, there are two options. In explicit WP, the parameter is transmitted within the slice header, and in implicit WP, the parameter is derived based on the picture - order count (POC) number signaled within the slice header. In this document, we focus only on explicit WP and how this method can be used to improve motion - compensation performance. Note that in HEVC and VVC, PB is used in the same way as macroblock partitioning in AVC.
[0159] For a single list explicit WP in a P slice or a B slice, the prediction block is derived from a single reference. Let p denote the sample value within prediction block P. If weighted prediction is not used, the final inter-prediction sample is f = p. Otherwise, the prediction sample is
[0160]
Number
[0161] as follows. Terms w x and o x are the WP gain and offset parameters for reference list x. Term logWD is transmitted in the bitstream and controls the mathematical precision of the weighted prediction process. For logWD ≥ 1, the above expression rounds up the decimal part. Similarly, for dual prediction, two prediction blocks, one for each reference list, are considered. Let p0 and p1 denote the samples within each of the two prediction blocks P0 and P1. If weighted prediction is not used, the prediction is f=(p0 + p1 + 1) >> 1 executed as. For weighted dual prediction, the prediction is f=((p0 × w0 + p1 × w1 + 2 logWD )>>(logWD + 1))+((o0 + o1 + 1) >> 1) executed as. It is worth noting that weighted prediction can compensate for illumination changes such as fade-in, fade-out, or cross-fade.
[0162] At high levels in VVC, weighted prediction is signaled within the SPS, PPS, and slice headers. In the SPS, the following syntax elements are used for that purpose. An sps_weighted_pred_flag equal to 1 specifies that weighted prediction may be applied to a P slice that references the SPS. An sps_weighted_pred_flag equal to 0 specifies that weighted prediction is not applied to a P slice that references the SPS. An sps_weighted_bipred_flag equal to 1 specifies that explicit weighted prediction may be applied to a B slice that references the SPS. An sps_weighted_bipred_flag equal to 0 specifies that explicit weighted prediction is not applied to a B slice that references the SPS.
[0163] In the PPS, the following syntax elements are used for that purpose. A pps_weighted_pred_flag equal to 0 specifies that weighted prediction is not applied to a P slice that references the PPS. A pps_weighted_pred_flag equal to 1 specifies that weighted prediction is applied to a P slice that references the PPS. When sps_weighted_pred_flag is equal to 0, the value of pps_weighted_pred_flag shall be equal to 0. A pps_weighted_bipred_flag equal to 0 specifies that explicit weighted prediction is not applied to a B slice that references the PPS. A pps_weighted_bipred_flag equal to 1 specifies that explicit weighted prediction is applied to a B slice that references the PPS. When sps_weighted_bipred_flag is equal to 0, the value of pps_weighted_bipred_flag shall be equal to 0.
[0164] Within the slice header, the weighted prediction parameters are structured as in Table 1 and signaled as pred_weight_table( ) containing the following elements.
[0165] luma_log2_weight_denom is the base-2 logarithm of the denominator for all luma weighting factors. Assume that the value of luma_log2_weight_denom is within the range including both ends from 0 to 7.
[0166] delta_chroma_log2_weight_denom is the difference of the base-2 logarithm of the denominator for all chroma weighting factors. When delta_chroma_log2_weight_denom does not exist, it is presumed to be equal to 0.
[0167] The variable ChromaLog2WeightDenom is derived to be equal to luma_log2_weight_denom + delta_chroma_log2_weight_denom, and assume that its value is within the range including both ends from 0 to 7.
[0168] luma_weight_l0_flag[ i ] equal to 1 specifies that there are weighting factors for the luma component of list 0 prediction using RefPicList
[0000] [ i ]. luma_weight_l0_flag[ i ] equal to 0 specifies that these weighting factors do not exist.
[0169] chroma_weight_l0_flag[ i ] equal to 1 specifies that there are weighting factors for the chroma prediction value of list 0 prediction using RefPicList
[0000] [ i ]. chroma_weight_l0_flag[ i ] equal to 0 specifies that these weighting factors do not exist. When chroma_weight_l0_flag[ i ] does not exist, it is presumed to be equal to 0.
[0170] delta_luma_weight_l0[ i ] is the difference of the weighting factors applied to the luma prediction value for list 0 prediction using RefPicList
[0000] [ i ].
[0171] The variable LumaWeightL0[i] is derived to be equal to (1 << luma_log2_weight_denom) + delta_luma_weight_l0[i]. When luma_weight_l0_flag[i] is equal to 1, the value of delta_luma_weight_l0[i] shall be within the range including both ends from -128 to 127. When luma_weight_l0_flag[i] is equal to 0, LumaWeightL0[i] is presumed to be equal to 2 luma_log2_weight_denom is presumed to be equal to
[0172] luma_offset_l0[i] is an additive offset applied to the luma prediction value for list 0 prediction using RefPicList
[0000] [i]. The value of luma_offset_l0[i] shall be within the range including both ends from -128 to 127. When luma_weight_l0_flag[i] is equal to 0, luma_offset_l0[i] is presumed to be equal to 0.
[0173] delta_chroma_weight_l0[i][j] is the difference of the weighting coefficients applied to the chroma prediction value for list 0 prediction using RefPicList
[0000] [i] having j equal to 0 for Cb and j equal to 1 for Cr.
[0174] The variable ChromaWeightL0[i][j] is derived to be equal to (1 << ChromaLog2WeightDenom) + delta_chroma_weight_l0[i][j]. When chroma_weight_l0_flag[i] is equal to 1, the value of delta_chroma_weight_l0[i][j] shall be within the range including both ends from -128 to 127. When chroma_weight_l0_flag[i] is equal to 0, ChromaWeightL0[i][j] is presumed to be equal to 2ChromaLog2WeightDenom is presumed to be equal to.
[0175] delta_chroma_offset_l0[ i ][ j ] is the difference of the additive offset applied to the chroma prediction value for list 0 prediction using RefPicList
[0000] [ i ] having j equal to 0 for Cb and j equal to 1 for Cr.
[0176] The variable ChromaOffsetL0[ i ][ j ] is derived as follows. ChromaOffsetL0[ i ][ j ] = Clip3( -128, 127, (128 + delta_chroma_offset_l0[ i ][ j ] - ( (128 * ChromaWeightL0[ i ][ j ] ) >> ChromaLog2WeightDenom ) ) )
[0177] The value of delta_chroma_offset_l0[ i ][ j ] is assumed to be within the range including both ends from -4 * 128 to 4 * 127. When chroma_weight_l0_flag[ i ] is equal to 0, ChromaOffsetL0[ i ][ j ] is presumed to be equal to 0.
[0178] luma_weight_l1_flag[ i ], chroma_weight_l1_flag[ i ], delta_luma_weight_l1[ i ], luma_offset_l1[ i ], delta_chroma_weight_l1[ i ][ j ], and delta_chroma_offset_l1[ i ][ j ] have the same semantics as luma_weight_l0_flag[ i ], chroma_weight_l0_flag[ i ], delta_luma_weight_l0[ i ], luma_offset_l0[ i ], delta_chroma_weight_l0[ i ][ j ], and delta_chroma_offset_l0[ i ][ j ] respectively, where l0, L0, list 0, and List0 are replaced by l1, L1, list 1, and List1 respectively.
[0179] The variable sumWeightL0Flags is derived to be equal to the sum of luma_weight_l0_flag[ i ] + 2 * chroma_weight_l0_flag[ i ] for i = 0..NumRefIdxActive
[0000] - 1.
[0180] When slice_type is equal to B, the variable sumWeightL1Flags is derived to be equal to the sum of luma_weight_l1_flag[ i ] + 2 * chroma_weight_l1_flag[ i ] for i = 0..NumRefIdxActive
[0001] - 1.
[0181] It is a requirement for bitstream compliance that sumWeightL0Flags be less than or equal to 24 when slice_type is equal to P, and that the sum of sumWeightL0Flags and sumWeightL1Flags be less than or equal to 24 when slice_type is equal to B.
[0182]
Table 1A
Table 1B
[0183] In the submission JVET-O0244 (V. Seregin et al., "AHG17: On zero delta POC in reference picture structure", the 15th JVET meeting, Ytterboholm, Sweden), it was pointed out that in the current VVC specification draft, the reference picture is signaled within the reference picture structure (RPS), and abs_delta_poc_st represents the delta POC value that can be equal to 0. The RPS can be signaled within the SPS and slice header. This function is required to signal different weights for the same reference picture and is potentially required if layered scalability is supported where the same POC value is used across layers within an access unit. It is stated therein that it is not necessary to repeat the reference picture when weighted prediction is not enabled. In particular, this submission proposes not to allow a zero delta POC value when weighted prediction is not enabled.
[0184]
Table 2A
Table 2B
Table 2C
Table 2D
Table 2E
[0185]
Table 3
[0186] The ref_pic_list_struct( listIdx, rplsIdx ) syntax structure may be present in the SPS or in the slice header. Depending on whether the syntax structure is included in the slice header or in the SPS, the following applies. - If present in the slice header, the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure specifies the reference picture list listIdx of the current picture (the picture containing that slice). - Otherwise (if present in the SPS), the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure specifies a candidate for the reference picture list listIdx, and the term "current picture" in the semantics specified in the rest of this section refers to each picture that (1) has one or more slices including ref_pic_list_idx[ listIdx ] equal to the index into the list of ref_pic_list_struct( listIdx, rplsIdx ) syntax structures included in the SPS, and (2) is in the CVS that references the SPS.
[0187] num_ref_entries[ listIdx ][ rplsIdx ] specifies the number of entries in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure. The value of num_ref_entries[ listIdx ][ rplsIdx ] shall be in the range from 0 to sps_max_dec_pic_buffering_minus1 + 14, inclusive.
[0188] ltrp_in_slice_header_flag[listIdx][rplsIdx] equal to 0 specifies that the POC LSB of the LTRP entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure exists in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure. ltrp_in_slice_header_flag[listIdx][rplsIdx] equal to 1 specifies that the POC LSB of the LTRP entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure does not exist in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure.
[0189] inter_layer_ref_pic_flag[listIdx][rplsIdx][i] equal to 1 specifies that the i-th entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure is an ILRP entry. inter_layer_ref_pic_flag[listIdx][rplsIdx][i] equal to 0 specifies that the i-th entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure is not an ILRP entry. When not present, the value of inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is assumed to be equal to 0.
[0190] st_ref_pic_flag[listIdx][rplsIdx][i] equal to 1 specifies that the i-th entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure is a STRP entry. st_ref_pic_flag[listIdx][rplsIdx][i] equal to 0 specifies that the i-th entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure is an LTRP entry. When inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is equal to 0 and st_ref_pic_flag[listIdx][rplsIdx][i] does not exist, the value of st_ref_pic_flag[listIdx][rplsIdx][i] is inferred to be equal to 1.
[0191] The variable NumLtrpEntries[listIdx][rplsIdx] is derived as follows. for( i = 0, NumLtrpEntries[listIdx][rplsIdx] = 0; i < num_ref_entries[listIdx][rplsIdx]; i++ ) if(!inter_layer_ref_pic_flag[listIdx][rplsIdx][i] &&!st_ref_pic_flag[listIdx][rplsIdx][i] ) NumLtrpEntries[listIdx][rplsIdx]++
[0192] abs_delta_poc_st[listIdx][rplsIdx][i] specifies the value of the variable AbsDeltaPocSt[listIdx][rplsIdx][i] as follows. if( sps_weighted_pred_flag || sps_weighted_bipred_flag ) AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] = abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] else AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] = abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] + 1
[0193] It is assumed that the value of abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] is within the range including both ends from 0 to 215 - 1.
[0194] A value of strp_entry_sign_flag[ listIdx ][ rplsIdx ][ i ] equal to 1 specifies that the i-th entry in the syntax structure ref_pic_list_struct( listIdx, rplsIdx ) has a value of 0 or greater. A value of strp_entry_sign_flag[ listIdx ][ rplsIdx ][ i ] equal to 0 specifies that the i-th entry in the syntax structure ref_pic_list_struct( listIdx, rplsIdx ) has a value less than 0. When it does not exist, the value of strp_entry_sign_flag[ listIdx ][ rplsIdx ][ i ] is assumed to be equal to 1.
[0195] The list DeltaPocValSt[ listIdx ][ rplsIdx ] is derived as follows. for( i = 0; i < num_ref_entries[ listIdx ][ rplsIdx ]; i++ ) if( !inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] && st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] ) DeltaPocValSt[ listIdx ][ rplsIdx ][ i ] = ( strp_entry_sign_flag[ listIdx ][ rplsIdx ][ i ] ) ? AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] : 0 - AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ]
[0196] rpls_poc_lsb_lt[ listIdx ][ rplsIdx ][ i ] specifies the value of the picture order count of the picture referred to by the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure, modulo MaxPicOrderCntLsb. The length of the rpls_poc_lsb_lt[ listIdx ][ rplsIdx ][ i ] syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits.
[0197] ilrp_idc[ listIdx ][ rplsIdx ][ i ] specifies the index of the ILRP of the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure to the list of directly dependent layers. The value of ilrp_idc[ listIdx ][ rplsIdx ][ i ] shall be in the range from 0 to GeneralLayerIdx[ nuh_layer_id ] - 1, inclusive.
[0198] The present invention is a method for joint signaling of high-level syntax (HLS) weighted prediction parameters and a reference picture list, where the reference picture list may comprise reference pictures having the same picture order count (POC) value. These reference pictures correspond to the same original picture to be coded, but are coded using different parameters, for example, when different parameters of weighted prediction are used. The signaling of the weighted prediction flag may depend on whether the reference list contains such an entry.
[0199] In one embodiment of the invention, when the weighted prediction flag is equal to 1, the reference picture list is restricted to non-zero values. However, in current state-of-the-art video coding, the weighted prediction parameters are signaled after the reference picture list signaling. The following table proposes reordering these syntax elements and restricting the binarization of the delta POC syntax element based on the value of the weighted prediction flag.
[0200]
Table 4A
Table 4B
[0201] And the value of the delta POC (variable AbsDeltaPocSt) is conditionally restored at the decoder side as follows.
[0202] abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] specifies the value of the variable AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] as follows. if( sps_weighted_pred_flag || sps_weighted_bipred_flag ) AbsDeltaPocSt[listIdx][rplsIdx][i] = abs_delta_poc_st[listIdx][rplsIdx][i] else AbsDeltaPocSt[listIdx][rplsIdx][i] = abs_delta_poc_st[listIdx][rplsIdx][i] + 1
[0203] The flowchart in FIG. 7 illustrates the method described above. In step 701, weighted prediction parameters (specifically, sps_weighted_pred_flag and sps_weighted_bipred_flag) are signaled. Depending on their values, the signaling 702 of the reference picture list is performed differently. In particular, when sps_weighted_pred_flag or sps_weighted_bipred_flag is true, AbsDeltaPocSt is allowed to have a value of 0. Otherwise, AbsDeltaPocSt is restored from the bitstream using the incremented value of abs_delta_poc_st that does not allow an AbsDeltaPocSt value of 0.
[0204] In another disclosed embodiment, the weighted prediction flags sps_weighted_pred_flag and sps_weighted_bipred_flag are signaled only if at least one reference picture list ref_pic_list_struct has at least one AbsDeltaPocSt value equal to 0.
[0205]
Table 5A
Table 5B
[0206]
Table 6
[0207] Figure 8 illustrates the method disclosed in this embodiment. According to the coding order defined in Table 5, the reference picture list 801 is signaled before the signaling of the weighted prediction parameter 803. The weighted prediction parameter 803 is signaled only when the reference picture list includes at least one element having AbsDeltaPocSt equal to 0. This check is performed in step 802 by a variable RestrictWPFlag that is initialized to false and set to true when a 0-valued AbsDeltaPocSt appears during the check of each element of the reference picture list.
[0208] The following is an explanation of the encoding method, the decoding method as represented in the above-described embodiment, and the application of the system using them.
[0209] Figure 9 is a block diagram showing a content supply system 3100 for realizing a content delivery service. This content supply system 3100 includes a capture device 3102 and a terminal device 3106, and optionally includes a display 3126. The capture device 3102 communicates with the terminal device 3106 over a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination of these types, or the like.
[0210] The capture device 3102 can generate data and encode the data by the encoding method as represented in the above embodiments. Alternatively, the capture device 3102 may deliver the data to a streaming server (not shown in the figure), and the server encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or a tablet, a computer or a laptop, a video conferencing system, a PDA, a vehicle-mounted device, or any combination thereof, or the like. For example, the capture device 3102 may include the source device 12 as described above. When the data includes video, the video encoder 20 included in the capture device 3102 can actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 can actually perform audio encoding processing. For some practical scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. For other practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.
[0211] In the content supply system 3100, the terminal device 310 receives and plays back the encoded data. The terminal device 3106 can be a smartphone or tablet 3108, computer or laptop 3110, network video recorder (NVR) / digital video recorder (DVR) 3112, TV 3114, set top box (STB) 3116, video conferencing system 3118, video surveillance system 3120, personal digital assistant (PDA) 3122, vehicle-mounted device 3124, or a combination thereof, or a device having data reception and restoration capabilities such as the like. For example, the terminal device 3106 may include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.
[0212] For the terminal device having the display, such as the smartphone or tablet 3108, computer or laptop 3110, network video recorder (NVR) / digital video recorder (DVR) 3112, TV 3114, personal digital assistant (PDA) 3122, or vehicle-mounted device 3124, the terminal device can supply the decoded data to the display. For the terminal device not equipped with a display, such as the STB 3116, video conferencing system 3118, or video surveillance system 3120, an external display 3126 is contacted there to receive and display the decoded data.
[0213] When each device in this system performs encoding or decoding, a picture encoding device or a picture decoding device can be used as represented in the embodiments described above.
[0214] FIG. 10 is a diagram showing the structure of an example of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol progress unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, the Real Time Streaming Protocol (RTSP), the Hyper Text Transfer Protocol (HTTP), the HTTP Live Streaming protocol (HLS), MPEG-DASH, the Real-time Transport protocol (RTP), the Real Time Messaging Protocol (RTMP), or any combination of these types, or the like.
[0215] After the protocol progress unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is transmitted to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0216] Via inverse multiplexing processing, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. A video decoder 3206 including a video decoder 30 as described in the above-described embodiments decodes the video ES by the decoding method represented in the above-described embodiments to generate a video frame, and supplies this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate an audio frame, and supplies this data to the synchronization unit 3212. Alternatively, the video frame may be stored in a buffer (not shown in FIG. 10) before supplying it to the synchronization unit 3212. Similarly, the audio frame may be stored in a buffer (not shown in FIG. 10) before supplying it to the synchronization unit 3212.
[0217] The synchronization unit 3212 synchronizes the video frame and the audio frame, and supplies the video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. The information may be coded in the syntax using time stamps related to the presentation of the coded audio and visual data, and time stamps related to the delivery of the data stream itself.
[0218] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video frame and the audio frame, and supplies the video / audio / subtitle to the video / audio / subtitle display 3216.
[0219] The present invention is not limited to the system described above, and any of the picture encoding device or the picture decoding device in the above-described embodiments can be incorporated into other systems, for example, an automotive system.
[0220] FIG. 11 illustrates an encoding method according to a first aspect of the present disclosure. The encoding method according to the first aspect is as follows: 1101. Determine the syntax elements to be coded, where the syntax elements include a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter; 1102. Code at least one HLS weighted prediction parameter; 1103. Code the reference picture list structure following the coding of at least one HLS weighted prediction parameter. The method comprises the steps.
[0221] FIG. 12 illustrates an encoding method according to a second aspect of the present disclosure. The encoding method according to the second aspect is as follows: 1201. Determine the syntax elements to be coded, where the syntax elements include a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter, and the reference picture list derived from the reference picture list structure comprises reference pictures having the same picture order count (POC) parameter; 1202. Code the syntax elements determined in coding order with a restriction on the binarization of the syntax elements having a later position in the coding order, where when at least one HLS weighted prediction parameter is coded after the reference picture list structure in the coding order, the restriction on the binarization of the syntax elements comprises coding at least one HLS weighted prediction parameter only when the reference picture list has at least one element having a delta POC value equal to 0. The method comprises the steps.
[0222] FIG. 13 illustrates a decoding method according to a third aspect of the present disclosure. The decoding method according to the third aspect is as follows: 1301. Receive a bitstream; 1302. Entropy-decode a bitstream to obtain syntax elements, where the syntax elements include a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter. Among the syntax elements, at least one HLS weighted prediction parameter is entropy-decoded before the reference picture list structure. 1303. Execute prediction based on the obtained syntax elements to obtain a predicted block. 1304. Reconstruct a reconstructed block based on the predicted block. 1305. Obtain a decoded picture based on the reconstructed block. Comprising the above.
[0223] FIG. 14 illustrates a decoding method by a decoder according to a fourth aspect of the present disclosure. The decoding method according to the fourth aspect is as follows: 1401. Receive a bitstream. 1402. Entropy-decode the bitstream to obtain syntax elements, where the syntax elements include a reference picture list structure and a preset flag. The value of the preset flag indicates whether the syntax elements include at least one high-level syntax (HLS) weighted prediction parameter. 1403. Execute prediction based on the obtained syntax elements to obtain a predicted block. 1404. Reconstruct a reconstructed block based on the predicted block. 1405. Obtain a decoded picture based on the reconstructed block. Comprising the above.
[0224] FIG. 15 illustrates a decoder according to an eighth aspect of the present disclosure. The decoder 1500 includes one or more processors 1501 and a non-transitory computer-readable storage medium 1502 coupled to the one or more processors 1502 and storing programming for execution by the processor 1501. When executed by the processor 1501, the programming configures the decoder 1500 to execute a method according to any one of the third aspect, the first implementation form of the third aspect, the fourth aspect, or the first to fourth implementation forms of the fourth aspect.
[0225] FIG. 16 illustrates a decoder according to a ninth aspect of the present disclosure. The decoder 1600 includes receiving means 1601 for receiving a bitstream, entropy decoding means 1602 for entropy decoding the bitstream to obtain syntax elements, where the syntax elements include a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter, and among the syntax elements, at least one HLS weighted prediction parameter is entropy decoded before the reference picture list structure, entropy decoding means 1602, prediction means 1603 for performing prediction based on the obtained syntax elements to obtain a predicted block, reconstruction means 1604 for reconstructing a reconstructed block based on the predicted block, and acquisition means 1605 for obtaining a decoded picture based on the reconstructed block.
[0226] FIG. 17 illustrates an encoder according to a tenth aspect of the present disclosure. The encoder 1700 includes one or more processors 1701 and a non-transitory computer-readable storage medium 1702 coupled to the processor 1701 and storing programming for execution by the processor 1701. When executed by the processor 1701, the programming configures the encoder 1700 to execute a method according to any one of the first aspect, any one of the first to seventh implementation forms of the first aspect, or the second aspect.
[0227] FIG. 18 illustrates an encoder according to an eleventh aspect of the present disclosure. The encoder 1800 includes a determining means 1801 for determining a syntax element to be coded, the syntax element including a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter, and a coding means 1802 for coding at least one HLS weighted prediction parameter and for coding a reference picture list structure following the coding of at least one HLS weighted prediction parameter.
[0228] The present disclosure provides the following further exemplary embodiments.
[0229] 1. Exemplary embodiment: A method of joint signaling of a high-level syntax (HLS) weighted prediction parameter and a reference picture list, the reference picture list comprising reference pictures having the same picture order count (POC) parameter, the method comprising: determining a syntax element to be signaled, the syntax element including a reference picture list and at least one HLS weighted prediction parameter; signaling the syntax element determined in coding order with a restriction on the binarization of syntax elements having a later position in the coding order.
[0230] 2. Exemplary embodiment: The method of exemplary embodiment 1, wherein at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted single prediction.
[0231] 3. Exemplary embodiment: The method of exemplary embodiment 1 or 2, wherein at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted bi-prediction.
[0232] 4. Exemplary Embodiment: When at least one HLS weighted prediction parameter is signaled before the reference picture list in coding order, the restriction on the binarization of the syntax element is When the sequence parameter set flag for weighted single prediction is set to 0, signaling a modified delta POC value for an element of the reference picture list, wherein the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, the method of Exemplary Embodiment 2.
[0233] 5. Exemplary Embodiment: When at least one HLS weighted prediction parameter is signaled before the reference picture list in coding order, the restriction on the binarization of the syntax element is When at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted bi-prediction is set to 0, signaling a modified delta POC value for an element of the reference picture list, wherein the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, said signaling, or When at least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted bi-prediction and a sequence parameter set flag for weighted single prediction, and at least one of the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted single prediction is set to 0, signaling a modified delta POC value for an element of the reference picture list, wherein the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, said signaling, or At least one HLS weighted prediction parameter includes a sequence parameter set flag for weighted bi-prediction and a sequence parameter set flag for weighted uni-prediction, and when both the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction are set to 0, signaling a modified delta POC value for an element of the reference picture list, wherein the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, said signaling The method of exemplary embodiment 3 comprising.
[0234] 6. Exemplary embodiment: The method of exemplary embodiment 4 or 5, wherein the modified delta POC value is 1 smaller than the delta POC value used in the coding process.
[0235] 7. Exemplary embodiment: When at least one HLS weighted prediction parameter is signaled after the reference picture list in coding order, the restriction on the binarization of the syntax element is Signaling at least one HLS weighted prediction parameter only when the reference picture list has at least one element with a delta POC value equal to 0, the method of any one of exemplary embodiments 1 to 3 comprising.
[0236] 8. Exemplary embodiment: Receiving a bitstream; Entropy decoding the bitstream to obtain syntax elements, wherein the syntax elements include a reference picture list and at least one HLS weighted prediction parameter, and among the elements, at least one HLS weighted prediction parameter is presented before the reference picture list; Performing a prediction based on the obtained syntax elements to obtain a predicted block; Reconstructing the reconstructed block based on the prediction block; Obtaining the decoded picture based on the reconstructed block; A decoding method by a decoder comprising:
[0237] 9. Exemplary embodiment: The method of exemplary embodiment 8, wherein at least one HLS weighted prediction parameter includes at least one of a sequence parameter set flag for weighted uni-prediction and a sequence parameter set flag for weighted bi-prediction.
[0238] 10. Exemplary embodiment: Receiving a bitstream; Entropy decoding the bitstream to obtain a syntax element, the syntax element including a reference picture list and a preset flag, the value of the preset flag indicating whether the syntax element includes at least one HLS weighted prediction parameter; Performing prediction based on the obtained syntax element to obtain a prediction block; Reconstructing the reconstructed block based on the prediction block; Obtaining the decoded picture based on the reconstructed block; A decoding method by a decoder comprising:
[0239] 11. Exemplary embodiment: The method of exemplary embodiment 10, wherein the value of the preset flag corresponds to whether the reference picture list has at least one element having a delta POC value equal to 0.
[0240] 12. Exemplary embodiment: When the value of the preset flag corresponding to the reference picture list has at least one element having a delta POC value equal to 0, the syntax element includes at least one HLS weighted prediction parameter, or When there is no element having a delta POC value equal to 0 for a preset flag value corresponding to a reference picture list, the method of exemplary embodiment 11 where the syntax element does not include at least one HLS weighted prediction parameter.
[0241] 13. Exemplary embodiment: A method according to any one of exemplary embodiments 10 to 12, where at least one HLS weighted prediction parameter includes at least one of a sequence parameter set flag for weighted single prediction and a sequence parameter set flag for weighted bi-prediction.
[0242] 14. Exemplary embodiment: A method according to any one of exemplary embodiments 10 to 13, where the preset flag is RestrictWPFlag as defined in the specification.
[0243] 15. Exemplary embodiment: An encoder (20) comprising a processing circuit for performing the method according to any one of exemplary embodiments 1 to 7.
[0244] 16. Exemplary embodiment: A decoder (30) comprising a processing circuit for performing the method according to any one of exemplary embodiments 8 to 14.
[0245] 17. Exemplary embodiment: A computer program product comprising program code for performing the method according to any one of exemplary embodiments 1 to 14.
[0246] 18. Exemplary embodiment: One or more processors, A non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming configuring the decoder to perform a method according to any one of exemplary embodiments 8 to 14 when executed by the processor.
[0247] 19. Exemplary embodiments: One or more processors, a non - transitory computer - readable storage medium coupled to the processor and storing programming for execution by the processor, the programming configuring an encoder to execute a method according to any one of exemplary embodiments 1 to 7 when executed by the processor.
[0248] 20. Exemplary embodiments: A non - transitory computer - readable medium carrying program code that causes a computer device to execute a method according to any one of exemplary embodiments 1 to 14 when executed by the computer device.
[0249] Mathematical operators The mathematical operators used in this application are similar to those used in the C programming language. However, the results of integer division and arithmetic shift operations are more precisely defined, and additional operations such as exponentiation and real - valued division are defined. The numbering and counting conventions generally start from 0. For example, "the first" is equivalent to the 0th, "the second" is equivalent to the 1st, and so on.
[0250] Arithmetic operators The following arithmetic operators are defined as follows. + Addition. - Subtraction (as a two - argument operator) or sign inversion (as a unary prefix operator). * Multiplication, including matrix multiplication. x y Exponentiation. Specifies x to the power of y. In other contexts, such notation is used to make superscripts that are not intended for interpretation as exponents. / Integer division with truncation of the result to 0. For example, 7 / 4 and - 7 / -4 are truncated to 1, and - 7 / 4 and 7 / -4 are truncated to - 1. ÷ Used to denote division in a mathematical formula where truncation or rounding is not intended.
[0251]
Number
[0252] It is used to represent division in a mathematical formula where truncation or rounding is not intended.
[0253]
Number
[0254] The sum of f(i) where i takes all integer values from x to y, including y. x % y method. The remainder when x is divided by y, defined only for integers x and y where x >= 0 and y > 0.
[0255] Logical operators The following logical operators are defined as follows. x && y The Boolean logical 'AND' of x and y. x || y The Boolean logical 'OR' of x and y. ! The Boolean logical 'NOT'. x? y : z If x is TRUE, i.e., not equal to 0, then evaluate to the value of y, otherwise evaluate to the value of z.
[0256] Relational operators The following relational operators are defined as follows. > Greater than. >= Greater than or equal to. < Less than. <= Less than or equal to. == Equal to. != Not equal to.
[0257] When a relational operator is applied to a syntax element or variable to which the value 'na' (not applicable) is assigned, the value 'na' is treated as a special value for that syntax element or variable. The value 'na' is considered not equal to any other value.
[0258] Bitwise operator The following bitwise operators are defined as follows. & Bitwise "logical AND". When operating on integer arguments, it operates on the two's complement representation of the integer value. When operating on a binary argument containing fewer bits than the other argument, the shorter argument is extended by adding higher-order bits equal to 0. | Bitwise "logical OR". When operating on integer arguments, it operates on the two's complement representation of the integer value. When operating on a binary argument containing fewer bits than the other argument, the shorter argument is extended by adding higher-order bits equal to 0. ^ Bitwise "exclusive OR". When operating on integer arguments, it operates on the two's complement representation of the integer value. When operating on a binary argument containing fewer bits than the other argument, the shorter argument is extended by adding higher-order bits equal to 0. x >> y Arithmetic right shift of the two's complement integer representation of x by y binary digits. This function is defined only for non-negative integer values of y. The bit shifted into the most significant bit (MSB) as a result of the right shift has the same value as the MSB of x before the shift operation. x << y Arithmetic left shift of the two's complement integer representation of x by y binary digits. This function is defined only for non-negative integer values of y. The bit shifted into the least significant bit (LSB) as a result of the left shift has a value equal to 0.
[0259] Assignment operator The following arithmetic operators are defined as follows. = Assignment operator. ++ Increment, i.e., x++ is equivalent to x = x + 1 and, when used in an array index, evaluates to the value of the variable before the increment operation. -- The decrement, i.e., x--, is equivalent to x = x - 1 and, when used in an array index, evaluates to the value of the variable before the decrement operation. += Increment by the specified amount, i.e., x += 3 is equivalent to x = x + 3 and x += (-3) is equivalent to x = x + (-3). -= Decrement by the specified amount, i.e., x -= 3 is equivalent to x = x - 3 and x -= (-3) is equivalent to x = x - (-3).
[0260] Range notation The following notation is used to specify a range of values. x=y..z x takes integer values starting from y up to and including z, where x, y, and z are integers and z is greater than y.
[0261] Mathematical functions The following mathematical functions are defined.
[0262] [Number]
[0263] Asin(x) Performs the operation on an argument x within the range including both ends from -1.0 to 1.0, and has an output value within the range including both ends from -π÷2 to π÷2 in radians, which is the inverse sine function of trigonometry. Atan(x) Performs the operation on an argument x and has an output value within the range including both ends from -π÷2 to π÷2 in radians, which is the inverse tangent function of trigonometry.
[0264] [Number]
[0265] Ceil(x) The smallest integer greater than or equal to x. Clip1 Y (x) = Clip3(0, (1 << BitDepthY ) - 1, x ) Clip1 C ( x ) = Clip3( 0, ( 1 << BitDepth C ) - 1, x )
[0266] [Number]
[0267] Cos(x) The cosine function of trigonometry that operates on the argument x in radians. Floor(x) The largest integer less than or equal to x.
[0268] [Number]
[0269] Ln(x) The natural logarithm of x (logarithm with base e, where e is the natural logarithm base constant 2.718 281 828...). Log2(x) The logarithm of x with base 2. Log10(x) The logarithm of x with base 10.
[0270] [Number]
[0271] Round( x ) = Sign( x ) * Floor( Abs( x ) + 0.5 )
[0272] [Number]
[0273] Sin(x) The sine function of trigonometry that operates on the argument x in radians.
[0274] [Number]
[0275] Swap(x, y) = (y, x) Tan(x) is the trigonometric tangent function that operates on the argument x in radians.
[0276] Order of operation precedence When the order of precedence in an expression is not explicitly indicated by the use of parentheses, the following rules apply. - Operations with higher precedence are evaluated before any operations of lower precedence. - Operations of the same precedence are evaluated sequentially from left to right.
[0277] The following table specifies the precedence of operations from highest to lowest, with higher positions in the table indicating higher precedence.
[0278] For those operators that are also used in the C programming language, the order of precedence used in this specification is the same as that used in the C programming language.
[0279]
Table 7
[0280] Text description of logical operations In the text, statements of logical operations that will be mathematically described in the following form, i.e., if (condition0) statement0 else if (condition1) statement1 ... else / * explanatory note for remaining conditions * / statementn can be described in the following form. ...as follows / ...the following applies - If condition0, then statement0 - Otherwise, if condition 1, then statement 1 -... - Otherwise (explanatory note for the remaining conditions), statement n.
[0281] Each "if... then... else... then... else..." statement in the text is introduced by "if... then... as follows" or "if... then... the following applies" immediately after it. The last condition of "if... then... else... then... else..." is always "else...". The alternating "if... then... else... then... else..." statements can be identified by aligning "if... then... as follows" or "if... then... the following applies" with the final "else...".
[0282] In the text, logical operation statements that will be mathematically described in the following form, that is, if (condition 0a && condition 0b) statement 0 else if (condition 1a || condition 1b) statement 1 ... else statement n can be described in the following form. ... as follows / ... the following applies - If all of the following conditions are true, then statement 0: - condition 0a - condition 0b - Otherwise, if one or more of the following conditions are true, then statement 1: - condition 1a - condition 1b -... - Otherwise, then statement n
[0283] In this document, a logical operation statement that will be mathematically described in the following form, that is, if (condition 0) Statement 0 if (condition 1) Statement 1 can be described in the following form. When condition 0, Statement 0 When condition 1, Statement 1.
[0284] For example, the embodiments of the encoder 20 and the decoder 30, and the functions described herein with reference to the encoder 20 and the decoder 30, may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes in a computer-readable medium or transmitted over a communication medium and executed by a processing unit based on hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one location to another, for example, in accordance with a communication protocol. In this form, the computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0285] By way of example and without limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that is accessible by a computer. Also, any connection can properly be called a computer-readable medium. For example, if the instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transient tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc (registered trademark), optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray (registered trademark) disc, where disk typically magnetically reproduces data, while disc optically reproduces data using a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0286] The commands can be executed by one or more processors such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated logic circuits or discrete logic circuits. Thus, the term "processor" as used herein can refer to either the foregoing structures or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functions described herein can be provided within dedicated hardware and / or software modules configured to encode and decode, or incorporated within a combined codec. Also, the techniques can be fully realized within one or more circuits or logic elements.
[0287] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chip sets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to execute the disclosed techniques, but do not necessarily require implementation by different hardware units. Rather, as described above, various units can be combined within a codec hardware unit, or provided with suitable software and / or firmware by a set of interoperable hardware units including one or more processors as described above.
Description of the Signs
[0288] 10 Video coding system 12 Source device 13 Communication channel 14 Destination device 16 Picture source 17 Picture, picture data, unprocessed picture, unprocessed picture data 18 Preprocessor, preprocessing unit 19 Preprocessed picture, preprocessed picture data 20 Video encoder 21 Encoded picture data 22 Communication interface, communication unit 28 Communication interface, communication unit 30 Video decoder, short decoder 31 Decoded picture, decoded picture data 32 Postprocessor, post-processing unit 33 Post-processed picture, post-processed picture data 34 Display device 46 Processing circuit 201 Input, input interface 203 Picture block 204 Residual calculation unit 205 Residual block, residual 206 Transformation processing unit 207 Transformation coefficient 208 Quantization unit 209 Quantized coefficient, quantized transformation coefficient, quantized residual coefficient 210 Inverse quantization unit 211 Dequantized coefficient, dequantized residual coefficient 212 Inverse transformation processing unit 213 Reconstructed residual block, corresponding dequantized coefficient, transformation block 214 Reconstruction unit 215 Reconstructed block 220 Loop filter unit 221 Filtered block, filtered reconstructed block 230 Decoded picture buffer 231 Decoded picture 244 Inter prediction unit 254 Intra prediction unit 260 Mode selection unit 262 Partitioning unit 265 Prediction block, predictor 266 Syntax element 270 Entropy encoding unit 272 Output, output interface 304 Entropy decoding unit 309 Quantized coefficient 310 Inverse quantization unit 311 Transform coefficient, dequantized coefficient 312 Inverse transform processing unit 313 Reconstructed residual block, transform block 314 Reconstruction unit, adder 315 Reconstructed block 320 Loop filter unit 321 Filtered block, decoded video block of picture 330 Decoded picture buffer (DPB) 331 Decoded picture 332 Output 344 Inter prediction unit 354 Intra prediction unit 360 Mode application unit 365 Prediction block 400 Video coding device 410 Inlet port, input port 420 Receiver unit 430 Processor, logic unit, central processing unit 440 Transmitter unit 450 Outlet port, output port 460 Memory 470 Coding module 500 Device 502 Processor 504 Memory 506 Code and data 508 Operating system 510 Application program 512 Bus 514 Secondary storage 518 Display 610 Video analysis 611 Statistical value 631 Status value 660 Encoding engine 1500 Decoder 1501 Processor 1502 Non - transitory computer - readable storage medium 1600 Decoder 1601 Receiving means 1602 Entropy decoding means 1603 Prediction means 1604 Reconstruction means 1605 Acquisition means 1700 Encoder 1701 Processor 1702 Non - transitory computer - readable storage medium 1800 Encoder 1801 Decision means 1802 Coding means 3100 Content supply system 3102 Capture device 3104 Communication link 3106 Terminal device 3108 Smartphone / tablet 3110 Computer / laptop 3112 Network video recorder / digital video recorder 3114 TV 3116 Set - top box 3118 Video conferencing system 3120 Video surveillance system 3122 Portable information terminal 3124 Vehicle - mounted device 3126 Display 3202 Protocol progress unit 3204 Demultiplexing unit 3206 Video decoder 3208 Audio decoder 3210 Subtitle Decoder 3212 Synchronization Unit 3214 Video / Audio Display 3216 Video / Audio / Subtitle Display
Claims
1. An encoding method, comprising: - determining syntax elements to be coded, the syntax elements including a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter; - coding the at least one HLS weighted prediction parameter; and - coding the reference picture list structure following the coding of the at least one HLS weighted prediction parameter. A method comprising the above steps.
2. The method according to claim 1, wherein the reference picture list derived from the reference picture list structure comprises reference pictures having the same picture order count (POC) parameter.
3. The method according to claim 1 or 2, wherein the at least one HLS weighted prediction parameter comprises a sequence parameter set flag for weighted single prediction.
4. The method according to any one of claims 1 to 3, wherein the at least one HLS weighted prediction parameter comprises a sequence parameter set flag for weighted bi-prediction.
5. The method according to any one of claims 1 to 4, wherein the coding of the reference picture list structure comprises a restriction on the binarization of at least a part of the reference picture list structure.
6. The restriction on the binarization of at least a part of the reference picture list structure comprises: when the sequence parameter set flag for weighted single prediction is set to 0, coding a modified delta POC value for elements of the reference picture list, the modified delta POC value (abs_delta_poc_st) being smaller than the delta POC value (AbsDeltaPocSt) used in the coding process. The method according to claim 5.
7. The restriction on the binarization comprises: (i) When the at least one HLS weighted prediction parameter includes the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted bi-prediction is set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, wherein the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, said coding, or (ii) When the at least one HLS weighted prediction parameter includes the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction, and at least one of the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction is set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, wherein the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, said coding, or (iii) When the at least one HLS weighted prediction parameter includes the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction, and both the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction are set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, wherein the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, said coding The method according to claim 5, comprising. Claim 8 The method according to claim 6 or 7, wherein the corrected delta POC value is 1 less than the delta POC value used in the coding process. **Claim 9** An encoding method, comprising: - determining a syntax element to be coded, the syntax element including a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter, and the reference picture list derived from the reference picture list structure comprising reference pictures having the same picture order count (POC) parameter; coding the determined syntax element in the coding order with a restriction on the binarization of syntax elements having a later position in the coding order; and when the at least one HLS weighted prediction parameter is coded after the reference picture list structure in the coding order, the restriction on the binarization of the syntax element includes coding the at least one HLS weighted prediction parameter only when the reference picture list has at least one element having a delta POC value equal to 0. **Claim 10** A decoding method by a decoder, comprising: receiving a bitstream; decoding the bitstream to obtain a syntax element, the syntax element including a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter, and among the syntax elements, the at least one HLS weighted prediction parameter being decoded before the reference picture list structure; performing a prediction based on the obtained syntax element to obtain a predicted block; reconstructing a reconstructed block based on the predicted block; and obtaining a decoded picture based on the reconstructed block. **Claim 11** The method according to claim 10, wherein the at least one HLS weighted prediction parameter includes at least one of a sequence parameter set flag for weighted single prediction and a sequence parameter set flag for weighted bi-prediction. **Claim 12** The method according to claim 10 or 11, wherein the decoding of the bitstream to obtain syntax elements is performed by entropy decoding.
13. The step of performing a prediction based on the obtained syntax elements to obtain a prediction block includes: obtaining a value of delta POC based on the at least one HLS weighted prediction parameter and syntax elements in the reference picture list structure; performing a prediction based on the value of delta POC; The method according to any one of claims 10 to 12.
14. The step of obtaining the value of delta POC based on the at least one HLS weighted prediction parameter includes: determining whether the value of delta POC is allowed to have a value of 0 based on the value of the at least one HLS weighted prediction parameter; when it is determined that the value of delta POC is not allowed to have a value of 0, restoring the value of delta POC using the incremented value of the syntax element in the reference picture list structure; The method according to claim 13.
15. The syntax element in the reference picture list structure is abs_delta_poc_st, and the value of delta POC is obtained as follows: if( sps_weighted_pred_flag || sps_weighted_bipred_flag ) AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] = abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] else AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] = abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] + 1 where AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] is the absolute value of delta POC, and abs_delta_poc_st[ listIdx ][ rplsIdx [ i ] is the syntax element in the reference picture list structure. The method according to claim 13 or 14.
16. A decoding method by a decoder, Receiving a bitstream; Entropy-decoding the bitstream to obtain syntax elements, where the syntax elements include a reference picture list structure and a preset flag, and the value of the preset flag indicates whether the syntax elements include at least one high-level syntax (HLS) weighted prediction parameter; Performing prediction based on the obtained syntax elements to obtain a prediction block; Reconstructing a reconstructed block based on the prediction block; Obtaining a decoded picture based on the reconstructed block; A decoding method comprising the above steps.
17. The method according to claim 16, wherein the value of the preset flag corresponds to whether the reference picture list derived from the reference picture list structure has at least one element with a delta POC value equal to 0.
18. When the value of the preset flag corresponding to the reference picture list has at least one element with a delta POC value equal to 0, the syntax elements include at least one HLS weighted prediction parameter; Or When the value of the preset flag corresponding to the reference picture list has no element with a delta POC value equal to 0, the syntax elements do not include the at least one HLS weighted prediction parameter. The method according to claim 17.
19. The method according to any one of claims 16 to 18, wherein the at least one HLS weighted prediction parameter includes at least one of a sequence parameter set flag for weighted single prediction and a sequence parameter set flag for weighted bi-prediction.
20. The method according to any one of claims 16 to 19, wherein the preset flag is the RestrictWPFlag that is set to true in the coding process when a delta POC value (AbsDeltaPocSt) of 0 appears during the inspection of each element of the reference picture list.
21. An encoder (20) comprising a processing circuit for performing the method according to any one of claims 1 to 9.
22. A decoder (30) comprising a processing circuit for executing the method according to any one of claims 10 to 20.
23. A computer program product comprising program code for executing the method according to any one of claims 1 to 20.
24. A decoder, comprising one or more processors, and a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming configuring the decoder to execute the method according to any one of claims 10 to 20 when executed by the processor.
25. A decoder, comprising receiving means for receiving a bitstream, decoding means for decoding the bitstream to obtain syntax elements, the syntax elements comprising a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter, and wherein, among the syntax elements, the at least one HLS weighted prediction parameter is entropy decoded before the reference picture list structure, prediction means for performing a prediction based on the obtained syntax elements to obtain a predicted block, reconstruction means for reconstructing a reconstructed block based on the predicted block, and obtaining means for obtaining a decoded picture based on the reconstructed block A decoder comprising.
26. The decoder according to claim 25, wherein the value of the preset flag corresponds to whether the reference picture list derived from the reference list structure has at least one element having a delta POC value equal to 0.
27. When the value of the preset flag corresponding to the reference picture list has at least one element having a delta POC value equal to 0, the syntax element includes at least one HLS weighted prediction parameter, or When the value of the preset flag corresponding to the reference picture list has no element having a delta POC value equal to 0, the syntax element does not include the at least one HLS weighted prediction parameter. The decoder according to claim 26.
28. The decoder according to any one of claims 25 to 27, wherein the at least one HLS weighted prediction parameter includes at least one of a sequence parameter set flag for weighted single prediction and a sequence parameter set flag for weighted bi-prediction.
29. The decoder according to any one of claims 25 to 28, wherein the preset flag is the RestrictWPFlag that is set to true in the coding process when a delta POC value (AbsDeltaPocSt) of 0 appears during the inspection of each element of the reference picture list.
30. The decoder according to any one of claims 25 to 29, wherein the decoding means is further for decoding the bitstream by entropy decoding to obtain syntax elements.
31. Performing prediction based on the obtained syntax elements to obtain a prediction block, obtaining a value of delta POC based on the at least one HLS weighted prediction parameter and a syntax element in the reference picture list structure, and performing prediction based on the value of delta POC The decoder according to any one of claims 25 to 30, comprising:
32. Obtaining the value of delta POC based on the at least one HLS weighted prediction parameter, determining whether the value of delta POC is allowed to have a value of 0 based on the value of the at least one HLS weighted prediction parameter, and when it is determined that the value of delta POC is not allowed to have a value of 0, restoring the value of delta POC using an incremented value of the syntax element in the reference picture list structure The decoder according to claim 31, comprising:
33. The syntax element in the reference picture list structure is abs_delta_poc_st, and the value of delta POC is as follows, that is, if( sps_weighted_pred_flag || sps_weighted_bipred_flag ) AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] = abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] else AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] = abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] + 1 is obtained as such, where AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] is the absolute value of the delta POC, and abs_delta_poc_st[ listIdx ][ rplsIdx [ i ] is the syntax element in the reference picture list structure, the decoder according to claim 31 or 32.
34. An encoder, comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming configuring the encoder to perform the method according to any one of claims 1 to 9 when executed by the processor.
35. An encoder, comprising: determining means for determining a syntax element to be coded, the syntax element including a reference picture list structure and at least one high-level syntax (HLS) weighted prediction parameter; coding means for coding the at least one HLS weighted prediction parameter and for coding the reference picture list structure following the coding of the at least one HLS weighted prediction parameter An encoder comprising.
36. The encoder according to claim 35, wherein the reference picture list derived from the reference picture list structure comprises reference pictures having the same picture order count (POC) parameter.
37. The encoder according to claim 35 or 36, wherein the at least one HLS weighted prediction parameter comprises a sequence parameter set flag for weighted single prediction.
38. The encoder according to any one of claims 35 to 37, wherein the at least one HLS weighted prediction parameter comprises a sequence parameter set flag for weighted bi-prediction.
39. The encoder according to any one of claims 35 to 38, wherein the coding of the reference picture list structure comprises a restriction on the binarization of at least a part of the reference picture list structure.
40. The restriction on the binarization of at least a part of the reference picture list structure is when the sequence parameter set flag for weighted uni-prediction is set to 0, signaling a modified delta POC value for an element of the reference picture list, the modified delta POC value (abs_delta_poc_st) being smaller than the delta POC value (AbsDeltaPocSt) used in the coding process. The encoder according to claim 39.
41. The restriction on the binarization is (i) when the at least one HLS weighted prediction parameter includes the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted bi-prediction is set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, the modified delta POC value (abs_delta_poc_st) being smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, the coding, or (ii) When the at least one HLS weighted prediction parameter includes the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction, and at least one of the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction is set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, wherein the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, said coding; or (iii) When the at least one HLS weighted prediction parameter includes the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction, and both the sequence parameter set flag for weighted bi-prediction and the sequence parameter set flag for weighted uni-prediction are set to 0, coding a modified delta POC value for an element of the reference picture list derived from the reference picture list structure, wherein the modified delta POC value (abs_delta_poc_st) is smaller than the delta POC value (AbsDeltaPocSt) used in the coding process, said coding The encoder according to claim 39, comprising. [
42. ] The encoder according to claim 40 or 41, wherein the modified delta POC value is 1 smaller than the delta POC value used in the coding process. [
43. ] A non-transitory computer-readable medium carrying program code for causing a computer device to execute the method according to any one of claims 1 to 20 when executed by the computer device.
Citation Information
Patent Citations
Weighted prediction parameter coding
US20130259130A1
Method and apparatus for efficient signaling of weighted prediction in advanced coding schemes
US20140056356A1
Weighted prediction parameter signaling for video coding
US20150103898A1