High-level syntax for video encoding and decoding
By constraining information signaling in the bitstream syntax to either the picture header or slice header, the complexity of the VVC standard is reduced, enhancing decoder simplicity and maintaining encoding efficiency.
Patent Information
- Application Number
- JP2025069889
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-04-20
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2041-03-05
AI Technical Summary
The high-level syntax changes in the Versatile Video Coding (VVC) standard have increased complexity without improving encoding performance, particularly for applications like 360-degree video and high-dynamic range (HDR) video.
Implementing constraints in the bitstream syntax structure by ensuring that certain information is signaled only in the picture header or slice header, reducing redundancy and simplifying decoding and encoding processes without degrading compression efficiency.
Simplifies decoder implementation and reduces signaling redundancy by allowing information to be present only in the picture header or slice header, maintaining encoding efficiency.
Smart Images

Figure 2025111598000043 
Figure 2025111598000044 
Figure 2025111598000045
Abstract
Description
Technical Field
[0001] The present invention relates to video encoding and decoding, and more particularly to high-level syntax used in bitstreams.
Background Art
[0002] Recently, the Joint Video Expert Team (JVET), a joint team formed by MPEG and the Video Coding Experts Group (VCEG) of ITU-T Study Group 16, has started work on formulating a new video coding standard called Versatile Video Coding (VVC). The goal of VVC is to significantly improve the compression performance compared to the existing HEVC standard (typically twice that of the previous one), and it is scheduled to be completed in 2020. The main target applications and services include, but are not limited to, 360-degree video and high-dynamic range (HDR) video. JVET evaluated responses from 32 organizations through formal subjective tests by independent test labs. In some proposals, an improvement in compression efficiency of typically more than 40% was shown compared to using HEVC. In particular, it was shown to be effective for test materials of ultra-high definition (UHD) video. Therefore, an improvement in compression efficiency far exceeding the target of 50% for the final standard is expected.
[0003] The JVET Exploration Model (JEM) uses all HEVC tools and introduces many new tools. These changes have required changes to the high-level syntax that can affect the structure of the bitstream, particularly the overall bitrate of the bitstream.
Summary of the Invention
[0004] The present invention relates to an improvement in the high-level syntax structure that reduces complexity without degrading the encoding performance.
[0005] According to a first aspect of the present invention, a method for decrypting video data from a bitstream is provided. The bitstream includes video data corresponding to one or more slices. The bitstream includes a picture header including a plurality of syntax elements to be used when decrypting one or more slices, and a slice header including a plurality of syntax elements to be used when decrypting one slice. The decryption includes imposing (e.g., imposing a constraint) that the picture header is not within the slice header when information that can be signaled within the picture header or within the slice header is signaled within the picture header, and decrypting the bitstream using the plurality of syntax elements.
[0006] Optionally, the decryption further includes parsing a first syntax element indicating whether the information is signaled within the picture header, and based on the first syntax element, permitting the parsing of the information that can be signaled within the slice header and within the picture header only in either the slice header or the picture header.
[0007] Optionally, the first syntax element is information within a picture parameter set flag or information within a picture header flag.
[0008] Optionally, when the first syntax element indicates that the information is signaled within the picture header, parsing of the information within the slice header is not permitted.
[0009] Optionally, the method further includes parsing a second syntax element indicating whether the picture header is within the slice header. When the first syntax element indicates that the information is signaled within the picture header, it is a requirement for bitstream compliance that the second syntax element indicates that the picture header is not within the slice header.
[0010] Optionally, the information includes one or more of quantization parameter value information, reference picture list information, deblocking filter information, sample adaptation offset (SAO) information, weighted prediction information, and adaptation loop filter (ALF) information.
[0011] Optionally, the information includes all information that can be signaled within a picture header and within a slice header.
[0012] Optionally, the reference picture list information includes one or more of slice_collocated_from_l0_flag, slice_collocated_ref_idx, ph_collocated_from_l0_flag, and ph_collocated_ref_idx.
[0013] According to a second aspect of the present invention, there is provided a method of encoding video data into a bitstream, the video data corresponding to one or more slices, the bitstream including a picture header including a plurality of syntax elements to be used when decoding one or more slices, and a slice header including a plurality of syntax elements to be used when decoding one slice, the encoding including signaling that the picture header is not within the slice header when information that can be signaled within the picture header or within the slice header is signaled within the picture header, and encoding the video data using the plurality of syntax elements.
[0014] Optionally, the encoding further includes encoding a first syntax element indicating whether the information is signaled within the picture header, and based on the first syntax element, permitting encoding of the information that can be signaled within the slice header and within the picture header in only one of the slice header or the picture header.
[0015] Optionally, the first syntax element is information within a picture parameter set flag or information within a picture header flag.
[0016] Optionally, if the first syntax element indicates that the information is signaled within the picture header, encoding of the information within the slice header is not permitted.
[0017] Optionally, the method further includes parsing a second syntax element indicating whether the picture header is within the slice header, and if the first syntax element indicates that the information is signaled within the picture header, it is a bitstream conformity requirement that the second syntax element indicates that the picture header is not within the slice header.
[0018] Optionally, the information includes one or more of quantization parameter value information, reference picture list information, deblocking filter information, sample adaptive offset (SAO) information, weighted prediction information, and adaptive loop filter (ALF) information.
[0019] Optionally, the information includes all information that can be signaled within the picture header and within the slice header.
[0020] Optionally, the reference picture list information includes one or more of slice_collocated_from_l0_flag, slice_collocated_ref_idx, ph_collocated_from_l0_flag, and ph_collocated_ref_idx.
[0021] In an alternative aspect of the present invention, a method for decoding video data from a bitstream is provided, the bitstream including video data corresponding to one or more slices, the bitstream including a picture header including syntax elements to be used when decoding one or more slices and a slice header including syntax elements to be used when decoding a slice, the decoding including, when information that can be signaled in the picture header or slice header is signaled in the slice header, imposing that it is not signaled in the slice header and decoding the bitstream using the syntax elements.
[0022] According to another aspect of the present invention, a method for decoding video data from a bitstream is provided, the bitstream including video data corresponding to one or more slices, the bitstream including a picture header including syntax elements to be used when decoding one or more slices and a slice header including syntax elements to be used when decoding a slice, the decoding including treating as inapplicable in combination (a) a syntax element indicating that tool information is signaled in the picture header rather than the slice header and (b) a syntax element indicating that it is signaled in the slice header, and decoding the bitstream using the syntax elements. The decoding is not performed if there are syntax elements (a) and (b) that are treated as inapplicable in combination.
[0023] According to related aspects of the present invention, a bitstream includes video data corresponding to one or more slices, a picture header including syntax elements used when decoding the one or more slices, and a slice header including syntax elements used when decoding a slice. The bitstream has a constraint that (a) a syntax element indicating that tool information is signaled in the picture header rather than in the slice header and (b) a syntax element indicating that the picture header is signaled in the slice header shall not exist there in combination. The tool information may be any one of quantization parameter value information, reference picture list information, deblocking filter information, sample adaptive offset (SAO) information, weighted prediction information, and adaptive loop filter (ALF) information. In related aspects, a method for decoding a bitstream is provided. In another related aspect, a decoder configured to decode a bitstream is provided. The bitstream may be constrained to conform to a video coding standard. In one embodiment, the video coding standard is a multi-purpose video coding standard. The constraint may be systematically applied to the entire bitstream. For example, in an embodiment, the constraint is applied to any or all of sequences, pictures, and slices within the bitstream.
[0024] According to another aspect of the present invention, a method for decoding video data from a bitstream is provided. The bitstream includes video data corresponding to one or more slices, and the bitstream includes a picture header including syntax elements to be used when decoding one or more slices, and a slice header including syntax elements to be used when decoding a slice. Decoding includes (a) treating as inapplicable a combination of a syntax element indicating that tool information is signaled in the slice header rather than the picture header and a syntax element indicating that the picture header is signaled in the slice header, and (b) decoding the bitstream using the syntax elements. Decoding is not performed if there are syntax elements (a) and (b) that are treated as inapplicable in combination.
[0025] According to another aspect of the present invention, a method for decoding video data from a bitstream is provided. The bitstream includes video data corresponding to one or more slices, and the bitstream includes a picture header including syntax elements to be used when decoding one or more slices, and a slice header including syntax elements to be used when decoding a slice. The picture header is to be signaled in the slice header, and parsing of information that can be signaled only in one of the slice header or the picture header is restricted. Decoding includes, when a syntax element (xxx_info_in_ph_flag) indicates that there is tool information in the picture header (for example, when xxx_info_in_ph_flag = 1), not signaling in the slice header (for example, forcing picture_header_in_slice_header_flag to 0).
[0026] According to related aspects of the present invention, a bitstream includes video data corresponding to one or more slices, a picture header including syntax elements used when decoding the one or more slices, and a slice header including syntax elements used when decoding a slice. The bitstream has a constraint that (a) a syntax element indicating that tool information is signaled in a slice header rather than in a picture header and (b) a syntax element indicating that a picture header is signaled in a slice header shall not be present there in combination.
[0027] According to other aspects of the present invention, a method for decoding video data from a bitstream is provided. The bitstream includes video data corresponding to one or more slices. The bitstream includes a picture header including syntax elements to be used when decoding the one or more slices and a slice header including syntax elements used when decoding a slice. The picture header is to be signaled in the slice header, and parsing of information that can be signaled only in one of the slice header or the picture header is restricted. Decoding includes not signaling in the slice header (e.g., forcing picture_header_in_slice_header_flag to 0) when a syntax element (xxx_info_in_ph_flag) indicates that there is no tool information in the picture header (e.g., when xxx_info_in_ph_flag = 0).
[0028] According to a first further aspect of the present invention, there is provided a method for decoding video data from a bitstream, the bitstream including video data corresponding to one or more slices, the bitstream including a picture header including syntax elements to be used when decoding one or more slices, and a slice header including syntax elements to be used when decoding a slice, the decoding including, when the picture header is signaled in the slice header, permitting parsing of information that may be signaled in only one of the slice header or the picture header and decoding the bitstream using the syntax elements.
[0029] If the picture header is within the slice header, this means that there is only one slice for the current picture. Therefore, even if information is transmitted or made transmissible for both the slice and the picture, since the parameters are the same, the flexibility of the encoder or decoder is not improved. That is, if there is information in the picture header, the corresponding information in the slice header becomes redundant. Similarly, if there is information in the slice header, the corresponding information in the picture header becomes redundant. When the picture header is in the slice header, by permitting only the information in the picture header or the slice header, redundancy in signaling can be limited and the implementation of the decoder can be simplified. Therefore, syntax analysis can be simplified without degrading the coding efficiency.
[0030] The decoding may further include parsing a first syntax element indicating whether or not to signal a picture header in the slice header, and permitting parsing of information that may be signaled in only one of the slice header or the picture header based on the first syntax element. The first syntax element may be a picture header flag in the slice header.
[0031] Optionally, parsing of the information in the slice header is not permitted if the first syntax element indicates that a picture header is signaled within the slice header. A second syntax element indicating whether the information is in the picture header may be parsed, and it is a requirement for bitstream conformance that if the first syntax element indicates that a picture header is signaled in the slice header, the second syntax element indicates that the information is signaled in the picture header.
[0032] Alternatively, if the first syntax element indicates that a picture header is signaled within the slice header, signaling of the information in the picture header is not permitted. This method may further include parsing a second syntax element indicating whether the information is in the picture header, and it is a requirement for bitstream conformance that if the first syntax element indicates that a picture header is signaled in the slice header, the second syntax element indicates that the information is in the picture header. The second syntax element may be a picture parameter set flag, where if the flag is set the information is in the picture header and if it is not set the information is in the slice header or does not exist.
[0033] According to an embodiment, the information can include one or more of quantization parameter value information, reference picture list information, deblocking filter information, sample adaptive offset (SAO) information, weighted prediction information, and adaptive loop filter (ALF) information. For example, it can include all of quantization parameter value information, reference picture list information, deblocking filter information, sample adaptive offset (SAO) information, weighted prediction information, and adaptive loop filter (ALF) information.
[0034] Optionally, this information includes all information that may be signaled in both the picture header and the slice header.
[0035] The reference picture list information can include one or more of the following syntax elements: slice_collocated_from_l0_flag, slice_collocated_ref_idx, ph_collocated_from_l0_flag, ph_collocated_ref_idx.
[0036] When signaling a picture header within a slice header, the number of weights for parseable weighted prediction can be restricted.
[0037] According to a second further aspect of the present invention, there is provided a method of encoding video data into a bitstream, wherein the video data corresponds to one or more slices, and the bitstream includes a picture header including syntax elements used when decoding one or more slices, and a slice header including syntax elements used when decoding a slice, and the encoding includes, when a picture header is signaled within a slice header, permitting encoding only in either the slice header or the picture header for information that can be signaled in both the slice header and the picture header, and encoding the video data using the syntax elements.
[0038] If the picture header is within the slice header, this means that there is only one slice for the current picture. Therefore, even if information is sent or made available for both the slice and the picture, the flexibility of the encoder or decoder is not improved because the parameters are the same. That is, if there is information in the picture header, the corresponding information in the slice header is redundant. Similarly, if there is information in the slice header, the corresponding information in the picture header is redundant. By allowing only the information in the picture header or the slice header when the picture header is within the slice header, redundancy in signaling can be limited and the implementation of the decoder can be simplified. Therefore, encoding can be simplified without sacrificing encoding efficiency, and signaling cost can be reduced (because related information is included only once in the bitstream).
[0039] Encoding may further include encoding a first syntax element indicating whether to signal the picture header within the slice header, and allowing encoding of information that can be signaled in only one of the slice header or the picture header based on the first syntax element.
[0040] The first syntax element may be the picture header of the slice header flag.
[0041] When the first syntax element indicates that the picture header is signaled within the slice header, encoding of information in the slice header may not be permitted.
[0042] A second syntax element indicating whether the information is in the picture header may be encoded, and when the first syntax element indicates that the picture header is signaled in the slice header, it is a requirement for bitstream conformity that the second syntax element indicates that the information is in the picture header.
[0043] Alternatively, if the first syntax element indicates that the picture header is signaled within the slice header, signaling of the picture header information is not permitted.
[0044] A second syntax element indicating whether the information is in the picture header may be encoded, and when the first syntax element indicates that the picture header is signaled within the slice header, it is a requirement for bitstream conformance that the second syntax element indicates that the information is within the picture header.
[0045] The second syntax element may be information of a picture parameter set flag or a picture header flag. When the flag is set, information is signaled within the picture header, and when it is not set, information is signaled within the slice header or does not exist.
[0046] The information may include one or more of quantization parameter value information, reference picture list information, deblocking filter information, sample adaptive offset (SAO) information, weighted prediction information, and adaptive loop filter (ALF) information. Optionally, the information includes all information that can be signaled in the picture header and the slice header. For example, it is all of quantization parameter value information, reference picture list information, deblocking filter information, sample adaptive offset (SAO) information, weighted prediction information, and adaptive loop filter (ALF) information.
[0047] The reference picture list information includes one or more of slice_collocated_from_l0_flag, slice_collocated_ref_idx, ph_collocated_from_l0_flag, and ph_collocated_ref_idx.
[0048] Optionally, when the picture header is signaled in the slice header, the number of weights for weighted prediction may be limited.
[0049] According to a third further aspect of the present invention, there is provided a method for decoding a bitstream including video data corresponding to one or more slices, the bitstream including a picture header including syntax elements to be used when decoding one or more slices, and a slice header including syntax elements to be used when decoding a slice, the method comprising parsing, in the slice header, a syntax element indicating whether the picture header is signaled within the slice header, wherein the APS_ID related syntax element of ALF is parsed before a syntax element indicating whether the picture header is signaled within the slice header. The APS_ID related information of ALF may be parsed in the vicinity or at the beginning of the slice header.
[0050] According to a fourth further aspect of the present invention, there is provided a method for encoding video data including one or more slices into a bitstream, the bitstream including a picture header including syntax elements to be used when decoding one or more slices, and a slice header including syntax elements to be used when decoding a slice, the method comprising parsing, in the slice header, a syntax element indicating whether the picture header is signaled in the slice header, wherein the APS_ID related syntax element of ALF is encoded prior to a syntax element indicating whether the picture header is signaled within the slice header. The APS_ID related information of ALF may be encoded in the vicinity or at the beginning of the slice header.
[0051] According to a fifth further aspect of the present invention, a method for decoding video data from a bitstream is provided. The bitstream includes video data corresponding to one or more slices. The bitstream includes a picture header including syntax elements used when decoding one or more slices, and a slice header including syntax elements used when decoding a slice. Decoding includes, when to be signaled in the slice header, restricting the number of weights signaled for the weighted prediction mode, and decoding the bitstream using the syntax elements. When signaling the picture header within the slice header, parsing of information that can be signaled within the slice header and within the picture header may be permitted only in either the slice header or the picture header.
[0052] According to a sixth further aspect of the present invention, a method for encoding video data from a bitstream is provided. The bitstream includes video data corresponding to one or more slices. The bitstream includes a picture header including syntax elements used when decoding one or more slices, and a slice header including syntax elements used when decoding a slice. Encoding includes, when signaling in the slice header, restricting the number of weights encoded for the weighted prediction mode, and encoding the bitstream using the syntax elements. When signaling the picture header within the slice header, encoding of information that can be signaled within the slice header and within the picture header may be permitted only in either the slice header or the picture header.
[0053] According to a seventh further aspect of the present invention, a decoder for decoding video data from a bitstream is provided. The decoder is configured to execute the method according to any of the first, third, or fifth further aspects.
[0054] According to an eighth further aspect of the present invention, there is provided an encoder for encoding video data into a bitstream, the encoder being configured to perform the method of any of the second, fourth or sixth further aspects.
[0055] According to a ninth further aspect of the present invention, there is provided a computer program which, when executed, causes a method according to any of the first to sixth further aspects to be performed. The program may be provided on its own, or may be carried on, by or in a carrier medium. The carrier medium may be non-transitory, for example a storage medium, in particular a computer-readable storage medium. The carrier medium may also be transitory, for example a signal or other transmission medium. The signal may be transmitted via any suitable network including the Internet. Further features of the present invention are characterized by the independent and dependent claims.
[0056] According to a first further alternative aspect of the present invention, there is provided a method of decoding video data from a bitstream, the bitstream including video data corresponding to one or more slices, the bitstream including a picture header including syntax elements used when decoding one or more slices, and a slice header including syntax elements used when decoding a slice. When the bitstream includes a first syntax element having a value indicating that information that can be signaled in the picture header or slice header is signaled in the picture header, the bitstream is constrained to also include a second syntax element having a value indicating that there is no slice header, and the method includes decoding the bitstream using the syntax elements. The bitstream may be constrained to conform to a video coding standard. In an embodiment, the video coding standard is a multi-purpose video coding standard. The second syntax element may be a picture header in the slice header syntax element. The first syntax element may be a flag indicating coding within one or more of quantization parameter value information, reference picture list information, deblocking filter information, sample adaptive offset (SAO) information, weighted prediction information, and adaptive loop filter (ALF) information within a picture header. The bitstream constraints may be applied systematically. For example, in an embodiment, the constraints are applied to any or all of sequences, pictures, and slices within the bitstream.
[0057] According to a second further aspect of the present invention, there is provided a method for encoding or decoding video data into or from a bitstream, the method including applying a constraint related to whether a picture header is permitted in a slice header based on whether information that can be signaled in a picture header or slice header is signaled in the picture header. According to a third further aspect of the present invention, there is provided an apparatus configured to execute the method of the second further aspect. According to a fourth further aspect of the present invention, there is provided a computer program including instructions which, when executed, cause the method of the second further aspect to be executed.
[0058] Any feature in one aspect of the present invention may be applied, in any suitable combination, to other aspects of the present invention. In particular, aspects of the method may be applied to aspects of the apparatus and vice versa.
[0059] Furthermore, features implemented in hardware may be implemented in software and vice versa. References herein to software and hardware features are to be construed accordingly.
[0060] Any apparatus feature described herein may also be provided as a method feature and vice versa. As used herein, means-plus-function features may alternatively be expressed from the perspective of their corresponding structures, such as a properly programmed processor and associated memory.
[0061] Also, it is to be understood that particular combinations of the various features described and defined in any aspect of the present invention can be implemented and / or supplied and / or used independently.
Brief Description of the Drawings
[0062] Next, by way of example, reference is made to the accompanying drawings.
[0063]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0064] FIG. 1 relates to an encoding structure used in the High Efficiency Video Coding (HEVC) video standard. Video sequence 1 is composed of consecutive digital images i, and each such digital image is represented by one or more matrices. The coefficients of the matrix represent pixels.
[0065] The image 2 of the sequence can be divided into a plurality of slices 3. The slices may, in some cases, constitute the entire image. These slices are divided into a plurality of non-overlapping coding tree units (CTUs). A coding tree unit (CTU) is the basic processing unit of the High Efficiency Video Coding (HEVC) video standard and conceptually structurally corresponds to the macroblock unit used in some previous video standards. A CTU may also be called a largest coding unit (LCU). A CTU has luma and chroma component parts, and each of the component parts is called a coding tree block (CTB). These different color components are not shown in FIG. 1.
[0066] A CTU generally has a size of 64 pixels × 64 pixels. Each CTU can be repeatedly divided into a plurality of coding units (CUs) 5 of smaller variable sizes using a quadtree (quad tree) decomposition.
[0067] A coding unit is a basic coding element and is composed of two types of sub-units called a prediction unit (PU) and a transform unit (TU). The maximum size of a PU or TU is equal to the size of the CU. The prediction unit corresponds to the division of the CU for predicting pixel values. When dividing a CU into PUs, various divisions are possible, such as dividing it into four square PUs or two rectangular PUs as shown at 606. The transform unit is the basic unit that undergoes spatial transformation using DCT. A CU can be divided into a plurality of TUs based on a quadtree representation 607.
[0068] Each slice is embedded in one Network Abstraction Layer (NAL) unit. Further, the encoding parameters of the video sequence are stored in a dedicated NAL unit called a parameter set. In HEVC and H.264 / AVC, two types of parameter set NAL units are employed: First, the Sequence Parameter Set (SPS) NAL unit collects all parameters that do not change throughout the video sequence. Typically, it deals with the encoding profile, the size of multiple video frames, and other multiple parameters. Next, the Picture Parameter Set (PPS) NAL unit contains multiple parameters that may vary from one picture (or frame) of the sequence to another. HEVC also includes a Video Parameter Set (VPS) NAL unit that contains multiple parameters describing the overall structure of the bitstream. The VPS is a new type of parameter set defined in HEVC and is applied to all layers of the bitstream. One layer may contain multiple sub-layers, and the version 1 bitstream is restricted to all being in one layer. HEVC has specific layer extensions for scalability and multi-view, which are backward-compatible version 1 base layers that enable multiple layers.
[0069] FIG. 2 is a diagram showing a data communication system in which one or more embodiments of the present invention can be implemented. The data communication system is composed of a transmitting device (in this case, server 201) operable to transmit data packets of a data stream to a receiving device (in this case, client terminal 202) via a data communication network 200. The data communication network 200 may be a wide area network (WAN) or a local area network (LAN). Such a network may be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a hybrid network including a plurality of different networks. In a particular embodiment of the present invention, the data communication system may be a digital television broadcast system in which the server 201 transmits the same data content to a plurality of clients.
[0070] The data stream 204 provided by the server 201 may be composed of multimedia data representing video data and audio data. The audio and video data streams may be acquired by the server 201 using a microphone and a camera, respectively, in some embodiments of the present invention. In some embodiments, the data stream may be stored in the server 201, or the server 201 may receive it from another data provider, or it may be generated by the server 201. The server 201 is provided with an encoder, particularly for encoding video and audio streams, and is for providing a compressed bit stream for transmission, which is a more compact representation of the data presented as input to the encoder.
[0071] In order to obtain a better ratio of the quality of the transmitted data to the amount of the transmitted data, the compression of the video data can be performed, for example, according to the HEVC format or the H.264 / AVC format.
[0072] Client 202 receives the transmitted bitstream, decodes the reconstructed bitstream, plays the video image on the display device, and plays the audio data on the loudspeaker.
[0073] Although a streaming scenario is considered in the example of FIG. 2, it will be understood that in some embodiments of the present invention, the data communication between the encoder and the decoder may be performed using a media storage device such as an optical disk.
[0074] In one or more embodiments of the present invention, the video image is transmitted together with representative data of a compensation offset applied to a plurality of reconstructed pixels of the image to provide a plurality of filtered pixels in the final image.
[0075] FIG. 3 schematically shows a processing device 300 configured to implement at least one embodiment of the present invention. The processing device 300 may be a device such as a microcomputer, a workstation, or a lightweight portable device. The device 300 includes - a central processing unit 311 such as a microprocessor (denoted as CPU); - a read-only memory 306 (denoted as ROM) for storing a computer program for implementing the present invention; - a random access memory 312 (denoted as RAM) for storing executable code of the method of the embodiment of the present invention, as well as registers adapted to record variables and parameters necessary for implementing a method of encoding a sequence of digital images and / or decoding a bitstream according to an embodiment of the present invention; and - a communication interface 302 connected to a communication network 303 through which the digital data to be processed is transmitted and received and includes a communication bus 313 connected thereto.
[0076] Optionally, the device 300 may include the following components: - A computer program for implementing the method of one or more embodiments of the present invention, and a data storage means 304 (hard disk) for storing data used or generated during the implementation of one or more embodiments of the present invention; - A disk drive 305 for a disk 306, the disk drive being adapted to read data from or write data to the disk 306; - A screen 309 for displaying data and / or functioning as a graphical interface with the user by means of a keyboard 310 or any other pointing means.
[0077] The apparatus 300 can be connected to various peripheral devices such as, for example, a digital camera 320 and a microphone 308, each of which is connected to an input / output card (not shown) so as to supply multimedia data to the apparatus 300.
[0078] The communication bus provides communication and interoperability between various elements included in or connected to the apparatus 300. The representation of the bus is not limiting, and in particular, the central processing unit is operable to transmit instructions directly to any element of the apparatus 300 or through another element of the apparatus 300.
[0079] The disk 306 can be replaced by any information medium such as, for example, a compact disk (CD-ROM), ZIP disk, memory card, whether rewritable or not, and generally is an information storage means readable by a microcomputer or microprocessor, removable regardless of whether it is incorporated in the apparatus, and capable of storing one or more programs for implementing a method of encoding a sequence of digital images according to the present invention and / or a method of decoding a bitstream, and can be any medium as long as it is suitable for such storage.
[0080] The executable code can be stored either in the read-only memory 306, on the hard disk 304, or on a removable digital medium such as, for example, the disk 306. According to a variant, the executable code of the program can be received by means of the communication network 303 via the interface 302 so as to be stored in one of the storage means of the device 300, for example the hard disk 304, before being executed.
[0081] The central processing unit 311 is adapted to control and direct the execution of the instructions of the program or the program's software code according to the invention, or of the instructions stored in one of the aforementioned storage means. At power-on, the program or programs stored, for example, on the hard disk 304 or in the non-volatile memory within the read-only memory 306 are transferred to the random access memory 312, and then the registers for storing the program or the executable code of the program, and the variables and parameters necessary for implementing the present invention are stored.
[0082] In this embodiment, the device is a programmable device that uses software to implement the present invention. However, alternatively, the present invention can be implemented in hardware (for example, in the form of an application-specific integrated circuit, i.e., an ASIC).
[0083] FIG. 4 is a block diagram of an encoder according to at least one embodiment of the present invention. The encoder is represented by a plurality of connected modules, each module being adapted to implement at least one corresponding step of a method for encoding an image of a sequence of images according to at least one embodiment of the present invention, in the form of programming instructions executed, for example, by the CPU 311 of the device 300.
[0084] The original sequence of a plurality of digital images (i0~in) 401 is received as input by an encoder 400. Each digital image is represented by a set of samples known as pixels.
[0085] After the encoding process is performed, a bitstream 410 is output by the encoder 400. The bitstream 410 includes a plurality of encoded units or slices, and each slice includes a slice header for transmitting the encoded value of the encoding parameters used for the encoding of the slice, and a slice body including the encoded video data.
[0086] The input digital images (i0~in) 401 are divided into pixel blocks by a module 402. The blocks correspond to image portions and may have variable sizes (e.g., 4x4, 8x8, 16x16, 32x32, 64x64, 128x128 pixels and some rectangular block sizes can also be considered). The encoding mode is selected for each input block. There are two families of encoding modes: an encoding mode based on spatial prediction (intra prediction) and an encoding based on temporal prediction (inter encoding, merge, SKIP). The possible encoding modes are tested.
[0087] A module 403 executes an intra prediction process, and a predetermined block to be encoded is predicted by a predictor calculated from a plurality of pixels in the vicinity of the block to be encoded. The difference between the display of the selected intra predictor and the given block and the predictor is encoded to provide a residual if intra encoding is selected.
[0088] Time prediction is performed by the motion estimation module 404 and the motion compensation module 405. First, one reference image is selected from a set of reference images 416, and a portion of the reference image (also called a reference region or image portion, which is the region closest to a predetermined block to be encoded) is selected by the motion estimation module 404. Then, the motion compensation module 405 uses the selected region to predict the block to be encoded. The difference between the selected reference region and the given block is also called the residual block and is calculated by the motion compensation module 405. The selected reference region is indicated by a motion vector.
[0089] Thus, in both cases (spatial prediction and time prediction), the residual is calculated by subtracting the predicted value from the original block.
[0090] In intra prediction performed by module 403, the prediction direction is encoded. In time prediction, at least one motion vector is encoded. In inter prediction performed by modules 404, 405, 416, 418, 417, at least one motion vector or data for identifying such a motion vector is encoded for time prediction.
[0091] When inter prediction is selected, the relative information of the motion vector and the residual block is encoded. Further, to reduce the bit rate, assuming that the motion is homogeneous, the motion vector is encoded by the difference with respect to the motion vector predictor. The motion vector predictors of the set of motion information predictors are obtained from the motion vector field 418 by the motion vector prediction and encoding module 417.
[0092] The symbolizer 400 further comprises a selection module 406 for selecting an encoding mode by applying an encoding cost criterion such as a rate-distortion criterion. To further reduce redundancy, a transformation (such as DCT) is applied to the residual block by the transformation module 407, and the resulting transformed data is then quantized by the quantization module 408 and entropy-encoded by the entropy encoding module 409. Finally, the encoded residual block of the currently encoded block is inserted into the bitstream 410.
[0093] Also, the symbolizer 400 performs decoding of the encoded image to generate a reference image for motion estimation of subsequent images. Thereby, the symbolizer and the decoder that receive the bitstream can have the same reference frame. The inverse quantization module 411 performs inverse quantization of the quantized data and then performs inverse transformation by the inverse transformation module 412. The inverse intra prediction module 413 determines a predictor to be used for a given block using prediction information, and the inverse motion compensation module 414 actually adds the residual obtained by the module 412 to the reference region obtained from the reference image set 416.
[0094] Thereafter, post-filtering is applied by the module 415 to filter the frame of the reconstructed pixels. In an embodiment of the present invention, an SAO loop filter is used and a compensation offset is added to the pixel value of the reconstructed pixels of the reconstructed image.
[0095] FIG. 5 is a block diagram of a decoder 60 that can be used to receive data from an encoder according to an embodiment of the present invention. The decoder is represented by a plurality of connected modules, and each module is adapted to perform the corresponding steps of the method implemented by the decoder 60, for example, in the form of programming instructions executed by the CPU 311 of the device 300.
[0096] Decoder 60 receives a bitstream 61 including a plurality of encoded units each including a header containing information on encoding parameters and a body containing encoded video data. The structure of the bitstream in VVC will be described in detail below with reference to FIG. 6. As described with respect to FIG. 4, the encoded video data is entropy encoded, and the index of the motion vector predictor is encoded with a predetermined number of bits for a predetermined block. The received encoded video data is entropy decoded by module 62. Then, the residual data is inverse quantized by module 63, and then inverse transformed by module 64 to obtain pixel values.
[0097] Also, mode data indicating the encoding mode is entropy decoded, and based on the mode, intra-type decoding or inter-type decoding is performed on the encoded image data block.
[0098] In the intra mode, an intra predictor is determined by intra inverse prediction module 65 based on the intra prediction mode specified in the bitstream.
[0099] In the inter mode, motion prediction information is extracted from the bitstream to find the reference area used by the encoder. The motion prediction information is composed of a reference frame index and a motion vector residual. The motion vector prediction information is added to the motion vector residual in order to obtain a motion vector by motion vector decoding module 70.
[0100] The motion vector decoding module 70 applies motion vector decoding to each current block encoded by motion prediction. When the index of the motion vector predictor for the current block is obtained, the actual value of the motion vector associated with the current block can be decoded and used to apply inverse motion compensation by module 66. The reference picture portion indicated by the decoded motion vector is extracted from the reference picture 68 to apply inverse motion compensation 66. The motion vector field data 71 is updated with the decoded motion vectors for use in inverse prediction of subsequent decoded motion vectors.
[0101] Finally, the decoded block is obtained. Post-filtering is applied by the post-filtering module 67. The decoder 60 finally provides the decoded video signal 69.
[0102] FIG. 6 is a diagram showing the configuration of a bitstream in an exemplary encoding system VVC described in JVET-Q2001-vD.
[0103] The bitstream 61 according to the VVC encoding method is composed of an ordered sequence of syntax elements and encoded data. The syntax elements and encoded data are arranged in network abstraction layer (NAL) units 601 to 608. There are various types of NAL units. The network abstraction layer provides the ability to encapsulate the bitstream in different protocols such as RTP / IP (Real-Time Protocol / Internet Protocol) and ISO base media file format. Also, the network abstraction layer provides a framework for packet loss recovery.
[0104] NAL units are divided into video coding layer (VCL) NAL units and non-VCL_NAL units. The VCL_NAL units contain the actually encoded video data. The non-VCL_NAL units contain additional information. This additional information may be parameters necessary for decoding the encoded video data or supplementary data that may improve the usability of the decoded video data. NAL unit 606 corresponds to a slice and constitutes the VCL_NAL unit of the bitstream.
[0105] The different NAL units 601 - 605 correspond to different parameter sets, and these NAL units are non-VCL_NAL units. The decoder parameter set (DPS) NAL unit 301 contains parameters that are constant for a given decoding process. The video parameter set (VPS) NAL unit 602 contains parameters defined for the entire video and thus for the entire bitstream. The DPS_NAL unit may define parameters that are more static than those of the VPS. That is, the parameters of the DPS are less frequently changed than those of the VPS.
[0106] The sequence parameter set (SPS) NAL unit 603 contains parameters defined for the video sequence. In particular, the SPS_NAL unit can define the sub-picture layout of the video sequence and related parameters. The parameters related to each sub-picture specify the encoding constraints applied to the sub-picture. In particular, a flag indicating that the temporal prediction between sub-pictures is restricted to data coming from the same sub-picture is included. Another flag can enable or disable the loop filter across sub-picture boundaries.
[0107] Picture Parameter Set (PPS) NAL unit 604. The PPS contains parameters defined for a picture or a group of pictures. The Adaptation Parameter Set (APS) NAL unit 605 typically contains parameters for an adaptation loop filter (ALF) or a reshape model (or a luma mapping with chroma scaling (LMCS) model) or a scaling matrix used at the slice level.
[0108] The syntax of the PPS proposed in the current version of VVC consists of syntax elements that specify the size of the picture in luma samples and further divide each picture into tiles and slices.
[0109] The PPS contains syntax elements for determining the slice positions within a frame. Since sub-pictures form rectangular regions within a frame, it is possible to determine the set of slices, parts of tiles, and tiles to which a sub-picture belongs in the parameter set NAL unit. The PPS has an ID mechanism similar to the APS and limits the transmission volume of the same PPS.
[0110] The main difference between the PPS and the picture header lies in their transmission. The PPS is generally transmitted for a group of pictures, while the PH is systematically transmitted for each picture. Therefore, the PPS contains parameters that are constant across multiple pictures compared to the PH.
[0111] In addition, the bitstream may contain Supplemental Enhancement Information (SEI) NAL units (not shown in FIG. 6). The period of appearance of these parameter sets in the bitstream is variable. The VPS defined for the entire bitstream may appear only once in the bitstream. Conversely, the APS defined for a slice may occur only once for each slice of each picture. In practice, different slices may depend on the same APS, and thus, generally, there are fewer APSs than slices in each picture. In particular, the APS is defined in the picture header. However, the APS of ALF can be refined in the slice header.
[0112] The Access Unit Delimiter (AUD) NAL unit 607 separates two access units. An access unit is a set of NAL units that can constitute one or more coded pictures with the same decoding timestamp. This optional NAL unit contains only one syntax element, pic_type, in the current VVC specification. This syntax element indicates the values of slice_type for all slices of the coded pictures within the AU. If pic_type is 0, the AU contains only intra slices. If it is 1, P and I slices are included. If it is 2, it includes any of B, P, and intra slices. This NAL unit contains only one syntax element, pic-type.
[0113] [Table 1]
[0114] In JVET-Q2001-vD, pic_type is defined as follows: "pic_type indicates that the slice_type values of all slices in the AU containing the AU delimiter NAL unit are members of the set indicated by the pic_type values in Table 2. The value of pic_type must be 0, 1, or 2 in a bitstream compliant with this version of this standard. Other values of pic_type are reserved for future use by ITU-T and ISO / IEC. Decoders compliant with this specification shall ignore reserved values of pic_type."
[0115] rbsp_trailing_bits() is a function that adds bits to align at the end of a byte. Therefore, when this function is executed, the amount of the bitstream to be parsed becomes an integer number of bytes.
[0116] [Table 2]
[0117] The PH_NAL unit 608 is a picture header NAL unit that groups parameters common to a set of slices of one coded picture. A picture can refer to one or more APSs to indicate the AFL parameters, resharper model, and scaling matrix used by the slices of the picture.
[0118] Each of the VCL_NAL units 606 contains a slice. A slice may correspond to the entire picture or a sub-picture, a single tile or multiple tiles or a part of a tile. For example, the slice in Figure 3 contains multiple tiles 620. A slice consists of a slice header 610 and a raw byte sequence payload (RBSP) 611 containing coded pixel data coded as coded blocks 640.
[0119] The syntax of the PPS proposed in the current version of VVC is composed of syntax elements that specify the picture size in luma samples and further divide each picture into tiles and slices.
[0120] The PPS contains syntax elements for determining the slice positions within a frame. Since a sub-picture forms a rectangular region within a frame, it is possible to determine the set of slices, parts of tiles, or tiles to which a sub-picture belongs in a parameter set NAL unit.
[0121] NAL unit slice As shown in Table 3, the NAL unit slice layer contains a slice header and slice data.
[0122]
Table 3
[0123] APS The adaptation parameter set (APS) NAL unit 605 is defined in Table 4 showing syntax elements.
[0124] As shown in Table 4, there are three types of APS given by the aps_params_type syntax element: · ALF_AP: For ALF parameters · LMCS_APS: For LMCS parameters · SCALING_APS: For scaling list relative parameters
[0125]
Table 4
[0126] These three types of APS parameters will be described in order below.
[0127] APS of ALF The parameters of the ALF are described in the data syntax elements of the adaptive loop filter (Table 5). First, four flags specify the presence or absence of the luma and chroma ALF filters, and the presence or absence of the CC-ALF (cross-component - adaptive loop filter) for the Cb and Cr components. When the luma filter flag is valid, another flag is decoded (alf_luma_clip_flag) to know whether the clip value is signaled. Next, the number of signaled filters is decoded using the alf_luma_num_filters_signalled_minus1 syntax element. If necessary, the syntax element "alf_luma_coeff_delta_idx" representing the ALF coefficient delta is decoded for each valid filter. Thereafter, the absolute value and sign of each coefficient of each filter are decoded.
[0128] When alf_luma_clip_flag is valid, the clip index of each coefficient of each valid filter is decoded.
[0129] Similarly, the chroma coefficients of the ALF are decoded as necessary.
[0130] When the CC-ALF is valid for Cr or Cb, the number of filters is decoded (alf_cc_cb_filters_signalled_minus1 or alf_cc_cr_filters_signalled_minus1), and the relevant coefficients are decoded (alf_cc_cb_mapped_coeff_abs and alf_cc_cb_coeff_sign or alf_cc_cr_mapped_coeff_abs and alf_cc_cr_coeff_sign respectively).
[0131]
Table 5
[0132] LMCS Syntax Elements for Both Luma Mapping and Chroma Scaling Table 6 below gives all LMCS syntax elements that are encoded in the adaptation parameter set (APS) syntax structure when the aps_params_type parameter is set to 1 (LMCS_APS). Up to four LMCS_APSs can be used in the encoded video sequence, but only a single LMCS_APS can be used for a given picture.
[0133] These parameters are used to construct the luma forward and inverse mapping functions and the chroma scaling function.
[0134] [Table 6]
[0135] Scaling List APS The scaling list provides the possibility to update the quantization matrix used for quantization. In VVC, this scaling matrix is signaled in the APS as described by the scaling list data syntax element (scaling list data syntax in Table 7). The first syntax element specifies whether the scaling matrix is used for the LFNST (Low Frequency Non-Separable Transform) tool based on the flag scaling_matrix_for_lfnst_disabled_flag. The second is specified when a scaling list is used for the chroma component (scaling_list_chroma_present_flag). Then, the syntax elements necessary to construct the scaling matrix are decoded (scaling_list_copy_mode_flag,scaling_list_pred_mode_flag,scaling_list_pred_id_delta,scaling_list_dc_coef,scaling_list_delta_coef).
[0136]
Table 7
[0137] Picture Header The picture header is transmitted at the beginning of each picture, prior to other slice data. This is very large compared to the headers in previous draft standards. A complete description of all these parameters is given in JVET-Q2001-vD. Table 9 shows these parameters in the current picture header decoding syntax.
[0138] The relevant syntax elements that can be decoded are as follows: · How this picture is used, whether it is a reference frame or not · The type of the picture · Output frame · Picture number · Sub-picture usage (if necessary) · List of reference pictures (if necessary) · Color plane (if necessary) · Partition update when the overwrite flag is valid · Delta QP parameter (if necessary) · Motion information parameter (if necessary) · ALF parameter (if necessary) · SAO parameter (if necessary) · Quantization parameter (if necessary) · LMCS parameter (if necessary) · Scaling list parameter (if necessary) · Picture header extension (if necessary) · And so on...
[0139] Picture “Type” The initial flag is gdr_or_irap_pic_flag, which indicates whether the current picture is a resynchronization picture (IRAP or GDR). If this flag is true, the gdr_pic_flag is decoded to know whether the current picture is an IRAP or GDR picture.
[0140] After that, the ph_inter_slice_allowed_flag is decoded to identify that inter-slice is allowed.
[0141] If allowed, the flag ph_intra_slice_allowed_flag is decoded to know whether intra-slice is allowed in the current picture.
[0142] Next, the non_reference_picture_flag, the ph_pic_parameter_set_id indicating the PPS_ID, and the picture order count ph_pic_order_cnt_lsb are decoded. The picture order count indicates the number of the current picture.
[0143] If the picture is a GDR or IRAP picture, the no_output_of_prior_pics_flag is decoded. Also, if the picture is a GDR, the recovery_poc_cnt is decoded. And if necessary, the ph_poc_msb_present_flag and poc_msb_val are decoded.
[0144] ALF After these parameters that describe important information about the current picture, if ALF is enabled at the SPS level and if ALF is enabled at the picture header level, the set of APS_ID syntax elements of ALF is decoded. ALF is enabled at the SPS level by the sps_alf_enabled_flag flag. Also, ALF signaling is enabled at the picture header level because alf_info_in_ph_flag is 1, and in other cases (alf_info_in_ph_flag is 0), ALF signaling is signaled at the slice level.
[0145] The alf_info_in_ph_flag is defined as follows: "An alf_info_in_ph_flag of 1 specifies that the ALF information exists in the PH syntax structure and does not exist in the slice header that refers to a PPS that does not contain the PH syntax structure. An alf_info_in_ph_flag of 0 specifies that the ALF information does not exist in the PH syntax structure and may exist in the slice header that refers to a PPS that does not contain the PH syntax structure."
[0146] First, the ph_alf_enabled_present_flag is decoded to determine whether the ph_alf_enabled_flag should be decoded. If the ph_alf_enabled_flag is valid, ALF is enabled for all slices of the current picture.
[0147] If ALF is enabled, the amount of APS_id of ALF for luma is decoded using the pic_num_alf_aps_ids_luma syntax element. For each APS_ID, the APS_ID value for luma is decoded as "ph_alf_aps_id_luma".
[0148] For chroma, the syntax element ph_alf_chroma_idc is decoded to determine whether ALF is valid for chroma, valid only for Cr, or valid only for Cb. If it is valid, the value of the APS_ID for chroma is decoded using the ph_alf_aps_id_chroma syntax element. The APS_ID of the CC-ALF method is decoded if necessary for the Cb and / or Cr components.
[0149] LMCS If LMCS is valid at the SPS level, the set of APS_ID syntax elements of LMCS is then decoded. First, ph_lmcs_enabled_flag is decoded to determine whether LMCS is valid for the current picture. If LMCS is valid, the ID value ph_lmcs_aps_id is decoded. For chroma only, ph_chroma_residual_scale_flag is decoded to enable or disable the method for chroma.
[0150] Scaling list If the scaling list is valid at the SPS level, the set of APS_IDs of the scaling list is then decoded. ph_scaling_list_present_flag is decoded to determine whether the scaling matrix is valid for the current picture. Then, the value of the APS_ID, ph_scaling_list_aps_id, is decoded.
[0151] Subpicture The subpicture parameters are valid when they are enabled in the SPS and subpicture ID signaling is disabled. They also include some information about the virtual boundaries. Eight syntax elements are defined for the subpicture parameters. ·ph_virtual_boundaries_present_flag ·ph_num_ver_virtual_boundaries ·ph_virtual_boundaries_pos_x[i] ·ph_num_hor_virtual_boundaries ·ph_virtual_boundaries_pos_y[i]
[0152] Output flag After these sub-picture parameters, if it exists, pic_output_flag follows.
[0153] Reference picture list When the reference picture list is signaled in the picture header (rpl_info_in_ph_flag is 1), the parameters of the reference picture list ref_pic_lists() are decoded and contain the following syntax elements: ·rpl_sps_flag[] ·rpl_idx[] ·poc_lsb_lt[][] ·delta_poc_msb_present_flag[][] ·delta_poc_msb_cycle_lt[][]
[0154] Partition The set of partition parameters is decoded if necessary and contains the following syntax elements: ·partition_constraints_override_flag ·ph_log2_diff_min_qt_min_cb_intra_slice_luma ·ph_max_mtt_hierarchy_depth_intra_slice_luma ·ph_log2_diff_max_bt_min_qt_intra_slice_luma ·ph_log2_diff_max_tt_min_qt_intra_slice_luma ·ph_log2_diff_min_qt_min_cb_intra_slice_chroma ·ph_max_mtt_hierarchy_depth_intra_slice_chroma ·ph_log2_diff_max_bt_min_qt_intra_slice_chroma ·ph_log2_diff_max_tt_min_qt_intra_slice_chroma ·ph_log2_diff_min_qt_min_cb_inter_slice ·ph_max_mtt_hierarchy_depth_inter_slice ·ph_log2_diff_max_bt_min_qt_inter_slice ·ph_log2_diff_max_tt_min_qt_inter_slice
[0155] Weighted prediction The weighted prediction parameter pred_weight_table() is decoded if the weighted prediction method is valid at the PPS level and the weighted prediction parameter is signaled in the picture header (wp_info_in_ph_flag is 1).
[0156] pred_weight_table() contains the weighted prediction parameters for list L0 and, if bidirectional weighted prediction is valid, the weighted prediction parameters for list L1. When the weighted prediction parameters are transmitted in the picture header, the number of weights for each list is transmitted explicitly as depicted in the pred_weight_table() syntax table (Table 8).
[0157]
Table 8
[0158] Delta QP When the picture is intra, ph_cu_qp_delta_subdiv_intra_slice and ph_cu_chroma_qp_offset_subdiv_intra_slice are decoded as necessary. Also, when inter slices are permitted, ph_cu_qp_delta_subdiv_inter_slice and ph_cu_chroma_qp_offset_subdiv_inter_slice are decoded as necessary. Finally, the picture header extension syntax elements are decoded as necessary.
[0159] All parameters of alf_info_in_ph_flag, rpl_info_in_ph_flag, qp_delta_info_in_ph_flag, sao_info_in_ph_flag, dbf_info_in_ph_flag, wp_info_in_ph_flag are signaled in the PPS.
[0160]
Table 9
[0161] Slice header The slice header is sent at the start of each slice. The slice header contains approximately 65 syntax elements. This is very large compared to the slice headers of previous video coding standards. A complete description of all slice header parameters is given in JVET-Q2001-vD. Table 10 shows these parameters in the current slice header decoding syntax.
[0162]
Table 10
[0163] First, decode picture_header_in_slice_header_flag to know whether picture_header_structure() exists in the slice header.
[0164] Next, if necessary, slice_subpic_id is decoded to determine the sub-picture ID of the current slice. Next, slice_address is decoded to determine the address of the current slice. Next, if the number of tiles in the current picture is greater than 1, num_tiles_in_slice_minus1 is decoded.
[0165] After that, slice_type is decoded.
[0166] If ALF is enabled at the SPS level (sps_alf_enabled_flag) and ALF is signaled in the slice header (alf_info_in_ph_flag is 0), the ALF information is decoded. This includes a flag (slice_alf_enabled_flag) indicating that ALF is enabled for the current slice. If it is enabled, the number of ALF_IDs of APS for luma (slice_num_alf_aps_ids_luma) is decoded, and then the APS_IDs are decoded (slice_alf_aps_id_luma[i]). Next, slice_alf_chroma_idc is decoded to know whether ALF is enabled for the chroma component and for which chroma components. Next, the APS_ID for chroma slice_alf_aps_id_chroma is decoded if necessary. Similarly, to know whether the CC_ALF mode is enabled, slice_cc_alf_cb_enabled_flag is decoded if necessary. If CC_ALF is enabled for Cr and / or Cb, the relevant APS_IDs for Cr and / or Cb are decoded.
[0167] When multiple color planes are transmitted independently (separate_colour_plane_flag is 1), the color_plane_id is decoded. When the reference picture list is not transmitted in the picture header (rpl_info_in_ph_flag is 0), and when the NAL unit is not an IDR or the reference picture list is transmitted for an IDR picture (sps_idr_rpl_present_flag is 1), the reference picture list parameters are decoded; these are the same parameters as those in the picture header.
[0168] When the reference picture list is transmitted in the picture header (rpl_info_in_ph_flag is 1), or when the NAL unit is not an IDR or the reference picture list is transmitted for an IDR picture (sps_idr_rpl_present_flag is 1), and when the number of references for at least one list is greater than 1, the override flag num_ref_idx_active_override_flag is decoded. When this flag is valid, the reference indices for each list are decoded.
[0169] When the slice type is not intra, and optionally, the cabac_init_flag is decoded. When the reference picture list is transmitted in the slice header, slice_collocated_from_l0_flag and slice_collocated_ref_idx are decoded. These data are related to CABAC encoding and motion vector placement. Similarly, when the slice type is not intra, the parameters of the weighted prediction pred_weight_table() are decoded.
[0170] slice_qp_delta is decoded when delta QP information is signaled in the slice header (qp_delta_info_in_ph_flag is 0). If necessary, the syntax elements of slice_cb_qp_offset, slice_cr_qp_offset, slice_joint_cbcr_qp_offset, and cu_chroma_qp_offset_enabled_flag are also decoded.
[0171] When SAO information is signaled in the slice header (sao_info_in_ph_flag is 0) and is enabled at the SPS level (sps_sao_enabled_flag), the enable flags for SAO in both luma and chroma are decoded: slice_sao_luma_flag, slice_sao_chroma_flag. Next, the deblocking filter parameters are decoded if they are signaled in the slice header (dbf_info_in_ph_flag is 0).
[0172] The flag slice_ts_residual_coding_disabled_flag is systematically decoded to know whether the transform skip residual coding method is enabled for the current slice.
[0173] If LMCS was enabled in the picture header (ph_lmcs_enabled_flag is 1), the flag slice_lmcs_enabled_flag is decoded.
[0174] Similarly, if the scaling list was enabled in the picture header (phpic_scaling_list_presentenabled_flag is 1), the flag slice_scaling_list_present_flag is decoded.
[0175] After that, other parameters are decoded as necessary.
[0176] Picture header within the slice header In a specific signaling method, as depicted in Figure 7, the picture header 708 can be signaled within the slice header 710. In that case, there is no NAL unit that contains only the picture header 608. Units 701, 702, 703, 704, 705, 706, 707, 720, 740 correspond to 601, 602, 603, 604, 605, 606, 607, 620, 640 in Figure 6, and thus can be understood from the previous description. This can be enabled in the slice header thanks to the flag picture_header_in_slice_header_flag. Furthermore, when the picture header is signaled within the slice header, the picture is assumed to contain only one slice. Therefore, there is always only one picture header in one picture. Additionally, the flag picture_header_in_slice_header_flag is assumed to have the same value for all pictures in the CLVS (Coded Layer Video Sequence). This means that all pictures between two IRAPs including the first IRAP consist of only one slice per picture.
[0177] The flag picture_header_in_slice_header_flag is defined as follows: "When picture_header_in_slice_header_flag is 1, it indicates that the PH syntax structure exists in the slice header. When picture_header_in_slice_header_flag is 0, it indicates that the PH syntax structure does not exist in the slice header. The requirement for bitstream compliance is that the value of picture_header_in_slice_header_flag is the same for all coded slices within the CLVS. When picture_header_in_slice_header_flag is 1 in a coded slice, it is a compliance requirement that there is no VCL_NAL unit with nal_unit_type being PH_NUT within the CLVS. When picture_header_in_slice_header_flag is 0, all coded slices of the current picture have picture_header_in_slice_header_flag equal to 0, and the current PU has a PH_NAL unit. picture_header_structure() includes the syntax elements of picture_rbsp() excluding the staff bits rbsp_trailing_bits().
[0178] Interaction between picture header in slice header and tool signaling in picture header and slice header QP delta information, reference picture list parameters, deblocking filter parameters, sample adaptation offset parameters, weighted prediction parameters, and ALF parameters can be transmitted in the picture header or slice header according to their respective flags: qp_delta_info_in_ph_flag rpl_info_in_ph_flag dbf_info_in_ph_flag sao_info_in_ph_flag wp_info_in_ph_flag alf_info_in_ph_flag These flags are transmitted within the PPS.
[0179] As shown in Table 11 (Summary of XXX signaling according to the flags picture_header_in_slice_header_flag and xxx_info_in_ph_flag), considering "xxx" for one of the mentioned tools, if xxx_info_in_ph_flag is set to 0 in the current syntax, xxx can be signaled within the slice header. That is, in the following explanations, the abbreviations XXX or xxx are used to refer to the information types that can be signaled in both the picture header and the slice header (examples of which are shown above).
[0180]
Table 11
[0181] If the picture header is within the slice header, this means that there is only one slice for the current picture. Therefore, even if information is sent or made available for both the slice and the picture, since the parameters are the same, the flexibility of the encoder or decoder is not improved. That is, if there is information in the picture header, the corresponding information in the slice header is redundant. Similarly, if the information is within the slice header, the corresponding information in the picture header is redundant. The embodiments described herein simplify the implementation of the decoder by limiting redundancy in the same coding signaling. In particular, in the embodiments, when the picture header is signaled within the slice header, the information is allowed to be in the picture header but not in the slice header.
[0182] Streaming application Some streaming applications extract only specific parts of the bitstream. These extractions can be spatial (as sub-pictures) or temporal (sub-parts of the video sequence). And these extracted parts can be merged with other bitstreams. There are also those that reduce the frame rate by extracting only some frames. Generally, the main purpose of these streaming applications is to provide the maximum quality to the end user using the maximum allowable bandwidth.
[0183] In VVC, in order to reduce the frame rate, the numbering of APS_ID is restricted so that the new APS_ID number of a certain frame cannot be used for the frame at the upper level of the temporal hierarchy. However, in a streaming application that extracts a part of the bitstream, since the frame (as IRAP) does not reset the numbering of APS_ID, it is necessary to track which APS should be retained for a part of the bitstream.
[0184] LMCS (Luma Mapping with Chroma Scaling) The luma mapping with chroma scaling (LMCS) technology is a method for converting sample values applied to a block before applying a loop filter in a video decoder such as VVC.
[0185] LMCS can be divided into two sub-tools. As described later, the first tool is applied to the luma block, and the second tool is applied to the chroma block: 1) The first sub-tool is the in-loop mapping of the luma component based on an adaptive piecewise linear model. The in-loop mapping of the luma component adjusts the dynamic range of the input signal by redistributing the codewords over the entire dynamic range to improve the compression efficiency. In luma mapping, a forward mapping function to the "mapping region" and a corresponding inverse mapping function to return to the "input region" are used. 2) The second sub-tool is related to the chroma component, and chroma-residual scaling that depends on luma is applied. Chroma-residual scaling is designed to compensate for the interaction between the luma signal and its corresponding chroma signal. Chroma-residual scaling depends on the average value of the reconstructed neighboring (above and / or left) luma samples of the current block.
[0186] Similar to other tools for video coding such as VVC, LMCS can be enabled / disabled at the sequence level using SPS flags. Whether chroma residual scaling is enabled is also signaled at the slice level. When luma mapping is enabled, an additional flag indicating whether luma-dependent chroma residual scaling is enabled is signaled. When luma mapping is not used, luma-dependent chroma residual scaling is completely disabled. Also, when the size of the chroma block is 4 or less, luma-dependent chroma residual scaling is always disabled.
[0187] Figure 8 shows the principle of LMCS described above for the luma mapping sub-tool. The hatched blocks in Figure 8 are the functional blocks of the new LMCS that include the forward and backward mapping of the luma signal. When using LMCS, it should be noted that some decoding operations are applied in the "mapped region". These operations are represented by the dashed blocks in this Figure 8. These usually correspond to the reconstruction steps including inverse quantization, inverse transform, luma intra prediction, addition of luma prediction and luma residual. Conversely, the solid blocks in Figure 8 indicate where the decoding process is applied in the original (i.e., unmapped) region, which includes loop filter processes such as deblocking, ALF, SAO, motion compensation prediction, and storage of the decoded image as a reference picture (DPB).
[0188] Figure 9 is a figure similar to Figure 8, but now it is a figure of the chroma scaling sub-tool of the LMCS tool. The hatched blocks in Figure 9 are the functional blocks of the new LMCS that include the luma-dependent chroma scaling process. However, in the case of chroma, there are some important differences compared to the case of luma. Here, for chroma samples, only inverse quantization and the inverse transform represented by the dashed block are executed in the "mapped region". All other steps of chroma intra prediction, motion compensation, and loop filtering are executed in the original region. As shown in Figure 9, only the scaling process exists, and there is no forward and inverse transform process such as luma mapping.
[0189] Luma mapping by piecewise linear model. The luma mapping subtool uses a piecewise linear model. That is, the piecewise linear model separates the dynamic range of the input signal into 16 equal sub-ranges, and for each sub-range, the parameters of its linear mapping are represented by the number of codewords assigned to that range.
[0190] Semantics of luma mapping lmcs_min_bin_idx specifies the minimum bin index to be used in the construction process of luma mapping with chroma scaling (LMCS). The value of lmcs_min_bin_idx ranges from 0 to 15.
[0191] lmcs_delta_max_bin_idx specifies the difference value between the maximum bin index LmcsMaxBinIdx used in the construction process of luma mapping with chroma scaling and 15. The value of lmcs_delta_max_bin_idx must range from 0 to 15. The value of LmcsMaxBinIdx is set to 15 - lmcs_delta_max_bin_idx. The value of LmcsMaxBinIdx must be greater than or equal to lmcs_min_bin_idx.
[0192] The syntax element lmcs_delta_cw_prec_minus1 plus 1 specifies the number of bits to be used in the representation of the syntax lmcs_delta_abs_cw[i].
[0193] The syntax element lmcs_delta_abs_cw[i] specifies the absolute delta codeword value of the i-th bin.
[0194] The syntax element lmcs_delta_sign_cw_flag[i] specifies the sign of the variable lmcsDeltaCW[i]. If lmcs_delta_sign_cw_flag[i] does not exist, it is inferred to be 0.
[0195] LMCS Intermediate Variable Calculation for Luma Mapping Several intermediate variables and data arrays are required to apply the forward and inverse luma mapping processes.
[0196] First, the variable OrgCW is derived as follows: OrgCW = (1 << BitDepth) / 16
[0197] Next, the variables lmcsDeltaCW[i] (i = lmcs_min_bin_idx....LmcsMaxBinIdx) are calculated as follows: lmcsDeltaCW[i] = (1 - 2 * lmcs_delta_sign_cw_flag[i]) * lmcs_delta_abs_cw[i]
[0198] The new variable lmcsCW[i] is derived as follows: When i = 0...lmcs_min_bin_idx - 1, lmcsCW[i] is set to 0. When i = lmcs_min_bin_idx...LmcsMaxBinIdx, the following applies: lmcsCW[i] = OrgCW + lmcsDeltaCW[i] The value of lmcsCW[i] must be in the range from (OrgCW >> 3) to (OrgCW << 3 - 1). For i = LmcsMaxBinIdx + 1...15, lmcsCW[i] is set to 0.
[0199] The variable InputPivot[i] (i = 0...16) is derived as follows: InputPivot[i] = i * OrgCW
[0200] The variables LmcsPivot[i] (i = 0..16), the variable ScaleCoeff[i] and InvScaleCoeff[i] (i = 0..15) are calculated as follows: LmcsPivot[0] = 0; for (i = 0; i <= 15; i++) { LmcsPivot[i + 1] = LmcsPivot[i] + lmcsCW[i] ScaleCoeff[i] = (lmcsCW[i] * (1 << 11) + (1 << (Log2(OrgCW) - 1))) >> (Log2(OrgCW)) if (lmcsCW[i] == 0) InvScaleCoeff[i] = 0 else InvScaleCoeff[i] = OrgCW * (1 << 11) / lmcsCW[i]
[0201] Forward Luma Mapping As shown in Figure 8, when LMCS is applied to luma, luma remap samples predMapSamples[i][j] can be obtained from the prediction samples predSamples[i][j].
[0202] predMapSamples[i][j] is calculated as follows: First, the index idxY at position (i, j) is calculated from the prediction sample predSamples[i][j]. idxY = predSamples[i][j] >> Log2(OrgCW) Then, using the intermediate variables idxY, LmcsPivot[idxY], and InputPivot[idxY] in the 0th section, predMapSamples[i][j] is derived as follows: predMapSamples[i][j] = LmcsPivot[idxY] + (ScaleCoeff[idxY] * (predSamples[i][j] - InputPivot[idxY]) + (1 << 10)) >> 11
[0203] Luma Reconstruction Samples Reconstruction processing is obtained from the predicted luma samples predMapSample[i][j] and the residual luma samples resiSamples[i][j].
[0204] The reconstituted luma image sample recSamples[i][j] is obtained only by adding predMapSample[i][j] to resiSamples[i][j] as follows: recSamples[i][j]=Clip1(predMapSamples[i][j]+resiSamples[i][j])
[0205] In the above relationship, the Clip1 function is a clipping function for ensuring that the reconstituted sample is between 0 and 1<<BitDepth-1.
[0206] Inverse luma mapping When applying inverse luma mapping according to FIG. 8, the following operations are applied to each sample recSample[i][j] of the current block being processed:
[0207] First, an index idxY is calculated from the reconstituted sample recSamples[i][j] at position (i,j). idxY=recSamples[i][j]>>Log2(OrgCW) The inverse mapped luma sample invLumaSample[i][j] is derived as follows: invLumaSample[i][j]=InputPivot[idxYInv]+(InvScaleCoeff[idxYInv]*(recSample[i][j]-LmcsPivot[idxYInv])+(1<<10))>>11 Then, a clipping operation is performed to obtain the final sample: finalSample[i][j]=Clip1(invLumaSample[i][j])
[0208] Chrominance scaling LMCS semantics for chrominance scaling The syntax element lmcs_delta_abs_crs in Table 6 specifies the absolute value of the codeword of the variable lmcsDeltaCrs. The value of lmcs_delta_abs_crs ranges from 0 to 7. If it does not exist, lmcs_delta_abs_crs is inferred to be 0.
[0209] The syntax element lmcs_delta_sign_crs_flag specifies the sign of the variable lmcsDeltaCrs. If it does not exist, lmcs_delta_sign_crs_flag is inferred to be 0.
[0210] LMCS Intermediate Variable Calculation for Chroma Scaling Several intermediate variables are required to apply the chroma scaling process. The variable lmcsDeltaCrs is derived as follows: lmcsDeltaCrs=(1 - 2 * lmcs_delta_sign_crs_flag) * lmcs_delta_abs_crs
[0211] The variable ChromaScaleCoeff[i] (i = 0...15) is derived as follows: if(lmcsCW[i] == 0) ChromaScaleCoeff[i]=(1 << 11) else ChromaScaleCoeff[i]=OrgCW * (1 << 11) / (lmcsCW[i]+lmcsDeltaCrs)
[0212] Chroma Scaling Process In the first step, the variable invAvgLuma is derived to calculate the average luma value of the reconstructed luma samples around the current corresponding chroma block. The average luma is calculated from the left and upper luma blocks surrounding the corresponding chroma block. If there are no samples, the variable invAvgLuma is set as follows: invAvgLuma=1 << (BitDepth - 1)
[0213] Then, based on the intermediate array LmcsPivot[] in the 0th section, the variable idxYInv is derived as follows: For(idxYInv = lmcs_min_bin_idx; idxYInv <= LmcsMaxBinIdx; idxYInv++){ if(invAvgLuma < LmcsPivot[idxYInv + 1]) break } IdxYInv = Min(idxYInv, 15)
[0214] The variable varScale is derived as follows: varScale = ChromaScaleCoeff[idxYInv]
[0215] When the current chroma block is transformed, the reconstructed chroma image sample array recSamples is derived as follows: recSamples[i][j] = Clip1(predSamples[i][j] + Sign(resiSamples[i][j]) * ((Abs(resiSamples[i][j]) * varScale + (1 << 10)) >> 11)) If no transformation is applied to the current block, it is as follows: recSamples[i][j] = Clip1(predSamples[i][j])
[0216] Consideration of the Encoder The basic principle of the LMCS encoder is that first, more codewords are assigned to the range with a lower codeword for the segment of the dynamic range than the average variance. As an alternative form of this, the main target of LMCS is to assign fewer codewords to the dynamic range segment with a higher codeword than the average variance. In this method, the smooth regions of the image are encoded with more codewords than the average, and vice versa.
[0217] All parameters of the LMCS tool stored in the APS (see Table 6) are determined on the encoder side. The LMCS encoder algorithm is based on the evaluation of local luma variance and optimizes the determination of LMCS parameters according to the above basic principle. Then, for the final reconstructed samples of a given block, optimization is performed to obtain the best PSNR metric.
[0218] An embodiment in which information is signaled in the picture header instead of the slice header In an embodiment, the signaling of information that can be signaled within the picture header or the slice header is signaled within the picture header when the picture header is signaled within the slice header and not signaled within the slice header. Also, when signaling the signaling of information that can be signaled within the picture header or the slice header within the slice header, there is an equivalent method that the picture header is not signaled within the slice header. In another equivalent method, when the signaling of information that can be signaled within the picture header or within the slice is signaled within the picture header, the picture header is signaled within the slice header. Table 12 is a table showing an embodiment in which the tool name is replaced with XXX. In this table, when the picture header is signaled within the slice header, parameter signaling is not permitted within the slice header.
[0219] [Table 12]
[0220] In one embodiment, in the semantics of xxx_info_in_ph_flag, the following conditions are added: "When the slice header referring to the PPS contains the PH syntax structure, it is a bitstream compliance requirement that XXX_info_in_ph_flag is 1." And / or 「When XXX_info_in_ph_flag is 0, picture_header_in_slice_header_flag is 0.」
[0221] If the picture header is within the slice header, this means that there is only one slice for the current picture. Therefore, even if information is sent or made available for both the slice and the picture, since the parameters are the same, the flexibility of the encoder or decoder is not improved. That is, if there is information in the picture header, the corresponding information in the slice header is redundant. Similarly, if there is information in the slice header, the corresponding information in the picture header is redundant. Thus, when the picture header is in the slice header, by forcing the condition that the information is in the picture header and not in the slice header, the redundancy of signaling can be limited and the implementation of the decoder can be simplified.
[0222] QP (Quantization Parameter) Delta In an embodiment, when the picture header is signaled within the slice header, QP delta signaling is avoided within the slice header. In an equivalent way, when QP delta signaling is signaled within the slice header, the picture header is not signaled within the slice header. In another equivalent way, when QP delta is signaled in the picture header, the picture header is signaled within the slice header. That is, the above tool XXX is the QP delta.
[0223] Table 13 shows a way in which this can be implemented such that QP delta information signaling is not approved (i.e., not permitted) when the picture header is signaled within the slice header.
[0224]
Table 13
[0225] If the picture header is within the slice header (which means there is only one slice for the current picture), since the parameters are the same, the information that can be transmitted in the slice and picture headers does not increase the flexibility of the encoder or decoder. Therefore, in order to reduce the complexity of decoder implementation, it is more preferable to have only one possible signaling QP delta information in either the slice or picture header to encode the same encoding possibility (in this case, the QP delta parameter (qp_delta_info)).
[0226] In the implementation of the first embodiment, the following conditions may be added to the semantics of qp_delta_info_in_ph_flag: "When the slice header referring to the PPS contains the PH syntax structure, it is a bitstream compliance requirement that qp_delta_info_in_ph_flag is 1." And / or "When qp_delta_info_in_ph_flag is 0, picture_header_in_slice_header_flag is 0."
[0227] In another embodiment, the decoding of the QP delta parameter in the slice header is only permitted when the value of the flag picture_header_in_slice_header_flag is set to 0, as depicted in Table 14. The syntax changes are underlined. In this table, the slice_qp_delta information of the slice header can be decoded when the QP delta information is signaled at the slice level (qp_delta_info_in_ph_flag is 0) and the picture header is not transmitted within the slice header (picture_header_in_slice_header_flag is 0).
[0228]
Table 14
[0229] In an embodiment, the decoding of the QP delta parameter of the picture header is systematically permitted when the value of the flag picture_header_in_slice_header_flag is set to 1, as depicted in Table 15. According to this table, the QP delta information in the slice header can be decoded only when the QP delta information is signaled in the picture header (qp_delta_info_in_ph_flag is 1) or when the picture header is transmitted within the slice header (picture_header_in_slice_header_flag is 1).
[0230] [Table 15]
[0231] Reference Picture List (RPL) In one embodiment, when the picture header is signaled within the slice header, the signaling of the reference picture list is avoided within the slice header. In an equivalent method, the picture header is not signaled in the slice header when the signaling of the reference picture list is signaled in the slice header. In another equivalent method, when the signaling of the reference picture list is performed in the picture header, the picture header is signaled within the slice header. In other words, the above tool XXX is the RPL.
[0232] Table 16 shows an example of this embodiment where RPL signaling is not permitted when signaling the picture header within the slice header.
[0233] [Table 16]
[0234] When the picture header is within the slice header (which means there is only one slice for the current picture), the fact that information can be transmitted in both the slice header and the picture header does not increase the flexibility for the encoder or decoder because the parameters are the same. According to this embodiment, for encoding the same encoding possibility (in this case, reference picture list (RPL) information (rpl_info)), it is better to have only possible signaling of the information in one of the slice or picture headers, so the complexity of decoder implementation is reduced.
[0235] In the embodiment, the following conditions are added in the semantics of rpl_info_in_ph_flag: "When a slice header that refers to the PPS contains a PH syntax structure, it is a bitstream compliance requirement that rpl_info_in_ph_flag be 1." And / or "When rpl_info_in_ph_flag is 0, picture_header_in_slice_header_flag is 0."
[0236] In the embodiment, the decoding of the reference picture list parameters in the slice header is permitted only when the value of the flag picture_header_in_slice_header_flag is set to 0, as depicted in Table 17. In this table, when the reference picture list information is signaled at the slice level (rpl_info_in_ph_flag is 0) and the picture header is not transmitted within the slice header (picture_header_in_slice_header_flag is 0), the ref_pic_lists() information of the slice header can be decoded.
[0237] In the embodiment, as depicted in Table 17, when the picture header is within the slice header, the time parameters slice_collocated_from_l0_flag and slice_collocated_ref_idx cannot be decoded.
[0238]
Table 17
[0239] In the embodiment, the decoding of the RPL parameters of the picture header is systematically permitted when the value of the flag picture_header_in_slice_header_flag is set to 1, as depicted in Table 18. In this table, the RPL information of the picture header can be decoded only when the RPL information is signaled at the picture level (rpl_info_in_ph_flag is 1) or when the picture header is transmitted in the slice header (picture_header_in_slice_header_flag is 1).
[0240] In the embodiment, as depicted in Table 18, when the picture header is within the slice header, the time parameters ph_collocated_from_l0_flag and ph_collocated_ref_idx can be decoded.
[0241]
Table 18
[0242] Deblocking Filter (DBF) In an embodiment, when the signaling of deblocking filter parameters is signaled in a slice header, it is avoided in the slice header when the picture header is signaled within the slice header. In an equivalent method, when the signaling of deblocking filter parameters is signaled within the slice header, the picture header is not signaled within the slice header. In another equivalent method, when the signaling of deblocking filter parameters is performed in the picture header, the picture header is signaled within the slice header. That is, the tool XXX described above is a DBF.
[0243] Table 19 is a diagram showing an implementation according to this embodiment that does not permit DBF signaling when signaling a picture header within a slice header.
[0244]
Table 19
[0245] When the picture header is within the slice header, since there is only one slice for the current picture, the information transmitted in the slice or picture has the same parameters and does not increase the flexibility of the encoder or decoder. Therefore, in order to reduce the implementation complexity of the decoder, it is better that there is only one possible encoding of the DBF information for encoding the same encoding possibility.
[0246] In the embodiment, the following conditions are added to the semantics of dbf_info_in_ph_flag: "When the slice header referring to the PPS includes a PH syntax structure, it is a bitstream compliance requirement that dbf_info_in_ph_flag is 1." And / or "When dbf_info_in_ph_flag is 0, picture_header_in_slice_header_flag is 0."
[0247] In the embodiment, as depicted in Table 20, the DBF parameter of the slice header is authorized only when the value of the flag picture_header_in_slice_header_flag is set to 0. In this table, when the DBF information is signaled at the slice level (dbf_info_in_ph_flag is 0) and the picture header is not transmitted in the slice header (picture_header_in_slice_header_flag is 0), the slice_deblocking_filter_override_flag flag of the slice header can be decoded.
[0248]
Table 20
[0249] In the embodiment, as depicted in Table 21, when the value of the flag picture_header_in_slice_header_flag is set to 1, the DBF parameter of the picture header is systematically authorized. In this table, the DBF information in the slice header can be decoded only when the DBF information is signaled in the picture header (dbf_info_in_ph_flag is 1) or when the picture header is transmitted in the slice header (picture_header_in_slice_header_flag is 1).
[0250]
Table 21
[0251] SAO (Sample Adaptive Offset) In one embodiment, when the SAO signaling is signaled in the slice header, the SAO signaling is avoided in the slice header. In an equivalent method, when the SAO signaling is signaled in the slice header, the picture header is not signaled in the slice header. In another equivalent method, when the SAO signaling is signaled in the picture header, the picture header is signaled in the slice header. That is, the above tool XXX is SAO.
[0252] Table 22 shows this embodiment that does not permit SAO signaling when signaling the picture header in the slice header.
[0253]
Table 22
[0254] When the picture header is in the slice header, since there is only one slice for the current picture, the information transmitted in the slice or picture has the same parameters and does not increase the flexibility for the encoder or decoder. Therefore, in order to reduce the complexity of decoder implementation, it is good to have only one possible encoding of SAO information for encoding the same encoding possibility.
[0255] In the embodiment, the following conditions are added in the semantics of sao_info_in_ph_flag: "When the slice header referring to the PPS contains the PH syntax structure, it is a bitstream compliance requirement that sao_info_in_ph_flag is 1." And / or "When sao_info_in_ph_flag is 0, picture_header_in_slice_header_flag is 0."
[0256] In the embodiment, the decoding of the SAO parameters of the slice header is permitted only when the value of the flag picture_header_in_slice_header_flag is set to 0, as illustrated below.
[0257] Table 23. In this table, when the SAO information is signaled at the slice level (sao_info_in_ph_flag is 0) and the picture header is not transmitted within the slice header (picture_header_in_slice_header_flag is 0), the slicice_sao_luma_flag of the slice header can be decoded.
[0258] [Table 23]
[0259] In the embodiment, the decoding of the SAO parameters of the picture header is systematically approved when the value of the flag picture_header_in_slice_header_flag is set to 1, as depicted in Table 24. In this table, the SAO information of the slice header can be decoded only when the SAO information is signaled in the picture header (sao_info_in_ph_flag is 1) or when the picture header is transmitted in the slice header (picture_header_in_slice_header_flag is 1).
[0260] [Table 24]
[0261] Weighted Prediction (WP) In one embodiment, when the picture header is signaled within the slice header, the signaling of weighted prediction is avoided in the slice header. In an equivalent method, when the signaling of weighted prediction is signaled within the slice header, the picture header is not signaled within the slice header. In another equivalent method, when the signaling of weighted prediction is signaled within the picture header, the picture header is signaled within the slice header. That is, the tool XXX described above is WP.
[0262] As shown in Table 25, the WP parameter can be signaled in the slice header when wp_info_in_ph_flag is set to 0 in the current syntax.
[0263] Table 25 shows an example of this embodiment where WP signaling is not permitted when signaling the picture header within the slice header.
[0264]
Table 25
[0265] When the picture header is within the slice header, since there is only one slice for the current picture, the information transmitted in the slice or picture does not increase the flexibility of the encoder or decoder because the parameters are the same. Therefore, in order to reduce the complexity of decoder implementation, it is good to have only one possible encoding of WP information for encoding the same encoding possibility.
[0266] In the embodiment, the following conditions are added to the semantics of wp_info_in_ph_flag: "When the slice header referring to the PPS contains the PH syntax structure, it is a bitstream compliance requirement that wp_info_in_ph_flag is 1." And / or 「When wp_info_in_ph_flag is 0, picture_header_in_slice_header_flag is 0.」
[0267] In an embodiment, the decoding of the WP parameter of the slice header is permitted only when the value of the flag picture_header_in_slice_header_flag is set to 0, as depicted in Table 26. In this table, when the WP information is signaled at the slice level (wp_info_in_ph_flag is 0) and the picture header is not transmitted within the slice header (picture_header_in_slice_header_flag is 0), the pred_weight_table() function including the weighted prediction parameter can be decoded.
[0268]
Table 26
[0269] In an embodiment, the decoding of the WP parameter of the picture header is systematically permitted when the value of the flag picture_header_in_slice_header_flag is set to 1, as depicted in Table 27. In this table, the WP information can be decoded only when the WP information is signaled within the picture header (wp_info_in_ph_flag is 1) or when the picture header is transmitted within the slice header (picture_header_in_slice_header_flag is 1).
[0270]
Table 27
[0271] ALF In one embodiment, when the ALF signaling is signaled in the slice header, the picture header is avoided in the slice header. In an equivalent method, when the ALF signaling is signaled in the slice header, the picture header is not signaled in the slice header. In another equivalent method, when the ALF signaling is signaled in the picture header, the picture header is signaled in the slice header. That is, the above tool XXX is ALF.
[0272] In fact, currently, when the picture header is signaled in the slice header (picture_header_in_slice_header_flag is 1) and the ALF is signaled in the slice header (alf_info_in_ph_flag is 0), all parameters of the picture header should be parsed before obtaining the APS_ID of the ALF. As a result, all variables such as PPS, SPS, and picture header have to be kept in memory to parse the APS_ID of the ALF, which increases the complexity of parsing for some streaming applications.
[0273] Furthermore, when the picture header is in the slice header, since there is only one slice for the current picture, the information transmitted in the slice or picture has the same parameters, so it does not increase the flexibility of the encoder or decoder. Therefore, in order to reduce the implementation complexity of the decoder, it is better to keep only one signaling with the same encoding possibility. Table 28 shows this embodiment that does not permit ALF signaling when signaling the picture header in the slice header.
[0274]
Table 28
[0275] In the embodiment, the following conditions are added in the semantics of alf_info_in_ph_flag: When the slice header referring to PPS contains the PH syntax structure, it is a requirement for bitstream compliance that alf_info_in_ph_flag be 1. and / or When alf_info_in_ph_flag is 0, picture_header_in_slice_header_flag is 0.
[0276] In an embodiment, the decoding of the ALF parameters of the slice header is permitted only when the value of the flag picture_header_in_slice_header_flag is set to 0, as depicted in Table 29. In this table, the ALF information in the slice header can be decoded only when the ALF is enabled at the SPS level (sps_alf_enabled_flag is 1), the ALF information is signaled at the slice level (alf_info_in_ph_flag is 0), and the picture header is not transmitted in the slice header (picture_header_in_slice_header_flag is 0).
[0277]
Table 29
[0278] In an embodiment, the decoding of the ALF parameters of the picture header is systematically permitted when the value of the flag picture_header_in_slice_header_flag is set to 1, as depicted in Table 30. In this table, the ALF information in the slice header can be decoded only when the ALF is enabled at the SPS level (sps_alf_enabled_flag is 1), the ALF information is signaled at the picture level (alf_info_in_ph_flag is 1), or the picture header is transmitted in the slice header (picture_header_in_slice_header_flag is 1).
[0279]
Table 30
[0280] All tools / parameters In one embodiment, when the picture header is signaled within the slice header, all tools (and / or parameters) that can be signaled within the picture header or within the slice header are restricted to be signaled within the picture header. As mentioned in the above description, in an embodiment, the related tools are as follows. QP delta information, reference picture list, deblocking filter, SAO weighted prediction, ALF. However, other tools are also possible and they can be signaled in both the slice and picture headers.
[0281] This can be expressed by adding the following constraints: "When at least one of the flags rpl_info_in_ph_flag, dbf_info_in_ph_flag, sao_info_in_ph_flag, alf_info_in_ph_flag, wp_info_in_ph_flag, qp_delta_info_in_ph_flag is set to 0, the value of picture_header_in_slice_header_flag shall be 0." And / or, add the following: "When picture_header_in_slice_header_flag is 1, the flags rpl_info_in_ph_flag, dbf_info_in_ph_flag, sao_info_in_ph_flag, wp_info_in_ph_flag, qp_delta_info_in_ph_flag shall be 1."
[0282] Also, add the following constraints to each XXX_info_in_ph_flag: "When the slice header referring to the PPS contains the PH syntax structure, it is a bitstream compliance requirement that XXX_info_in_ph_flag is 1."
[0283] If all of these parameters are signaled with the same header, it is better to have only one signaling for encoding the same encoding possibilities, so that the complexity of decoder implementation can be reduced.
[0284] Embodiment of Signaling Order In one alternative embodiment related to ALF, as depicted in Table 31, the information related to the APS_ID of ALF in the slice header is set before the structure of the picture header. In this embodiment, when ALF is signaled within the slice header and when the intra picture header is signaled in the slice header, the APS_ID can be quickly obtained without parsing all the parameters of the picture header. [[ID=##]] [[ID=##]]
[0285] [[ID=##]] [[ID=##]] [[ID=##]][Table 31] [[ID=##]] [[ID=##]] [[ID=##]] [[ID=##]] [[ID=##]]
[0286] [[ID=##]] [[ID=##]]Embodiment of Weight Number [[ID=##]] [[ID=##]]In the embodiment where the picture header is within the slice header, as depicted in the part of Table 32 of the syntax element, the number of weights of the weighted prediction for each list L0 L1 is decoded. Therefore, the signaling of the weight number can be limited to the picture header. [[ID=##]] [[ID=##]]
[0287] [[ID=##]] [[ID=##]] [[ID=##]][Table 32] [[ID=##]] [[ID=##]] [[ID=##]] [[ID=##]] [[ID=##]]
[0288] [[ID=##]] [[ID=##]]Overview of the Embodiment of Notifying Information within the Slice Header and Not Notifying within the Picture Header [[ID=##]] In one embodiment, for the signaling of information that can be signaled within the picture header or within the slice header, when the picture header is signaled within the slice header, it is signaled in the slice header and not in the picture header. Also, when signaling information that can be signaled within the picture header or within the slice header within the picture header, it is not signaled within the slice header. In another equivalent method, when the signaling of information that can be signaled within the picture header or within the slice header is signaled within the slice header, the picture header is signaled within the slice header. Table 33 shows this embodiment with the tool (or parameter) name replaced by XXX. In this table, when the picture header is signaled within the slice header, signaling of parameters in the picture header is not permitted.
[0289]
Table 33
[0290] In one embodiment, in the semantics of xxx_info_in_ph_flag, the following conditions are added: "When the slice header referring to the PPS contains the PH syntax, it is a requirement for bitstream conformance that XXX_info_in_ph_flag be 0." And / or "When XXX_info_in_ph_flag is 1, picture_header_in_slice_header_flag is 0."
[0291] If the picture header is within the slice header, this means that there is only one slice for the current picture. Thus, even if information is transmitted or made available for both the slice and the picture, since the parameters are the same, the flexibility of the encoder or decoder is not improved. That is, if there is information in the picture header, the corresponding information in the slice header is redundant. Similarly, if the information is within the slice header, the corresponding information in the picture header is redundant. The embodiments described herein simplify the implementation of the decoder by limiting the redundancy in the same coding signaling. In particular, in an embodiment, when the picture header is signaled within the slice header, it is allowed for the information to be within the slice header, but not within the picture header.
[0292] QP Delta In one embodiment, the tool XXX is the QP Delta. In one embodiment, as depicted in Table 33, when the value of the flag picture_header_in_slice_header_flag is set to 1, the decoding of the QP Delta parameter of the slice header is permitted. Also, as shown in Table 33, picture_header_in_slice_header_flag is set to a value equal to 0 when the QP Delta parameter is signaled within the picture header. In another equivalent way, as depicted in Table 33, when the QP Delta parameter is signaled within the slice header, picture_header_in_slice_header_flag is set to 1. In this table, when the QP Delta information is signaled at the slice level (qp_delta_info_in_ph_flag is 0), or when the picture header is transmitted in the slice header (picture_header_in_slice_header_flag is 1), the slice_qp_delta information of the slice header can be decoded.
[0293] This is obtained by adding the following to the semantics: "When the slice header referring to PPS contains PH syntax, it is a bitstream compliance requirement that qp_delta_info_in_ph_flag be 0." and / or "When qp_delta_info_in_ph_flag is 1, picture_header_in_slice_header_flag is 0."
[0294] In an embodiment, as depicted in Table 35, when the value of the flag picture_header_in_slice_header_flag is set to 1, the decoding of the QP delta parameter of the picture header is systematically avoided. In this table, the QP delta information of the picture header can be decoded only when the QP delta information is signaled in the picture header (qp_delta_info_in_ph_flag is 1) and when picture_header_in_slice_header_flag is set to 0.
[0295] RPL (Reference Picture List) In one embodiment, the tool XXX is a reference picture list. In one embodiment, as depicted in Table 33, the decoding of the parameters of the reference picture list in the slice header is authorized only when the value of the flag picture_header_in_slice_header_flag is set to 1. Also, picture_header_in_slice_header_flag is set to 0 when the list parameters of the reference picture are signaled within the picture header, as shown in Table 33. Also, picture_header_in_slice_header_flag is set to 1 when notifying the parameters of the reference picture within the slice header, as shown in Table 33. In this table, when the reference picture list information is signaled at the slice level (rpl_info_in_ph_flag is 0) and the picture header is transmitted in the slice header (picture_header_in_slice_header_flag is 1), the ref_pic_lists() information of the slice header can be decoded.
[0296] In an embodiment, as depicted in Table 34, when the picture header is within the slice header, the time parameters slice_collocated_from_l0_flag and slice_collocated_ref_idx can be decoded.
[0297] In an embodiment, as depicted in Table 35, when the value of the flag picture_header_in_slice_header_flag is set to 1, the decoding of the RPL parameters of the picture header is systematically avoided. In this table, it is possible to decode the RPL information of the picture header only when the RPL information is signaled at the picture level (rpl_info_in_ph_flag is 1) and picture_header_in_slice_header_flag is set to 0.
[0298] In an embodiment, as depicted in Table 35, when the picture header is within the slice header, the time parameters ph_collocated_from_l0_flag and ph_collocated_ref_idx cannot be decoded.
[0299] Deblocking Filter (DBF) In one embodiment, tool XXX is a Deblocking Filter (DBF). In an alternative or additional embodiment, as depicted in Table 33, when the value of the flag picture_header_in_slice_header_flag is set to 1, the DBF parameters of the slice header are approved. Also, as shown in Table 33, when picture_header_in_slice_header_flag is set to 0, the DBF parameters are signaled within the picture header. Also, as shown in Table 33, when the DBF parameters are stored within the slice header, picture_header_in_slice_header_flag is set to 1. In this table, when the DBF information is signaled at the slice level (dbf_info_in_ph_flag is 0), or when the picture header is transmitted within the slice header (picture_header_in_slice_header_flag is 1), the slice_deblocking_filter_override_flag flag of the slice header can be decoded.
[0300] In an embodiment, as depicted in Table 35, when the value of the flag picture_header_in_slice_header_flag is set to 1, the DBF parameters of the picture header are systematically avoided. In this table, the DBF information of the picture header can be decoded only when the DBF information is signaled within the picture header (dbf_info_in_ph_flag is 1) and picture_header_in_slice_header_flag is 0.
[0301] Sample Adaptive Offset (SAO) In one embodiment, the tool XXX is SAO (Sample Adaptive Offset). In certain alternative or additional embodiments, the SAO parameters of the slice header are only recognized if the value of the flag picture_header_in_slice_header_flag is set to 1, as depicted in Table 33. Also, as shown in Table 33, picture_header_in_slice_header_flag is set to 0 when the SAO parameters are signaled in the picture header. In another equivalent method, picture_header_in_slice_header_flag is set to 1 when the SAO parameters are signaled in the slice header, as depicted in Table 33. In this table, when the SAO information is signaled at the slice level (sao_info_in_ph_flag is 0) and the picture header is transmitted in the slice header (picture_header_in_slice_header_flag is 1), the slice_sao_luma_flag of the slice header can be decoded.
[0302] In an embodiment, as depicted in Table 35, the SAO parameters of the picture header are systematically avoided when the value of the flag picture_header_in_slice_header_flag is set to 1. In this table, the SAO information of the picture header can only be decoded if the SAO information is signaled within the picture header (sao_info_in_ph_flag is 1) and picture_header_in_slice_header_flag is set to 0.
[0303] WP (Weighted Prediction) In one embodiment, the tool XXX is WP (weighted prediction). In one embodiment, the WP parameters in the slice header are only authorized when the value of the flag picture_header_in_slice_header_flag is set to 1, as depicted in Table 33. Also, as shown in Table 33, picture_header_in_slice_header_flag is set to 0 when the WP parameters are signaled within the picture header. Also, as shown in Table 33, when notifying the WP parameters within the slice header, picture_header_in_slice_header_flag is set to 1. In this table, when the WP information is signaled at the slice level (wp_info_in_ph_flag is 0), or when the picture header is transmitted in the slice header (picture_header_in_slice_header_flag is 1), the pred_weight_table() function including the weighted prediction parameters can be decoded.
[0304] In an additional embodiment, as depicted in Table 35, when the value of the flag picture_header_in_slice_header_flag is set to 1, the decoding of the WP parameters within the picture header is systematically avoided. In this table, the WP information can be decoded only when the WP information is signaled within the picture header (wp_info_in_ph_flag is 1) and when picture_header_in_slice_header_flag is set to 0.
[0305] ALF (Adaptive Loop Filter) In one embodiment, the tool XXX is an ALF (Adaptive Loop Filter). In one embodiment, the decoding of the ALF parameters of the slice header is permitted only when the value of the flag picture_header_in_slice_header_flag is set to 1, as depicted in Table 33. Also, picture_header_in_slice_header_flag is set to 0 when the ALF parameters are signaled within the picture header (Table 33). In another equivalent method, picture_header_in_slice_header_flag is set to 1. In this table, the ALF information in the slice header can be decoded only when the ALF is enabled at the SPS level (sps_alf_enabled_flag is 1), the ALF information is signaled at the slice level (alf_info_in_ph_flag is 0), or the picture header is transmitted within the slice header (picture_header_in_slice_header_flag is 1).
[0306] In an embodiment, as depicted in Table 35, when the value of the flag picture_header_in_slice_header_flag is set to 1, the decoding of the ALF parameters of the picture header is systematically avoided. In this table, the ALF information in the slice header can be decoded only if the ALF is enabled at the SPS level (sps_alf_enabled_flag is 1), the ALF information is signaled at the picture level (alf_info_in_ph_flag is 1), and picture_header_in_slice_header_flag is set to 0.
[0307]
Table 34
[0308]
Table 35
[0309] All tools / parameters In an embodiment, when a picture header is signaled within a slice header, all tools (and / or parameters) that can be signaled within the picture header or within the slice header are signaled in the slice header. As mentioned in the above description in the embodiment, the relevant tools are as follows. QP delta information, reference picture list, deblocking filter, SAO weighted prediction, and ALF. However, other tools may be possible if they can be signaled in both the slice and picture headers.
[0310] This can be expressed by adding the following constraints: "When at least one of the flags rpl_info_in_ph_flag, dbf_info_in_ph_flag, sao_info_in_ph_flag, alf_info_in_ph_flag, wp_info_in_ph_flag, qp_delta_info_in_ph_flag is 1, picture_header_in_slice_header_flag shall be 0." And / or, add the following: "When picture_header_in_slice_header_flag is 1, the flags rpl_info_in_ph_flag, dbf_info_in_ph_flag, sao_info_in_ph_flag, wp_info_in_ph_flag, qp_delta_info_in_ph_flag shall be 0."
[0311] Also, add the following constraints to each XXX_info_in_ph_flag: "When a slice header referring to PPS contains the PH syntax, it is a bitstream compliance requirement that XXX_info_in_ph_flag is 0."
[0312] If all of these parameters are signaled within the same header, and if the tools or parameters necessarily have the same values or characteristics, it is better for there to be only one signaling, so as to reduce the complexity of decoder implementation.
[0313] Implementation FIG. 10 is a diagram showing systems 191, 195 including at least one of an encoder 150 or a decoder 100 according to an embodiment of the present invention and a communication network 199. According to an embodiment, system 195 is a system for processing and providing content (e.g., video and audio content for display / output or streaming) to a user who accesses decoder 100 via, for example, a user interface of a user terminal that constitutes decoder 100 or a user terminal capable of communicating with decoder 100. Such a user terminal may be a computer, a mobile phone, a tablet, or any other type of device capable of providing / displaying the content (provided / streamed to the user). System 195 acquires / receives bitstream 101 via communication network 199 in the form of a continuous stream or signal (e.g., while previous video / audio is being displayed / output). According to an embodiment, system 191 is for processing content and storing the processed content, e.g., video and audio content processed for display / output / streaming at a later time. System 191 acquires / receives content including the original image sequence 151 received and processed (including filtering by the deblocking filter according to the present invention) by encoder 150, and encoder 150 generates bitstream 101 that will be transmitted to decoder 100 via communication network 191. Bitstream 101 is communicated to decoder 100 in several ways. For example, it is pre-generated by encoder 150 and stored as data in a storage device (e.g., a server or cloud storage) within communication network 199 until the user requests the content (i.e., bitstream data) from the storage device, at which point the data may be communicated / streamed from the storage device to decoder 100.In addition, the system 191 may include a content providing device that provides / streamlines content information of the content stored in the storage device (e.g., the title of the content and other meta / storage location data for identifying, selecting, and requesting the content) to the user (e.g., by communicating the data of the user interface displayed on the user terminal), and receives and processes the user request for the content so that the requested content can be delivered / streamlined from the storage device to the user terminal. Alternatively, when the user requests the content, the encoder 150 generates a bitstream 101 and communicates / streamlines it directly to the decoder 100. Thereafter, the decoder 100 receives the bitstream 101 (or signal), performs filtering by the deblocking filter according to the present invention, obtains / generates a video signal 109 and / or an audio signal, and the user terminal uses this to provide the requested content to the user.
[0314] Any step of the method / process according to the present invention or the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the step / function may be stored on or transmitted onto one or more instructions or codes or programs, or a computer-readable medium, and may be executed by one or more hardware-based processing devices such as a programmable computer, which may be a PC (personal computer). It may be a DSP (digital signal processor), a circuit, a processor and a memory, a general-purpose microprocessor or a central processing unit, a microcontroller, an ASIC (application-specific integrated circuit), a field-programmable logic array (FPGA), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein may refer to any of the foregoing structures or other structures suitable for implementing the technology described herein.
[0315] Embodiments of the present invention can also be implemented by a variety of devices or apparatuses including wireless handsets, integrated circuits (ICs) or IC sets (e.g., chip sets). Although various components, modules, or units are described herein to explain the functional aspects of the devices / apparatuses configured to execute those embodiments, it is not necessarily required to be implemented by different hardware units. Rather, they can be provided by combining various modules / units into codec hardware units or by an assembly of interoperable hardware units including one or more processors in cooperation with appropriate software / firmware.
[0316] Embodiments of the present invention can be realized by a computer of a system or apparatus that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium to execute one or more modules / units / functions of the above-described embodiments and / or includes one or more processing units or circuits for executing one or more functions of the above-described embodiments. For example, it can be realized by a method executed by a computer of a system or apparatus that reads and executes computer-executable instructions from a storage medium to execute one or more functions of the above-described embodiments and / or controls one or more processing units or circuits for executing one or more functions of the above-described embodiments. The computer may include a separate computer or a network of separate processing units for reading and executing computer-executable instructions. The computer-executable instructions may be provided to the computer from a computer-readable medium such as a communication medium via a network or a tangible storage medium. The communication medium may be a signal / bitstream / carrier wave. The tangible storage medium is a "non-transitory computer-readable storage medium" and may include, for example, a hard disk, random access memory (RAM), read-only memory (ROM), storage of a distributed computing system, an optical disk (such as a compact disk (CD), digital versatile disk (DVD), Blu-ray disk (BD) (trademark), etc.), a flash memory device, a memory card, or one or more of the like. Also, at least a part of the steps / functions may be implemented in hardware by a machine or a dedicated component such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).
[0317] FIG. 11 is a schematic block diagram of a computing device 2000 for implementing one or more embodiments of the present invention. The computing device 2000 may be a device such as a microcomputer, a workstation, or a lightweight portable device. The computing device 2000 includes a central processing unit (CPU) 2001, such as a microprocessor, connected to a communication bus; a random access memory (RAM) 2002 for storing registers adapted to record executable code of the method of one or more embodiments of the present invention and variables and parameters necessary for implementing a method of encoding or decoding at least a part of an image according to one or more embodiments of the present invention (the memory capacity can be extended, for example, by optional RAM connected to an expansion port); a read only memory (ROM) 2003 for storing a computer program for implementing one or more embodiments of the present invention; a network interface (NET) 2004 connected to a communication network through which digital data to be processed is usually transmitted or received. The network interface (NET) 2004 may be a single network interface or may be composed of a set of different network interfaces (for example, a wired interface and a wireless interface, or different types of wired or wireless interfaces). Data packets are written to the network interface for transmission or read from the network interface for reception under the control of a software application executed by the CPU 2001. A user interface (UI) 2005 may be used to receive input from a user or display information to the user. A hard disk (HD) 2006 may be provided as a mass storage device. An input / output module (IO) 2007 may be used to receive / send data from / to an external device such as a video source or a display, etc. The executable code can be stored in any of the ROM 2003, the HD 2006, or a removable digital medium such as a disk.According to a modification example, the executable code of the program can be received by the communication network means via NET2004 in order to be stored in one of the storage means of the communication device 2000 such as HD2006 before being executed. The CPU2001 is adapted to control and direct the execution of the instructions or portions of the program or software code according to an embodiment of the present invention, and these instructions are stored in one of the aforementioned storage means. After power-on, the CPU2001 can execute instructions from the main RAM memory 2002 related to the software application, for example, after those instructions are loaded from the program ROM 2003 or HD2006. When such a software application is executed by the CPU2001, it causes the steps of the method according to the present invention to be executed.
[0318] Also, according to another embodiment of the present invention, it is understood that the decoder according to the aforementioned embodiment is provided in a computer, a mobile phone (cellular phone), a user terminal such as a table, or another type of device (for example, a display device) that can provide / display content to the user. According to yet another embodiment, the encoder according to the aforementioned embodiment is provided in an imaging device including a camera, a video camera, or a network camera (for example, a closed-circuit television or a video surveillance camera), and captures and provides the content to be encoded by the encoder. Two such examples are provided below with reference to FIGS. 11 and 12.
[0319] · Network camera FIG. 11 is a diagram for explaining a network camera system 2100 including a network camera 2102 and a client device 2104. The network camera 2102 includes an imaging unit 2106, an encoding unit 2108, a communication unit 2110, and a control unit 2112. The network camera 2102 and the client device 2104 are communicably connected to each other via the network 200. The imaging unit 2106 includes a lens and an image sensor (e.g., a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS)), captures an image of an object, and generates image data based on the image. This image may be a still image or a moving image. The encoding unit 2108 encodes the image data using the above-described encoding method. The communication unit 2110 of the network camera 2102 transmits the encoded image data encoded by the encoding unit 2108 to the client device 2104. Furthermore, the communication unit 2110 receives a command from the client device 2104. The command includes a command for setting parameters for encoding by the encoding unit 2108. The control unit 2112 controls other units within the network camera 2102 according to the command received by the communication unit 2110. The client device 2104 includes a communication unit 2114, a decoding unit 2116, and a control unit 2118. The communication unit 2114 of the client device 2104 transmits a command to the network camera 2102. Furthermore, the communication unit 2114 of the client device 2104 receives the encoded image data from the network camera 2102. The decoding unit 2116 decodes the encoded image data using the above-described decoding method. The control unit 2118 of the client device 2104 controls other units within the client device 2104 according to a user operation or a command received by the communication unit 2114. The control unit 2118 of the client device 2104 controls the display device 2120 to display the image decoded by the decoding unit 2116. Also, the control unit 2118 of the client device 2104 controls the display device 2120 to display a GUI (graphical user interface) for specifying the values of the parameters of the network camera 2102, including the parameters for encoding by the encoding unit 2108. In addition, the control unit 2118 of the client device 2104 controls other units within the client device 2104 in response to user operation inputs to the GUI displayed on the display device 2120. The control unit 2119 of the client device 2104 controls the communication unit 2114 of the client device 2104 to send a command specifying a parameter value for the network camera 2102 to the network camera 2102 in response to a user's operation input to the GUI displayed on the display device 2120.
[0320] ·Smartphone FIG. 12 is a diagram for explaining the smartphone 2200. The smartphone 2200 includes a communication unit 2202, a decoding unit 2204, a control unit 2206, a display unit 2208, an image recording device 2210, and a sensor 2212. The communication unit 2202 receives encoded image data via the network 200. The decoding unit 2204 decodes the encoded image data received by the communication unit 2202. The decoding unit 2204 decodes the encoded image data using the decoding method described above. The control unit 2206 controls other units within the smartphone 2200 according to user operations and commands received by the communication unit 2202. For example, the control unit 2206 controls the display unit 2208 to display the image decoded by the decoding unit 2204.
[0321] Although the present invention has been described with reference to embodiments, it will be understood that the present invention is not limited to the disclosed embodiments. It will be understood by those skilled in the art that various changes and modifications can be made without departing from the scope of the invention as defined in the appended claims. All features disclosed in this specification (including the appended claims, abstract and drawings), and / or all steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. Each feature disclosed in this specification (including the appended claims, summary and drawings) may be replaced by an alternative feature serving the same, equivalent or similar purpose, unless expressly stated otherwise. Thus, unless expressly stated otherwise, each feature disclosed is only an example of a general series of equivalent or similar features.
[0322] Also, any result of the above comparisons, determinations, evaluations, selections, executions, performings, or considerations (e.g., a selection made during an encoding or filtering process) can be indicated or determinable / referenceable in the data (e.g., a flag or data indicating a result) in the bitstream, and it is understood that the indicated or determined / referenceable result can be used in processing, for example, instead of actually performing the comparisons, determinations, evaluations, selections, executions, performings, or considerations during a decoding process.
[0323] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be advantageously used.
[0324] The reference numerals recited in the claims are for illustrative purposes only and do not limit the claims.
Claims
1. A method for decoding video data from a bitstream, comprising: The bitstream includes a picture parameter set, a picture header syntax structure including a plurality of syntax elements used when decoding one or more slices, and a slice header including a plurality of syntax elements used when decoding a slice; The method includes: decoding a plurality of syntax elements including a first flag indicating whether the picture header syntax structure is in the slice header and a second flag in the picture parameter set; decoding the video data from the bitstream using the decoded plurality of syntax elements; and when the second flag has a first value, the second flag indicates that adaptive loop filter (ALF) information may be present in the picture header syntax structure and is not present in the slice header that refers to the picture parameter set not including the picture header syntax structure; when the second flag has a second value different from the first value, the second flag indicates that the ALF information is not present in the picture header syntax structure and may be present in the slice header that refers to the picture parameter set; when the second flag has the first value, the first flag having a value indicating that the picture header syntax structure is not in the slice header is constrained to be decoded; The ALF information includes an adaptation parameter set id for an adaptive loop filter (ALF), and the adaptation parameter set (APS) indicated by the adaptation parameter set id includes a third flag indicating whether a plurality of clipping indexes related to the ALF are to be decoded. A method characterized by the above.
2. The one or more slices may include an image portion of any size of 4x4, 8x8, 16x16, 32x32, 64x64, 128x128. The method according to claim 1, characterized by the above.
3. The third flag is alf_luma_clip_flag, and each of the plurality of clipping indexes is alf_luma_clip_idx. The method according to claim 1, characterized by the above.
4. The picture header syntax structure corresponds to picture_header_structure(). The method according to claim 1, characterized in that.
5. The adaptation parameter set (APS) indicated by the adaptation parameter set id includes information corresponding to the number of filters for luma. The method according to claim 1, characterized in that.
6. A method for encoding video data into a bitstream, wherein the bitstream includes a picture parameter set, a picture header syntax structure including a plurality of syntax elements used when decoding one or more slices, and a slice header including a plurality of syntax elements used when decoding a slice, the method includes encoding the video data using a plurality of syntax elements including a first flag indicating whether the picture header syntax structure is in the slice header and a second flag in the picture parameter set, when the second flag has a first value, the second flag indicates that the adaptive loop filter (ALF) information may exist in the picture header syntax structure and does not exist in the slice header referring to the picture parameter set not including the picture header syntax structure, when the second flag has a second value different from the first value, the second flag indicates that the ALF information does not exist in the picture header syntax structure and may exist in the slice header referring to the picture parameter set, when the second flag has the first value, the first flag always has a value indicating that the picture header syntax structure is not in the slice header, the ALF information includes an adaptation parameter set id for an adaptive loop filter (ALF), and the adaptation parameter set (APS) indicated by the adaptation parameter set id includes a third flag indicating whether a plurality of clipping indexes related to the ALF are decoded A method characterized by the above.
7. The one or more slices may include an image portion of any size of 4x4, 8x8, 16x16, 32x32, 64x64, 128x128 The method according to claim 6, characterized in that...
8. wherein the third flag is alf_luma_clip_flag, and each of the plurality of clipping indexes is alf_luma_clip_idx The method according to claim 6, characterized in that...
9. A decoding apparatus for decoding video data from a bitstream, wherein the bitstream includes a picture parameter set, a picture header syntax structure including a plurality of syntax elements used when decoding one or more slices, and a slice header including a plurality of syntax elements used when decoding a slice, the decoding apparatus includes means for decoding a plurality of syntax elements including a first flag indicating whether the picture header syntax structure is in the slice header and a second flag in the picture parameter set; means for decoding the video data from the bitstream using the decoded plurality of syntax elements; and has when the second flag has a first value, the second flag indicates that the adaptive loop filter (ALF) information may exist in the picture header syntax structure and does not exist in the slice header referring to the picture parameter set not including the picture header syntax structure; when the second flag has a second value different from the first value, the second flag indicates that the ALF information does not exist in the picture header syntax structure and may exist in the slice header referring to the picture parameter set; when the second flag has the first value, the first flag having a value indicating that the picture header syntax structure is not in the slice header is constrained to be decoded; the ALF information includes an adaptation parameter set id for an adaptive loop filter (ALF), and the adaptation parameter set (APS) indicated by the adaptation parameter set id includes a third flag indicating whether a plurality of clipping indexes related to the ALF are decoded; A decoding apparatus characterized in that...
10. An encoding apparatus for encoding video data into a bitstream, The bitstream includes a picture parameter set, a picture header syntax structure including a plurality of syntax elements used when decoding one or more slices, and a slice header including a plurality of syntax elements used when decoding a slice. The encoding device has means for encoding the video data using a plurality of syntax elements including a first flag indicating whether the picture header syntax structure is in the slice header and a second flag in the picture parameter set. When the second flag has a first value, the second flag indicates that the adaptive loop filter (ALF) information may be present in the picture header syntax structure and is not present in the slice header that refers to the picture parameter set not including the picture header syntax structure. When the second flag has a second value different from the first value, the second flag indicates that the ALF information is not present in the picture header syntax structure and may be present in the slice header that refers to the picture parameter set. When the second flag has the first value, the first flag always has a value indicating that the picture header syntax structure is not in the slice header. The ALF information includes an adaptation parameter set id for an adaptive loop filter (ALF), and the adaptation parameter set (APS) indicated by the adaptation parameter set id includes a third flag indicating whether a plurality of clipping indexes related to the ALF are to be decoded. An encoding device characterized by the above.
11. A computer program for causing a computer to execute the method according to any one of Claims 1 to 8.
Citation Information
Patent Citations
Modified Adaptive Loop Filter Temporal Prediction for Temporal Scalability Support.
JP2020503801A
Image encoding / decoding method and device for signaling information about sub-pictures and picture headers, and method for transmitting bitstreams
JP2023510576A
Conditional signaling of syntax elements in picture headers
JP2023515186A