Encoder, decoder, and corresponding method using adaptive loop filters
Adaptive loop filtering with fixed-length codes addresses the challenge of enhancing video coding efficiency by improving compression ratios while maintaining picture quality through efficient encoding and decoding processes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-12-17
- Publication Date
- 2026-05-22
AI Technical Summary
Existing video coding technologies face challenges in achieving improved compression ratios with minimal sacrifice in picture quality, particularly in the context of limited network bandwidth and memory resources.
The implementation of adaptive loop filtering (ALF) using fixed-length codes for signaling clipping indices and coefficients, allowing for efficient encoding and decoding of video data by limiting sample value differences and modifying target sample values based on these indices.
This approach enhances coding efficiency and maintains picture quality by simplifying the signaling of clipping parameters, thereby improving the compression ratio without significant degradation.
Smart Images

Figure 0007864174000060 
Figure 0007864174000061 
Figure 0007864174000062
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This patent application claims priority to U.S. Provisional Patent Application No. 62 / 843,431, filed on 4 May 2019. The disclosure of the aforementioned patent application is incorporated herein by reference in its entirety.
[0002] Technical field Embodiments of the present application (disclosure) relate broadly to the field of picture processing, and more specifically to filtering samples of blocks within a picture. [Background technology]
[0003] Video coding (video encoding and decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the internet and mobile networks, real-time conversation applications like video chat and video conferencing, DVD and Blu-ray discs, video content acquisition and editing systems, and camcorders in security applications.
[0004] The amount of video data required to depict a relatively short video can be substantial, which can cause difficulties when the data is streamed or otherwise transmitted over a communication network with limited bandwidth. Therefore, video data is generally compressed before being transmitted over modern telecommunication networks. Video size can also be an issue when video is stored in storage, as memory resources may be limited. Video compressors often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompressor that decodes the video data. Given limited network resources and the ever-increasing demand for higher video quality, improved compression and decompression techniques that improve the compression ratio with little to no sacrifice of picture quality are desirable. [Overview of the Initiative] [Means for solving the problem]
[0005] Embodiments of the present application provide an apparatus and method for encoding and decoding, as defined by an independent claim. The above and other objectives are achieved by the subject matter of the independent claims. Further implementations are evident from the dependent claims, specification and drawings.
[0006] According to the first aspect of this disclosure, a coding method is provided which is implemented by a decoding device, the method being: The process includes: a step of acquiring a bitstream, wherein at least one bit in the bitstream represents a syntactic element for the current block, the syntactic element specifying a clipping index for a clipping value for an adaptive loop filter (ALF); a step of parsing the bitstream to acquire the value of the syntactic element for the current block, wherein the syntactic element is coded using a fixed-length code; and a step of applying adaptive loop filtering to the current block based on the value of the syntactic element for the current block. Here, a fixed-length code means that all possible values of the syntactic element are signaled using the same number of bits. This provides a simpler way to signal the clipping parameters. Furthermore, coding efficiency is improved.
[0007] In one possible implementation of the method by the first aspect itself, the value of the syntactic element for the block is obtained by using only the at least one bit.
[0008] In any possible implementation of the method by any prior implementation or by the first aspect itself, the at least one bit is 2 bits.
[0009] In any possible implementation of the method by any preceding implementation or by the first aspect itself, the at least one bit in the bitstream represents the value of the syntactic element.
[0010] In any possible implementation of the method by any preceding implementation or by the first aspect itself, the syntactic elements are for a chroma-adaptive loop filter or a luma-adaptive loop filter.
[0011] In any possible implementation of the method by any prior implementation or by the first aspect itself, the clipping value is used to determine the clipping range used to limit (or clip) the difference between a target sample value and a neighboring sample value, and the limited sample value difference (or clipped sample value difference) is used to modify the target sample value in the ALF process.
[0012] In any possible implementation of the method by any prior implementation or by the first aspect itself, the step of applying adaptive loop filtering to a current block based on the value of the syntactic element includes: obtaining a clipping value based on the value of the syntactic element; using the clipping value to limit (or clip) the difference between a target sample value and a neighboring sample value of the current block; multiplying the difference of the limited sample values (or the clipped sample value difference) by a coefficient of the adaptive loop filter (ALF); and using the result of the multiplication to modify the target sample value.
[0013] In any possible implementation of the method by any prior implementation or by the first aspect itself, the clipping value is determined by using the clipping index specified by the syntactic element and a mapping between the clipping index and the clipping value.
[0014] In any possible implementation of the method by any prior implementation or by the first aspect itself, the fixed-length code includes a binary representation of an unsigned integer using the at least one bit. In other words, the at least one bit is a binary representation of the value of the syntactic element, and the value of the syntactic element is an unsigned integer.
[0015] In any possible implementation of the method by any preceding implementation or by the first aspect itself, the syntactic element is applied to a set of blocks, where now a block is one block in the set of blocks.
[0016] In any possible implementation of the method by any preceding implementation or by the first aspect itself, the syntactic element is at the slice level.
[0017] According to a second aspect of this disclosure, a coding method is provided which is implemented by a decoding device, the method being: The process includes: obtaining a bitstream in which at least one bit of the bitstream represents a syntactic element for the current block, the syntactic element being an adaptive loop filter (ALF) clipping value index and / or an ALF coefficient parameter; parsing the bitstream to obtain the value of the syntactic element for the current block, the value of the syntactic element for the current block being obtained by using only the at least one bit of the syntactic element; and applying adaptive loop filtering to the current block based on the value of the syntactic element for the current block.
[0018] In one possible implementation of the method by the second aspect itself, the syntactic elements are coded using fixed-length code.
[0019] In one possible implementation of the method described in the preceding implementation, the fixed-length code includes a binary representation of an unsigned integer using the at least one bit. In other words, the at least one bit is a binary representation of the value of the syntactic element, and the value of the syntactic element is an unsigned integer.
[0020] In one possible implementation of the second aspect itself or any preceding implementation of the method, the syntactic element itself defines the value of the syntactic element.
[0021] In one possible implementation of the method by the second aspect itself or any preceding implementation, the at least one bit in the bitstream represents the value of the syntactic element.
[0022] In a possible implementation of the method according to the second aspect itself or any of its preceding implementations, the ALF clipping value index specifies the clipping index of the clipping values for an adaptive loop filter (ALF).
[0023] In a possible implementation of the method according to the second aspect itself or any of its preceding implementations, an ALF coefficient parameter is used to obtain the coefficients of the ALF.
[0024] In a possible implementation of the method according to the second aspect itself or any of its preceding implementations, the fact that the value of the syntax element for the current block is obtained by using only at least one bit of the syntax element means that the value of the syntax element is defined by the syntax element itself. [[ID=[]]
[0025] In a possible implementation of the method according to the second aspect itself or any of its preceding implementations, the syntax element is applied to a set of blocks, and the current block is one of the blocks in the set of blocks.
[0026] In a possible implementation of the method according to the second aspect itself or any of its preceding implementations, the syntax element is at the slice level.
[0027] In a possible implementation of the method according to the second aspect itself or any of its preceding implementations, the ALF coefficient parameter is used to determine the ALF coefficients.
[0028] In a possible implementation of the method according to the first aspect or the second aspect itself, or any of its preceding implementations, the syntax element is the ALF clipping value index, and the at least one bit representing the syntax element is 2 bits.
[0029] In one possible implementation of the method described above, the ALF clipping value index identifies one of the four clipping values.
[0030] In a possible implementation of the first aspect or the second aspect itself, or any prior implementation thereof, the value of the ALF clipping value index is used to determine the clipping range, which is used in the adaptive loop filtering process.
[0031] A third aspect of this disclosure provides a coding method implemented by an encoding device. The method is: The present invention includes the steps of: determining the value of a syntactic element for a block, wherein the syntactic element specifies the clipping index of the clipping value for an adaptive loop filter (ALF); and generating a bitstream based on the value of the syntactic element, wherein at least one bit in the bitstream represents the syntactic element, and the syntactic element is coded using a fixed-length code.
[0032] In one possible implementation of the method by the third aspect itself, the at least one bit of the syntactic element is obtained by using only the value of the syntactic element for the current block.
[0033] In one possible implementation of the third aspect itself or any prior implementation of the method, the value of the syntactic element corresponds to the smallest difference (e.g., mean squared error or rate distortion cost) between the reconstructed (or filtered) block of the current block and the original signal of the current block, the reconstructed (or filtered) block being the result of using the value of the syntactic element, the smallest difference being smaller than any other difference corresponding to any other possible value of the syntactic element.
[0034] In one possible implementation of the third aspect itself or any preceding implementation of the method, at least one bit in the bitstream represents the value of the syntactic element.
[0035] In one possible implementation of the third aspect itself or any prior implementation of the method, the clipping value is used to determine a clipping range used to limit (or clip) the difference between a target sample value and a neighboring sample value, and the limited sample value difference (or clipped sample value difference) is used to modify the target sample value in the ALF process.
[0036] In one possible implementation of the third aspect itself or any prior implementation of the method, the fixed-length code includes a binary representation of an unsigned integer using the at least one bit. In other words, the at least one bit is a binary representation of the value of the syntactic element, and the value of the syntactic element is an unsigned integer.
[0037] In one possible implementation of the third aspect itself or any prior implementation of the method, the syntactic element is applied to a set of blocks, where now a block is one block in the set of blocks.
[0038] In one possible implementation of the third aspect itself or any preceding implementation of the method, the syntactic element is at the slice level.
[0039] A fourth aspect of this disclosure provides a coding method implemented by an encoding device. The method is: The process includes: a step of determining the value of a syntactic element for a current block, wherein the syntactic element is an adaptive loop filter (ALF) clipping value index and / or an ALF filter coefficient parameter; and a step of generating a bitstream based on the value of the syntactic element, wherein at least one bit in the bitstream represents the syntactic element, and the at least one bit of the syntactic element is obtained by using only the value of the syntactic element for the current block.
[0040] In one possible implementation of the fourth aspect itself or any preceding implementation of the method, the syntactic elements are coded using fixed-length code.
[0041] In one possible implementation of the method described in the preceding implementation, the fixed-length code includes a binary representation of an unsigned integer using the at least one bit. In other words, the at least one bit is a binary representation of the value of the syntactic element, and the value of the syntactic element is an unsigned integer.
[0042] In one possible implementation of the fourth aspect itself or any preceding implementation of the method, the at least one bit in the bitstream represents the value of the syntactic element.
[0043] In one possible implementation of the fourth aspect itself or any preceding implementation of the method, the syntactic element is applied to a set of blocks, where now a block is one block in the set of blocks.
[0044] In one possible implementation of the fourth aspect itself or any preceding implementation of the method, the syntactic element is at the slice level.
[0045] In one possible implementation of the fourth aspect itself or any prior implementation of the method, the ALF coefficient parameter is used to determine the ALF coefficient.
[0046] In one possible implementation of the third aspect or the fourth aspect itself, or any prior implementation thereof, the syntactic element is the ALF clipping value index, and the at least one bit representing the syntactic element is 2 bits.
[0047] In one possible implementation of the method immediately preceding the fourth aspect, the ALF clipping value index identifies one of the four clipping values.
[0048] In one possible implementation of the third aspect or the fourth aspect itself, or any prior implementation thereof, the value of the ALF clipping value index is used to determine the clipping range, which is used in the adaptive loop filtering process.
[0049] According to a fifth aspect of this disclosure, a decoder is provided comprising a processing circuit for performing the first or second aspect or any implementation thereof.
[0050] According to the sixth aspect of this disclosure, an encoder is provided comprising a processing circuit for performing the third or fourth aspect or any implementation thereof.
[0051] According to the seventh aspect of this disclosure, a computer program product is provided which includes program code for performing a method according to any of the first to fourth aspects or any implementation thereof.
[0052] According to the eighth aspect of this disclosure, a non-temporary computer-readable medium is provided which carries program code that, when executed by a computer device, causes the computer device to execute a method according to any of the first to fourth aspects or any implementation thereof.
[0053] A ninth aspect of the present disclosure provides a decoder having one or more processors; and a non-temporary computer-readable storage medium coupled to the processors and storing a program for execution by the processors. The program is configured to perform, when executed by the processors, a method according to the first or second aspect or any implementation thereof.
[0054] According to a tenth aspect of the present disclosure, an encoder is provided having one or more processors; and a non-temporary computer-readable storage medium coupled to the processors and storing a program for execution by the processors. The program is configured to perform a third or fourth aspect or a method according to any implementation thereof when executed by the processors.
[0055] According to the eleventh aspect of this disclosure, a decoder is provided having an entropy decode unit configured to acquire a bitstream, At least one bit in the bitstream represents a syntactic element for the current block, the syntactic element specifying the clipping index of the clipping value for the adaptive loop filter (ALF); The entropy decoding unit is configured to parse the bitstream to obtain the values of the syntax elements for the current block, the syntax elements being coded using fixed-length code; and the filtering unit is configured to apply adaptive loop filtering to the current block based on the values of the syntax elements for the current block.
[0056] According to a twelfth aspect of this disclosure, a decoder is provided having an entropy decoding unit configured to acquire a bitstream, At least one bit in the bitstream represents a syntactic element for the current block, which is an adaptive loop filter (ALF) clipping value index or an ALF coefficient parameter; The entropy decode unit is further configured to parse the bitstream to obtain the value of the syntax element for the current block, which is obtained by using only the at least one bit of the syntax element; and a filtering unit is configured to apply adaptive loop filtering to the current block based on the value of the syntax element for the current block.
[0057] According to the 13th aspect of this disclosure, an encoder is provided, the encoder is: The current configuration includes a decision unit configured to determine the value of a syntactic element for a block, wherein the syntactic element specifies the clipping index of the clipping value for an adaptive loop filter (ALF); and an entropy encoding unit configured to generate a bitstream based on the value of the syntactic element, wherein at least one bit in the bitstream represents the syntactic element, and the syntactic element is encoded using a fixed-length code.
[0058] According to a fourteenth aspect of this disclosure, an encoder is provided, which is: A decision unit configured to determine the value of a syntactic element for a current block, wherein the syntactic element is an ALF clipping value index or an adaptive loop filter (ALF) coefficient parameter; and an entropy encode unit configured to generate a bitstream based on the value of the syntactic element, wherein at least one bit in the bitstream represents the syntactic element, and the at least one bit of the syntactic element is obtained by using only the value of the syntactic element for the current block.
[0059] According to the 15th aspect of this disclosure, a coding method is provided which is implemented by a decoding device, the method being: A step of acquiring a bitstream, wherein n bits in the bitstream represent syntactic elements specifying the clipping index of the clipping value for the adaptive loop filter (ALF), where n is a non-negative integer; The process includes: parsing the bitstream to obtain the value of the syntax element for the current block, wherein the value of the syntax element is a binary representation of an unsigned integer using the n bits; and applying adaptive loop filtering to the current block based on the value of the syntax element for the current block.
[0060] In one possible implementation of the method according to the 15th aspect, the syntactic element may be a slice-level syntactic element.
[0061] According to the sixteenth aspect of this disclosure, a coding method is provided which is implemented by an encoding device, the method being: The steps include: determining the value of a syntax element that specifies the clipping index of the clipping value for an adaptive loop filter (ALF), where n is a non-negative integer; and generating a bitstream containing n bits based on the value of the syntax element, where the binary representation of an unsigned integer using the n bits is the value of the syntax element.
[0062] In one possible implementation of the 16th aspect method, the syntactic element may be a slice-level syntactic element.
[0063] According to the 17th aspect of this disclosure, a decoder is provided, and said decoder is: Entropy decoding unit configured to acquire a bitstream, wherein n bits in the bitstream represent a slice-level syntactic element specifying the clipping index of the clipping value for an adaptive loop filter (ALF), where n is a non-negative integer, and the entropy decoding unit is further configured to parse the bitstream to obtain the value of the syntactic element for the current block, where the value of the syntactic element is a binary representation of an unsigned integer using the n bits; and filtering unit configured to apply adaptive loop filtering to the current block based on the value of the syntactic element for the current block.
[0064] According to the 18th aspect of this disclosure, an encoder is provided, which is: A decision unit configured to determine the value of a slice-level syntactic element specifying the clipping index of the clipping value for an adaptive loop filter (ALF), wherein n is a non-negative integer; and an entropy encode unit configured to generate a bitstream containing n bits based on the value of the syntactic element, wherein the binary representation of an unsigned integer using the n bits is the value of the syntactic element.
[0065] According to the 19th aspect of this disclosure, a decoder is provided comprising a processing circuit for performing the method of the 15th aspect or any implementation thereof.
[0066] According to the 20th aspect of this disclosure, an encoder is provided comprising a processing circuit for performing the method of the 16th aspect or any implementation thereof.
[0067] According to the 21st aspect of this disclosure, a computer program product is provided which includes program code for performing the 15th aspect, the 16th aspect, or any implementation thereof.
[0068] According to the 22nd aspect of this disclosure, a non-temporary computer-readable medium is provided which carries program code that, when executed by a computer device, causes the computer device to execute a method according to the 15th aspect or the 16th aspect or any implementation thereof.
[0069] According to a 23rd aspect of the present disclosure, a decoder is provided having one or more processors; and a non-temporary computer-readable storage medium coupled to the processors and storing a program for execution by the processors. The program is configured to perform the 15th aspect or any implementation thereof when executed by the processors.
[0070] According to a 24th aspect of the present disclosure, an encoder is provided having one or more processors; and a non-temporary computer-readable storage medium coupled to the processors and storing a program for execution by the processors. The program, when executed by the processors, is configured to perform a method according to a 16th aspect or any implementation thereof.
[0071] According to a 25th aspect of the present disclosure, a non-temporary storage medium is provided which includes a bitstream containing n bits, wherein the binary representation of an unsigned integer using the n bits is the value of a syntactic element, the syntactic element specifies the clipping index of the clipping value for an adaptive loop filter (ALF), where n is a non-negative integer.
[0072] According to a 26th aspect of the present disclosure, a non-temporary storage medium is provided, comprising a bitstream, wherein at least one bit in the bitstream represents a syntactic element, the syntactic element is coded using a fixed-length code, and specifies the clipping index of the clipping value for an adaptive loop filter (ALF).
[0073] In one implementation of the method using the 26th aspect itself, the syntactic element itself defines the value of the syntactic element.
[0074] According to a 27th aspect of the present disclosure, a non-temporary storage medium is provided which includes a bitstream, wherein at least one bit in the bitstream represents a syntactic element, which is an adaptive loop filter (ALF) clipping value index or an ALF filter coefficient parameter, and which is obtained by using only the value of the syntactic element.
[0075] According to the 28th aspect of this disclosure, a non-temporary storage medium is provided which includes a bitstream encoded by any aspect or any implementation thereof.
[0076] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, purposes, and advantages will be apparent from the specification, drawings, and claims. [Brief explanation of the drawing]
[0077] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings and figures. [Figure 1A] This is a block diagram showing an example of a video coding system configured to implement an embodiment of the present invention. [Figure 1B] This is a block diagram showing another example of a video coding system configured to implement embodiments of the present invention. [Figure 2] This is a block diagram showing an example of a video encoder configured to implement an embodiment of the present invention. [Figure 3] This is a block diagram illustrating the structure of a video decoder configured to implement embodiments of the present invention. [Figure 4] This is a block diagram showing examples of encoding or decoding devices. [Figure 5] This is a block diagram showing another example of an encoding or decoding device. [Figure 6] The ALF filter shapes are shown: Chroma 5x5 diamond, Luma 7x7 diamond. [Figure 7] This shows the subsampled ALF block classification. [Figure 8] This shows the signal transmission of the VTM-5.0 ALF luma and chroma clipping parameters. [Figure 9] This shows the signal transmission of the modified VTM-5.0 ALF luma and chroma clipping parameters, where the clipping parameters are transmitted using a 2-bit fixed-length code. [Figure 10] This is a block diagram showing the method according to the first aspect of this disclosure. [Figure 11] This is a block diagram showing the method according to the second aspect of this disclosure. [Figure 12] This is a block diagram showing the method according to the third aspect of this disclosure. [Figure 13] This is a block diagram showing the method according to the fourth aspect of this disclosure. [Figure 14] A block diagram of a decoder according to the fifth aspect of this disclosure. [Figure 15] This is a block diagram of an encoder according to the sixth aspect of this disclosure. [Figure 16] A block diagram of a decoder according to the ninth aspect of this disclosure. [Figure 17] This is a block diagram of an encoder according to the tenth aspect of this disclosure. [Figure 18] A block diagram of a decoder according to the eleventh aspect of this disclosure. [Figure 19] A block diagram of a decoder according to the twelfth aspect of this disclosure. [Figure 20] This is a block diagram of an encoder according to the thirteenth aspect of this disclosure. [Figure 21] This is a block diagram of an encoder according to the 14th aspect of this disclosure. [Figure 22] This is a block diagram showing an exemplary structure of a content supply system 3100 that realizes a content distribution service. [Figure 23] A block diagram showing the structure of an example terminal device.
[0078] In the following, unless otherwise explicitly specified, the same reference numeral refers to the same or at least functionally equivalent feature. [Modes for carrying out the invention]
[0079] The following description refers to the accompanying drawings, which constitute part of this disclosure and, as an example, illustrate specific aspects of embodiments of the present invention or specific aspects in which embodiments of the present invention may be used. Embodiments of the present invention may be used in other aspects and may include structural or logical modifications not shown in the drawings. Therefore, the following detailed description should not be construed as limiting, and the scope of the invention is defined by the appended claims.
[0080] For example, disclosure relating to a described method may also apply to a corresponding apparatus or system configured to perform that method, and vice versa. For example, if one or more specific method steps are described, the corresponding apparatus may include one or more units, e.g., functional units (e.g., one unit that performs the one or more steps, or multiple units that each perform one or more of the steps), even if such one or more units are not explicitly described or illustrated. On the other hand, if a particular apparatus is described based on one or more units, e.g., functional units, the corresponding method may include one step for performing the function of the one or more units (e.g., one step that performs the function of the one or more units, or multiple steps that each perform the function of one or more of the units), even if such one or more steps are not explicitly described or illustrated. Furthermore, it is understood that the various exemplary embodiments and / or aspect features described herein may be combined with each other unless otherwise specified.
[0081] Video coding typically refers to the processing of a series of pictures that make up a video or video sequence. The terms “frame” or “image” are sometimes used synonymously in the field of video coding instead of “picture.” Video coding (or coding in general) consists of two parts: video encoding and video decoding. Video encoding is performed on the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves the reverse processing compared to encoding, for the purpose of reconstructing the video picture. Embodiments referring to “coding” a video picture (or picture in general) are understood to relate to “encoding” or “decoding” a video picture or its respective video sequence. The combination of the encoding and decoding parts is also called a codec (coding and decoding).
[0082] In lossless video coding, the original video picture can be reconstructed. That is, the reconstructed video picture has the same quality as the original video picture (assuming there is no transmission loss or other data loss during storage or transmission). In lossy video coding, further compression is performed, for example by quantization, to reduce the amount of data representing the video picture, and the video picture cannot be fully reconstructed in the decoder. That is, the quality of the reconstructed video picture is lower or worse than the quality of the original video picture.
[0083] Several video coding standards belong to the group of “lossy hybrid video codecs” (i.e., combining spatial and temporal prediction in the sample domain with 2D transform coding to apply quantization in the transform domain). Each picture in a video sequence is typically divided into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in an encoder, video is typically processed, or encoded, at the block (video block) level. This is done, for example, by generating predicted blocks using spatial (in-picture) and / or temporal (inter-picture) predictions, subtracting the predicted blocks from the current blocks (the blocks currently being processed / to be processed) to obtain residual blocks, transforming the residual blocks, and quantizing the residual blocks in the transform domain to reduce (compress) the amount of data to be transmitted. In a decoder, the reverse process compared to the encoder is applied to the encoded or compressed blocks to reconstruct the current blocks for representation. Furthermore, the encoder duplicates the decoder processing loop, so both generate the same predictions (e.g., intra-predictions and inter-predictions) and / or subsequent blocks for processing, i.e., coding.
[0084] Embodiments of the video coding system 10, video encoder 20, and video decoder 30 are described below with reference to Figures 1 to 3.
[0085] Figure 1A is a schematic block diagram showing an exemplary coding system 10, for example, a video coding system 10 (or simply coding system 10), in which the technology of the present application can be utilized. The video encoder 20 (or simply encoder 20) and video decoder 30 (or simply decoder 30) of the video coding system 10 represent examples of devices that may be configured to perform the technology described in the various examples of the present application.
[0086] As shown in Figure 1A, the coding system 10 has a source device 12 configured to provide encoded picture data 21 to a destination device 14 that decodes the encoded picture data 13, for example.
[0087] The source device 12 has an encoder 20 and may additionally, i.e., optionally, have a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22.
[0088] The picture source 16 may have, or be, any kind of picture capturing device, e.g., a camera for capturing real-world pictures, and / or any kind of picture generating device, e.g., a computer graphics processor for generating computer-animated pictures, or any kind of other device for acquiring and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). The picture source may also have any kind of memory or storage for storing any of the above pictures.
[0089] To distinguish it from the processing performed by the preprocessor 18 and the preprocessing unit 18, the picture or picture data 17 may be referred to as the raw picture or raw picture data 17.
[0090] The preprocessor 18 is configured to receive (raw) picture data 17, perform preprocessing on the picture data 17, and obtain a preprocessed picture 19 or preprocessed picture data 19. The preprocessing performed by the preprocessor 18 may include, for example, cropping, color format conversion (e.g., RGB to YCbCr), color correction, or noise reduction. It is understood that the preprocessing unit 18 may be an optional component.
[0091] The video encoder 20 is configured to receive pre-processed picture data 19 and provide encoded picture data 21 (further details will be described later, for example, based on Figure 2).
[0092] The communication interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) through the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.
[0093] The destination device 14 has a decoder 30 (for example, a video decoder 30) and may additionally, or optionally, have a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.
[0094] The communication interface 28 of the destination device 14 is configured to receive encoded picture data 21 (or a further processed version thereof) from, for example, the source device 12 directly, or from any other source, such as a storage device, such as an encoded picture data storage device, and to provide the encoded picture data 21 to the decoder 30.
[0095] Communication interfaces 22 and 28 may be configured to transmit or receive encoded picture data 21 or encoded data 13 via a direct communication link between the source device 12 and the destination device 14, for example, via a direct wired or wireless connection, or via any type of network, for example, a wired or wireless network or any combination thereof, or any type of private and public network or any combination thereof.
[0096] The communication interface 22 may be configured, for example, to package the encoded picture data 21 into an appropriate format, such as a packet, and / or to process the encoded picture data using any kind of transmission encoding or processing for transmission over a communication link or communication network.
[0097] The communication interface 28, which is the counterpart to the communication interface 22, may be configured, for example, to receive transmitted data and process the transmitted data using any type of corresponding transmission decoding or processing and / or depackaging to obtain encoded picture data 21.
[0098] Both communication interfaces 22 and 28 can be configured as one-way or two-way communication interfaces, as indicated by the arrows in Figure 1A for the communication channel 13 pointing from the source device 12 to the destination device 14, and may be configured, for example, to send and receive messages, for example, to set up a connection, to receive and confirm and exchange any other information related to the communication link and / or data transmission, such as encoded picture data transmission.
[0099] The decoder 30 is configured to receive encoded picture data 21 and provide decoded picture data 31 or decoded picture 31 (further details will be described later, for example, based on Figure 3 or Figure 5).
[0100] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also called reconstructed picture data), for example, the decoded picture 31, to obtain post-processed picture data 33, for example, the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., YCbCr to RGB), color correction, cropping, or resampling, or any other processing to prepare the decoded picture data 31 for display, for example, by the display device 34.
[0101] The display device 34 of the destination device 14 is configured to receive post-processed picture data 33 for displaying the picture, for example, to a user or viewer. The display device 34 may be any type of display for representing the reconstructed picture, for example, an integrated or external display or monitor, or may include one. The display may include, for example, a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a plasma display, a projector, a microLED display, a liquid crystal on silicon (LCoS), a digital optical processor (DLP), or any other type of display.
[0102] Figure 1A depicts the source device 12 and the destination device 14 as separate devices, but the embodiment of the device may include both or both functions, such as the source device 12 or its corresponding function and the destination device 14 or its corresponding function. In such embodiments, the source device 12 or its corresponding function and the destination device 14 or its corresponding function may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.
[0103] As will be obvious to those skilled in the art based on the foregoing description, the functions of different units or the presence and (strict) division of functions within the source device 12 and / or destination device 14 as shown in Figure 1A may vary depending on the actual device and application.
[0104] The encoder 20 (for example, a video encoder 20), or the decoder 30 (for example, a video decoder 30), or both the encoder 20 and the decoder 30, may be implemented via processing circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof, as shown in Figure 1B. The encoder 20 may be implemented via processing circuit 46 to embody various modules discussed with respect to the encoder 20 of Figure 2 and / or any other encoder systems or subsystems described herein. The decoder 30 may be implemented via processing circuit 46 to embody various modules discussed with respect to the decoder 30 of Figure 3 and / or any other decoder systems or subsystems described herein. The processing circuits may be configured to perform various operations, as shown later, as shown in Figure 5. If the technology is partially implemented in software, the device may store instructions for the software in a suitable non-temporary computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of the present disclosure. Either the video encoder 20 or the video decoder 30 may also be integrated within a single device as part of a combined encoder / decoder (codec), for example, as shown in Figure 1B.
[0105] The source device 12 and destination device 14 may include any of a wide range of devices, including any type of handheld or stationary device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or content delivery server), a broadcast receiver, a broadcast transmitter, etc., and may or may not use an operating system.
[0106] In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices. In some cases, the video coding system 10 shown in Figure 1A is merely an example, and the technology of the present application may be applied to video coding scenarios (e.g., video encoding or video decoding) that do not necessarily involve data communication between the encoding device and the decoding device. In other examples, the data may be retrieved from local memory, streamed over a network, etc. A video encoding device can encode data and store it in memory, and / or a video decoding device can retrieve the data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other, but simply encode the data into memory and / or retrieve the data from memory and decode it.
[0107] For convenience of description, embodiments of the present invention are described herein with reference to, for example, High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC) reference software, next-generation video coding standards developed by the ITU-T Video Coding Expert Group (VCEG) and the ISO / IEC Motion Picture Expert Group (MPEG) Joint Collaboration Team on Video Coding (JCT-VC). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC.
[0108] Encoder and encoding method Figure 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the technology of the present invention. In the example of Figure 2, the video encoder 20 has an input 201 (or input interface 201), a residual calculation unit 204, a transformation unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transformation unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output 272 (or output interface 272). The mode selection unit 260 may include an inter-prediction unit 244, an intra-prediction unit 254, and a partitioning unit 262. The inter-prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in Figure 2 may be referred to as a hybrid video encoder or a video encoder with a hybrid video codec.
[0109] The residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, and the mode selection unit 260 may be said to form the forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decode picture buffer (DPB) 230, the inter-prediction unit 244, and the intra-prediction unit 254 may be said to form the reverse signal path of the video encoder 20. The reverse signal path of the video encoder 20 corresponds to the signal path of the decoder (see video decoder 30 in Figure 3). The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the decode picture buffer (DPB) 230, the inter-prediction unit 244, and the intra-prediction unit 254 may also be said to constitute the “built-in decoder” of the video encoder 20.
[0110] Picture and picture partitioning (picture and block) The encoder 20 may be configured to receive, for example, a picture 17 (or picture data 17) via input 201, for example, a picture of a sequence of pictures forming a video or video sequence. The received picture or picture data may be a pre-processed picture 19 (or pre-processed picture data 19). For simplicity, the following description refers to picture 17. Picture 17 may also be referred to as the current picture or the picture to be coded (particularly in video coding, to distinguish the current picture from other pictures, for example, previously encoded and / or decoded pictures of the same video sequence, i.e., the video sequence that also contains the current picture).
[0111] A (digital) picture can be considered, or may be considered, a two-dimensional array or matrix of samples with intensity values. Samples in the array may be referred to as pixels (short for picture element) or picture elements. The number of samples in the horizontal and vertical (or axis) directions of the array or picture defines the size and / or resolution of the picture. Typically, three color components are used for color representation; that is, a picture can be represented by, or contain, three sample arrays. In the RGB format or color space, a picture contains corresponding red, green, and blue sample arrays. However, in video coding, each pixel is typically represented in a luminance and chrominance format or color space, for example, YCbCr, which includes a luminance component represented by Y (sometimes L is used instead) and two chrominance components represented by Cb and Cr. The luminance (or luma for short) component Y represents brightness or gray level intensity (for example, as in a grayscale picture), while the two chrominance (or chroma for short) components Cb and Cr represent chromaticity or color information components. Thus, a picture in YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in RGB format can be converted to or from YCbCr format, and vice versa. This process is also known as color conversion or transformation. If the picture is monochrome, it may contain only a luminance sample array. Thus, a picture can be, for example, in 4:2:0, 4:2:2, and 4:4:4 color formats, an array of luma samples or arrays of luma samples in a monochrome format and two corresponding arrays of chroma samples.
[0112] Embodiments of the video encoder 20 may include a picture partitioning unit (not shown in Figure 2) configured to partition a picture 17 into a plurality of (typically non-overlapping) picture blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), coding tree blocks (CTBs), or coding tree units (CTUs) (H.265 / HEVC and VVC). The picture partitioning unit may be configured to use the same block size and corresponding grid defining said block size for all pictures in a video sequence, or to change the block size between pictures, or between subsets or groups of pictures, and partition each picture into a corresponding block.
[0113] In further embodiments, the video encoder may be configured to directly receive a block 203 of picture 17, for example, one, some, or all of the blocks that make up picture 17. The picture block 203 may also be referred to as the current picture block or the picture block to be coded.
[0114] Similar to picture 17, picture block 203 is also considered, or can be considered, as a two-dimensional array or matrix of samples having intensity values (sample values), although it is smaller in dimensions than picture 17. In other words, block 203 can contain, for example, one sample array (e.g., a luma array in the case of monochrome picture 17, or a luma or chroma array in the case of a color picture), or three sample arrays (e.g., a luma array and two chroma arrays in the case of color picture 17), or any other number and / or type of arrays depending on the applied color format. The number of samples in the horizontal and vertical (or axis) directions of block 203 defines the size of block 203. Thus, the block may be, for example, an M×N (M columns × N rows) array of samples, or an M×N array of conversion coefficients.
[0115] An embodiment of the video encoder 20, as shown in Figure 2, may be configured to encode the picture 17 block by block, for example, encoding and prediction may be performed for each block 203.
[0116] Residual calculation The residual calculation unit 204 may be configured to calculate a residual block 205 (also referred to as residual 205) based on the picture block 203 and the prediction block 265 (further details about the prediction block 265 will be described later). This can be done, for example, by subtracting the sample values of the prediction block 265 from the sample values of the picture block 203 sample by sample (pixel by pixel) to obtain the residual block 205 in the sample region.
[0117] conversion The conversion processing unit 206 may be configured to apply a conversion, such as a discrete cosine transform (DCT) or discrete sine transform (DST), to the sample values of the residual block 205 to obtain conversion coefficients 207 in the conversion region. The conversion coefficients 207 may also be called conversion residual coefficients and represent the residual block 205 in the conversion region.
[0118] The conversion processing unit 206 may be configured to apply an integer approximation of the DCT / DST, such as the conversion specified for H.265 / HEVC. Compared to the orthogonal DCT conversion, such an integer approximation is typically scaled by a factor. An additional scaling factor is applied as part of the conversion process to preserve the norm of the residual blocks processed by the forward and inverse conversions. The scaling factor is typically selected based on certain constraints, such as the scaling factor being a power of 2 for the shift operation, the bit depth of the conversion coefficients, and a trade-off between precision and implementation cost. A specific scaling factor may be specified, for example, for the inverse conversion by the inverse conversion processing unit 212 (and the corresponding inverse conversion by the inverse conversion processing unit 312 in the video decoder 30), and a corresponding scaling factor for the forward conversion by the conversion processing unit 206 in the encoder 20 may be specified accordingly.
[0119] Embodiments of the video encoder 20 (specifically the conversion processing unit 206) may be configured to output conversion parameters, such as one or more conversion types, encoded or compressed, for example, directly or via the entropy encoding unit 270, so that, for example, the video decoder 30 may receive and use the conversion parameters for decoding.
[0120] quantization The quantization unit 208 may be configured to quantize the transformation coefficient 207, for example by applying scalar quantization or vector quantization, to obtain the quantized coefficient 209. The quantized coefficient 209 may also be referred to as the quantized transformation coefficient 209 or the quantized residual coefficient 209.
[0121] The quantization process can reduce the bit depth associated with some or all of the 207 conversion coefficients. For example, n-bit conversion coefficients may be rounded to m-bit conversion coefficients during quantization, where n is greater than m. The degree of quantization may be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different scaling may be applied to achieve finer or coarser quantization. Smaller quantization step sizes correspond to finer quantization, while larger quantization step sizes correspond to coarser quantization. Applicable quantization step sizes may be indicated by the quantization parameter (QP), which may be, for example, an index to a predefined set of applicable quantization step sizes. For example, small quantization parameters may correspond to finer quantization (smaller quantization step sizes), large quantization parameters may correspond to coarser quantization (larger quantization step sizes), or vice versa. Quantization may involve division by the quantization step size, and the corresponding and / or inverse dequantization by, for example, the inverse quantization unit 210 may involve multiplication by the quantization step size. Some standards, such as the HEVC embodiment, may be configured to use a quantization parameter to determine the quantization step size. In general, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of the expression involving division. Additional scaling factors may be introduced for quantization and dequantization to restore the norm of the residual block, which may be modified for the scaling used in the fixed-point approximation of the expression for the quantization parameter and quantization step size. In one example implementation, the scaling of the inverse transform and dequantization may be combined. Alternatively, a customized quantization table may be used to transmit signals from encoder to decoder, for example, in a bitstream. Quantization is a lossy operation, and the loss increases with increasing quantization step size.
[0122] Embodiments of the video encoder 20 (specifically the quantization unit 208) may be configured to output quantization parameters (QP) that are encoded, for example, directly or via the entropy encoding unit 270, so that, for example, the video decoder 30 can receive and apply the quantization parameters for decoding.
[0123] inverse quantization The inverse quantization unit 210 is configured to obtain dequantized coefficients 211 by applying the inverse quantization of the quantization unit 208 to the quantized coefficients, for example, by applying the reciprocal of the quantization scheme applied by the quantization unit 208, either based on or using the same quantization step size as the quantization unit 208. The dequantized coefficients 211 may also be called dequantized residual coefficients 211 and typically correspond to the transformation coefficients 207, although they are not identical to the transformation coefficients due to quantization losses.
[0124] Inverse Transform The inverse transform processing unit 212 is configured to apply the inverse transform of the transform applied by the transform processing unit 206, such as the inverse discrete cosine transform (DCT) or the inverse discrete sine transform (DST), or other inverse transforms, to obtain a reconstructed residual block 213 (or the corresponding dequantized coefficient 213) in the sample region. The reconstructed residual block 213 is sometimes referred to as the transform block 213.
[0125] Reconstruction The reconstruction unit 214 (for example, an adder or summer 214) is configured to add the transformed block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain the reconstructed block 215 in the sample region. This is done, for example, by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265 sample by sample.
[0126] filtering The loop filter unit 220 (or simply "loop filter" 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or more generally, to filter the reconstructed sample to obtain a filtered sample. The loop filter unit is configured, for example, to smooth pixel transitions or to otherwise improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a cooperative filter, or any combination thereof. Although the loop filter unit 220 is shown as an in-loop filter in Figure 2, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 is sometimes referred to as the filtered reconstructed block 221.
[0127] Embodiments of the video encoder 20 (specifically the loop filter unit 220) may be configured to output loop filter parameters (such as sample-adaptive offset information) encoded, for example, directly or via the entropy encoding unit 270, so that, for example, the decoder 30 can receive the same loop filter parameters or the respective loop filters and apply them for decoding.
[0128] Decode picture buffer The decode picture buffer (DPB) 230 may be a memory that stores a reference picture or generally reference picture data for encoding video data by the video encoder 20. The DPB 230 may be made up of any of a variety of memory devices, such as dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decode picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decode picture buffer 230 may further be configured to store other previously filtered blocks of the same current picture or a different picture, for example, a previously reconstructed picture, for example, a previously reconstructed and filtered block 221, and may provide a complete previously reconstructed, i.e., decoded picture (and corresponding reference blocks and samples) and / or a partially reconstructed current picture (and corresponding reference blocks and samples), for example, for interpretation. The decode picture buffer (DPB) 230 may be configured to store, for example, one or more unfiltered reconfigured blocks 215, or generally unfiltered reconfigured samples, or any other further processed versions of the reconfigured blocks or samples, if the reconfigured blocks 215 are not filtered by the loop filter unit 220.
[0129] Mode selection (partitioning and prediction) The mode selection unit 260 includes a partitioning unit 262, an inter-prediction unit 244, and an intra-prediction unit 254, and is configured to receive or acquire original picture data, for example, the original block 203 (the current block 203 of the current picture 17), and reconstructed picture data, for example filtered and / or unfiltered reconstructed samples or blocks, from the same (current) picture and / or from one or more previously decoded pictures, from, for example, a decoded picture buffer 230 or other buffer (for example, a line buffer, not shown). The reconstructed picture data is used as reference picture data for predictions, for example, inter-prediction or intra-prediction, in order to obtain a prediction block 265 or predictor 265.
[0130] The mode selection unit 260 may be configured to determine or select a partitioning (including no partitioning) and prediction mode (e.g., intra or inter-prediction mode) for the current block prediction mode, and to generate the corresponding prediction blocks 265 used for calculating the residual blocks 205 and reconstructing the reconstructed blocks 215.
[0131] Embodiments of the mode selection unit 260 may be configured to select a partitioning and prediction mode (for example, from those supported or available to the mode selection unit 260) that provides the best match, i.e., the smallest residual (smallest residual means better compression for transmission or storage) or the smallest signal transmission overhead (smallest signal transmission overhead means better compression for transmission or storage), or that considers or balances both. The mode selection unit 260 may also be configured to determine the partitioning and prediction mode based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides the smallest rate distortion. In this context, terms such as “best,” “smallest,” and “optimal” do not necessarily refer to an overall “best,” “smallest,” and “optimal,” but may also refer to the achievement of selection criteria, or other constraints, such as termination or a value being above or below a threshold, which could potentially lead to a “suboptimal selection” but reduce complexity and processing time.
[0132] In other words, the partitioning unit 262 may be configured to partition block 203 into smaller block partitions or subblocks (which also form blocks) using, for example, quadtree partitioning (QT), binary partitioning (BT), or ternary partitioning (TT), or any combination thereof, and to perform predictions for each of the block partitions or subblocks, for example, where mode selection involves selecting the tree structure of the partitioned block 203, and prediction modes are applied to each of the block partitions or subblocks.
[0133] The following describes in more detail the splitting (e.g., by the splitting unit 260) and prediction processing (by the inter-prediction unit 244 and the intra-prediction unit 254) performed by the exemplary video encoder 20.
[0134] Partitioning The partitioning unit 262 can now partition (or divide) block 203 into smaller partitions, for example, smaller blocks of square or rectangular size. These smaller blocks (which may also be called subblocks) may be further partitioned into even smaller partitions. This is also called tree partitioning or hierarchical tree partitioning, and may be recursively partitioned, for example, the root block at root tree level 0 (hierarchical level 0, depth 0), and then partitioned into two or more blocks at the next lower tree level, for example, a node at tree level 1 (hierarchical level 1, depth 1), and these blocks may again be partitioned into two or more blocks at the next lower level, for example, tree level 2 (hierarchical level 2, depth 2), and so on, until partitioning is terminated because a termination criterion is met, for example, the maximum tree depth or the minimum block size is reached. Blocks that are not further partitioned are also called leaf blocks or leaf nodes of the tree. A tree that uses partitioning into two partitions is called a binary tree (BT), a tree that uses partitioning into three partitions is called a ternary tree (TT), and a tree that uses partitioning into four partitions is called a quadary tree (QT).
[0135] As stated herein, the term “block” may refer to a part of a picture, particularly a square or rectangle. For example, with reference to HEVC and VVC, a block may be a coding tree unit (CTU), coding unit (CU), prediction unit (PU), and transformation unit (TU), and / or a corresponding block, such as a coding tree block (CTB), coding block (CB), transformation block (TB), or prediction block (PB), or may correspond to these.
[0136] For example, a coding tree unit (CTU) may be, or include, a CTB of a lumen sample of a picture having three sample arrays, two corresponding CTBs of a chroma sample, or a CTB of a sample of a picture coded using three separate color planes and syntax structures used to code a monochrome picture or sample. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some value of N, such that the division of a component into a CTB is a partition division. A coding unit (CU) may be, or include, a coding block of a lumen sample of a picture having three sample arrays, two corresponding coding blocks of a chroma sample, or a coding block of a sample of a picture coded using three separate color planes and syntax structures used to code a monochrome picture or sample. Correspondingly, a coding block (CB) may be an M×N block of samples for some values of M and N, such that the division of a CTB into a coding block is a partition division.
[0137] In embodiments, for example, according to HEVC, a coding tree unit (CTU) may be partitioned into CUs by using a quadtree structure referred to as a coding tree. The decision of whether to code a picture region using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. Each CU can further be partitioned into one, two, or four PUs, depending on the PU partitioning type. Within a single PU, the same prediction process is applied, and relevant information is transmitted to the decoder for each PU. After obtaining residual blocks by applying the prediction process based on the PU partitioning type, the CU can be partitioned into transformation units (TUs) according to another quadtree structure similar to the coding tree for that CU.
[0138] In embodiments, according to a modern video coding standard currently under development, for example, called Multipurpose Video Coding (VVC), quadtree and binary tree (QTBT) partitioning is used to partition coding blocks. In the QTBT block structure, CUs can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quadtree structure. The quadtree leaf nodes are further partitioned by a binary or ternary (or trident) structure. The partitioned tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transformation processing without further partitioning. This means that in the QTBT coding block structure, CUs, PUs, and TUs have the same block size. In parallel, it has also been proposed that multi-partitioning, such as ternary tree partitioning, be used in conjunction with the QTBT block structure.
[0139] In one example, the mode selection unit 260 of the video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.
[0140] As described above, the video encoder 20 is configured to determine or select the best or most optimal prediction mode from a set of (predetermined) prediction modes. The set of prediction modes may include, for example, an intra-prediction mode and / or an inter-prediction mode.
[0141] Intra Prediction The set of intra-prediction modes may include 35 different intra-prediction modes, such as DC (or mean) modes and non-directional modes like planar modes, or directional modes, as defined, for example in HEVC, or it may include 67 different intra-prediction modes, such as DC (or mean) modes and non-directional modes like planar modes, or directional modes, as defined, for example in VVC.
[0142] The intra-prediction unit 254 is configured to use reconfigured samples from neighboring blocks of the same current picture to generate an intra-prediction block 265 according to a certain intra-prediction mode from a set of intra-prediction modes.
[0143] The intra-prediction unit 254 (or, generally, the mode selection unit 260) is further configured to output intra-prediction parameters (or, generally, information indicating the selected intra-prediction mode for that block) to the entropy encoding unit 270 in the form of syntactic elements 266, to be included in the encoded picture data 21. Thereafter, for example, the video decoder 30 can receive the prediction parameters and use them for decoding.
[0144] Interpretation The (or possible) interpretation modes of the set depend on the available reference picture (i.e., a previously at least partially decoded picture stored, for example, in DBP 230) and other interpretation parameters. These other interpretation parameters include, for example, whether the entire reference picture is used to find the best-matching reference block, or only a portion of the reference picture, for example, a search window area around that region of the current block, and / or whether, for example, pixel interpolation, for example, half / semi-pixel and / or quarter-pixel interpolation is applied.
[0145] In addition to the prediction modes described above, skip mode and / or direct mode may also be applied.
[0146] The interpretation unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (neither of which are shown in Figure 2). The motion estimation unit may be configured to receive or acquire for motion estimation a picture block 203 (the current picture block 203 of the current picture 17) and a decoded picture 231, or at least one or more previously reconstructed blocks, for example, one or more other / different reconstructed blocks of a previously decoded picture 231. For example, a video sequence may include the current picture and a previously decoded picture 231, or in other words, the current picture and a previously decoded picture 231 may be part of a sequence of pictures that make up the video sequence, or may make up such sequence.
[0147] The encoder 20 may be configured, for example, to select a reference block from multiple reference blocks of the same or different pictures among several other pictures, and to provide the reference picture (or reference picture index) and / or the offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as an interprediction parameter to the motion estimation unit. This offset is also called the motion vector (MV).
[0148] The motion compensation unit is configured to obtain, for example, interprediction parameters, and to perform interprediction based on or using said interprediction parameters to obtain interprediction blocks 265. Motion compensation performed by the motion compensation unit may include taking or generating prediction blocks by performing interpolation, possibly to sub-pixel precision, based on the motion / block vector determined by motion estimation. Interpolation filtering may generate additional pixel samples from known pixel samples, thus potentially increasing the number of candidate prediction blocks that can be used to code picture blocks. Now receiving motion vectors for the picture block PU, the motion compensation unit can locate the prediction block that the motion vectors point to in one of the reference picture lists.
[0149] The motion compensation unit may also generate syntactic elements related to blocks and video slices for use by the video decoder 30 when decoding picture blocks of video slices.
[0150] Entropy coding The entropy encoding unit 270 is configured to obtain encoded picture data 21 by applying, for example, an entropy encoding algorithm or scheme (e.g., variable-length coding (VLC) scheme, context-adaptive VLC scheme (CAVLC), arithmetic coding scheme, binarization, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding, or other entropy encoding methods or techniques) or bypass (uncompressed) to quantized coefficients 209, inter-prediction parameters, intra-prediction parameters, loop filter parameters and / or other syntactic elements. The encoded picture data 21 can be output via output 272, for example, in the form of an encoded bitstream 21, so that, for example, a video decoder 30 can receive its parameters and use them for decoding. The encoded bitstream 21 may be sent to the video decoder 30 or stored in memory for later transmission or retrieval by the video decoder 30.
[0151] Other structural variations of the video encoder 20 can be used to encode the video stream. For example, a non-conversion-based encoder 20 can quantize the residual signal directly for certain blocks or frames without a conversion processing unit 206. In another implementation, the encoder 20 may have a single unit that combines a quantization unit 208 and an inverse quantization unit 210.
[0152] Decoder and decoding method Figure 3 shows an example of a video decoder 30 configured to implement the technology of the present invention. The video decoder 30 is configured to receive encoded picture data 21 (e.g., encoded bitstream 21), which has been encoded by, for example, an encoder 20, in order to obtain a decoded picture 331. The encoded picture data or bitstream includes information for decoding the encoded picture data, such as data representing picture blocks and associated syntactic elements of an encoded video slice.
[0153] In the example in Figure 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transformation unit 312, a reconstruction unit 314 (e.g., an adder 314), a loop filter 320, a decode picture buffer (DBP) 330, an interpretation unit 344, and an intraprediction unit 354. The interpretation unit 344 may be or may include a motion compensation unit. In some examples, the video decoder 30 can perform a decoding path that is generally the reverse of the encoding path described with respect to the video encoder 100 from Figure 2.
[0154] As described with respect to encoder 20, the inverse quantization unit 210, inverse processing unit 212, reconstruction unit 214, loop filter 220, decode picture buffer (DPB) 230, inter-prediction unit 344, and intra-prediction unit 354 are also referred to as the “built-in decoder” of video encoder 20. Thus, the inverse quantization unit 310 may be functionally identical to the inverse quantization unit 110, the inverse processing unit 312 may be functionally identical to the inverse processing unit 212, the reconstruction unit 314 may be functionally identical to the reconstruction unit 214, the loop filter 320 may be functionally identical to the loop filter 220, and the decode picture buffer 330 may be functionally identical to the decode picture buffer 230. Accordingly, the descriptions given for each unit and function of the video encoder 20 apply correspondingly to each unit and function of the video decoder 30.
[0155] Entropy Decoding The entropy decode unit 304 is configured to parse the bitstream 21 (or generally, the encoded picture data 21), perform entropy decoding on the encoded picture data 21, for example, to obtain any or all of the following: quantized coefficients 309 and / or decoded coding parameters (not shown in Figure 3), such as inter-prediction parameters (e.g., reference picture index and motion vector), intra-prediction parameters (e.g., intra-prediction mode or index), transformation parameters, quantization parameters, loop filter parameters, and / or other syntactic elements. The entropy decode unit 304 may be configured to apply a decoding algorithm or scheme corresponding to the encoding scheme described with respect to the entropy encoding unit 270 of the encoder 20. The entropy decode unit 304 may be further configured to provide the inter-prediction parameters, intra-prediction parameters, and / or other syntactic elements to the mode selection unit 360, and the other parameters to other units of the decoder 30. The video decoder 30 may receive syntactic elements at the video slice level and / or video block level.
[0156] inverse quantization The inverse quantization unit 310 may be configured to receive quantization parameters (QP) (or generally, information about inverse quantization) and quantized coefficients from the encoded picture data 21 (for example, by parsing and / or decoding by the entropy decoding unit 304), and to apply inverse quantization to the decoded quantized coefficients 309 based on the quantization parameters to obtain dequantized coefficients 311. The dequantized coefficients 311 may also be referred to as transformed coefficients 311. The inverse quantization process may use the quantization parameters determined by the video encoder 20 for each video block in the video slice to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied.
[0157] Inverse Transform The inverse transformation processing unit 312 may be configured to receive dequantized coefficients 311, also called transformation coefficients 311, and to apply a transformation to the dequantized coefficients 311 in order to obtain a reconstructed residual block 213 in the sample region. The reconstructed residual block 213 may also be referred to as the transformation block 313. The transformation may be an inverse transformation, such as an inverse DCT, inverse DST, inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 may further be configured to receive transformation parameters or corresponding information from the encoded picture data 21 (for example, by parsing and / or decoding by the entropy decoding unit 304) to determine which transformation to apply to the dequantized coefficients 311.
[0158] Reconstruction The reconstruction unit 314 (for example, an adder or summer 314) may be configured to add the reconstructed residual block 313 to the prediction block 365 to obtain the reconstructed block 315 in the sample region. This is done, for example, by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365.
[0159] filtering The loop filter unit 320 (either within or after the coding loop) is configured to filter the reconstructed block 315 to obtain a filtered block 321, for example, to smooth pixel transitions or otherwise improve video quality. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a cooperative filter, or any combination thereof. Although the loop filter unit 320 is shown as an in-loop filter in Figure 3, in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.
[0160] Decode picture buffer Next, the decoded video block 321 of the picture is stored in the decoded picture buffer 330. The decoded picture buffer 330 stores the decoded picture 331 as a reference picture for subsequent motion compensation for other pictures and / or for their respective output or display.
[0161] The decoder 30 is configured to output the decoded picture 311, for example via output 312, for presentation or viewing to the user.
[0162] prediction The inter-prediction unit 344 may be identical to the inter-prediction unit 244 (in particular, the motion compensation unit), and the intra-prediction unit 354 may be functionally identical to the inter-prediction unit 254, and performs partitioning or partitioning decision and prediction based on partitioning and / or prediction parameters or respective information received from the encoded picture data 21 (for example by parsing and / or decoding by the entropy decoding unit 304). The mode selection unit 360 may be configured to perform block-by-block prediction (intra-prediction or inter-prediction) based on the reconstructed picture, block, or each sample (filtered or unfiltered) to obtain a predicted block 365.
[0163] When a video slice is coded as an intra-coded (I) slice, the intra-prediction unit 354 of the mode selection unit 360 is configured to generate a prediction block 365 for the picture block of the current video slice based on the signaled intra-prediction mode and data from previously decoded blocks of the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, the inter-prediction unit 344 (e.g., a motion compensation unit) of the mode selection unit 360 is configured to generate a prediction block 365 for the video block of the current video slice based on motion vectors and other syntactic elements received from the entropy decode unit 304. For inter-prediction, the prediction block may be generated from one of the reference pictures in one of the list of reference picture lists. The video decoder 30 can construct the reference frame lists, list 0 and list 1, using default construction techniques based on reference pictures stored in the DPB 330.
[0164] The mode selection unit 360 is configured to determine predictive information about the video blocks of the current video slice by parsing motion vectors and other syntactic elements, and to use this predictive information to generate a predictive block for the current video block to be decoded. For example, the mode selection unit 360 uses some of the received syntactic elements to determine the predictive mode used to code the video blocks of the video slice (e.g., intra-predictive or inter-predictive), the inter-predictive slice type (e.g., B-slice, P-slice, or GPB-slice), construction information for one or more lists of reference picture lists for the slice, motion vectors for each inter-encoded video block of the slice, the inter-predictive status for each intercoded video block of the slice, and other information for decoding the video blocks in the current video slice.
[0165] Other variations of the video decoder 30 can be used to decode the encoded picture data 21. For example, the decoder 30 can generate an output video stream without a loop filtering unit 320. For example, a non-transformation-based decoder 30 can dequantize the residual signal directly for certain blocks or frames without an inverse processing unit 312. In another embodiment, the video decoder 30 may have a combination of an inverse quantization unit 310 and an inverse processing unit 312 in a single unit.
[0166] The encoder 20 and decoder 30 should understand that the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.
[0167] It should be noted that further operations can be applied to the currently derived motion vectors of a block (including, but not limited to, affine mode control point motion vectors, affine, planar, and ATMVP mode subblock motion vectors, and temporal motion vectors). For example, the value of a motion vector is constrained to a predefined range according to its representation bits. If the representation bits of a motion vector are bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1. For example, if bitDepth is set to 16, the range is -32768 to 32767, and if bitDepth is set to 18, the range is -131072 to 131071. For example, the value of a derived motion vector (e.g., the MV of four 4x4 subblocks in one 8x8 block) is constrained such that the maximum difference between the integer parts of the four 4x4 subblock MVs is less than or equal to N pixels, for example, less than or equal to 1 pixel. Here, we provide two methods for constraining the motion vector according to bitDepth.
[0168] Method 1: Remove the overflow MSB (most significant bit) using the following calculation.
number
number
[0169] Method 2: Remove the overflow MSB by clipping the value.
number
number
[0170] Figure 4 is a schematic diagram of a video coding device 400 according to one embodiment of the present disclosure. The video coding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video coding device 400 may be a decoder, such as the video decoder 30 in Figure 1A, or an encoder, such as the video encoder 20 in Figure 1A.
[0171] The video coding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit (CPU) 430 for processing data; a transmitter unit (Tx) 440 and an exit port 450 (or output port 450) for transmitting data; and memory 460 for storing data. The video coding device 400 may also have optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the inlet port 410, receiver unit 420, transmitter unit 440, and exit port 450 for the entry and exit of optical or electrical signals.
[0172] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (for example, as a multi-core processor), FPGAs, ASICs, and DSPs. The processor 430 communicates with the inlet port 410, the receiver unit 420, the transmitter unit 440, the exit port 450, and the memory 460. The processor 430 has a coding module 470. The coding module 470 implements the embodiments disclosed above. For example, the coding module 470 implements, processes, prepares, or provides various coding operations. Thus, by including the coding module 470, the functionality of the video coding device 400 is substantially improved and conversion of the video coding device 400 to different states is realized. Alternatively, the coding module 470 may be implemented as instructions stored in the memory 460 and executed by the processor 430.
[0173] Memory 460 may include one or more disks, tape drives, and solid-state drives, and may be used as an overflow data storage device to store programs when the program is selected for execution, and to store instructions and data to be read during program execution. Memory 460 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), tertiary associative memory (TCAM), and / or static random access memory (SRAM).
[0174] Figure 5 is a simplified block diagram of a device 500 that may be used as either or both of the source device 12 and destination device 14 from Figure 1, according to an exemplary embodiment.
[0175] The processor 502 within the device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or multiple devices capable of manipulating or processing information that currently exists or will be developed in the future. The disclosed implementation can be carried out using a single processor, for example, processor 502, as shown in the figure, but advantages in speed and efficiency can be achieved by using multiple processors.
[0176] The memory 504 within the device 500 may, in some implementations, be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as memory 504. Memory 504 may contain code and data 506 accessed by the processor 502 using the bus 512. Memory 504 may further contain an operating system 508 and an application program 510, the application program 510 containing at least one program that allows the processor 502 to perform the method described herein. For example, the application program 510 may contain applications 1 to N, which further contain a video coding application that performs the method described herein.
[0177] The device 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines the display with a touch-sensitive element that can operate to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512.
[0178] Although shown here as a single bus, the bus 512 of device 500 can consist of multiple buses. Furthermore, the secondary storage 514 can be directly coupled to other components of device 500 or accessed via a network, and can include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, device 500 can be implemented in a wide variety of configurations.
[0179] In-loop filtering VTM3 has a total of three in-loop filters. In addition to the deblocking filter and SAO (the two loop filters in HEVC), VTM3 applies an adaptive loop filter (ALF). The filtering process in VTM3 is in the order of deblocking filter, SAO, and ALF.
[0180] ALF In VTM5, an Adaptive Loop Filter (ALF) with block-based filter adaptation is applied. For the luma component, one of 25 filters is selected for each 4x4 block based on the direction and function of the local gradient.
[0181] Filter shape: In JEM, two diamond filter shapes (shown in Figure 6) are used for the lumern component. A 7x7 diamond shape is applied to the lumern component, and a 5x5 diamond shape is applied to the chromatic component.
[0182] Block classification: For the luma components, each 4x4 block is categorized into one of 25 classes. The classification index C is the quantized value of its direction D and function.
number
number
number
[0183] To reduce the complexity of block classification, a subsampled 1-D Laplacian calculation is applied. As shown in Figure 7, the same subsampled locations are used for gradient calculations in all directions.
[0184] Next, the maximum and minimum values for the horizontal and vertical slopes are set as follows:
number
number
[0185] To derive the directional value D, these values are compared with each other and with two thresholds t1 and t2:
number
[0186] The working value A is calculated as follows:
number
number
[0187] No classification method is applied to the chroma components within the picture. That is, a single set of ALF coefficients is applied to each chroma component.
[0188] Geometric transformation of filter coefficients Before filtering each 4x4 lumen block, geometric transformations such as rotation or diagonal and vertical inversions are applied to the filter coefficients f(k,l), depending on the gradient value calculated for that block. This is equivalent to applying these transformations to the samples within the filter support region. The idea is to make different blocks to which ALF is applied more similar by aligning their orientations.
[0189] Three geometric transformations are introduced, including diagonal, vertical inversion, and rotation:
number
[0190] Signal transmission of filter parameters In VTM3, ALF filter parameters are transmitted in the slice header. Up to 25 lumar filter coefficients can be transmitted. To reduce bit overhead, filter coefficients from different classifications can be merged.
[0191] The filtering process can be controlled at the CTB level. A flag indicating whether ALF is applied to the chroma CTB is always signaled. For each chroma CTB, a flag indicating whether ALF is applied to the chroma CTB may also be signaled, depending on the value of alf_chroma_ctb_present_flag.
[0192] The filter coefficients are quantized using a norm equal to 128. To further constrain the complexity of the multiplication, the coefficient values at the central position are between 0 and 2. 8 It is within the range, and the coefficient value for the remaining positions is -2 7 From 2 7 The bitstream conformance requirement applies, which stipulates that the bitstream must be within the range of -1 (including both ends). Filtering process On the decoder side, when ALF is enabled for a given CTB, each sample R(i,j) in the CU is filtered, resulting in the sample value R'(i,j) shown below. Here, L represents the filter length, and f m,n represents the filter coefficients, and f(k,l) represents the decoded filter coefficients.
number
[0193] Alternatively, filtering can also be expressed as follows:
number
number
[0194] From VTM5 onward (ITU JVET-N0242), ALF is executed in a non-linear manner. Equation 21 can be formulated as follows:
number
number
[0195] The filter is further modified by introducing nonlinearity to make ALF more efficient. This is done by using a clipping function to reduce the influence of neighboring sample values when they are too different from the current sample value being filtered (I(x,y)).
[0196] In VTM5, the ALF filter is modified as follows:
number
[0197] To limit signal transmission costs and encoder complexity, the evaluation of clipping values is reduced to a small set of possible values. In VTM5, only four possible fixed values are used, which are the same for inter and intra-tile groups.
[0198] Because the variance of local differences is often greater for "luma" than for "chroma," two different sets of clipping values are used: one for the "luma" filter and one for the "chroma" filter. The clipping values also include the maximum sample value in each set (1024 here for a 10-bit depth), so clipping can be disabled if not needed.
[0199] Table 2 shows the set of clipping values used in VTM5. The four values are selected by dividing the entire range of sample values (coded in 10 bits) for the lumens and the range from 4 to 1024 for the chromens in roughly equal proportions in the logarithmic domain. More precisely, the lumens table of clipping values is obtained by the following formula:
number
number
[0200] Another filter in a loop VVC has a total of three in-loop filters. In addition to the deblocking filter and SAO (the two loop filters in HEVC), an adaptive loop filter (ALF) is applied. The ALF includes the lumar ALF, chromar ALF, and cross-component ALF (CC-ALF). The ALF filtering process is designed so that the lumar ALF, chromar ALF, and CC-ALF can run in parallel. The order of the filtering process in VVC is deblocking filter, SAO, and ALF. The SAO in VVC is the same as that in HEVC.
[0201] VVC adds a new process called luma mapping with chroma scaling (formerly known as adaptive in-loop reshaper). LMCS corrects pre-encoded and post-reconstruction sample values by redistributing codewords across the entire dynamic range. This new process is performed before deblocking.
[0202] Adaptive Loop Filter In VVC, an Adaptive Loop Filter (ALF) with block-based filter adaptation is applied. For the luma component, one of 25 filters is selected for each 4x4 block based on the direction and effect of the local gradient.
[0203] Filter shape: Two diamond filter shapes (shown in Figure 6) are used. A 7x7 diamond shape is applied to the lumens component, and a 5x5 diamond shape is applied to the chromatic component.
[0204] Block classification: For the ruma components, each 4x4 block is categorized into one of 25 classes. The classification index C is derived based on its direction D and the quantized value of its action ^A as follows:
number
[0205] To calculate D and ^A, first, the horizontal, vertical, and two diagonal gradients are calculated using the 1-D Laplacian:
number
[0206] To reduce the complexity of block classification, subsampled 1-D Laplacian calculation is applied. As shown in FIG. 7, the same subsampled positions are used for gradient calculations in all directions.
[0207] Next, the maximum and minimum values of the horizontal and vertical gradients D are set as follows:
Equation
Equation
[0208] To derive the directional value D, these values are compared with each other and with two threshold values t1 and t2:
Equation
[0209] The working value A is calculated as follows:
Equation
Equation
[0210] For the chroma components in the picture, the classification method is not applied. That is, for each chroma component, a single set of ALF coefficients is applied.
[0211] Geometric transformation of filter coefficients and clipping values Before filtering each 4x4 lumen block, geometric transformations such as rotation or diagonal and vertical inversion are applied to the filter coefficients f(k,l) and the corresponding filter clipping values c(k,l), depending on the gradient values calculated for that block. This is equivalent to applying these transformations to the samples within the filter support region. The idea is to make different blocks to which ALF is applied more similar by aligning their orientations.
[0212] Three geometric transformations are introduced, including diagonal, vertical inversion, and rotation:
number
[0213] Signal transmission of filter parameters ALF filter parameters are transmitted in an Adaptive Parameter Set (APS). A single APS can transmit up to 25 sets of lumern filter coefficients and clipping value indices, and up to 8 sets of chroma filter coefficients and clipping value indices. To reduce bit overhead, filter coefficients from different classifications of lumern components can be merged. In the slice header, the indices of the APS used for the current slice are transmitted.
[0214] The clipping value index decoded from the APS allows determining the clipping value using a table of clipping values for both the lumen and chroma components. These clipping values depend on the internal bit depth. More precisely, the clipping value is obtained by the following formula:
number
[0215] The slice header can signal up to seven APS indices to specify the rumor filter set used for the current slice. The filtering process can be further controlled at the CTB level. A flag is always signaled indicating whether the ALF is applied to the rumor CTB. The rumor CTB can select a filter set from 16 fixed filter sets and the aforementioned filter sets from the APS. A filter set index is signaled for the rumor CTB to indicate which filter set is applied. The 16 fixed filter sets are predefined and hardcoded into both the encoder and decoder.
[0216] For chroma components, the APS index is transmitted in the slice header to indicate the set of chroma filters currently used in the slice. At the CTB level, if the APS has multiple chroma filter sets, the filter index is transmitted for each chroma CTB.
[0217] The filter coefficients are quantized using a norm equal to 128. To constrain the complexity of the multiplication, the coefficient values for non-central positions are -2 7 From 2 7Bitstream compliance is applied which must be within the range of -1 (including both ends). The central position coefficient is not signaled in the bitstream and is considered equal to 128.
[0218] Filtering process On the decoder side, when ALF is enabled for a certain CTB, each sample R(i,j) within the CU is filtered, and as a result, sample values R'(i,j) as shown below are obtained.
Equation
[0219] The selected clipping value is coded in the "alf_data" syntax element using the Golomb encoding method corresponding to the index of the clipping value in Table 1 above. This encoding method is the same as the encoding method for the filter index. alf_data may be within adaptation_parameter_set_rbsp(), and adaptation_parameter_set_rbsp() may be referenced by the slice header.
[0220] The details of the syntax are shown in the following table.
Table 4-1
Table 4-2
[0221] The meanings of the newly introduced syntactic elements are as follows: an alf_luma_clip equal to 0 specifies that a linear adaptive loop filter is applied to the lumen component. An alf_luma_clip equal to 1 specifies that a nonlinear adaptive loop filter may be applied to the lumen component.
[0222] If alf_chroma_clip is equal to 0, it specifies that a linear adaptive loop filter is applied to the chroma components, and if alf_chroma_clip is equal to 1, it specifies that a nonlinear adaptive loop filter is applied to the chroma components. If it does not exist, alf_croma_clip is assumed to be 0.
[0223] Adding 1 to alf_luma_clip_min_eg_order_minus1 specifies the minimum order of the exponential Golomb code for Luma Clipping Index signal transmission. The value of alf_luma_clip_min_eg_order_minus1 must be in the range of 0 to 6 (including both ends).
[0224] If alf_luma_clip_eg_order_increase_flag[i] is equal to 1, it specifies that the minimum order of the exponential Golomb code for luma clipping index signaling is incremented by 1. If alf_luma_clip_eg_order_increase_flag[i] is equal to 0, it specifies that the minimum order of the exponential Golomb code for luma clipping index signaling is not incremented by 1. The exponential Golom code order expGoOrderYClip[i] used to decode the value of alf_luma_clip_idx[sigFiltIdx][j] is derived as follows: expGoOrderYClip[i] = alf_luma_clip_min_eg_order_minus1 + 1 + alf_luma_clip_eg_order_increase_flag[i]
[0225] alf_luma_clip_idx[sigFiltIdx][j] specifies the clipping index of the clipping value to be used before multiplying the j-th coefficient of the signal-transmitted luma filter indicated by sigFiltIdx. If alf_luma_clip_idx[sigFiltIdx][j] does not exist, it is assumed to be equal to 0 (no clipping). The degree k of the exponential Golomb binary evolution uek(v) is derived as follows: golombOrderIdxYClip[] = {0, 0, 1, 0, 0, 1, 2, 1, 0, 0, 1, 2} k = expGoOrderYClip[golombOrderIdxYClip[j]]
[0226] Assuming sigFiltIdx = 0..alf_luma_num_filters_signalled_minus1 and j = 0..11, the variable filterClips[signFiltIdx][j] is initialized as follows:
Table 5
[0227] Adding 1 to alf_chroma_clip_min_eg_order_minus1 specifies the minimum order of the exponential Golomb code for chroma clipping index signal transmission. The value of alf_chroma_clip_min_eg_order_minus1 must be in the range of 0 to 6 (including both ends).
[0228] If alf_chroma_clip_eg_order_increase_flag[i] is equal to 1, it specifies that the minimum order of the exponential Golomb code for chroma clipping index signaling is incremented by 1. If alf_chroma_clip_eg_order_increase_flag[i] is equal to 0, it specifies that the minimum order of the exponential Golomb code for chroma clipping index signaling is not incremented by 1. The exponential Golom coding order expGoOrderC[i] used to decode the value of alf_chroma_clip_idx[j] is derived as follows: expGoOrderC[i]=alf_chroma_clip_min_eg_order_minus1+1+alf_chroma_clip_eg_order_increase_flag[i]
[0229] alf_chroma_clip_idx[j] specifies the clipping index of the clipping value to use before multiplying by the j-th coefficient of the chroma filter. If alf_chroma_clip_idx[j] does not exist, it is assumed to be 0 (no clipping). The degree k of the exponential Golomb binary uek(v) is derived as follows: golombOrderIdxC[]={0,0,1,0,0,1} k=expGoOrderC[golombOrderIdxC[j]] Set j=0..5 to the element AlfClip C Chroma filter clipping value AlfClip with [j] C This is derived as follows: [Table 6]
[0230] ALF syntax specification compliant with VVC specification Adaptive Loop Filter Process 1.1 General The inputs to this process are reconfigured picture sample arrays recPictureL, recPictureCb, and recPictureCr, which precede an adaptive loop filter. The output of this process is the modified and reconstructed picture sample arrays alfPictureL, alfPictureCb, and alfPictureCr after adaptive loop filtering.
[0231] The sample values in the modified and reconstructed picture sample arrays alfPictureL, alfPictureCb, and alfPictureCr after adaptive loop filtering are initially set to be equal to the sample values in the reconstructed picture sample arrays recPictureL, recPictureCb, and recPictureCr prior to adaptive loop filtering, respectively.
[0232] If the value of tile_group_alf_enabled_flag is equal to 1, then for all coding tree units with a ruma coding tree block position (rx,ry), assuming rx=0..PicWidthInCtbs-1 and ry=0..PicHeightInCtbs-1, the following process is applied: When the value of alf_ctb_flag[0][rx][ry] is equal to 1, the coding tree block filtering process for the luma samples specified in section 1.2 is called. The luma coding tree block position (xCtb, yCtb) set equal to recPictureL, alfPictureL, and (rx<<CtbLog2SizeY,ry<<CtbLog2SizeY) is input, and the output is the modified filtered picture alfPictureL. When the value of alf_ctb_flag[1][rx][ry] is equal to 1, the coding tree block filtering process for the chroma samples specified in section 1.1 is called. The recPicture set equal to recPictureCb, the alfPicture set equal to alfPictureCb, and the chroma coding tree block position (xCtbC, yCtbC) set equal to (rx<<(CtbLog2SizeY-1),ry<<(CtbLog2SizeY-1)) are input, and the output is the modified filtered picture alfPictureCb. When the value of alf_ctb_flag[2][rx][ry] is equal to 1, the coding tree block filtering process for the chroma samples specified in section 1.4 is called. The recPicture set equal to recPictureCr, the alfPicture set equal to alfPictureCr, and the chroma coding tree block position (xCtbC, yCtbC) set equal to (rx<<(CtbLog2SizeY-1),ry<<(CtbLog2SizeY-1)) are input, and the output is the modified filtered picture alfPictureCr.
[0233] 1.2 Coding Tree Block Filtering Process for Ruma Samples The inputs to this process are as follows: Reconfigured luma picture sample array recPictureL prior to adaptive loop filtering process, Filtered and reconstructed Luma Picture sample array alfPictureL, The rumor position (xCtb, yCtb) that specifies the top-left sample of the current rumor coding tree block relative to the top-left sample of the current picture.
[0234] The output of this process is the modified, filtered, and reconstructed lumens picture sample array alfPictureL.
[0235] The derivation process for the filter index in Section 1.3 is invoked. The input is the position (xCtb, yCtb) and the reconstructed lumen picture sample array recPictureL, and the output is filtIdx[x][y] and transposeIdx[x][y], where x, y = 0..CtbSizeY-1.
[0236] To derive the filtered and reconstructed ruma sample alfPictureL[x][y], each reconstructed ruma sample in the current ruma coding tree block recPictureL[x][y] is filtered as follows, with x,y=0..CtbSizeY-1 The array f[j] of lumar filter coefficients corresponding to the filter specified by filtIdx[x][y] is derived as follows, with j=0..12: f[j]=AlfCoeffL[filtIdx[x][y]][j] The array c[j] of lumar filter clipping values corresponding to the filter specified by filtIdx[x][y] is derived as follows, with j=0..11: c[j]=AlfClipL[filtIdx[x][y]][j]
[0237] The Luma filter coefficients filterCoeff are derived as follows, depending on transposeIdx[x][y]: [Table 7]
[0238] The positions (hx, vy) for each corresponding ruma sample (x, y) within a given ruma sample sequence recPicture are derived as follows: [Table 8]
[0239] The variable sum is derived as follows: [Table 9-1] [Table 9-2]
[0240] The modified, filtered, and reconstructed Luma picture sample alfPictureL[xCtb+x][yCtb+y] is derived as follows: [Table 10]
[0241] 1.3 Derivation process for ALF transpose and filter index for lumar samples The inputs for this process are as follows: The rumor position (xCtb, yCtb) specifies the top-left sample of the current rumor coding tree block relative to the top-left sample of the current picture. Reconfigured luma picture sample array recPictureL prior to adaptive loop filtering process
[0242] The output of this process is as follows: Classification filter index array filtIdx[x][y], where x,y = 0..CtbSizeY-1 The transpose index array transposeIdx[x][y], where x,y = 0..CtbSizeY-1
[0243] The positions (hx, vy) for each corresponding ruma sample (x, y) within a given ruma sample sequence recPicture are derived as follows: hx=Clip3(0,pic_width_in_luma_samples-1,x) vy=Clip3(0,pic_height_in_luma_samples-1,y)
[0244] The classification filter index array filtIdx and the transpose index array transposeIdx are derived by the following ordered steps: Assuming x,y = -2..CtbSizeY+1, the variables filtH[x][y], filtV[x][y], filtD0[x][y], and filtD1[x][y] are derived as follows: If both x and y are even numbers, or if both x and y are uneven numbers, then the following applies: [Table 11] Otherwise, filtH[x][y], filtV[x][y], filtD0[x][y], and filtD1[x][y] are set to 0.
[0245] Assuming x,y=0..(CtbSizeY-1)>>2, the variables varTempH1[x][y], varTempV1[x][y], varTempD01[x][y], varTempD11[x][y], and varTemp[x][y] are derived as follows: [Table 12]
[0246] The variables hv1, hv0, and dirHV are derived as follows: [Table 13]
[0247] The variables d1, d0, and dirD are derived as follows: [Table 14]
[0248] The variables hvd1 and hvd0 are derived as follows: [Table 15]
[0249] The variables dirS[x][y], dir1[x][y], and dir2[x][y] are derived as follows: [Table 16]
[0250] Assuming x, y = 0..CtbSizeY-1, the variables avgVar[x][y] are derived as follows: [Table 17]
[0251] Assuming x=y=0..CtbSizeY-1, the classification filter index array filtIdx[x][y] and the transpose index array transposeIdx[x][y] are derived as follows: [Table 18] If dirS[x][y] is not equal to 0, then filtIdx[x][y] is modified as follows: [Table 19]
[0252] 1.4 Coding tree block filtering process for chroma samples The inputs for this process are as follows: Reconstructed chroma-picture sample array recPicture prior to adaptive loop filtering process, Filtered and reconstructed chroma picture sample array alfPicture, The chroma position (xCtbC, yCtbC) specifies the top-left sample of the current chroma coding tree block relative to the top-left sample of the current picture.
[0253] The output of this process is a modified, filtered, and reconstructed chroma-picture sample array, alfPicture.
[0254] The current size of the chroma coding tree block ctbSizeC is derived as follows: ctbSizeC = CtbSizeY / SubWidthC
[0255] To derive the filtered and reconstructed chroma sample alfPicture[x][y], each reconstructed chroma sample in the current chroma coding tree block recPicture[x][y] is filtered as follows, with x,y = 0..ctbSizeC-1:
[0256] The position (hx, vy) for each corresponding chromatic sample (x, y) in a given array recPicture is derived as follows: [Table 20]
[0257] The variable sum is derived as follows: [Table 21]
[0258] The modified, filtered, and reconstructed chroma picture sample alfPicture[xCtbC+x][yCtbC+y] is derived as follows: [Table 22]
[0259] As described above and shown in Figure 8, the ALF luma and chroma clipping parameters are transmitted using a K-order exponential Golomb code similar to the ALF filter coefficients.
[0260] The use of K-order exponential Golomb codes for clipping parameters can be inefficient in terms of coding efficiency. This is because the clipping parameters being signaled are simply indices to a table of clipping values (see Table 2 above). The index values range from 0 to 3. Therefore, signaling index values from 0 to 3 using K-order exponential Golomb codes, similar to the ALF filter coefficients, requires determining the value K (the order of the exponential Golomb code used) using additional syntactic elements alf_luma_clip_min_eg_order_min_minus1, alf_luma_clip_eg_order_increase_flag[i], and then the syntactic element alf_luma_clip_idx is signaled using the K-order exponential Golomb code. Consequently, this method of signaling is complex and inefficient in terms of coding efficiency. Therefore, a simpler method for signaling clipping parameters is desired.
[0261] In one embodiment of the proposed solution (Solution 1), as shown in Figure 9, the clipping parameters are signaled using a fixed-length code, and thus the syntactic elements alf_luma_clip_min_eg_order_minus1 and alf_luma_clip_eg_order_increase_flag[i] are not used. The syntactic element alf_luma_clip_idx is signaled using a 2-bit fixed-length code. This method has the advantage of improved coding efficiency because the clipping parameters are signaled in a very simple manner and most of the syntactic elements related to the K-order exponential Golomb code are no longer signaled.
[0262] The corrected alf_data syntax is as follows: [Table 23-1] [Table 23-2] [Table 23-3]
[0263] As an embodiment of the alternative solution (Solution 2), a truncated unary coding may be used to signal the clipping parameter index.
[0264] As an embodiment of the alternative solution (Solution 3), if the number of clipping parameters is changed from a fixed value of 4 to a different number greater than 4, the fixed-length code (v) "value v" increases accordingly. For example, if the number of clipping parameters increases from 4 to 5 or 6, the fixed-length code uses 3 bits to signal the clipping parameters.
[0265] As an embodiment of the alternative solution (Solution 4), the ALF filter coefficients are transmitted using a fixed-length code instead of a K-order exponential Golomb code.
[0266] Figure 10 is a block diagram illustrating a method according to a first aspect of the present disclosure. The method includes the steps of: acquiring a bitstream in which at least one bit in the bitstream represents a syntactic element for the current block (1001), the syntactic element being an adaptive loop filter (ALF) clipping value index specifying a clipping index of a clipping value to be used before multiplying by a coefficient of ALF; parsing the bitstream to obtain the value of the syntactic element for the current block (1002); and applying adaptive loop filtering to the current block based on the value of the syntactic element for the current block (1003).
[0267] Figure 11 is a block diagram illustrating a method according to a second aspect of the present disclosure. According to the second aspect, a coding method is provided which is implemented by a decoding device. The method includes the steps of: acquiring a bitstream in which at least one bit in the bitstream represents a syntactic element for the current block (1101), the syntactic element being an adaptive loop filter (ALF) clipping value index or an ALF coefficient parameter; parsing the bitstream to obtain a value of the syntactic element for the current block, the value of the syntactic element for the current block is obtained by using only the at least one bit of the syntactic element (1102); and applying adaptive loop filtering to the current block based on the value of the syntactic element for the current block (1103).
[0268] Figure 12 is a block diagram illustrating a method according to a third aspect of the present disclosure. According to the third aspect, a coding method is provided which is implemented by an encoding device. The method includes: (1201) determining a value for a current block, wherein the syntactic element specifies the clipping index of the clipping value to be used before multiplying by coefficients of an adaptive loop filter (ALF); and (1202) generating a bitstream based on the value of the syntactic element, wherein at least one bit in the bitstream represents the syntactic element, and the syntactic element is coded using a fixed-length code.
[0269] Figure 13 is a block diagram illustrating a method according to a fourth aspect of the present disclosure. According to the fourth aspect, a coding method is provided which is implemented by an encoding device. The method is: a step (1301) of determining a value of a syntactic element for a current block, wherein the syntactic element is an adaptive loop filter (ALF) clipping value or an ALF filter coefficient parameter; and a step of generating a bitstream based on the value of the syntactic element, wherein at least one bit in the bitstream represents the syntactic element, and the at least one bit of the syntactic element is obtained by using only the value of the syntactic element for a current block.
[0270] Figure 14 is a block diagram illustrating a decoder according to a fifth aspect of the present disclosure. According to the fifth aspect of the present disclosure, a decoder 1400 is provided, comprising a processing circuit 1401 for performing the method according to the first or second aspect, or an implementation of either thereof.
[0271] Figure 15 is a block diagram illustrating an encoder according to a sixth aspect of the present disclosure. According to the sixth aspect of the present disclosure, an encoder 1500 is provided, comprising a processing circuit 1501 for performing a method according to the third or fourth aspect or an implementation of either of them.
[0272] Figure 16 is a block diagram illustrating a decoder according to the ninth aspect of the present disclosure. According to the ninth aspect of the present disclosure, a decoder 1600 is provided, comprising one or more processors 1601; and a non-temporary computer-readable storage medium 1602 coupled to the processors 1601 and storing a program for execution by the processors 1601, wherein the program, when executed by the processors 1601, configures the decoder 1600 to perform a method according to the first or second aspect or an implementation of either of them.
[0273] Figure 17 is a block diagram illustrating an encoder according to the tenth aspect of the present disclosure. According to the tenth aspect of the present disclosure, an encoder 1700 is provided, comprising one or more processors 1701; and a non-temporary computer-readable storage medium 1702 coupled to the processors 1701 and storing a program for execution by the processors 1701, wherein the program, when executed by the processors 1701, configures the encoder 1700 to perform a method according to the third or fourth aspect or an implementation of either of them.
[0274] Figure 18 is a block diagram illustrating a decoder according to the eleventh aspect of the present disclosure. According to the eleventh aspect of the present disclosure, a decoder 1800 is provided, the decoder comprising an entropy decode unit 1801 (which may be an entropy decode unit 304) configured to acquire a bitstream 1811, wherein at least one bit in the bitstream 1811 represents a syntactic element for the current block, the syntactic element specifying the clipping index of the clipping value to be used before multiplying by a coefficient of an adaptive loop filter (ALF); the entropy decode unit 1801 is further configured to parse the bitstream 1811 to acquire a value 1812 of the syntactic element for the current block, the syntactic element being coded using a fixed-length code; and a filtering unit 1803 (which may be a loop filter 320) configured to apply adaptive loop filtering to the current block based on the value 1812 of the syntactic element for the current block.
[0275] Figure 19 is a block diagram illustrating a decoder according to a twelfth aspect of the present disclosure. According to a twelfth aspect of the present disclosure, a decoder 1900 is provided, comprising an entropy decode unit 1901 (which may be an entropy decode unit 304) configured to acquire a bitstream 1911, wherein at least one bit in the bitstream 1911 represents a syntactic element for a current block, the syntactic element being an ALF clipping value index or an ALF coefficient parameter, and the entropy decode unit 1801 is further configured to parse the bitstream to acquire a value 1912 of the syntactic element for a current block, the value of the syntactic element for a current block is obtained by using only the at least one bit of the syntactic element; and a filtering unit (which may be a loop filter 320) configured to apply adaptive loop filtering to a current block based on the value 1912 of the syntactic element for a current block.
[0276] Figure 20 is a block diagram illustrating an encoder according to the thirteenth aspect of this disclosure. According to the thirteenth aspect of this disclosure, an encoder 2000 is provided, which is: The current configuration includes a decision unit 2001 (which may be a loop filter 220) configured to determine a value 2012 for a syntactic element about a block, wherein the syntactic element specifies the clipping index of the clipping value to be used before multiplying by the coefficients of an adaptive loop filter (ALF); and an entropy encode unit 2002 (which may be an entropy encode unit 270) configured to generate a bitstream 2011 based on the value 2012 for the syntactic element, wherein at least one bit in the bitstream 2011 represents the syntactic element, and the syntactic element is encoded using a fixed-length code.
[0277] Figure 21 is a block diagram illustrating an encoder according to a fourteenth aspect of the present disclosure. According to a fourteenth aspect of the present disclosure, an encoder 2100 is provided, the encoder having: a decision unit 2101 (which may be a loop filter 220) configured to determine a value 2112 of a syntactic element for a current block, wherein the syntactic element is an ALF clipping value index or an ALF coefficient parameter; and an entropy encode unit 2102 (which may be an entropy encode unit 270) configured to generate a bitstream 2111 based on the value 2112 of the syntactic element, wherein at least one bit in the bitstream 2111 represents the syntactic element, wherein the at least one bit of the syntactic element is obtained by using only the value of the syntactic element for a current block.
[0278] This disclosure provides the following further embodiments.
[0279] Embodiment 1. A coding method implemented by a decoding device, comprising: a step of acquiring a bitstream, wherein at least one bit in the bitstream corresponds to a syntactic element for a current block (or a set of blocks in which one block is the current block); a step of parsing the bitstream to obtain a value of the syntactic element for a current block, wherein the value of the syntactic element for a current block refers only to the at least one bit; and a step of filtering the current block based on the value of the syntactic element for a current block.
[0280] Embodiment 2. The method according to Embodiment 1, wherein the value of the syntactic element is coded according to a fixed-length code (where a fixed-length code means that all possible values of the syntactic element are signaled using the same number of bits).
[0281] Embodiment 3. The method according to Embodiment 1, wherein the value of the syntactic element is coded according to a truncated unary code (a truncated unary code means that the most frequently occurring value of the given syntactic element is signaled using the fewest number of bits, and the least frequently occurring value of the syntactic element is signaled using the most number of bits).
[0282] Embodiment 4. Any one of Embodiments 1 to 3, wherein the syntactic element is an adaptive loop, filter, clipping index, or parameter.
[0283] Embodiment 5. A method according to any one of Embodiments 1 to 3, wherein the syntactic element is an adaptive loop filter coefficient parameter.
[0284] Embodiment 6. Any one of Embodiments 1 to 5, wherein the value of the syntactic element is used to determine a filter coefficient, and the filter coefficient is used in the filtering process.
[0285] Embodiment 7. Any one of Embodiments 1 to 5, wherein the value of the syntactic element is used to determine a clipping range, and the clipping range is used in the filtering process (the clipping range is used to limit the amount of modification allowed for a given sample by its neighboring samples).
[0286] Embodiment 8. A decoder (30) comprising a processing circuit for performing a method according to any one of Embodiments 1 to 7.
[0287] Embodiment 9. A computer program product comprising program code for performing the method described in any of Embodiments 1 to 7.
[0288] Embodiment 10. One or more processors; A decoder comprising a non-temporary computer-readable storage medium coupled to the processor and storing a program for execution by the processor, wherein the program, when executed by the processor, configures the decoder to perform a method according to any one of Embodiments 1 to 7.
[0289] The following is a description of the encoding and decoding methods shown in the above embodiments, as well as the systems using them.
[0290] Figure 22 is a block diagram showing a content delivery system 3100 for realizing a content delivery service. This content delivery system 3100 includes a capture device 3102, a terminal device 3106, and optionally a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, Wi-Fi, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination thereof.
[0291] The capture device 3102 may generate data and encode the data by encoding methods as shown in the embodiments described above. Alternatively, the capture device 3102 may deliver the data to a streaming server (not shown in the figure), which encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or tablet, a computer or laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include the source device 12 as described above. If the data includes video, the video encoder 20 included in the capture device 3102 may actually perform the video encoding process. If the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 may actually perform the audio encoding process. In some practical scenarios, the capture device 3102 delivers the encoded video and audio data by multiplexing them together. For example, in other practical scenarios in a video conferencing system, the encoded audio data and encoded video data are not multiplexed. The capture device 3102 delivers the encoded audio data and encoded video data separately to the terminal device 3106.
[0292] In the content supply system 3100, the terminal device 310 receives and plays back the encoded data. The terminal device 3106 may be a device with data receiving and restoration capabilities that can decode the encoded data described above, such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, a vehicle-mounted device 3124, or any combination thereof. For example, the terminal device 3106 may include the destination device 14 as described above. If the encoded data includes video, the video decoder 30 included in the terminal device is preferred for performing video decoding. If the encoded data includes audio, the audio decoder included in the terminal device is preferred for performing audio decoding.
[0293] For terminal devices with a display, such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or a vehicle-mounted device 3124, the terminal device can provide the decoded data to its display. For terminal devices without a display, such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is made contact thereto to receive and display the decoded data.
[0294] When each device in this system performs encoding or decoding, a picture encoding device or a picture decoding device as shown in the embodiments described above can be used.
[0295] Figure 23 shows the structure of an example terminal device 3106. After the terminal device 3106 receives the stream from the capture device 3102, the protocol progression unit 3202 analyzes the transmission protocol of the stream. This protocol includes, but is not limited to, Real-Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-Time Transport Protocol (RTP), Real-Time Messaging Protocol (RTMP), or any combination thereof.
[0296] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the multiplexing / decomposition unit 3204. The multiplexing / decomposition unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As mentioned above, in some practical scenarios, such as in a video conferencing system, the encoded audio data and encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and audio decoder 3208 without passing through the multiplexing / decomposition unit 3204.
[0297] Through multiplexing, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. A video decoder 3206, including a video decoder 30 as described in the embodiments described above, decodes the video ES by the decoding method shown in the embodiments described above to generate video frames and provides this data to the synchronization unit 3212. An audio decoder 3208 decodes the audio ES to generate audio frames and provides this data to the synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in Figure 23) before being provided to the synchronization unit 3212. Similarly, the audio frames may be stored in a buffer (not shown in Figure 23) before being provided to the synchronization unit 3212.
[0298] The synchronization unit 3212 synchronizes video frames and audio frames and supplies video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. The information may be coded in syntax using timestamps for the presentation of coded audio and visual data and for the delivery of the data stream itself.
[0299] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video and audio frames, and supplies the video / audio / subtitles to the video / audio / subtitle display 3216.
[0300] The present invention is not limited to the systems described above, and any of the picture encoding or picture decoding devices in the embodiments described above can be incorporated into other systems, such as automotive systems.
[0301] While embodiments of this disclosure have been described primarily in relation to video coding, it should be noted that embodiments of the coding system 10, encoder 20, and decoder 30 (and correspondingly system 10), as well as other embodiments described herein, may be configured for the processing or coding of still pictures, i.e., for the processing or coding of individual pictures independent of any preceding or consecutive pictures, as in video coding. Generally, when picture processing coding is limited to a single picture 17, only the interpretation units 244 (encoder) and 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and video decoder 30 may be equally applicable to still picture processing. Other functions include, for example, residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, partitioning 262 / 362, intra prediction 254 / 354, and / or loop filtering 220, 320, entropy coding 270, and entropy decoding 304.
[0302] For example, embodiments of encoder 20 and decoder 30, and functions described herein with reference to encoder 20 and decoder 30, may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted through a communication medium and executed by a hardware-based processing unit. The computer-readable medium may include computer-readable storage media corresponding to tangible media such as data storage media, or communication media including any medium that facilitates the transfer of computer programs from one location to another, for example, according to a communication protocol. Thus, the computer-readable medium may generally correspond to (1) non-transient tangible computer-readable storage media, or (2) communication media such as signals or carrier waves. The data storage medium may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, codes and / or data structures for implementations of the technologies described herein. A computer program product may include computer-readable media.
[0303] For example, and not limited to, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio waves, and microwaves, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio waves, and microwaves are included in the definition of a medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other temporary media, but instead refer to non-temporary, tangible storage media. As used herein, the terms "disk" and "disc" include compact discs (CDs), laser discs, optical discs, digital multipurpose discs (DVDs), floppy disks, and Blu-ray discs, where a "disk" typically reproduces data magnetically, while a "disc" reproduces data optically using a laser. Any combination of the above should also be included within the scope of computer-readable media.
[0304] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term “processor” as used herein may refer to any of the aforementioned structures or any other structure suitable for implementing the technologies described herein. Furthermore, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. These technologies may also be fully implemented by one or more circuits or logic elements.
[0305] The technology of this disclosure can be implemented in a wide variety of devices or apparatus, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). While various components, modules, or units are described in this disclosure to highlight the functional aspects of apparatus configured to perform the disclosed technology, these are not necessarily required to be implemented by different hardware units. Rather, as described above, the various units may be combined within a codec hardware unit, or they may be provided by a collection of interoperable hardware units that include one or more processors as described above in association with suitable software and / or firmware.
Claims
1. A coding method implemented by a decoding device, the method being: The first step involves parsing the bitstream to obtain the ALF flag (adaptive loop filter flag), the ALF flag indicating whether adaptive loop filtering is applied to the luma component; If the value of the ALF flag is true, the step is to parse the bitstream to obtain syntactic elements, wherein the syntactic elements specify the clipping index for the clipping value for ALF; The step includes applying adaptive loop filtering based on the value of the syntactic element, method.
2. The method according to claim 1, wherein the syntactic elements are coded using fixed-length codes.
3. The method according to claim 2, wherein the fixed-length code includes a binary representation of an unsigned integer using at least one bit.
4. The method according to claim 1 or 2, wherein the syntactic element is applied to a set of blocks.
5. The method according to claim 1 or 2, wherein the syntactic element is at the slice level.
6. A coding method implemented by an encoding device, the method being: The step of determining the value of the ALF flag (adaptive loop filter flag), wherein the ALF flag indicates whether ALF is applied to the luma component; If the value of the ALF flag is true, the step of determining the value of a syntactic element, wherein the syntactic element is an ALF filter coefficient parameter; A step of generating a bitstream containing the value of the syntactic element and the ALF flag, wherein at least one bit in the bitstream represents the syntactic element, method.
7. The method according to claim 6, wherein the syntactic elements are coded using fixed-length codes.
8. The method according to claim 6 or 7, wherein the syntactic element is applied to a set of blocks.
9. The method according to claim 6 or 7, wherein the syntactic element is at the slice level.
10. A decoder (30) comprising a processing circuit for performing the method according to any one of claims 1 to 5.
11. An encoder (20) comprising a processing circuit for performing the method according to any one of claims 6 to 9.
12. A non-temporary computer-readable medium carrying an instruction code that, when executed by a computer device, causes the computer device to perform the method according to any one of claims 1 to 9.
13. One or more processors; A decoder comprising a non-temporary computer-readable storage medium coupled to the processor and storing a program for execution by the processor, wherein the program, when executed by the processor, configures the decoder to perform the method according to any one of claims 1 to 5. decoder.
14. One or more processors; An encoder comprising a non-temporary computer-readable storage medium coupled to the processor and storing a program for execution by the processor, wherein the program, when executed by the processor, configures the encoder to perform the method according to any one of claims 6 to 9. Encoder.
15. A method of storing a bitstream: The stage of receiving a bitstream through a communication interface; The steps include storing the bitstream in one or more storage media, wherein the bitstream includes an ALF flag (adaptive loop filter flag), the ALF flag indicating whether adaptive loop filtering is applied to the luma component; if the value of the ALF flag is true, the bitstream further includes a syntactic element, the syntactic element specifying the clipping index of the clipping value for ALF, and the value of the syntactic element being used for adaptive loop filtering. method.
16. A method for sending a bitstream: A step of storing at least one bitstream in at least one storage medium; The step of transmitting at least one bitstream The bitstream includes an ALF flag (adaptive loop filter flag), which indicates whether adaptive loop filtering is applied to the luma component; if the value of the ALF flag is true, the bitstream further includes a syntactic element, which specifies the clipping index of the clipping value for ALF, and the value of the syntactic element is used for adaptive loop filtering. method.
17. A device for storing a bitstream, the device having at least one storage medium and at least one communication interface, The at least one communication interface is configured to receive the bitstream; The at least one storage medium is configured to store the bitstream, the bitstream including an ALF flag (adaptive loop filter flag), the ALF flag indicating whether adaptive loop filtering is applied to the luma component; if the value of the ALF flag is true, the bitstream further includes a syntactic element, the syntactic element specifying the clipping index of the clipping value for ALF, the value of the syntactic element being used for adaptive loop filtering. Device.
18. A device for transmitting a bitstream: A storage medium configured to store at least one bitstream; A transmitter for transmitting the aforementioned at least one bitstream and The bitstream includes an ALF flag (adaptive loop filter flag), which indicates whether adaptive loop filtering is applied to the luma component; if the value of the ALF flag is true, the bitstream further includes a syntactic element, which specifies the clipping index of the clipping value for ALF, and the value of the syntactic element is used for adaptive loop filtering. Device.