Access unit delimiter and use of adaptive parameter set

By deriving in-loop filter parameters from access units, the decoder reduces latency and enhances video quality in low-latency environments by applying filters adaptively within the access unit.

JP2025106558AActive Publication Date: 2025-07-15FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025067429
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-08-19
Filing Date
2025-04-16
Publication Date
2025-07-15
Estimated Expiration
2040-08-18

AI Technical Summary

Technical Problem

In low-latency environments, modern video encoding standards face challenges in efficiently transmitting in-loop filter parameters due to the need for complete picture encoding before parameter signaling, leading to decoding delays and reduced subjective quality.

Method used

A video decoder derives necessary parameters from an access unit, using motion compensation and transform-based residual decoding to reconstruct decoded pictures, and applies in-loop filters with adaptive parameter sets within the access unit, allowing early decoding and reduced latency.

Benefits of technology

This approach enables early decoding of video data by deriving in-loop filter parameters within the access unit, reducing latency and improving subjective quality by applying filters efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106558000001_ABST
    Figure 2025106558000001_ABST
Patent Text Reader

Abstract

To provide a decoder and a method for deriving a necessary parameter from an access unit.SOLUTION: A video decoder 20 comprises an in-loop filter 90 which filters a reconfiguration version of a decoded picture, and a parameterizing unit. The parameterizing unit reads in-loop filter control information for parameterizing the in-loop filter from a parameter set located in an access unit of a decoded picture following a video encoding unit along the sequence of a data stream 14 and / or a part of a video encoding unit following data included in a video encoding unit carrying block-based prediction parameter data and predicted residue data along the sequence of the data stream, and also parameterizes the in-loop filter so as to filter the reconfiguration version of the encoded picture by a method corresponding to the in-loop filter control information.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the use of access delimiters and adaptive parameter sets for signaling encoding parameters.

Background Art

[0002] Modern video encoding standards utilize in-loop filters such as an Adaptive Loop Filter (ALF), Sample Adaptive Offset (SAO), and a Deblocking Filter.

[0003] In-loop filters are placed within the decoder loop of an encoder. During all video encoding stages, and particularly in the lossy compression performed at the quantization stage, the subjective quality of a video sequence may be reduced as a result of the appearance of blocking, ringing, or blurred artifacts. To remove these artifacts and improve the subjective and objective quality of the reconstructed sequence, a set of in-loop filters is used. The in-loop filter within the encoder estimates the optimal filter parameters that will most improve the objective quality of the frame. These parameters are then sent to the decoder, and the in-loop filter of the decoder uses these parameters to optimally filter the reconstructed frame and achieve the same quality improvement as achieved for the reconstructed frame within the encoder.

[0004] The Deblocking Filter aims to remove blocking artifacts that appear at the edges of CUs (Coding Units), specifically PUs (Prediction Units) and TUs (Transform Units), as a result of using the block structure in the processing of all stages of the encoder.

[0005] The SAO filter aims to reduce unwanted visible artifacts such as ringing. The key idea of SAO is first to classify the reconstructed samples into different categories, obtain an offset for each category, and then reduce sample distortion by adding the offset to each sample in the category.

[0006] The key idea of ALF is to minimize the mean squared error between the original pixel and the decoded pixel using Wiener-based adaptive filter coefficients. ALF is placed at the last processing stage of each picture and can be regarded as a tool to capture and correct artifacts from the previous stage. Appropriate filter coefficients are determined by the encoder and explicitly signaled to the decoder. That is, ALF requires a set of parameters to be sent to the decoder, namely, appropriate filter coefficients. These parameters are sent in a high-level syntax structure, such as an adaptive parameter set (APS). The APS is a parameter set that is sent in the bitstream before the video coding layer (VCL) NAL (network abstraction layer) unit, i.e., the slice of the picture. ALF is applied to the complete picture after reconstruction. Also, in the encoder, ALF estimation is one of the last steps of the encoding process.

[0007] In low-latency environments, this causes problems because the encoder wants to start transmitting the processed part of the picture as early as possible, especially before finishing the encoding process of the picture. ALF cannot be optimally used in these environments because the APS with the filter parameters estimated for the encoded picture must be sent before the first slice of the picture.

[0008] Furthermore, a set of NAL units of a specified format is called an access unit, AU, and the decoding of each AU results in one decoded picture. Each AU includes a set of VCL NAL units that together constitute a primary coded picture. Also, an access unit delimiter (AUD) can be added at the beginning to help identify the start position of the AU.

[0009] The AUD is used to separate AUs within the bitstream and can optionally contain information about subsequent pictures, such as the permitted slice types (I, P, B).

[0010] In VVC (Versatile Video Coding), several different parameter sets can be referenced within a picture, namely, a video parameter set (VPS), a decoder parameter set (DPS), a sequence parameter set (SPS), multiple picture parameter sets (PPS), different types of adaptive parameter sets (APS), and one or more. To enable a decoder to decode a picture, all parameter sets must be available.

[0011] Different slices of a picture can reference different PPSs and APSs. Therefore, it can be difficult for a decoder to determine whether all the necessary parameter sets are available because it needs to analyze all the slice headers of the picture and which parameters are being referenced.

Summary of the Invention

Problems to be Solved by the Invention

[0012] The object of the subject matter of this application is to provide a decoder that derives the necessary parameters from an access unit.

Means for Solving the Problems

[0013] This object is achieved by the subject matter of the claims of this application.

[0014] According to an embodiment of the present application, in order to obtain a reconstructed version (46a) of a decoded picture, a video decoder uses motion compensation prediction and transform-based residual decoding from one or more video coding units (100) within an access unit AU of a video data stream, such as a VCL NAL unit, to reconstruct a decoded picture, such as the currently decoded picture or a subsequent decoded picture. The video decoder includes a decoding core (94) configured to perform such reconstruction, a loop filter (90), such as an ALF, configured to filter the reconstructed version (46b) of the decoded picture to obtain a version (46b) of the decoded picture to be inserted into the decoded picture buffer DPB (92) of the video decoder, and a parameterizer configured to parameterize the loop filter. The parameterizer reads, for example, ALF coefficients (or parameters) and ALF flags for each CTU (coded tree unit), and parameterizes the loop filter to filter the reconstructed version of the decoded picture in a manner according to the loop filter control information. That is, the loop filter control information is derived for one or more video coding units, so it is possible to start decoding before receiving all the video coding units of a picture. Therefore, the decoding delay is reduced in a low-latency environment.

[0015] According to an embodiment of the present application, the in-loop filter control information includes one or more filter coefficients for parameterizing the in-loop filter with respect to the transfer function. That is, the ALF is, for example, a FIR (Finite Impulse Response) or IIR (Infinite Impulse Response) filter and a FIR or IIR coefficient of the filter coefficient that controls the transfer function of the filter.

[0016] According to an embodiment of the present application, the in-loop filter control information includes spatially selective in-loop filter control information for spatially varying the filtering of a decoded picture, for example, a reconstructed version of the currently decoded picture or a subsequent decoded picture, by the in-loop filter.

[0017] According to an embodiment of the present application, each video encoding unit (100) is continuously arithmetic encoded along the data stream order to the end of the portion (106), that is, up to the ALF. A predetermined parameter set (102) of one or more parameter sets follows each of the one or more video encoding units (100) in the data stream order and includes one or more filter coefficients for parameterizing the in-loop filter with respect to the transfer function.

[0018] According to an embodiment of the present application, one or more parameter sets (104) include, for each of the one or more video encoding units (100), a further predetermined parameter set that follows each video encoding unit (100) in the data stream order, and spatially selective in-loop filter control information for spatially varying the filtering of a decoded picture, for example, a reconstructed version of the currently decoded picture or a subsequent decoded picture, by the in-loop filter within the portion of the picture encoded by each video encoding unit (100).

[0019] According to an embodiment of the present application, each of one or more video encoding units (100) includes a filter information section (106) following in data stream order in the data section (108) of each video encoding unit (100), and the filter information section includes spatial selective in-loop filter control information for spatially varying the filtering of a decoded picture, for example, a reconstructed version of the currently decoded picture or a subsequent decoded picture, by an in-loop filter within a portion of a picture in which block-based prediction parameter data and prediction residual data are encoded in the data section of each video encoding unit (100).

[0020] According to an embodiment of the present application, the parameterizer is configured to arrange, within the access unit (AU), the one or more parameter sets (102, 104), for example, an ALF for each of an ALF APS and a CTU APS, of the decoded picture, for example, the currently decoded picture or a subsequent decoded picture, at a position following the one or more video encoding units (100) in data stream order, for example, individually with a VCL NALU or at a position following all of them, in the case of a predetermined indication in the video data stream assuming a first state, and at a different position within the access unit preceding all of the one or more video encoding units (100) in the case of the predetermined indication in the video data stream assuming a second state.

[0021] According to an embodiment of the present application, a portion (106) of one or more video encoding units (100), for example, the ALF for each CTU data, is located at a position following data (108) constituted by one or more video encoding units (100) along the data stream order, which carries block-based prediction parameter data and prediction residual data in the case of a predetermined instruction in a video data stream assuming a first state, and is located at different positions within one or more video encoding units where the block-based prediction parameter data and the prediction residual data are scattered in the case of a predetermined instruction in a video data stream assuming a second state.

[0022] According to an embodiment of the present application, a video decoder is configured to read a predetermined instruction from one or more video encoding units (100). The predetermined instruction indicates one or more parameter sets by one or more identifiers when assuming a first state, and indicates different one or more in-loop filter control information parameter settings when assuming a second state. The video decoder is configured to respond to the predetermined instruction for each access unit so as to perform different position identifications for different access units of the video data stream when the predetermined instruction is different for different access units. The parameterizer is configured to reconstruct the decoded picture using the in-loop filter control information included in the previously signaled access unit AU.

[0023] According to an embodiment of the present application, when detecting the boundary of an access unit AU, the video decoder interprets a video encoding unit that conveys in-loop filter control information, for example, ALF filter data, as not starting an access unit in the form of one or more parameter sets (102, 104), for example, a suffix APS, and therefrom, for example, ignores them in AU boundary detection, thereby detecting that there is no AU boundary, and interprets a video encoding unit that conveys in-loop filter control information not in the form of one or more parameter sets (102, 104), for example, a suffix APS, as starting an access unit from such a video encoding unit, and is configured to detect an AU boundary from such a video encoding unit, for example.

[0024] According to an embodiment of the present application, the video decoder decodes video from a video data stream by decoding a decoded picture, for example, the currently decoded picture or a subsequent decoded picture, in a parameterized manner using one or more predetermined encoding parameters from one or more video encoding units (100) within an access unit AU of the video data stream, derives a predetermined encoding parameter (122) from a plurality of parameter sets (120) scattered in the video data stream, and is configured to read out an identifier (200) for identifying a predetermined parameter set from a plurality of parameter sets including the predetermined encoding parameter from a predetermined unit (124) of the access unit AU. That is, the presence or absence of the encoding parameter is indicated by the identifier, and thus it is efficiently recognized which parameter set can be derived from the received video encoding unit. Further, since the identifier is included in a predetermined unit of the AU, it is easy to include different parameter sets for different video encoding units.

[0025] According to an embodiment of the present application, a predetermined unit of the AU includes a flag (204) indicating whether an identifier (200) exists within the predetermined unit. That is, it is possible to indicate with a flag which identifier is included in a predetermined unit of the AU, and for example, it is possible to indicate a delimiter of an access unit.

[0026] According to an embodiment of the present application, a plurality of parameter sets (120) are at different hierarchical levels, and one or more video encoding units include an identifier referring to a first predetermined parameter set (126) within one or more first predetermined hierarchical levels, for example, in a slice header. The first predetermined parameter set (126) within one or more first predetermined hierarchical levels includes an identifier referring to a second predetermined parameter set (128) within one or more second predetermined hierarchical levels. The first and second predetermined parameter sets are included by a predetermined parameter set (122). An identifier read from a predetermined unit (124) of the access unit AU identifies all predetermined parameter sets directly or indirectly referred to by one or more video encoding units of the access unit. Thereby, when all predetermined parameter sets identified by the identifier are available, the access unit is decodable.

[0027] According to an embodiment of the present application, a predetermined unit of the AU includes a flag (204) indicating (205) whether a predetermined identifier of an identifier (200) referring to a specific predetermined parameter set (126b), for example, a specific APS, exists within the predetermined unit (124), or whether a predetermined identifier referring to a specific predetermined parameter set (126b) exists within one or more video encoding units (100).

[0028] According to an embodiment of the present application, the first predetermined parameter set (126) includes a third predetermined parameter set (126a) referred to by an identifier in one or more video encoding units (100), and a fourth predetermined parameter set (126b) referred to by an identifier (200) existing in a predetermined unit (124), but not referred to by any identifier in one or more video encoding units (100) or by any predetermined parameter set.

[0029] According to an embodiment of the present application, a predetermined unit of an AU includes one or more identifiers of one or more adaptation parameter sets, APS, one or more identifiers of one or more picture parameter sets, PPS, an identifier of a video parameter set, VPS, an identifier of a decoder parameter set, DPS, and one or more identifiers of one or more sequence parameter sets, SPS. The plurality of parameter sets include a video parameter set, VPS, a decoder parameter set, DPS, a sequence parameter set, SPS, one or more picture parameter sets, PPS, and one or more adaptation parameter sets, APS.

[0030] According to an embodiment of the present application, a video decoder decodes video from a video data stream by decoding a picture from one or more video coding units (100) of an access unit (AU) of the video data stream, and reads one or more parameters from an access unit delimiter AUD arranged in the video data stream so as to form the start of the access unit AU. The one or more parameters control (300) whether separate access units are defined in the video data stream for pictures related to different layers at one instant of the video data stream, or whether pictures related to different layers at one instant of the video data stream are encoded in one of the access units, and / or indicate (302) that when an indication of a video coding type includes different video coding units within one access unit, the coding type of the video coding unit is included within the access unit assigned to the video coding unit within the access unit, and / or indicate (304) that the access unit is not referenced by other pictures, and / or indicate (306) a picture that is not output. That is, since the parameters required for decoding a picture are indicated by the AUD, the decoding of the slices of the pictures included in the AU can be started before acquiring all the parameter sets for decoding a complete picture. That is, since the parameter set required for each slice can be efficiently indicated by the AUD, the decoding speed can be improved.

[0031] According to an embodiment of the present application, one or more parameters form a deviation with respect to the parameters defined by the previous AUD. The AUD includes an indication of whether the parameters defined by the previous AUD should be adopted. The AUD includes an indication of whether one or more parameters are applied to all layers of the video data stream or only to a single layer thereof. The video coding type of the video coding unit is indicated by describing the random access characteristics of a plurality of pictures.

[0032] According to an embodiment of the present application, the video decoder decodes video from a video data stream by decoding a picture from one or more video coding units (100) of an access unit (AU) of the video data stream, and is configured to read one or more parameters (308) from an access unit delimiter (AUD) arranged in the data stream so as to form the start of the access unit (AU), and the one or more parameters indicate characteristics of the access unit and indicate whether the characteristics apply to all layers of the video data stream or only to a single layer thereof.

[0033] According to an embodiment of the present application, in order to obtain a reconstructed version of a decoded picture, one or more video coding units (100) within an access unit AU of a video data stream, for example, motion compensation prediction and transform-based residual decoding from a VCL NAL unit, are used to reconstruct the decoded picture, and an in-loop filter is used to filter the reconstructed version of the decoded picture to obtain a version of the decoded picture to be inserted into the decoded picture buffer DPB of the video decoder, and the in-loop filter is used to filter the reconstructed version of the decoded picture to obtain a version of the decoded picture to be inserted into the decoded picture buffer DPB of the video decoder, and the in-loop filter is parameterized, including parameterizing the in-loop filter by reading in-loop filter control information for parameterizing the in-loop filter from one or more parameter sets (102, 104) located within the access unit AU of the decoded picture subsequent to one or more video coding units (100) along the data stream order, for example, ALF for each of the ALF APS and CTU APS, and / or from a portion (106) of one or more video coding units (100) containing block-based prediction parameter data and prediction residual data along the data stream order, for example, ALF for each CTU data, and filtering the reconstructed version of the decoded picture in a manner according to the in-loop filter control information. A method is provided by doing so.

[0034] According to an embodiment of the present application, from one or more video encoding units (100) within an access unit AU of a video data stream, in a parameterized manner using one or more predetermined encoding parameters, by decoding a decoded picture, for example, the currently decoded picture or a subsequent decoded picture, decoding a video from the video data stream, deriving the predetermined encoding parameters from a plurality of parameter sets scattered in the video data stream, and reading an identifier (200) from a predetermined unit of the access unit AU that identifies a predetermined parameter set from the plurality of parameter sets including the predetermined encoding parameters, a method is provided.

[0035] According to an embodiment of the present application, decoding a video from a video data stream by decoding a picture from one or more video encoding units (100) of an access unit AU of the video data stream, and reading one or more parameters from an access unit delimiter AUD arranged in the video data stream so as to form the start of the access unit AU, wherein the one or more parameters control (300) whether a separate access unit is defined in the video data stream for pictures related to different layers related to one instant of the video data stream, or whether pictures related to different layers related to one instant of the video data stream are encoded in one of the access units, and when an indication of the video encoding type includes mutually different video encoding units within one access unit, the encoding type of the video encoding unit indicates (302) that it is included within the access unit assigned within the video encoding unit within one access unit, and / or indicates (304) a picture that is not referenced by other pictures, and / or indicates (306) a picture that is not output, a method is provided.

[0036] According to an embodiment of the present application, video is decoded from a video data stream by decoding a picture from one or more video encoding units (100) of an access unit (AU) of the video data stream, and one or more parameters (308) are read from an access unit delimiter (AUD) arranged in the data stream so as to form the start of the access unit (AU), and the one or more parameters indicate the characteristics of the access unit (AU), and the characteristics indicate whether they are applied to all layers of the video data stream or only to a single layer thereof. A method is provided.

Brief Description of Drawings

[0037] Preferred embodiments of the present application will be described below with reference to the drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6a

Figure 6b

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11a

Figure 11b

Figure 11c

Figure 12a

Figure 12b

Figure 13

BEST MODE FOR CARRYING OUT THE INVENTION

[0038] In the following description, the same or equivalent elements, or elements having the same or equivalent functions, are denoted by the same or equivalent reference numerals.

[0039] In the following description, numerous specific details are set forth in order to provide a more thorough explanation of embodiments of the present application. However, it will be apparent to those skilled in the art that embodiments of the present application may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail in order to avoid obscuring embodiments of the present application. Also, unless otherwise specified, the features of the different embodiments described below may be combined.

[0040] Introduction Note that the individual aspects described herein may be used separately or in combination. Thus, details may be added to each of the individual aspects without adding details to another one of the aspects.

[0041] Also, note that the present disclosure describes, either explicitly or implicitly, features that can be used in a video decoder (an apparatus for providing a decoded representation of a video signal based on an encoded representation). Thus, any of the features described herein can be used in the context of a video decoder.

[0042] Furthermore, the features and functionality disclosed herein in connection with a method can also be used in an apparatus (configured to perform such functionality). Additionally, any features and functions disclosed herein with respect to an apparatus can also be used in a corresponding method. In other words, the methods disclosed herein can be supplemented by any of the features and functionality described with respect to an apparatus.

[0043] The following description of the figures begins with an explanation of a block-based predictive codec for encoding a video picture to form an example of an encoding framework that can incorporate embodiments for a hierarchical video data stream codec. The video encoder and video decoder of the block-based predictive codec are described in connection with FIGS. 1-3. Hereinafter, an explanation of an embodiment of the concept of the hierarchical video data stream codec of the present application will be presented together with an explanation of how such a concept can be incorporated into the video encoder and decoder of FIGS. 1 and 2, respectively. However, the embodiments described below can also be used to form video encoders and video decoders that do not operate according to the encoding framework underlying the video encoder and video decoder of FIGS. 1 and 2.

[0044] FIG. 1 shows a block diagram of an apparatus for predictive encoding a video as an example of a video decoder capable of performing motion compensation prediction of an inter prediction block according to an embodiment of the present application. That is, FIG. 1 shows an apparatus for predictive encoding a video 11 consisting of a series of pictures 12 into a data stream 14. For this purpose, block-based predictive encoding is used. Further, transform-based residual encoding is exemplarily used. The apparatus, i.e., the encoder, is denoted by reference numeral 10.

[0045] FIG. 2 shows a block diagram of an apparatus for predictive decoding of video as an example of a video decoder capable of performing motion compensation prediction of an inter prediction block according to an embodiment of the present application. That is, FIG. 2 shows a corresponding decoder 20, i.e., an apparatus 20 configured to predictively decode a video 11' composed of a picture 12' within a picture block from a data stream 14, also using transform-based residual decoding illustratively here. The apostrophe is used to indicate that the picture 12' and video 11' reconstructed by the decoder 20 deviate from the original encoded picture 12 by the apparatus 10 with respect to the coding loss introduced by quantization of the prediction residual signal, respectively. FIGS. 1 and 2 illustratively use transform-based prediction residual coding, but the embodiments of the present application are not limited to this type of prediction residual coding. This also applies to other details described with respect to FIGS. 1 and 2, as outlined below.

[0046] The encoder 10 is configured to subject the prediction residual signal to a spatial-to-spectral transformation and encode the thus obtained prediction residual signal into the data stream 14. Similarly, the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and subject the thus obtained prediction residual signal to a spectral-to-spatial transformation.

[0047] Internally, the encoder 10 can comprise a prediction residual signal former 22 that generates a prediction residual 24 so as to measure the deviation of the original signal, i.e., the prediction signal 26 from the video 11 or the current picture 12. The prediction residual signal former 22 can be, for example, a subtractor that subtracts the prediction signal from the original signal, i.e., the current picture 12. Next, the encoder 10 further comprises a transformer 28 that subjects the prediction residual signal 24 to a spatial spectral transform and then obtains a spectral-domain prediction residual signal 24' that is quantized by a quantizer 32 constituted by the encoder 10. The prediction residual signal 24'' thus quantized is encoded in the data stream 14. For this purpose, the encoder 10 optionally includes an entropy encoder 34, and the entropy encoder performs entropy encoding so that the prediction residual signal is transformed and quantized into the data stream 14. Based on the prediction residual signal 24'' in the data stream 14 and decoded from the data stream 14, a prediction residual 26 is generated by the prediction stage 36 of the encoder 10. For this purpose, the prediction stage 36 may internally include an inverse quantizer 38 that inverse quantizes the prediction residual signal 24'' so as to obtain a spectral-domain prediction residual signal 24''' corresponding to the signal 24' excluding the quantization loss, and then includes an inverse transformer 40 that inverse-transforms, i.e., subjects the latter prediction residual signal 24'' to a spectral space transform, to obtain a prediction residual signal 24'''' corresponding to the original prediction residual signal 24 excluding the quantization loss. The synthesizer 42 of the prediction stage 36 then recombines, such as by adding, the prediction signal 26 and the prediction residual signal 24'''' so as to obtain a reconstructed signal 46a, i.e., a reconstruction (reconstructed version) of the original signal 12. The reconstructed signal 46a may correspond to the signal 12'.

[0048] The in-loop filter 90 filters the reconstructed signal 46a to obtain a version of the decoded picture, e.g., the decoded signal 46b of the currently decoded picture or a subsequent decoded picture, and inserts it into the decoded picture buffer DPB 92.

[0049] Next, the prediction module 44 in the prediction stage 36 generates a prediction signal 26 based on the signal 46b, for example, by using spatial prediction, i.e., intra prediction, and / or temporal prediction, i.e., inter prediction. Details regarding this will be described below.

[0050] The decoder 20 includes a decoding core 94 that includes an entropy decoder 50, an inverse quantizer 52, an inverse transformer 54, a combiner 56, a prediction module 58, an in-loop filter 90, and a DPB 94.

[0051] Similarly, the decoder 20 may correspond to the prediction stage 36 and may be internally configured from components interconnected in a corresponding manner. In particular, the entropy decoder 50 of the decoder 20 can entropy-decode the quantized spectral region prediction residual signal 24'' from the data stream, where the inverse quantizer 52, the inverse transformer 54, the combiner 56, and the prediction module 58 are interconnected and cooperate in the manner described above with respect to the modules of the prediction stage 36 to recover the signal reconstructed based on the prediction residual signal 24'', whereby, as shown in FIG. 3, the output of the combiner 56 yields the reconstructed signal, i.e., the video 11' or its current picture 12'.

[0052] Although not specifically described above, it is readily apparent that the encoder 10 can set some encoding parameters, such as prediction mode, motion parameters, etc., according to some optimization scheme, such as a method of optimizing certain rate and distortion related criteria, i.e., encoding cost, and / or a method of using some rate control. As will be described in more detail below, the encoder 10, decoder 20, and corresponding modules 44, 58 each support different prediction modes, such as an intra encoding mode and an inter encoding mode, which form a set or pool of a kind of primitive prediction mode in which the prediction of picture blocks is configured in a manner described in more detail below. The granularity at which the encoder and decoder switch between these prediction syntheses corresponds to the subdivision of each block of pictures 12 and 12'. Some of these blocks may simply be intra-encoded blocks, some may simply be inter-encoded blocks, and optionally, additional blocks may be obtained using both intra-encoding and inter-encoding, but note that the details will be described below. According to the intra encoding mode, the prediction signal for a block is obtained based on the spatial, already encoded / decoded neighborhood of each block. There may be several intra encoding sub-modes, among which certain intra prediction parameters are represented. There may also be a direction or angle intra encoding sub-mode in which the prediction signal for each block is filled by extrapolating the sample values of the neighborhood along a certain direction specified for each directional intra encoding sub-mode to each block.The intra-coding sub-mode may include one or more additional sub-modes such as, for example, a DC coding mode in which the prediction signal of each block assigns a DC value to all samples within each block, and / or a planar intra-coding mode in which the prediction signal of each block is approximated or determined to be a spatial distribution of sample values described by a two-dimensional linear function over the sample positions of each block, and the slope and offset of a plane defined by the two-dimensional linear function are derived based on adjacent samples. In contrast, according to the inter-prediction mode, for example, a prediction signal for a block can be obtained by temporally predicting the inside of the block. For parameterization of the inter-prediction mode, motion vectors may be signaled in the data stream, and the motion vectors indicate the spatial displacement of a portion of a previously encoded picture of the video 11 from which a previously encoded / decoded picture is sampled in order to obtain a prediction signal for each block. This means that, in addition to the residual signal coding constituted by the data stream 14, such as the entropy-coded transform coefficient levels representing the quantized spectral region prediction residual signal 24'', the data stream 14 may be encoded with prediction-related parameters for assigning block prediction modes, prediction parameters for the assigned prediction modes such as motion parameters for the inter-prediction mode, and optionally, additional parameters for controlling the construction of the final prediction signal for the block using the assigned prediction mode and prediction parameters, as will be described in more detail later. Additionally, the data stream may include parameters for controlling and signaling the subdivision of pictures 12 and 12' into blocks respectively. The decoder 20 uses these parameters to subdivide the picture in the same manner as the encoder did, assign the same prediction mode and parameters to the blocks, and perform the same prediction to obtain the same prediction signal.

[0053] FIG. 3 is a schematic diagram showing an example of the relationship among a prediction residual signal, a prediction signal, and a reconstructed signal, and shows possibilities such as the setting of sub-divisions defining the prediction signal and the handling of the prediction residual signal. That is, FIG. 3 shows the relationship between the reconstructed signal, i.e., the reconstructed picture 12', on the one hand, and the prediction residual signal 24'''' signaled in the data stream on the other hand, and the relationship with the prediction signal 26. As already described above, the combination may be additional. The prediction signal 26 is merely an example, and is obtained by dividing the picture area into blocks 80 of various sizes. The subdivision may be any subdivision that regularly divides the picture area into rows and columns of blocks, or may be a multi-tree subdivision that divides the picture 12 into leaf blocks of various sizes, such as quadtree subdivision, etc. The picture area may first be divided into rows and columns of tree root blocks, and then further divided according to the recursive multi-tree subdivision to form the blocks 80, or a mixture thereof.

[0054] Hereinafter, each aspect of the present invention will be described.

[0055] Suffix-APS According to one aspect of the present invention of the present application, the encoder can start transmitting a part (e.g., a slice) of a picture before finishing the encoding process of the entire picture while still using slices. This is achieved by enabling the adaptation parameter set (APS) to be transmitted after the encoded slice of the picture that moves for each CTU (Coding Tree Unit) ALF parameter behind the actual slice data.

[0056] Figure 4 shows a diagram of a state-of-the-art encoder. First, the entire picture is encoded (intra prediction, motion estimation, residual encoding, etc.), and then the ALF estimation process is started. The ALF filter coefficients are written into the APS, and then the slice data can be written including the ALF for each CTU parameter (with other parameters scattered). The picture can be transmitted only after the ALF encoding is completed.

[0057] Figure 5 shows an example of a low-latency encoder assumed by the present invention. Slices can be sent before the picture encoding process is completed. In particular, the APS carrying the coefficients is moved behind the slice data (VCL NAL unit).

[0058] In this process, the encoder can first send the encoded slices of the picture while collecting the estimated ALF parameters (filter coefficients, filter control information) and then collecting the APS including the ALF parameters after the encoded picture. The decoder can start syntax analysis and decoding of the slices of the picture as soon as the slices of the picture arrive. Since ALF is one of the last decoding steps, the ALF parameters can arrive after the encoded picture to which other decoding steps are applied.

[0059] The present invention includes the following aspects. - An APS type indicating that the APS belongs to the previous picture. - The APS type is encoded with syntax elements. Embodiments are as follows. 〇A new NAL unit type (e.g., suffix APS) encoded in the NAL unit header. 〇A new APS type (e.g., suffix ALF APS) encoded with the APS. 〇A flag indicating whether the APS is applied to the previous picture or the subsequent picture. - Instead of the suffix APS, only assume the start of a new access unit AU by the signaling prefix APS, a change to the decoding process. - Alternatively, the APS type (prefix or suffix) can be determined by the relative position in the bitstream with respect to the surrounding access unit delimiter (AUD) and the coded slice of the picture. 〇 If the APS is located between the last VCL NAL unit of the picture and the AUD, it applies to the previous picture. 〇 If the APS is located between the AUD and the first VCL NAL unit of the picture, it applies to the next picture. - In this case, the decoder only needs to determine the start of a new access unit by the position of the AUD. - In one embodiment, the slice header indicates that a suffix APS is used instead of a prefix APS, and as a result, the dependency is resolved at the end of the picture decoding process, for example, by the following equation. 〇 Signal the APS identifier at a position within the slice header according to its prefix or suffix characteristic. 〇 Signal a flag indicating that the referenced APS follows the slice coded in bitstream order.

[0060] Typically, not only is the derivation of the ALF parameters (filter coefficients) performed towards the end of the encoding process (based on the reconstructed sample values), but additional ALF control information (information regarding whether the coding tree unit CTU is filtered and how it is filtered) is also derived at this stage. The ALF control information is carried by several syntax elements for each coding_tree_unit within the slice payload, with block partitioning (e.g., as shown in Figure 6b), transform coefficients, etc. scattered therein. For example, the following syntax elements may exist. 〇 alf_ctb_flag: Specifies whether to apply the adaptive loop filter to the coding tree block CTB. 〇alf_ctb_use_first_APS_flag: Specifies whether to use the filter information of the APS where adaptive_parameter_set_id is equal to slice_alf_APS_id_luma[0]. 〇alf_use_APS_flag: Specifies whether to apply the filter set from the APS to the luma CTB. 〇Others.

[0061] All of this ALF control information depends on the derivation of the ALF filter parameters towards the end of the picture encoding process.

[0062] In one embodiment, the ALF control information is signaled in a separate loop over the CTUs of the slices at the end of each slice payload so that the encoder can finalize the first part of the slice payload (transform coefficients, block structure, etc.) before the ALF is executed. This embodiment is shown in FIGS. 6a and 6b.

[0063] As shown in FIG. 6a, the video coding unit (VCL NAL unit) 100 includes a slice header, slice data 108, and parts (ALF for each CUT APS) 106, and one or more parameter sets (ALF coefficients) 102 are signaled separately, i.e., as suffix APSs (within non-VCL-NAL units). That is, each video coding unit 100 is arithmetically encoded continuously across the data to the end of the part along the data stream order.

[0064] FIG. 6b shows that the slice data 108 is interspersed in the part 106. That is, the ALF for each CUT APS is interspersed in blocks, and one or more parameter sets 102 are signaled separately as suffix APSs.

[0065] In another embodiment, the slice header indicates that the ALF control information is not in the syntax elements within the coded slice payload, i.e., not within the CTU loop described above, but is within a suffix APS, i.e., for example, with reference to an APS which is of the suffix APS type, within a separate loop over all CTUs within each suffix APS.

[0066] In another embodiment, the slice header indicates that the ALF control information is not in the syntax elements within the coded slice payload, i.e., not within the CTU loop described above, but is within a new type of suffix APS different from the suffix APS that carries the ALF coefficients, i.e., for example, with reference to an APS which is of the suffix APS type, within a separate loop over all CTUs within each referenced APS. The data for each CTU can optionally be CABAC coded. This embodiment is shown in FIG. 7.

[0067] As shown in FIG. 7, the data unit is signaled in data stream order by a video coding unit (VCL NAL unit) 100 including a slice header and slice data 108, a parameter set 104 (Suffix ALF CTU - data APS: non - VCL NAL unit), a further video coding unit 100, a further parameter set 104, and a parameter set (filter control information) 102 (Suffix ALF coefficient APS: non - VCL NAL unit). That is, it is the reverse of the data stream shown in FIGS. 6a and 6b, and the filter coefficients do not need to be signaled following all video coding units. In other words, the filter coefficients may be sent collectively for two or more video coding units 100 in bit - stream order, or may be further used by a further video coding unit following the filter coefficients in bit - stream order.

[0068] In another embodiment, the slice header that refers to the suffix APS and all CTUs is presumed to have an adaptive loop filter to which the default values of the filter parameters and ALF control information signaled with the suffix APS are applied.

[0069] Signaling of reference to parameter set ID in AUD Hereinafter, another aspect of the present invention, that is, a method for facilitating access to a list of all parameter sets referred to within a picture will be described.

[0070] According to this aspect of the present invention of the present application, the decoder can easily determine whether all necessary parameter sets are available before starting decoding. - The list of all parameter sets used is included in a high-level syntax structure. - The list is composed of the following. 〇One VPS (Video Parameter Set) 〇One DPS (Decoder Parameter Set) 〇One or more SPSs (Sequence Parameter Sets) 〇One or more PPSs (Picture Parameter Sets) 〇One or more APSs (Adaptation Parameter Sets) (ordered by APS type) - Optionally, one or more syntax elements are present before each list, indicating which parameter set types are in the list (including the option to disable transmission). - The syntax structure that holds the information is included in the Access Unit Delimiter (AUD).

[0071] An example of the syntax is shown in FIG. 8, that is, a predetermined unit, that is, the AUD includes a plurality of identifiers 200. For example, the identifier "aud_vps_id" of the VPS, the identifier "aud_dps_id" of the DPS, the identifier "aud_sps_id" of the SPS, the identifier "aud_pps_id" of the PPS, and the like.

[0072] Signaling of APS ID within AUD only The APS is referenced by each slice of the picture. When combining bitstreams, it may be necessary to rewrite and / or combine different APSs.

[0073] To avoid rewriting the slice header, the APS ID is signaled by the access unit delimiter instead of the slice header. Therefore, in case of change, there is no need to rewrite the slice header. Rewriting the access unit delimiter is a much easier operation.

[0074] An example of the syntax is shown in Figure 9.

[0075] In another embodiment, the APS ID is transmitted only in an AUD conditioned on another syntax element. If the syntax element indicates that the APS ID does not exist in the AUD, the APS ID exists in the slice header. An example of the syntax is shown in Figure 10. That is, as shown in Figure 10, the AUD includes a flag 204, for example, the syntax "aps_ids_in_aud_enabled_flag".

[0076] Figures 11a to 11c are schematic diagrams of examples showing the relationship between a predetermined coding parameter and an AU according to the above embodiments shown in Figures 8 to 10.

[0077] As shown in FIG. 11a, the plurality of parameter sets 120 includes one or more first predetermined parameter sets 126 including, for example, APA and PPS, and one or more second predetermined parameter sets 128 including, for example, SPS, DPS, and VPS. The second predetermined parameter set 128 belongs to a higher hierarchical level than the first predetermined parameter set 126. As shown in FIG. 11a, the AU includes a plurality of slice data, for example, VCL0 to VCLn, and the first and second predetermined parameter sets are included in a predetermined parameter set 122, for example, a parameter set for VCL0.

[0078] The plurality of parameter sets 120 are stored in the AUD of the AU and signaled to the decoder.

[0079] As shown in FIG. 11b, when the flag 204 is included in the AUD, the flag 204 indicates whether a predetermined identifier of the identifier 200 referring to a specific predetermined parameter set 126b exists in the predetermined unit 124, or whether a predetermined identifier referring to a specific predetermined parameter set 126b exists in one or more video coding units 100 (205). That is, the flag 204 indicates whether the APS126b is in the AUD or in the VLC (as indicated by the arrow in FIG. 11b).

[0080] As shown in FIG. 11c, the first predetermined parameter set 126 includes a third predetermined parameter set 126a, for example, PPS, which is referred to by an identifier within one or more video coding units 100 (for example, AU), and an identifier 200 that exists within a predetermined unit 124 (for example, AUD) (shown in FIGS. 8 and 9), but is not referred to by any of the identifiers within one or more video coding units 100 (AU), and is not referred to by any of the predetermined parameter sets, a fourth predetermined parameter set 126b, for example, APS.

[0081] Signaling of Access Unit Characteristics to AUD Currently, the AUD indicates whether the following slices are of type I, B, or P. In most systems, this function is not very useful because it does not necessarily mean that an I picture has a random access point. The prioritization of AUs when they need to be dropped can usually be done in other ways, such as by analyzing the temporal ID and whether the picture is a disposable picture (not referenced by others).

[0082] Instead of indicating the picture type, it can indicate the NAL unit type and whether they are disposable pictures, etc. Furthermore, in the case of multiple layers, it may become more difficult to describe the characteristics. 〇Instead of specifying within the NAL unit header, the random access characteristics of the picture are specified by the overall NAL unit type (e.g., IDR, CRA, etc.) used for all VCL NAL units within the access unit. 〇Pictures within a layer can be discarded in one layer but are not collocated pictures in another layer. 〇Pictures within a layer are marked as no output (pic_output_flag) in one layer and are not collocated pictures in another layer.

[0083] Therefore, in the embodiment shown in FIG. 12a, the AUD indicates whether the information applies to a single layer or all layers. That is, in the AUD flag, information regarding the AU characteristics is referenced, for example, as shown below.

[0084] "layer_specific_aud_flag" 300: Controls whether separate access units are defined in the video data stream for pictures related to different layers that are instantaneous in the video data stream, and / or whether pictures related to different layers that are instantaneous in the video data stream are encoded in one of the access units, and / or "nal_unit_type_present_flag" 302: In the case of video coding type indication, it indicates the video coding type of the video coding units within an access unit that are assigned to the video coding units within one access unit. The video coding units included within one access unit are different from each other. That is, by indicating the presence of the syntax element of the NAL unit type, it indicates the NAL unit type, and / or "discardable_flag" 304: Indicates a picture that is not referenced by any other picture in the access unit, and / or "pic_output_flag" 306: Indicates a picture that is not output. In another embodiment, the AUD can indicate that the AUD is a dependent AUD. This means the following. 〇 Inherit parameters from the AUD before the dependent layer, but add some layer-specific information, and / or 〇 A new AU has not started.

[0085] Figure 12b shows the indication in the AUD. That is, the video coding type of the video coding unit is indicated by describing the random access characteristics of multiple pictures. That is, the syntax "random_access_info_present_flag" is specified by the overall NAL unit type (such as IDR, CRA, etc.) used for all VCL NAL units within the access unit, instead of being specified within the NAL unit header, as shown by "all_pics_in_au_random_access_flag" in Figure 12b, to indicate the random access characteristics of the pictures.

[0086] An example of the syntax according to an embodiment of the present invention is shown in Figure 13.

[0087] In this example, the parameter set 308, i.e., the implementation "layer_specific_aud_flag" indicating whether the information in the AUD is applied to all layers, "dependent_aud_flag" indicating whether the AUD starts a new global access unit, or whether it is only a "layer-access unit". In a dependent AUD, the inheritance from the base layer AUD is indicated by the "aud_inheritance_flag".

[0088] Although several aspects have been described in the context of an apparatus, it is clear that these aspects also represent corresponding method descriptions where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent descriptions of corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.

[0089] The data stream of the present invention can be stored in a digital storage medium or transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0090] Depending on specific implementation requirements, embodiments of the present application can be implemented in hardware or software. This implementation can be executed using a digital storage medium, such as a floppy (registered trademark) disk, DVD, Blu-ray (registered trademark), CD, ROM, PROM, EPROM, EEPROM, or flash memory, which stores electronically readable control signals that cooperate (or can cooperate) with a programmable computer system such that their respective methods are executed. Thus, the digital storage medium may be computer-readable.

[0091] Some embodiments according to the present invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system so that one of the methods described herein is executed.

[0092] In general, embodiments of the present application can be implemented as a computer program product having program code, and the program code is operable to execute one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, on a machine-readable carrier.

[0093] Other embodiments comprise a computer program for executing one of the methods described herein, stored on a machine-readable carrier.

[0094] In other words, an embodiment of the method of the present invention is thus a computer program having program code for executing one of the methods described herein when the computer program is executed on a computer.

[0095] Accordingly, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for executing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.

[0096] Accordingly, a further embodiment of the method of the present invention is a sequence of a data stream or signal representing a computer program for executing one of the methods described herein. The sequence of the data stream or signal may be configured to be transferred via a data communication connection, for example via the Internet.

[0097] Further embodiments comprise processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.

[0098] Further embodiments include a computer having installed thereon a computer program for performing one of the methods described herein.

[0099] Further embodiments according to the present invention comprise an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0100] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0101] The apparatuses described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0102] The apparatuses described herein, or any component of the apparatuses described herein, can be at least partially implemented in hardware and / or software.

Claims

1. A method for decrypting a video, comprising: reading a filter coefficient located in a suffix unit corresponding to the first picture of the video stream, wherein the filter coefficient is for a second picture following the first picture; reading a flag from the video data corresponding to the second picture of the video stream, the flag specifying that an adaptive filter is applied to a specific portion of the second picture; applying the adaptive filter to the specific portion of the second picture using the filter coefficient.

2. The method according to claim 1, wherein the suffix unit is the last unit of an access unit (AU) of the first picture.

3. The method according to claim 1, wherein the video data corresponding to the second picture includes one or more video coding layer (VCL) network abstraction layer (NAL) units representing slices.

4. The method according to claim 1, wherein the suffix unit corresponding to the first picture includes a non-VCL NAL unit following all VCL NAL units corresponding to the first picture.

5. The method according to claim 1, wherein the specific portion of the second picture includes a coding tree block (CTB) of the second picture.

6. The method according to claim 1, wherein the filter coefficient includes adaptive loop filter (ALF) parameters, and the flag includes an ALF CTB flag.

7. A method for encoding a video, comprising: encoding a filter coefficient located in a suffix unit corresponding to the first picture of the video stream, wherein the filter coefficient is for a second picture following the first picture; encoding a flag from the video data corresponding to the second picture of the video stream, the flag indicating that the adaptive filter is applied to a specific portion of the second picture, wherein the adaptive filter is applied to the specific portion of the second picture using the filter coefficient.

8. The method according to claim 7, wherein the suffix unit is the last unit of the access unit (AU) of the first picture.

9. The method according to claim 7, wherein the video data corresponding to the second picture includes one or more video coding layer (VCL) network abstraction layer (NAL) units representing slices.

10. The method according to claim 7, wherein the suffix unit corresponding to the first picture includes a non-VCL NAL unit following all the VCL NAL units corresponding to the first picture.

11. The method according to claim 7, wherein the specific portion of the second picture includes a coding tree block (CTB) of the second picture.

12. The method according to claim 7, wherein the filter coefficient includes an adaptive loop filter (ALF) parameter and the flag includes an ALF CTB flag.

13. A non-transitory computer-readable medium storing instructions that, when executed, cause a computer to execute the method according to any one of claims 1 to 12.

14. Comprising at least one processor and a memory, A video coding apparatus, wherein the at least one processor is configured to cooperate with the memory to execute the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Adaptation parameter sets for video coding

    US20130022104A1

  • Adaptation parameter sets (APS) for adaptive loop filter (ALF) parameters

    US20200344473A1

  • Modified adaptive loop filter temporal prediction for temporal scalability support

    WO2018129168A1