Non-uniform classification for adaptive loop filter classifiers
By using the non-uniform classification technique of adaptive loop filters, the problem of low coding efficiency caused by uneven filter classification in existing technologies is solved, and more efficient video encoding and decoding effects are achieved.
Patent Information
- Application Number
- CN202480026944.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-21
- Filing Date
- 2024-04-20
- Publication Date
- 2025-11-21
AI Technical Summary
Existing video encoding and decoding technologies struggle to effectively utilize non-uniformly distributed filters to optimize image quality when processing video data, resulting in poor encoding efficiency and decoding performance.
The non-uniform classification technique of Adaptive Loop Filter (ALF) is adopted to determine the filter category through non-uniform distribution within the dynamic range, and the filter coefficients are calculated based on the classification value to filter the video data.
It improves the efficiency of video encoding and decoding, enhances image quality, and optimizes encoding efficiency.
Smart Images

Figure CN121002871A_ABST
Abstract
Description
Related applications
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 461,225, filed April 21, 2023, entitled “Non-uniform Classification for Adaptive Loop Filter Classifier,” which is incorporated herein by reference in its entirety. Technical Field
[0002] This disclosure generally describes implementation methods involving video encoding and decoding. Background Technology
[0003] The background description provided herein is for the purpose of presenting the overall context of this disclosure. To the extent that the work of the currently attributed inventors is described in the background section, such work, and aspects of the description which at the time of submission may not otherwise be considered prior art, are neither expressly nor implicitly acknowledged as prior art to this disclosure.
[0004] Image / video compression can help transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial and temporal redundancy. For instance, a video codec can use a technique called intra-frame prediction, which can compress images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current image being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress images based on temporal redundancy. For example, inter-frame prediction can utilize motion compensation to predict samples in the current image based on previously reconstructed images. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention
[0005] This disclosure includes methods and apparatus for video encoding / decoding.
[0006] Some aspects of this disclosure provide a method for processing visual media data. This method includes processing a bitstream of visual media data according to format rules. The bitstream includes encoded information for one or more images. The format rules specify: determining that a non-uniform classification for an Adaptive Loop Filter (ALF) is applied to a current block in the current image, and calculating a classification value associated with a classification unit for applying the ALF, the classification unit being in the current block. The format rules also specify: determining a filter class for the classification unit from multiple filter classes based on the classification value in a dynamic range, the multiple filter classes being distributed according to a non-uniform distribution in the dynamic range, the dynamic range including at least a first filter class corresponding to a first value range in the dynamic range and a second filter class corresponding to a second value range in the dynamic range, the first value range and the second value range having a difference range size. The format rules further specify: determining filter coefficients associated with the filter class, and generating at least one filtered sample of the classification unit based on the filter coefficients of the ALF.
[0007] Some aspects of this disclosure provide a video decoding apparatus including a processing circuit system configured to receive encoded information of one or more images, the encoded information indicating the use of non-uniform classification for applying filters. The processing circuit system is configured to compute a classification value associated with a classification unit for applying a filter to the classification unit, the classification unit being in a current block within a current image. Furthermore, the processing circuit system is configured to determine a filter class for the classification unit from a plurality of filter classes distributed according to a non-uniform distribution in the dynamic range based on the classification value. Additionally, the processing circuit system is configured to obtain filter coefficients associated with the filter class determined for the classification unit, generate at least one filtered sample of the classification unit based on the filter coefficients, and reconstruct the current block based on the at least one filtered sample.
[0008] In some examples, the dynamic range includes at least a first filter category corresponding to a first value range in the dynamic range and a second filter category corresponding to a second value range in the dynamic range, wherein the first value range and the second value range have a difference in range size.
[0009] In some examples, the dynamic range includes sub-intervals with indices ranging from 0 to k-1, where k is the number of sub-intervals, and the size of each sub-interval is set to... The sub-interval has a corresponding number of categories. The sum of the corresponding class numbers in the sub-intervals equals the total number of classes in the multiple filter classes. The processing circuitry is configured to determine the filter category based on a specific sub-interval to which the classification value belongs and the number of specific categories within that sub-interval.
[0010] In some examples, the dynamic range is non-uniformly divided into sub-intervals, each associated with one of multiple filter categories, and the processing circuitry is configured to determine the filter category based on the specific sub-interval to which the classification value belongs.
[0011] In some examples, the dynamic range includes intervals with non-uniform sizes. The subintervals, and They do not have the same value. In the example, They have the same value.
[0012] In some examples, the dynamic range includes at least a first sub-interval and a second sub-interval with different interval ranges, and the first and second sub-intervals have the same number of categories.
[0013] In some examples, the dynamic range includes at least a first sub-interval and a second sub-interval with the same interval size, the first sub-interval having a first number of categories, the second sub-interval having a second number of categories, and the first number of categories being different from the second number of categories.
[0014] In some examples, the dynamic range includes sub-intervals with the same interval size, and the number of corresponding categories for each sub-interval. They are not the same number.
[0015] In some examples, the processing circuitry is configured to determine a specific sub-interval to which a classification value belongs, and to determine a specific category within that specific sub-interval for the classification value.
[0016] In some examples, the processing circuitry is configured to decode one or more signals for a non-uniform distribution based on one or more high-level syntax elements in the encoded information, wherein the one or more high-level syntax elements are at least one of the following: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, and Slice Header.
[0017] In the example, the number of subintervals, the size of the subinterval range, and the number of categories in the subinterval are predefined parameters for nonuniform distribution, and the processing circuitry is configured to decode the flag indicating whether nonuniform distribution is used based on the encoded information.
[0018] In another example, the processing circuitry is configured to decode a flag indicating whether a non-uniform distribution is used based on encoded information, and when the flag indicates the use of a non-uniform distribution, to decode at least one of the number of sub-intervals, the range size of the sub-intervals, and the number of categories in the sub-intervals based on the encoded information.
[0019] In some examples, the filter is an adaptive loop filter.
[0020] In some examples, the classification value is one of the following: sample value, residual value, and derived value from the window covering the classification unit.
[0021] Some aspects of this disclosure provide a video coding method. The method includes: determining the use of a non-uniform classification for an adaptive loop filter (ALF) to be applied to a current block in a current image; and calculating a classification value associated with a classification unit for applying the ALF, the classification unit being associated with the current block. The method also includes determining a filter class for the classification unit from a plurality of filter classes based on the classification value in a dynamic range, the plurality of filter classes being distributed according to a non-uniform distribution in the dynamic range, the dynamic range including at least a first filter class corresponding to a first value range in the dynamic range and a second filter class corresponding to a second value range in the dynamic range, the first value range and the second value range having a difference range size. The method further includes obtaining filter coefficients associated with the filter class determined for the classification unit; and generating at least one filtered sample of the classification unit based on the filter coefficients of the ALF.
[0022] In some examples, the dynamic range includes sub-intervals with indices ranging from 0 to k-1, where k is the number of sub-intervals, and the size of each sub-interval is set to... The sub-interval has a corresponding number of categories. The sum of the corresponding class numbers in the sub-intervals equals the total number of classes in the multiple filter classes. This method involves determining the filter category based on a specific sub-interval to which the classification value belongs and the number of specific categories within that sub-interval.
[0023] In some examples, the dynamic range includes intervals with corresponding non-uniform sizes. The subintervals, and They have the same value. In the example, the dynamic range includes sub-intervals with the same interval size and the corresponding number of categories in the sub-intervals. They are not the same number.
[0024] According to another aspect of this disclosure, an apparatus is provided. The apparatus includes a processing circuitry system. This processing circuitry system can be configured to perform any of the described methods for video decoding / encoding.
[0025] This disclosure also provides a non-transitory computer-readable medium for storing instructions that, when executed by a computer, cause the computer to perform any of the described methods for video decoding / encoding. Attached Figure Description
[0026] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:
[0027] Figure 1 It is a schematic illustration of an exemplary block diagram of a communication system (100).
[0028] Figure 2 This is a schematic illustration of an exemplary block diagram of a decoder.
[0029] Figure 3 This is a schematic illustration of an exemplary block diagram of an encoder.
[0030] Figure 4 A diagram illustrating filtering according to an embodiment of this disclosure is shown.
[0031] Figure 5 A diagram showing two filter shapes in some examples is provided.
[0032] Figures 6A to 6D Examples of subsampling locations used to compute gradients are shown in some examples.
[0033] Figure 7 Examples of filter shapes using residual samples as additional inputs are shown in some examples.
[0034] Figure 8 A flowchart outlining the decoding process according to some embodiments of this disclosure is shown.
[0035] Figure 9 A flowchart is shown that outlines the coding process according to some embodiments of this disclosure.
[0036] Figure 10 It is a schematic diagram of a computer system according to an implementation method. Detailed Implementation
[0037] Figure 1 Block diagrams of some example video processing systems (100) are shown. The video processing system (100) is an example of the application of the disclosed subject matter—video encoders and video decoders—in a streaming environment. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs (Compact Discs), DVDs (Digital Versatile Discs), memory sticks, etc.
[0038] The video processing system (100) includes a capture subsystem (113) that may include a video source (101), such as a digital camera device, which creates, for example, an uncompressed video picture stream (102). In the example, the video picture stream (102) includes samples captured by the digital camera device. The video picture stream (102) is depicted as a thick line to emphasize the high data volume when compared with encoded video data (104) (or encoded video bitstream), which may be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize its lower data volume when compared to the video picture stream (102). The encoded video data (104) (or encoded video bitstream) can be stored on a streaming server (105) for future use. One or more streaming client subsystems, for example... Figure 1 Client subsystems (106) and (108) can access a streaming server (105) to retrieve copies (107) and (109) of encoded video data (104). Client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of encoded video data and creates an outgoing video picture stream (111) that can be presented on a display (112) (e.g., a screen) or other presentation device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., video bitstreams) may be encoded according to certain video coding standards / video compression standards. Examples of such standards include ITU-T (International Telecommunication Union-Telecommunication Standardization Sector, ITU-T) Recommendation H.265. In the example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The topics that are exposed can be used in the context of VVC.
[0039] Note that electronic devices (120) and (130) may include other components (not shown). For example, electronic device (120) may include a video decoder (not shown), and electronic device (130) may also include a video encoder (not shown).
[0040] Figure 2An exemplary block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuitry system). The video decoder (210) may be used in place of... Figure 1 The video decoder (110) in the example.
[0041] The receiver (231) can receive, for example, one or more encoded video sequences to be decoded by the video decoder (210) included in a bitstream. In an implementation, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences can be received from a channel (201), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (231) can receive the encoded video data as well as other data such as encoded audio data and / or auxiliary data streams, which can be forwarded to their respective user entities (not depicted). The receiver (231) can separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) can be external to the video decoder (210) (not depicted). In other applications, a buffer memory (not depicted) may exist outside the video decoder (210) to prevent network jitter, for example, and another buffer memory (215) may exist inside the video decoder (210) to handle broadcast timing, for example. The buffer memory (215) may not be necessary, or it may be small, when the receiver (231) is receiving data from a store / forward device with sufficient bandwidth and controllability, or from an isochronous synchronization network. For optimal use on packet networks such as the Internet, the buffer memory (215) may be required; it may be relatively large and advantageously have an adaptive size, and may be implemented at least partially in the operating system or in a similar element (not depicted) outside the video decoder (210).
[0042] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. These symbols include: information for managing the operation of the video decoder (210), and potential information for controlling presentation devices such as presentation device (212) (e.g., a display screen), which is not part of the electronic device (230) but may be coupled to it, such as... Figure 2As shown. Control information for the presentation device can be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (220) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can conform to video coding techniques or standards and can follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) can extract a subgroup parameter set of at least one subgroup of pixels in the subgroups of the encoded video sequence for use in the video decoder based on at least one parameter corresponding to a group. Subgroups can include Group of Picture (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (220) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0043] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0044] Depending on the type of the encoded video picture or its portion (e.g., inter-frame picture and intra-frame picture, inter-frame block and intra-frame block) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (220). For clarity, the flow of such subgroup control information between the parser (220) and the multiple units below is not depicted.
[0045] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.
[0046] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantization transform coefficients as symbols (221) from the parser (220) and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output blocks containing sample values that can be input into the aggregator (255).
[0047] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to intra-coded blocks. Intra-coded blocks are blocks that do not use predictive information from previously reconstructed images, but can use predictive information from previously reconstructed portions of the current image. Such predictive information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information obtained from the current picture buffer (258) to generate blocks of the same size and shape as the blocks in the reconstruction. For example, the current picture buffer (258) buffers the partially reconstructed current image and / or the fully reconstructed current image. In some cases, the aggregator (255) adds the predictive information already generated by the intra-picture prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.
[0048] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-frame coded block and potentially to a motion-compensated block. In this case, the motion-compensated prediction unit (253) can access the reference image memory (257) to obtain samples for prediction. After motion compensation of the obtained samples according to the symbols (221) belonging to the block, these samples can be added to the output of the scaler / inverse transform unit (251) by the aggregator (255) (in this case, referred to as residual samples or residual signals) to generate output sample information. The address in the reference image memory (257) from which the motion-compensated prediction unit (253) obtains the predicted samples can be controlled by motion vectors, which are provided to the motion-compensated prediction unit (253) in the form of symbols (221), which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values obtained from the reference image memory (257) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0049] The output samples of the aggregator (255) can undergo various loop filtering techniques in the loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters available to the loop filter unit (256) as included in the encoded video sequence (also referred to as the encoded video bitstream) and as symbols (221) from the parser (220). Video compression may also respond to metadata acquired during decoding of previous portions of the encoded picture or encoded video sequence (in decoding order), as well as to previously reconstructed and loop-filtered sample values.
[0050] The output of the loop filter unit (256) can be a sample stream, which can be output to the presentation device (212) and stored in the reference image memory (257) for use in future inter-frame image prediction.
[0051] Once a certain encoded image has been fully reconstructed, it can be used as a reference image for future predictions. For example, once the encoded image corresponding to the current image has been fully reconstructed and that encoded image (by, for example, the parser (220)) is identified as the reference image, the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.
[0052] The video decoder (210) can perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T Recommendation H.265. The encoded video sequence can conform to the syntax specified by the video compression technology or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the configuration file as documented in the video compression technology or standard. Specifically, the configuration file can select certain tools from all available tools in the video compression technology or standard as tools usable only under said configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference picture size, etc. In some cases, the limitations set by the hierarchy can be further restricted by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented as a signal in the encoded video sequence.
[0053] In this implementation, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be, for example, in the form of temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0054] Figure 3 An exemplary block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit system). The video encoder (303) can be used in place of Figure 1 The video encoder (103) in the example.
[0055] The video encoder (303) can obtain data from the video source (301) (which is not...). Figure 3 In one example, a portion of the electronic device (320) receives video samples, and a video source (301) can capture video images to be encoded by a video encoder (303). In another example, the video source (301) is a portion of the electronic device (320).
[0056] The video source (301) can provide a sequence of source video samples in the form of a digital video sample stream to be encoded by the video encoder (303). This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCb, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera device capturing local image information as a video sequence. Video data can be provided as multiple individual pictures that are given motion when viewed sequentially. The pictures themselves can be organized as spatial pixel arrays, where each pixel can include one or more samples, depending on the sampling structure, color space, etc. The following description focuses on samples.
[0057] According to the implementation, the video encoder (303) can encode and compress images of the source video sequence into an encoded video sequence (343) in real time or under any other required time constraints. Implementing an appropriate encoding rate is a function of the controller (350). In some implementations, the controller (350) controls and is functionally coupled to other functional units as described below. For clarity, the coupling is not depicted. Parameters set by the controller (350) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of images (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions belonging to the video encoder (303) optimized for a specific system design.
[0058] In some implementations, the video encoder (303) is configured to operate within an encoding loop. As a simplified description, in this example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (333) embedded within the video encoder (303). The decoder (333) reconstructs the symbols, creating sample data in a manner similar to how a (remote) decoder would also create them. The reconstructed sample stream (sample data) is input to a reference image memory (334). Since decoding the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference image memory (334) are also bit-accurate between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same sample values that the decoder will "see" when using prediction during decoding. This basic principle of reference image synchronicity (and drift that occurs, for example, due to channel errors) is also used in some related techniques.
[0059] The operation of the "local" decoder (333) can be combined with what has already been done above. Figure 2 The operation of a "remote" decoder, such as a video decoder (210), is the same as described in the detailed description. However, a brief reference is also provided. Figure 2 Since symbols are available and the encoding of symbols into an encoded video sequence by the entropy encoder (345) and the decoding of symbols by the parser (220) can be lossless, the entropy decoding portion of the video decoder (210), which includes the buffer (215) and the parser (220), may not be fully implemented in the local decoder (333).
[0060] In the implementation, decoder techniques other than parsing / entropy decoding present in the decoder exist in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on decoder operation. Since the encoder technique is the inverse of the fully described decoder technique, the description of the encoder technique can be simplified. A more detailed description is provided below in certain sections.
[0061] In some examples, during operation, the source encoder (330) may perform motion-compensated predictive coding, referencing one or more previously encoded images designated as "reference images" from the video sequence to predictively encode the input image. In this way, the encoding engine (332) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which can be selected as predictive references for the input image.
[0062] The local video decoder (333) can decode encoded video data of a picture that can be designated as a reference picture based on symbols created by the source encoder (330). The operation of the encoding engine (332) can advantageously be lossy. When the encoded video data can be decoded by the video decoder (333), Figure 3 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process performed on the reference image by the video decoder and can store the reconstructed reference image in the reference image memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.
[0063] The predictor (335) can perform a prediction search against the encoding engine (332). That is, for a new image to be encoded, the predictor (335) can search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or certain metadata such as reference image motion vectors, block shapes, etc. that can be used as appropriate prediction references for the new image. The predictor (335) can operate on a pixel-by-pixel basis based on the sample blocks to find appropriate prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (334).
[0064] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, the setting of parameters and subgroup parameters for encoding video data.
[0065] The outputs of all the functional units mentioned above can be entropy encoded in the entropy encoder (345). The entropy encoder (345) converts the symbols generated by the various functional units into an encoded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0066] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission via a communication channel (360), which can be a hardware / software link to a storage device storing the encoded video data. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).
[0067] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a specific encoded image type to each encoded image, which may affect the encoding techniques that can be applied to the corresponding image. For example, images can typically be assigned to one of the following image types:
[0068] Intra-frame pictures (I-pictures) can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0069] Predictive images (P-images) are encoded and decoded using intra-frame or inter-frame predictions that utilize motion vectors and reference indices to predict sample values for each block.
[0070] Bidirectional predictive images (B-images) can be encoded and decoded using intra-frame or inter-frame predictions that utilize two motion vectors and a reference index to predict sample values for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata for the reconstruction of a single block.
[0071] Source images are typically spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, which are determined by the coding assignment of the corresponding images applied to the blocks. For example, blocks of image I can be non-predictively coded, or blocks of image I can be predictively coded (spatial prediction or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of image P can be predictively coded with reference to a previously coded reference image via spatial prediction or via temporal prediction. Blocks of image B can be predictively coded with reference to one or two previously coded reference images via spatial prediction or via temporal prediction.
[0072] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In its operation, the video encoder (303) can perform various compression operations, including predictive coding operations utilizing temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard in use.
[0073] In this implementation, the transmitter (340) may transmit additional data along with the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant images and slices, SEI messages, VUI parameter set fragments, etc.
[0074] Video can be captured in a time series as multiple source images (video images). Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In the example, a specific image during encoding / decoding—referred to as the current image—is segmented into blocks. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and in the case of multiple reference images, the motion vector can have a third dimension that identifies the reference images.
[0075] In some implementations, bidirectional prediction techniques can be used for inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. A block can be predicted using a combination of the first and second reference blocks.
[0076] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.
[0077] According to some embodiments of this disclosure, predictions such as inter-frame picture prediction and intra-frame picture prediction are performed on a block-by-block basis. For example, according to the HEVC (High Efficiency Video Coding) standard, pictures in a video picture sequence are segmented into Coding Tree Units (CTUs) for compression. The CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU comprises three Coding Tree Blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU can be recursively partitioned into one or more Coding Units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be segmented into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In the example, each CU is analyzed to determine the prediction type used for that CU, such as inter-frame prediction or intra-frame prediction. Based on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In implementations, prediction operations in encoding / decoding are performed on a block-by-block basis. Using a luma prediction block as an example, this block comprises a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0078] Note that any suitable technology can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In one implementation, one or more integrated circuits can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In another implementation, one or more processors that execute software instructions can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210).
[0079] Various aspects of this disclosure provide non-uniform classification techniques for filter classifiers such as adaptive loop filter (ALF) classifiers.
[0080] In some examples (e.g., VVC), adaptive loop filters (ALF) and cross component adaptive loop filters (CC-ALF) are used in video encoding and decoding.
[0081] In some examples, artifacts can be reduced by applying an ALF with block-based filter adaptation by the encoder / decoder. Two filter shapes for block-based ALF can be used in VVC. For example, a 7×7 rhombus shape is used for the luma component, and a 5×5 rhombus shape is used for the chroma component. In the examples, one of up to 25 filters is selected for each 4×4 block based on the directionality and activity of the local gradients. Each 4×4 block is classified and categorized into one of 25 classes based on the directionality and activity of the local gradients. Each class can have its own filter coefficient assignment. In some examples, geometric transformations, such as a 90-degree rotation, diagonal flip, or vertical flip, can be applied to the filter shape before filtering based on the gradient values computed for the block. Geometric transformations of the filter shape are equivalent to applying these transformations to samples in the filter's support region. Geometric transformations can be performed for each block by aligning the filter with the block's directionality.
[0082] In some examples, in addition to 4×4 block-level filter adaptation for luma, ALF also supports CTU-level filter adaptation. For example, each CTU can use a filter bank computed based on the current slice, or one of a filter bank represented by the signal at the encoded slice, or one of 16 offline-trained filter banks. Within each CTU, the selected filter bank can be applied to each 4×4 block. The filter coefficients and clipping indexes are carried in the ALF Adaptive Parameter Set (APS). The ALF APS can include up to eight chroma filters and a luma filter bank with up to 25 filters. In the example, each of the 25 luma categories also includes an index. By merging different categories, the number of bits in the filter coefficients can be reduced.
[0083] CC-ALF uses luminance sample values to refine chrominance sample values within the ALF processing. The linear filtering operation takes luminance samples as input and generates correction values for the chrominance sample values. Corrections are generated independently for each chrominance component.
[0084] Figure 4 A cross-component filter (e.g., CC-ALF) for generating chromaticity components is illustrated according to embodiments of this disclosure. In some examples, Figure 4 Filtering is illustrated for a first chromaticity component (e.g., first chromaticity CB), a second chromaticity component (e.g., second chromaticity CB), and a luminance component (e.g., luminance CB). The luminance component can be filtered by a Sample Adaptive Offset (SAO) filter (410) to generate a SAO-filtered luminance component (441). The SAO-filtered luminance component (441) can be further filtered by an ALF luminance filter (416) to become a filtered luminance CB (461) (e.g., "Y").
[0085] The first chromaticity component can be filtered by a SAO filter (412) and an ALF chromaticity filter (418) to generate a first intermediate component (452). Furthermore, the SAO-filtered luminance component (441) can be filtered by a cross-component filter (e.g., CC-ALF) (421) for the first chromaticity component to generate a second intermediate component (442). Subsequently, a filtered first chromaticity component (462) (e.g., 'Cb') can be generated based on at least one of the second intermediate component (442) and the first intermediate component (452). In the example, the filtered first chromaticity component (462) (e.g., 'Cb') can be generated by combining the second intermediate component (442) and the first intermediate component (452) using an adder (422). The cross-component adaptive loop filtering process for the first chromaticity component can include steps performed by the CC-ALF (421) and steps performed by, for example, the adder (422).
[0086] The above description can be adapted to the second chromaticity component. The second chromaticity component can be filtered by a SAO filter (414) and an ALF chromaticity filter (418) to generate a third intermediate component (453). Furthermore, the SAO-filtered luminance component (441) can be filtered by a cross-component filter (e.g., CC-ALF) (431) for the second chromaticity component to generate a fourth intermediate component (443). Subsequently, a filtered second chromaticity component (463) (e.g., 'Cr') can be generated based on at least one of the fourth intermediate component (443) and the third intermediate component (453). In the example, the filtered second chromaticity component (463) (e.g., 'Cr') can be generated by combining the fourth intermediate component (443) and the third intermediate component (453) using an adder (432). In the example, the cross-component adaptive loop filtering process for the second chromaticity component can include steps performed by a CC-ALF (431) and steps performed by, for example, an adder (432).
[0087] Cross-component filters (e.g., CC-ALF (421), CC-ALF (431)) can be operated by applying a linear filter with any suitable filter shape to the luminance component (or luminance channel) to refine each chrominance component (e.g., the first chrominance component, the second chrominance component).
[0088] ALFs can have any suitable shape and size. Figure 5 The diagram shows the shapes of two ALF filters in some examples. (Refer to...) Figure 5 ALF(510) to ALF(511) have a rhombus shape, such as a 5×5 rhombus shape in ALF(510) and a 7×7 rhombus shape in ALF(511). In ALF(510), elements (520) to (532) form a rhombus shape and can be used in filtering. Seven values (e.g., C0 to C6) can be used for elements (520) to (532). In ALF(511), elements (540) to (564) form a rhombus shape and can be used in filtering. Thirteen values (e.g., C0 to C12) can be used for elements (540) to (564).
[0089] Reference Figure 5 In some examples, two ALFs (510) to ALF (511) with diamond filter shapes are used. A 5×5 diamond-shaped filter (510) can be used for the chroma component (e.g., chroma block, chroma CB), and a 7×7 diamond-shaped filter (511) can be used for the luma component (e.g., luma block, luma CB). Other suitable shapes and sizes can be used in the ALF. For example, a 9×9 diamond-shaped filter can be used.
[0090] The filter coefficients at the locations indicated by values (e.g., C0 to C6 in (510) or C0 to C12 in (520)) can be nonzero. Furthermore, when the ALF includes a clipping function, the clipping value at those locations can be nonzero.
[0091] For block classification of the luminance component, a 4×4 block (or luminance block, luminance CB) can be classified or categorized into one of several (e.g., 25) categories. This can be based on the quantized values of the directional parameter D and the activity value A. Use equation (1) to derive the category index C.
[0092] Equation (1)
[0093] To calculate the directional parameter D and the quantization value The gradients g in the vertical, horizontal, and two diagonal directions (e.g., d1 and d2) can be calculated using a 1-dimensional Laplacian as follows. v g h g d1 and g d2 .
[0094] Equation (2)
[0095] Equation (3)
[0096] Equation (4)
[0097] Equation (5)
[0098] Among them, index and This refers to the coordinates of the top-left sample within the 4x4 block, and Indicator at coordinates The reconstructed sample at the location. Directions (e.g., d1 and d2) can refer to two diagonal directions.
[0099] To reduce the complexity of the block classification described above, a 1-dimensional Laplace calculation based on subsampling can be applied. Figures 6A to 6D The following diagrams show the methods for calculating the vertical direction ( Figure 6A ), horizontal direction ( Figure 6B ) and the two diagonal directions d1 ( Figure 6C ) and d2 ( Figure 6D The gradient g of ) v g h g d1 and g d2 Examples of subsampling positions. The same subsampling position can be used for gradient calculations in different directions. Figure 6AIn the diagram, the symbol "V" indicates the symbol used to calculate the vertical gradient g. v The subsampling location. Figure 6B In the diagram, the symbol "H" indicates the symbol used to calculate the horizontal gradient g. h The subsampling location. Figure 6C In the diagram, the label "D1" indicates the value used to calculate the diagonal gradient g of d1. d1 The subsampling location. Figure 6D In the diagram, the label "D2" indicates the value used to calculate the diagonal gradient g of d2. d2 The sub-sampling position.
[0100] The gradients g in the horizontal and vertical directions v and g h maximum value and minimum value It can be set to:
[0101] , Equation (6)
[0102] gradient g in the two diagonal directions d1 and g d2 maximum value and minimum value It can be set to:
[0103] , Equation (7)
[0104] The directional parameter D can be based on the above values and the following two thresholds. and Export.
[0105] Step 1. If (1) and (2) If true, then Set to 0.
[0106] Step 2. If If yes, proceed to step 3; otherwise, proceed to step 4.
[0107] Step 3. If Then Set to 2; otherwise, Set to 1.
[0108] Step 4. If Then Set to 4; otherwise, Set it to 3.
[0109] Activity Value It can be calculated as:
[0110] Equation (8)
[0111] The quantization can be further divided into a range of 0 to 4 (inclusive), and the quantized value is represented as... .
[0112] For the chromaticity components in the image, block classification is not applied, and therefore a single set of ALF coefficients can be applied for each chromaticity component.
[0113] Geometric transformations can be applied to filter coefficients and the corresponding filter limiting values (also known as limiting values). Before filtering a block (e.g., a 4×4 brightness block), the filtering can be performed, for example, based on gradient values (e.g., g) calculated for the block. v g h g d1 and / or g d2 ) for filter coefficients and the corresponding filter limiting value Apply geometric transformations such as rotation, diagonal flipping, and vertical flipping to the filter coefficients. and the corresponding filter limiting value Applying a geometric transformation is equivalent to applying a geometric transformation to samples within a region supported by a filter. By aligning their respective orientations, the geometric transformation can make different blocks to which ALF is applied more similar.
[0114] The three geometric transformations, including diagonal flip, vertical flip, and rotation, can be performed as described by equations (9) to (11), respectively.
[0115] , Equation (9)
[0116] , Equation (10)
[0117] , Equation (11)
[0118] in, It is the size of the ALF or filter, and , These are the coordinates of the coefficients. For example, position. At the top left corner of the filter f or the limiting matrix (or limiting matrix) c, and at the position At the bottom right corner of filter f or the limiting matrix (or limiting matrix) c. The filter coefficients can be adjusted based on the gradient values calculated for the block. and amplitude limit Apply the transformation. Table 1 summarizes examples of the relationship between the transformation and the four gradients.
[0119] Table 1: Mapping of gradients and transformations calculated for each block
[0120] gradient value Transformation <![CDATA[g d2 < g d1 And g h < g v ]]> No transformation <![CDATA[g d2 < g d1 And g v < g h ]]> diagonal flip <![CDATA[g d1 < g d2 And g h < g v ]]> Vertical flip <![CDATA[g d1 < g d2 And g v < g h ]]> Rotation
[0121] For filtering, on the decoder side, when ALF is enabled for CTB, for each sample within the CU Filtering is performed to obtain the sample values shown in equation (12). :
[0122]
[0123] Equation (12)
[0124] in, This represents the decoded filter coefficients. It is a limiting function, and This represents the decoded limiting parameters. Variables k and l are... and The value varies between these values, where L represents the filter length. Limiting function. Its corresponding function The clipping operation introduces non-linearity to make ALF more efficient by reducing the influence of neighboring sample values that differ too much from the current sample value.
[0125] In some examples (e.g., ECM), modifications can be made to the ALF. For instance, ALF gradient subsampling and ALF virtual boundary processing can be removed. In another example, the block size used for classification is reduced from 4×4 to 2×2. Furthermore, the filter size for both luma and chroma, representing the ALF coefficients as signals, is increased to 9×9.
[0126] In some examples, a technique known as ALF with fixed filters is used. To filter the luminance samples, three different classifiers (C0, C1, and C2) and three different filter groups (F0, F1, and F2) are used. Groups F0 and F1 contain fixed filters with coefficients trained on classifiers C0 and C1, respectively. The coefficients of the filters in F2 are represented by signals. For a given sample, the filters from group F1 are used... i Which filter in the classifier C is used by? i The category assigned to the sample Decide.
[0127] In some examples (e.g., ECM8), for classification, the directionality is based on the expression shown in equation (13). and activity Assign a category to each 2×2 block
[0128] Equation (13)
[0129] in, Indicates directionality The total number.
[0130] In some examples, a 1-dimensional Laplacian is used to compute the horizontal, vertical, and two diagonal gradients for each sample. The sum of the gradients of samples within a 4×4 window covering a 2×2 block of the target is used for classifier C0, and the sum of the gradients of samples within a 12×12 window is used for classifiers C1 and C2. The sum of the horizontal, vertical, and diagonal gradients are expressed as follows: , , and Directionality By comparing equation (14) and It is determined by a set of thresholds.
[0131] , Equation (14)
[0132] In VVC, for example, the directionality is derived using thresholds 2 and 4.5. .for and First, calculate the horizontal / vertical edge strength. and diagonal edge strength Using thresholds .exist In the case of edge strength =0; otherwise, In order to make The largest integer. In the case of edge strength =0; otherwise, In order to make The largest integer. In the case where horizontal / vertical edges are dominant, the results are derived using Table 2(a). Otherwise, the diagonal edges dominate, as derived using Table 2(b). .
[0133] Table 1- and arrive mapping
[0134] (a) (b)
[0135]
[0136] In order to obtain The sum of the vertical gradient and the horizontal gradient Mapped to the range 0 to n, where, for n equals 4, and for and n equals 15. In ALF_APS, up to 4 luminance filter banks are represented by signals, and each luminance filter bank can have up to 25 filters.
[0137] In some examples, an alternative 2×2 ALF classifier can be used. Classification in the ALF is further extended using an additional alternative classifier. For a luminance filter bank represented by a signal, a flag indicating whether an alternative classifier is applied is represented by a signal. No geometric transformation is applied to the alternative band-based classifier. When applying a band-based classifier, the sum of the sample values of the 2×2 luminance blocks is first calculated. Then the class index is calculated as follows:
[0138] Equation (15)
[0139] In some examples, for filtering, two fixed 13×13 rhombus-shaped filters, F0 and F1, are first applied to derive two intermediate samples. and Then, apply F2. , And neighboring samples, to derive the filtered samples as follows,
[0140] Equation (16)
[0141] in, Neighboring samples and current sample The difference in amplitude between them, and for The amplitude difference between the current sample and the current sample. Filter coefficients. Represented by signals.
[0142] In some examples (e.g., ECM), for ALF, the online trained filters include four types of filter taps: spatial taps, taps based on reconstruction before DBF, taps based on the extended fixed filter output, and residual taps.
[0143] Figure 7An example of an ALF filter shape using residual samples as additional input is shown. Following the spatial taps positioned in a cross-shape (i.e., taps #0 to #19) are eight offline filtering taps (i.e., taps #20 to #25, taps #28 to #29), three taps based on reconstruction before DBF (i.e., taps #26, #27, and #30), and two residual-based taps (i.e., taps #31 and #32). In this filter shape, the filtered samples are derived as shown in Equation (17):
[0144]
[0145] Equation (17)
[0146] in, Neighboring samples and current sample The difference in the amplitude limit between them For intermediate samples and current samples The difference in amplitude between them, and Neighbor samples before DBF and current sample The difference in amplitude limits between them. These are the nearest residual sample values after amplitude limiting, and These are residual samples filtered by a fixed filter and then clipped. For the residual samples, the fixed filter is reused after SAO for reconstruction training.
[0147] In the adaptive parameter set, a signal is used to indicate whether to use only residual-based taps for ALF, or to use both residual-based taps and taps based on residuals filtered by a fixed filter.
[0148] In some examples, a technique known as a residual-based classifier can be used. For example, a third classifier based on the luminance residual sample values can be used. For each 2×2 luminance block, the sum of the absolute values of the residual samples in the neighboring 8×8 windows is calculated, and the class index is derived as shown in equation (18):
[0149] Equation (18)
[0150] The value of classIdx ranges from 0 to 24, the same as in ECM-8.0. In APS, a signal representation is used for each luminance filter bank in the classifier.
[0151] Both band-based and residual-based classifiers uniformly classify values across the entire dynamic range. However, sampled or residual values may not have an ideally uniform distribution across the dynamic range, leading to coarse classification results and potentially inefficient coding.
[0152] Some aspects of this disclosure provide non-uniform classification techniques for ALF classifiers. In some examples, non-uniform classification can lead to better coding efficiency. In some examples, non-uniform classification means that the number of classes may differ for different value intervals, rather than using the same number of classes for all value intervals and performing uniform classification across the entire dynamic range. For example, the encoder / decoder can compute a classification value associated with a classification unit to apply a filter to the classification unit, which is in the current block. The encoder / decoder can determine a filter class for the classification unit from multiple filter classes distributed according to a non-uniform distribution across the dynamic range based on the classification value. Furthermore, the encoder / decoder can obtain filter coefficients associated with the filter class, generate at least one filtered sample of the classification unit based on the filter coefficients, and reconstruct the current block based on the at least one filtered sample.
[0153] For example, the dynamic range N is divided into k sub-intervals, each sub-interval has an index ranging from 0 to k-1, and the value range (size) of each sub-interval is set to... Each sub-interval has its corresponding number of categories. And the sum of the number of categories in all sub-intervals equals the total number of categories used for the ALF classifier. Classification is performed based on which sub-interval the value to be classified belongs to and the corresponding number of categories within that sub-interval.
[0154] In some examples, the dynamic range N is divided into k sub-intervals, where each sub-interval is associated with its own category, and where the sub-interval division is based on a specified... It is executed in a non-uniform manner. In other words, the value k in this implementation represents the total number of categories (all). (equal to 1), and classification is performed by checking which sub-interval the topic value belongs to.
[0155] In the example, for a given value The number of categories is determined by equation (19).
[0156] Equation (19)
[0157] argmin returns the index of the smallest element.
[0158] In some examples, the dynamic range N is divided into k sub-intervals, where each sub-interval is associated with its own category, and where the sub-interval division is based on a specified... It is executed in a non-uniform manner. These k sub-intervals have their corresponding number of categories. For each subinterval, the range of values... and number of categories Should meet They are not entirely the same.
[0159] In the example, all sub-intervals use the same number of categories, which should be [number of categories]. It should be noted that in this implementation, the number of categories... Some categories can have a number of zero.
[0160] In some examples, the dynamic range N is uniformly divided into k subintervals, each subinterval having a value range of... However, all sub-intervals use a different number of categories, that is, all... They are all different. To perform non-uniform classification, first, for the values to be classified... Export the sub-range index, and then generate the final category index based on the number of categories in the sub-range as follows: 1) By... 1) Generate sub-interval indices for the values; 2) Classify the values based on the corresponding number of categories, as shown in equation (20).
[0161]
[0162] Equation (20)
[0163] According to one aspect of this disclosure, non-uniform classification can be represented by signals using relevant syntax in HLS (Higher-Level Syntax) (VPS (Video parameter set, VPS), SPS (Sequence Parameter S, SPS), PPS (Picture Parameter Set, PPS), APS (Adaptive Parameter Set, APS), picture header, slice header).
[0164] In some examples, the number of subintervals, the range of values for each subinterval, the number of categories for each subinterval, and other relevant parameters all use predefined parameters from both the encoder and decoder. A signal is used to indicate whether non-uniform classification is used. If this signal is resolved to true at the decoder, the corresponding non-uniform classification is applied based on the predefined parameters.
[0165] In some examples, a signal is used to indicate whether non-uniform classification is used. If the signal is true, the remaining non-uniform classification-related syntax is further signaled to the decoder side. Based on the parsed parameters, non-uniform classification is applied at the decoder.
[0166] According to another aspect of this disclosure, non-uniform classification can be used by any classifier that directly uses values for classification. In some examples, the values used for classification are derived from sample values.
[0167] In some examples, the values used for classification are derived from the residual values.
[0168] In some examples, the values used for classification are derived from an N×N window covering an M×M target cell, where N≥M.
[0169] Figure 8 A flowchart outlining a process (800) according to an embodiment of this disclosure is shown. The process (800) can be used in a video decoder. In various embodiments, the process (800) is executed by a processing circuitry system, such as a processing circuitry system that performs the functions of a video decoder (110), a processing circuitry system that performs the functions of a video decoder (210), etc. In some embodiments, the process (800) is implemented as software instructions; therefore, when the processing circuitry system executes the software instructions, the processing circuitry system executes the process (800). The process begins at (S801) and proceeds to (S810).
[0170] At (S810), encoded information for one or more images is received. The encoded information indicates the use of non-uniform classification for the filter.
[0171] At (S820), a classification value associated with the classification unit is calculated to apply a filter to the classification unit, which is in the current block of the current image.
[0172] At (S830), a filter category is determined for a classification unit from multiple filter categories based on the classification value in the dynamic range, and the multiple filter categories are distributed in the dynamic range according to a non-uniform distribution.
[0173] At (S840), the filter coefficients associated with the filter category are determined.
[0174] At (S850), at least one filtered sample of the classification unit is generated based on the filter coefficients.
[0175] At (S860), the current block is reconstructed based on at least one filtered sample.
[0176] In some examples, the dynamic range includes at least a first filter category associated with a first value range in the dynamic range and a second filter category associated with a second value range in the dynamic range, the first value range and the second value range having a difference range size.
[0177] In some examples, the dynamic range includes sub-intervals with indices ranging from 0 to k-1, where k is the number of sub-intervals, and the size of each sub-interval is set to... The sub-interval has a corresponding number of categories. The sum of the corresponding class numbers in the sub-intervals equals the total number of classes in the multiple filter classes. Then, in some examples, the filter category is determined based on a specific sub-interval to which the classification value belongs and the number of specific categories within that sub-interval.
[0178] In some examples, the dynamic range is non-uniformly divided into sub-intervals, each associated with a corresponding filter category among multiple filter categories. The filter category is then determined based on the specific sub-interval to which the classification value belongs.
[0179] In some examples, the dynamic range includes intervals with corresponding non-uniform sizes. The subintervals, and They do not have the same value. In the example, They have the same value.
[0180] In some examples, the dynamic range includes at least a first sub-interval and a second sub-interval with different interval ranges, and the first and second sub-intervals have the same number of categories.
[0181] In some examples, the dynamic range includes at least a first sub-interval and a second sub-interval with the same interval size, the first sub-interval having a first number of categories, the second sub-interval having a second number of categories, and the first number of categories and the second number of categories are different.
[0182] In some examples, the dynamic range includes sub-intervals with the same interval size, and the number of corresponding categories for each sub-interval. They are not the same number.
[0183] In some examples, the sub-interval to which the categorical value belongs is determined, and a specific category within the sub-interval is determined for the categorical value.
[0184] In some examples, decoding is performed on one or more signals for a non-uniform distribution based on one or more high-level syntax elements in the encoded information. One or more high-level syntax elements include one of the following: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, and Slice Header.
[0185] In the example, the number of subintervals, the range size of the subintervals, and the number of classes in the subintervals are predefined parameters for a non-uniform distribution. A flag indicating whether a non-uniform distribution is used is decoded.
[0186] In another example, a flag indicating whether a non-uniform distribution is used is decoded. When the flag indicates the use of a non-uniform distribution, at least one of the following is decoded based on the encoded information: the number of subintervals, the size of the subinterval range, and the number of categories in the subinterval.
[0187] Note that the filter can be any suitable filter based on the classifier. In some examples, the filter is an adaptive loop filter.
[0188] In some examples, the classification value is one of the sample value, the residual value, and the derived value from the window covering the classification unit.
[0189] Note that the sorting unit can be any suitable size. In the example, the sorting unit is a 2×2 block. In another example, the sorting unit is a 4×4 block.
[0190] Then, the process proceeds to (S899) and terminates.
[0191] The process (800) can be adjusted as appropriate. Steps in the process (800) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.
[0192] Figure 9 A flowchart outlining a process (900) according to an embodiment of this disclosure is shown. The process (900) can be used in a video encoder. In various embodiments, the process (900) is executed by a processing circuitry system, such as a processing circuitry system that performs the functions of a video encoder (103), a processing circuitry system that performs the functions of a video encoder (303), etc. In some embodiments, the process (900) is implemented as software instructions; therefore, when the processing circuitry system executes the software instructions, the processing circuitry system executes the process (900). The process begins at (S901) and proceeds to (S910).
[0193] At (S910), the use of non-uniform classification for the adaptive loop filter (ALF) is determined to be applied to the current block in the current image.
[0194] At (S920), a classification value is determined that is associated with the classification unit used to apply ALF, and the classification unit is associated with the current block.
[0195] At (S930), a filter category is selected for a classification unit from multiple filter categories based on the classification value in the dynamic range. The multiple filter categories are distributed in the dynamic range according to a non-uniform distribution. The dynamic range includes at least a first filter category having a first value range in the dynamic range and a second filter category having a second value range in the dynamic range. The first value range and the second value range have a difference range size.
[0196] At (S940), the filter coefficients associated with the filter category are obtained.
[0197] At (S950), at least one filtered sample of the classification unit is generated based on the filter coefficients of ALF.
[0198] In some examples, the dynamic range includes sub-intervals with indices ranging from 0 to k-1, where k is the number of sub-intervals, and the size of each sub-interval is set to... The sub-interval has a corresponding number of categories. The sum of the corresponding class numbers in the sub-intervals equals the total number of classes in the multiple filter classes. In some examples, a specific sub-interval to which a categorical value belongs is determined, and the number of specific categories within that specific sub-interval is determined.
[0199] In some examples, the dynamic range includes intervals with corresponding non-uniform sizes. The subintervals, and They have the same value.
[0200] In some examples, the dynamic range includes sub-intervals with the same interval size, and the number of corresponding categories for each sub-interval. They are not the same number.
[0201] Then, the process proceeds to (S999) and terminates.
[0202] The process (900) can be adjusted as appropriate. Steps in the process (900) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.
[0203] According to one aspect of this disclosure, a method for processing visual media data is provided. In this method, a bitstream of visual media data is processed according to format rules. For example, the bitstream may be a bitstream decoded / encoded using any of the decoding and / or encoding methods described herein. Format rules may specify one or more constraints on the bitstream and / or one or more processes to be performed by the decoder and / or encoder.
[0204] In the example, the bitstream includes encoded information for one or more images, including the current image. The format rules specify that the use of a non-uniform classification method for an adaptive loop filter (ALF) is applied to the current block in the current image; a classification value is determined associated with a classification unit for which the ALF is applied, the classification unit being in the current block; a filter class is determined for the classification unit from multiple filter classes based on the classification value within a dynamic range, the multiple filter classes being distributed according to a non-uniform distribution within the dynamic range, the dynamic range including at least a first filter class corresponding to a first value range within the dynamic range and a second filter class corresponding to a second value range within the dynamic range, the first value range and the second value range having a difference range size; filter coefficients are determined associated with the filter class, and at least one filtered sample of the classification unit is generated based on the filter coefficients of the ALF.
[0205] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 10 A computer system (1000) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0206] Computer software can be coded using any suitable machine code or computer language. Machine code or computer language can be subjected to mechanisms such as assembly, compilation, and linking to create code that includes instructions. These instructions can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.
[0207] The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0208] Figure 10 The components of the computer system (1000) shown are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any one or a combination of the components shown in the exemplary embodiments of the computer system (1000).
[0209] The computer system (1000) may include certain human-machine interface input devices. Such human-machine interface input devices can respond to input made by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, tapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface devices can also be used to capture certain media that are not necessarily directly related to conscious input made by humans, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images acquired from still image capturing devices), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0210] The input human-machine interface device may include one or more of the following (only one of each is depicted): keyboard (1001), mouse (1002), touchpad (1003), touch screen (1010), data glove (not shown), joystick (1005), microphone (1006), scanner (1007), and camera device (1008).
[0211] The computer system (1000) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback via a touchscreen (1010), data gloves (not shown), or joystick (1005), but tactile feedback devices that are not used as input devices may also exist); audio output devices (e.g., speakers (1009), headphones (not depicted)); visual output devices (e.g., screens (1010), including CRT (Cathode Ray Tube) screens, LCD (Liquid Crystal Display) screens, plasma screens, OLED (Organic Light Emitting Diode) screens, each screen may or may not have touchscreen input capability, each screen may or may not have tactile feedback capability—some of the screens may be able to output two-dimensional visual output or more than three-dimensional output in a manner such as stereoscopic image output; virtual reality glasses (not depicted); holographic displays and smoke generators (not depicted)); and printers (not depicted).
[0212] The computer system (1000) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM (Read Only Memory, ROM) / RW (1020) having media such as CD / DVD (1021), thumb drives (1022), removable hard disk drives or solid-state drives (1023), conventional magnetic media such as magnetic tape and floppy disks (not depicted), devices based on dedicated ROM / ASIC (Application Specific Integrated Circuit, ASIC) / PLD (Programable Logic Device, PLD) such as security dongles (not depicted), etc.
[0213] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0214] The computer system (1000) may also include interfaces (1054) to one or more communication networks (1055). The networks may be, for example, wireless, wired, or optical. The networks may also be local area, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of networks include: local area networks (LANs) such as Ethernet; wireless LANs; cellular networks including GSM (Global System for Mobile Communications), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; cable or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus (Controller Area Network Bus), etc. Some networks typically require external network interface adapters that attach to certain general-purpose data ports or peripheral buses (1049) (such as, for example, the USB (Universal Serial Bus, USB) port of the computer system (1000); other networks are typically integrated into the core of the computer system (1000) via system buses that attach to, for example, an Ethernet interface to a PC (Personal Computer, PC) computer system or a cellular network interface to a smartphone computer system. Using any of these networks, the computer system (1000) can communicate with other entities. Such communication can be one-way receiving (e.g., broadcasting TV), one-way transmitting (e.g., to a CANBus device), or bidirectional, such as to other computer systems using local or wide-area digital networks. Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.
[0215] The human-machine interface device, human-accessible storage device and network interface mentioned above can be attached to the core (1040) of the computer system (1000).
[0216] The core (1040) may include one or more central processing units (CPU) (1041), graphics processing units (GPUs) (1042), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1043), hardware accelerators (1044) for certain tasks, graphics adapters (1050), etc. These devices, along with read-only memory (ROM) (1045), random access memory (1046), internal mass storage devices such as internal non-user-accessible hard disk drives, SSDs (SSDs), etc. (1047), may be connected via the system bus (1048). In some computer systems, the system bus (1048) may be accessed in the form of one or more physical plugs to allow for expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (1048) or may be attached to the core's system bus (1048) via a peripheral bus (1049). In the example, the screen (1010) can be connected to the graphics adapter (1050). The peripheral bus architecture includes PCI (Peripheral Component Interconnect), USB, etc.
[0217] The CPU (1041), GPU (1042), FPGA (1043), and accelerator (1044) can execute certain instructions, which, when combined, constitute the computer code mentioned above. This computer code can be stored in ROM (1045) or RAM (1046). Transient data can also be stored in RAM (1046), while permanent data can be stored, for example, in an internal mass storage device (1047). Fast storage and retrieval to any memory device in the memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1041), GPUs (1042), mass storage devices (1047), ROMs (1045), RAMs (1046), etc.
[0218] Computer-readable media may have computer code thereon for performing operations of various computer implementations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a type known and available to those skilled in the art of computer software.
[0219] By way of example and not limitation, a computer system (1000) having an architecture, and particularly a core (1040), can provide functionality by having a processor (including a CPU, GPU, FPGA, accelerator, etc.) execute software embodied in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with a user-accessible mass storage device as described above, and with certain non-transitory storage devices of the core (1040), such as a mass storage device (1047) or ROM (1045) within the core. Software implementing various embodiments of this disclosure can be stored in such a device and executed by the core (1040). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the core (1040), and particularly the processor therein (including a CPU, GPU, FPGA, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (1046) and modifying such data structures according to the processes defined by the software. Alternatively or as an alternative, the computer system may be provided with functionality by means of hard-wired logic or otherwise embodied in circuitry (e.g., an accelerator (1044)), which may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and references to logic may also include software. Where appropriate, references to a computer-readable medium may include circuitry storing software for execution (e.g., an integrated circuit (IC)), circuitry implementing logic for execution, or both. This disclosure includes any suitable combination of hardware and software.
[0220] The use of “at least one of…” or “one of…” in this disclosure is intended to include any one or a combination of the elements described. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A to C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include either A or B or (A and B). The use of “one of…” does not exclude any combination of the elements described where applicable, such as when the elements are not mutually exclusive.
[0221] While this disclosure has described several exemplary embodiments, variations, substitutions, and various alternative equivalents fall within the scope of this disclosure. Therefore, it will be appreciated that those skilled in the art will be able to conceive of many systems and methods that, while not expressly shown or described herein, implement the principles of this disclosure and are thus within its spirit and scope.
Claims
1. A method for processing visual media data, the method comprising: The bitstream of visual media data is processed according to format rules, where: The bitstream includes encoded information for one or more images, including the current image; and The formatting rules specify: The use of non-uniform classification for the Adaptive Loop Filter (ALF) is determined to be applied to the current block in the current image; Determine a classification value associated with a classification unit used to apply the adaptive loop filter, the classification unit being in the current block; In a dynamic range, a filter category is determined for the classification unit from multiple filter categories based on the classification value. The multiple filter categories are distributed in the dynamic range according to a non-uniform distribution. The dynamic range includes at least a first filter category corresponding to a first value range in the dynamic range and a second filter category corresponding to a second value range in the dynamic range. The first value range and the second value range have a difference range size. Determine the filter coefficients associated with the filter category; and At least one filtered sample for the classification unit is generated based on the filter coefficients of the adaptive loop filter.
2. A video decoding device, the device comprising a processing circuit system configured to perform the following operations: Receive encoded information of one or more images, the encoded information indicating the use of non-uniform classification for the filter; Calculate a classification value associated with a classification unit to apply the filter to the classification unit, which is in the current block of the current image; In the dynamic range, a filter category is determined for the classification unit from multiple filter categories based on the classification value, the multiple filter categories being distributed in the dynamic range according to a non-uniform distribution; Obtain the filter coefficients associated with the filter category determined for the classification unit; At least one filtered sample for the classification unit is generated based on the filter coefficients; as well as The current block is reconstructed based on the at least one filtered sample.
3. The device according to claim 2, wherein, The dynamic range includes at least a first filter category corresponding to a first value range in the dynamic range and a second filter category corresponding to a second value range in the dynamic range, wherein the first value range and the second value range have a difference range size.
4. The device according to any one of claims 2 to 3, wherein, The dynamic range includes sub-intervals with an index range from 0 to k-1, where k is the number of sub-intervals, and the size of each sub-interval is set to... The sub-interval has a corresponding number of categories. The sum of the corresponding category numbers in the sub-intervals is equal to the total number of categories in the plurality of filter categories. And the processing circuit system is configured to: The filter category is determined based on the specific sub-interval to which the classification value belongs and the number of specific categories within the specific sub-interval.
5. The device according to claim 4, wherein, The dynamic range is non-uniformly divided into sub-intervals, each sub-interval being associated with one of the plurality of filter categories, and the processing circuitry is configured to: The filter category is determined based on the specific sub-interval to which the classification value belongs.
6. The device according to claim 4, wherein, The dynamic range includes regions with non-uniform interval sizes. The subintervals, and They do not have the same value.
7. The device according to claim 6, wherein, They have the same value.
8. The device according to claim 4, wherein, The dynamic range includes at least a first sub-range and a second sub-range with different ranges, and the first sub-range and the second sub-range have the same number of categories.
9. The device according to claim 4, wherein, The dynamic range includes at least a first sub-interval and a second sub-interval with the same interval size, the first sub-interval having a first number of categories, the second sub-interval having a second number of categories, and the first number of categories being different from the second number of categories.
10. The device according to claim 4, wherein, The dynamic range includes sub-intervals with the same interval size, and the number of corresponding categories of the sub-intervals. They are not the same number.
11. The device according to claim 10, wherein, The processing circuit system is configured to: Determine the specific sub-interval to which the classification value belongs; and A specific category is determined within the specific sub-interval based on the classification value.
12. The device according to claim 4, wherein, The processing circuit system is configured to: Decoding is performed on one or more signals for the non-uniform distribution based on one or more high-level syntax elements in the encoded information, wherein the one or more high-level syntax elements are in the Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, and Slice Header.
13. The device according to claim 12, wherein, The number of sub-intervals, the size of the sub-intervals, and the number of categories in the sub-intervals are predefined parameters for the non-uniform distribution, and the processing circuit system is configured to: The flag indicating whether to use the non-uniform distribution is decoded based on the encoded information.
14. The device according to claim 12, wherein, The processing circuit system is configured to: Decoding the flag indicating whether to use the non-uniformly distributed flag based on the encoded information; and When the flag indicates the use of the non-uniform distribution, at least one of the number of sub-intervals, the range size of the sub-intervals, and the number of categories in the sub-intervals is decoded according to the encoding information.
15. A video coding method, comprising: Determine the use of non-uniform classification for the Adaptive Loop Filter (ALF) to apply to the current patch in the current image; Calculate the classification value associated with the classification unit used to apply the adaptive loop filter in the current block; In a dynamic range, a filter category is determined for the classification unit from multiple filter categories based on the classification value. The multiple filter categories are distributed in the dynamic range according to a non-uniform distribution. The dynamic range includes at least a first filter category corresponding to a first value range in the dynamic range and a second filter category corresponding to a second value range in the dynamic range. The first value range and the second value range have a difference range size. Obtain the filter coefficients associated with the filter category determined for the classification unit; as well as At least one filtered sample for the classification unit is generated based on the filter coefficients of the adaptive loop filter.