CCSO with offset-specific options
Through the cross-component offset filtering method, the chroma sample offset value is determined using the isometric and adjacent reconstruction samples of the brightness sample, which solves the video quality degradation problem caused by the brightness and chroma sample offset processing in the prior art, and achieves more efficient video encoding and reduces reconstruction errors.
Patent Information
- Application Number
- CN202480005442.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-09
- Filing Date
- 2024-05-20
- Publication Date
- 2025-07-18
AI Technical Summary
Existing video encoding and decoding technologies are difficult to effectively reduce reconstruction errors during compression, especially in the offset processing between brightness and chromaticity samples, resulting in video quality degradation.
The cross-component offset filtering method is used to determine the offset value of the chromaticity sample through a loop filter using the isometric reconstruction sample and adjacent reconstruction sample of the luminance sample. The precise reconstruction of the video frame is achieved independently of the luminance gradient.
Improve the quality of video reconstruction, reduce reconstruction errors, and enhance video encoding efficiency and compression ratio.
Smart Images

Figure CN120345244A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims the priority of U.S. Provisional Patent Application No. 63 / 544,410, entitled "CCSO with Band Offset Only Option", filed on October 16, 2023, and is a continuation - in - part of and claims the priority of U.S. Patent Application No. 18 / 660,058, entitled "CCSO with Band Offset Only Option", filed on May 9, 2024. All of these applications are hereby incorporated by reference in their entirety. Technical Field
[0002] The disclosed embodiments generally relate to video coding and decoding, including but not limited to systems and methods for loop filtering (e.g., cross - component offset filtering) of video data. Background Art
[0003] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices transmit and receive digital video data over a communication network or otherwise convey digital video data and / or store digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited memory resources of the storage device, video coding can be used to compress video data according to at least one video coding standard before transmitting or storing the video data. Video coding and decoding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services.
[0004] Video coding and decoding typically utilize prediction methods that exploit the redundancy inherent in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video coding and decoding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4). Versatile Video Coding (VVC / H.266) is a video compression standard designed to be a successor to HEVC. The ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (Version 1) and 2022 (Version 2). AOMedia Video 1 (AV1) is an open video coding and decoding format designed as an alternative to HEVC. On January 8, 2019, the verified Version 1.0.0 of the specification with Errata 1 was released. Summary of the Invention
[0005] As described above, encoding (compression) reduces the bandwidth and / or storage space requirements. As will be described in detail later, lossless compression and lossy compression can be employed. Lossless compression refers to a technique in which an exact copy of the original signal is reconstructed from the compressed original signal through the decoding process. Lossy compression refers to a process in which the original video information is not fully retained during the encoding process and not fully restored during the decoding process. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is reduced to a level sufficient for the reconstructed signal to be useful for the intended application. The amount of tolerable distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows for coding algorithms that produce higher losses and higher compression ratios.
[0006] The present disclosure describes methods, systems, and non-transitory computer-readable storage media for applying a loop filter for video (image) compression. A video codec includes multiple functional modules for at least one of intra / inter prediction, transform coding, quantization, entropy coding, and in-loop filtering. In-loop filtering techniques are used to adjust reconstructed image samples to further reduce reconstruction errors. A cross-component offset filtering method is implemented to apply co-located reconstructed samples of a first color component and associated adjacent reconstructed samples to derive an offset value to be added to a current sample of a second color component, thereby adjusting the reconstructed value of the current sample. An example of the first color component is a luminance color component, and an example of the second color component is a chrominance color component. In some implementations, the first color component and the second color component correspond to the same color component, such as luminance samples.
[0007] In various embodiments of the present application, samples of the first color component are processed by a cross-component offset filter in the loop filter to determine an offset value to be added to samples of the second color component. The cross-component offset filtering is implemented based on an edge-preserving loop filter that uses reconstructed chrominance samples to determine sample offsets for luminance and / or chrominance components. For example, a sample offset is determined based on luminance values of a first luminance sample and at least one adjacent luminance sample, e.g., independent of any sideband offset corresponding to a gradient between the first luminance sample and the associated adjacent luminance samples.
[0008] According to some embodiments, a video decoding method is provided. The method includes: receiving a video bitstream including a current image frame, where the video bitstream includes a first syntax element that indicates whether a first sample offset of a first chrominance sample is determined based on values of at least one luminance sample, e.g., independent of any associated luminance gradient of the at least one luminance sample. The method further includes: determining that the first syntax element has a first predefined value indicating an enabled offset-specific mode; based on the offset-specific mode, generating at least one quantization value based on at least one luminance sample including a first luminance sample co-located with the first chrominance sample in the current image frame; classifying the first chrominance sample of the current image frame based on the at least one quantization value to determine the first sample offset of the first chrominance sample; and reconstructing the current image frame at least by adjusting the first chrominance sample based on the first sample offset of the first chrominance sample.
[0009] According to some embodiments, a video encoding method is provided. The method includes receiving video data including a current picture frame and a first syntax element; encoding the current picture frame; determining, based on the first syntax element, that the band offset special mode is enabled to determine a first sample offset of a first chroma sample based on values of at least one luma sample, e.g., independent of any associated luma gradient of the at least one luma sample; transmitting the encoded current picture frame via a video bitstream; and signaling the first syntax element via the video bitstream to indicate that the band offset special mode is applied to reconstruct the first chroma sample juxtaposed with the first luma sample based on the first sample offset.
[0010] According to some embodiments, a bitstream conversion method is provided. The method includes obtaining a source video sequence including a current picture frame and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes the current picture frame and a first syntax element that indicates whether to determine a first sample offset of a first chroma sample based on values of at least one luma sample, e.g., independent of any associated luma gradient of the at least one luma sample. Based on the band offset special mode, a first sample offset of the first chroma sample is determined based on at least one luma sample, and the at least one luma sample includes a first luma sample collocated with the first chroma sample in the current picture frame.
[0011] According to some embodiments, a computing system, such as a streaming media system, a server system, a personal computer system, or other electronic device, is provided. The computing system includes a control circuit and a memory storing at least one set of instructions. The at least one set of instructions includes instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.
[0012] According to some embodiments, a non - volatile computer - readable storage medium is provided. The non - volatile computer - readable storage medium stores at least one set of instructions for execution by a computing system. The at least one set of instructions includes instructions for performing any of the methods described herein.
[0013] Thus, apparatuses and systems having methods for video encoding and decoding are disclosed. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, or systems for video encoding.
[0014] The features and advantages described in the specification are not necessarily all inclusive. In particular, in view of the figures, specification, and claims provided in the present disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Further, it should be noted that the language used in the specification has been principally selected for readability and instructional purposes and not necessarily to describe or limit the subject matter hereof. Description of the Drawings
[0015] For a more detailed understanding of the present disclosure, a more specific description can be made by referring to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the accompanying drawings only illustrate the relevant features of the present disclosure and are therefore not necessarily considered restrictive, as those skilled in the art will understand that there may be other valid features in this specification after reading the present disclosure.
[0016] Figure 1 FIG. is a block diagram of an exemplary communication system according to some embodiments.
[0017] Figure 2A FIG. is a block diagram of exemplary elements of an encoder component according to some embodiments.
[0018] Figure 2B FIG. is a block diagram of exemplary elements of a decoder component according to some embodiments.
[0019] Figure 3 FIG. is a block diagram of an exemplary server system according to some embodiments.
[0020] Figure 4 FIG. is a flowchart of an exemplary process of applying a band-offset dedicated mode in intra-loop filtering according to some embodiments.
[0021] Figure 5 FIG. is a flowchart of an exemplary process of applying cross-component sample offset in intra-loop filtering according to some embodiments.
[0022] Figure 6 FIG. is a histogram associated with multiple luminance samples of a current image frame according to some embodiments.
[0023] Figure 7 FIG. is a flowchart of a method for encoding and decoding video according to some embodiments.
[0024] According to conventional practice, the various features illustrated in the accompanying drawings need not be drawn to scale, and throughout the specification and the drawings, like reference numerals may be used to represent like features. DETAILED DESCRIPTION
[0025] The present disclosure describes methods, systems, and non-transitory computer-readable storage media for applying a loop filter to video (image) compression. Intra-loop filtering techniques are applied to adjust reconstructed picture samples to further reduce reconstruction errors. A cross-component offset filtering method is implemented to apply co-located reconstructed samples of a first color component and associated neighboring reconstructed samples to derive an offset value that is added to a current sample of a second color component, thereby adjusting the reconstructed value of the current sample. In various embodiments of the present application, a decoder receives a video bitstream from an encoder, the video bitstream including a current image frame and a first syntax element that indicates whether a first sample offset of a first chrominance sample is determined based on values of at least one luma sample. For example, independent of any associated luma gradient of the at least one luma sample. Sample values of the first color component (e.g., unassociated gradient values) are used in cross-component offset filtering to determine the offset value added to samples of the second color component. For example, luma samples are applied to generate an offset value for a first luma sample or a first chrominance sample co-located with the first luma sample, without involving gradient values of the luma samples.
[0026] More specifically, in some embodiments, a video decoder identifies a set of luma samples, e.g., based on a filter shape, the set of luma samples including a first luma sample and at least one neighboring luma sample of the first luma sample. The decoder does not determine a difference between the at least one neighboring luma sample and the first luma sample. For example, a scalar quantizer is used to quantize the luma samples to generate at least one quantization value, without involving any difference associated with the luma samples. The scalar quantizer can be specified by a quantization interval (e.g., a range of values assigned to the same integer) and a quantization level (e.g., the integer value to which the quantization interval is assigned). For example, a classifier classifies the first chrominance sample based on the at least one quantization value to determine a first sample offset of the first chrominance sample. The first chrominance sample is adjusted based on the first sample offset of the first chrominance sample, thereby achieving reconstruction of the current image frame.
[0027] Figure 1 A block diagram of a communication system 100 in accordance with some embodiments is shown. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via at least one network. In some embodiments, the communication system 100 is a streaming system, e.g., used with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0028] The source device 102 includes a video source 104 (e.g., a camera component or a media memory) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates at least one encoded video bitstream from the video bitstream. The video bitstream from the video source 104 can be of high data volume compared to the encoded video bitstream 108 generated by the encoder component 106. Since the encoded video bitstream 108 has a lower data volume (less data) compared to the video bitstream from the video source, the encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared to the video bitstream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video to the network 110).
[0029] The at least one network 110 represents any number of networks for transmitting information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wired (wired) and / or wireless communication networks. The at least one network 110 can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0030] The at least one network 110 includes a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as the encoded video bitstream from the source device 102). The server system 112 includes a codec component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec component 114 includes an encoder component and / or a decoder component. In various embodiments, the codec component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the codec component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video bitstream 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 can be configured to trim the encoded video bitstream 108 to customize potentially different bitstreams for at least one of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.
[0031] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video bitstream that can be presented on a display or other type of rendering device. In some embodiments, at least one electronic device 120 does not include a display component (e.g., is communicatively coupled to an external display device and / or includes a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0032] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or at least one electronic device 120 are examples of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.
[0033] In an example operation of the communication system 100, the source device 102 transmits the encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may decode and / or encode the encoded video bitstream 108 using the codec component 114. For example, the server system 112 may apply an encoding that is more suitable for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., at least one encoded video bitstream) to at least one of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.
[0034] Figure 2AA block diagram showing example elements of an encoder component 106 in accordance with some embodiments. The encoder component 106 receives video data such as a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that, when viewed in sequence, produce motion. The pictures themselves may be organized as a spatial array of pixels, where each pixel may include at least one sample depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can readily understand the relationship between pixels and samples.
[0035] The encoder component 106 is configured to encode and / or compress pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, the encoder component 106 is configured to perform a conversion between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Implementing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ (lambda) value of rate distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as they may relate to the encoder component 106 optimized for a particular system design.
[0036] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and at least one reference picture), and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since decoding the symbol stream results in a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the same sample values as the sample values that the decoder will interpret when using prediction during decoding as reference picture samples. This principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is known to those of ordinary skill in the art.
[0037] The operation of the decoder 210 can be the same as the operation of a remote decoder (such as the decoder component 122), which will be described in detail below in conjunction with Figure 2B For a brief reference to Figure 2B , since the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols into an encoded video sequence losslessly, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.
[0038] Except for parsing / entropy decoding, the decoder techniques described herein can exist in a corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder techniques can be brief as they can be the opposite of decoder techniques.
[0039] As part of its operation, the source encoder 202 can perform motion-compensated predictive coding, where the source encoder predictively encodes an input frame by referring to at least one previously encoded frame designated as a reference frame from a video sequence. In this way, the coding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of at least one reference frame that can be selected as at least one prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and sub-group parameters for encoding video data.
[0040] Decoder 210 decodes the encoded video data of a frame that can be designated as a reference frame based on symbols created by source encoder 202. The operation of encoding engine 212 can advantageously be a lossy process. When the encoded video data is decoded at the video decoder ( Figure 2A not shown herein), the reconstructed video sequence can be a copy of the source video sequence with some errors. Decoder 210 replicates the decoding process that can be performed by a remote video decoder on the reference frame and can cause the reconstructed reference frame to be stored in reference picture memory 208. In this way, encoder component 106 locally stores copies of the reconstructed reference frames that have the same content as the reconstructed reference frames that would be obtained by a remote video decoder (in the absence of transmission errors).
[0041] Predictor 206 can perform a prediction search on encoding engine 212. That is, for a new frame to be encoded, predictor 206 can search reference picture memory 208 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. Predictor 206 can operate on a per-pixel-block basis of sample blocks to find an appropriate prediction reference. As determined by the search results obtained by predictor 206, the input picture can have a prediction reference extracted from multiple reference pictures stored in reference picture memory 208.
[0042] The outputs of all the above functional units can be entropy encoded in entropy encoder 214. Entropy encoder 214 converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).
[0043] In some embodiments, the output of entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer at least one encoded video sequence created by entropy encoder 214 to make them ready for transmission via communication channel 218, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter can be configured to merge the encoded video data from source encoder 202 with other data to be transmitted (e.g., encoded audio data and / or an auxiliary data stream (not shown in the source)). In some embodiments, the transmitter can transmit additional data along with the encoded video. Source encoder 202 can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data (such as redundant pictures and slices), supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.
[0044] The controller 204 can manage the operation of the encoder component 106. During encoding, the controller 204 can assign a specific encoded picture type to each encoded picture, which can affect the encoding techniques applied to the corresponding picture. For example, a picture can be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). Intra pictures can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example independent decoder refresh (IDR) pictures. Those of ordinary skill in the art are aware of those variants of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. Predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures can be used to reconstruct a single block using more than two reference pictures and associated metadata.
[0045] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and encoded on a block-by-block basis. Blocks can be encoded predictively with reference to other (already encoded) blocks determined by the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture can be encoded non-predictively, or they can be encoded predictively with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be encoded non-predictively with reference to one previously encoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture can be encoded non-predictively with reference to one or two previously encoded reference pictures via spatial prediction or via temporal prediction.
[0046] Video can be captured as multiple source pictures (video pictures) in a time series. Intra picture prediction (commonly abbreviated as intra prediction) exploits the spatial correlation within a given picture, and inter picture prediction exploits the (temporal or other) correlation between pictures. In an example, a specific picture in encoding / decoding (which is referred to as the current picture) is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously encoded and still buffered reference picture in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, can have a third dimension identifying the reference picture.
[0047] The encoder component 106 may perform encoding operations according to a predetermined video coding technique or standard, such as any of the techniques or standards described herein. In its operation, the encoder component 106 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video coding technique or standard being used.
[0048] Figure 2B A block diagram showing example elements of a decoder component 122 according to some embodiments is shown. Figure 2B The decoder component 122 in is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to a loop filter 256 and is configured to send data to the display 124 (e.g., via a wired or wireless connection).
[0049] In some embodiments, the decoder component 122 includes a receiver that is coupled to the channel 218 and is configured to receive data from the channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive at least one encoded video sequence to be decoded by the decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from the channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not shown). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data with the encoded video. The additional data may be included as part of at least one encoded video sequence. The decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0050] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The decoder component 122 may be implemented at least partially in software.
[0051] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to counter network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 inside the decoder component 122 (e.g., the buffer memory 252 is configured to handle playout timing), a separate buffer memory is also provided outside the decoder component 122 (e.g., to counter network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory 252 may not be required, or the buffer memory 252 can be small. For use on a best-effort packet network such as the Internet, the buffer memory 252 may be necessary, the buffer memory 252 can be relatively large and / or can have an adaptive size, and can be implemented at least partially in an operating system or a similar element outside the decoder component 122.
[0052] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols can include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device can be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not shown). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technology or standard, and can follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the set. Subgroups can include group of pictures (GOP), pictures, tiles, stripes, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The parser 254 can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0053] Depending on the type of the encoded video picture or a part thereof (such as: inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols 270 may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information, which is parsed by the parser 254 from the encoded video sequence. For clarity, this subgroup control information flow between the parser 254 and the multiple units is not depicted below.
[0054] The decoder component 122 can be conceptually divided into multiple functional units. In some embodiments, many of these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, the conceptual division of the functional units is maintained here.
[0055] The scaler / inverse transform unit 258 receives the quantized transform coefficients and control information (such as which transform to use, block size, quantization factor, and / or quantization scaling matrix) as at least one symbol 270 from the parser 254. The scaler / inverse transform unit 258 can output a block including sample values, and these sample values can be input into the aggregator 268.
[0056] In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed pictures but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by the intra prediction unit 262. The intra prediction unit 262 can use the surrounding reconstructed information obtained from the current (partially reconstructed) picture in the current picture memory 264 to generate a block having the same size and shape as the block being reconstructed. The aggregator 268 can, on a per-sample basis, add the predictive information already generated by the intra prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258.
[0057] In other cases, the output samples of the scaler / inverse transform unit 258 belong to inter-coded and possibly motion-compensated blocks. In such cases, the motion compensation prediction unit 260 can access the reference picture memory 266 to obtain samples for prediction. After motion-compensating the obtained samples according to the symbol 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (which is called the residual sample or residual signal in this case) to generate the output sample information. The address in the reference picture memory 266 from which the motion compensation prediction unit 260 obtains the prediction samples can be controlled by the motion vector. The motion vector can be provided to the motion compensation prediction unit 260 in the form of a symbol 270, and the symbol 270 can have, for example, an X component, a Y component, and a reference picture component. Motion compensation can also include interpolation of the sample values obtained from the reference picture memory 266 when using sub-sampled accurate motion vectors, motion vector prediction mechanisms, etc.
[0058] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and provided as symbols 270 from parser 254 to loop filter unit 256, but can also respond to meta-information obtained during the decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to the sample values of previously reconstructed and loop-filtered samples. The output of loop filter unit 256 can be a sample stream that can be output to a rendering device such as display 124, and stored in reference picture memory 266 for use in future inter-frame prediction.
[0059] Once reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is reconstructed and the encoded picture has been identified (e.g., by parser 254) as a reference picture, the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting the reconstruction of subsequent encoded pictures.
[0060] Decoder component 122 can perform decoding operations according to predetermined video compression techniques that can be documented in standards such as any of the standards described herein. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that it follows the syntax of the video compression technique or standard as specified in the video compression technique documentation or standard, particularly in the profile documentation thereof. Additionally, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, and so on. In some cases, the limits set by the level can be further restricted by the hypothetical reference decoder (HRD) specifications and metadata signaled in the encoded video sequence for HRD buffer management.
[0061] Figure 3 A block diagram of server system 112 according to some embodiments is shown. Server system 112 includes control circuit 302, at least one network interface 304, memory 314, user interface 306, and at least one communication bus 312 for interconnecting these components. In some embodiments, control circuit 302 includes at least one processor (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes at least one field programmable gate array (FPGA), hardware accelerator, and / or at least one integrated circuit (e.g., application specific integrated circuit).
[0062] The network interface 304 can be configured to interface with at least one communication network (e.g., wireless, wired, and / or optical networks). The communication network can be a local area network, wide area network, metropolitan area network, vehicular network, industrial network, real-time network, delay-tolerant network, etc. Examples of communication networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANBus, etc. Such communication can be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., to other computer systems using local or wide area digital networks). Such communication can include communication with at least one cloud computing network.
[0063] The user interface 306 includes at least one output device 308 and / or at least one input device 310. The at least one input device 310 can include one or more of the following: keyboard, mouse, touchpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. The at least one output device 308 can include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or monitors), etc.
[0064] The memory 314 can include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as at least one disk storage device, optical disk storage device, flash device, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes at least one storage device located remotely from the control circuit 302. The memory 314 or optionally at least one non-volatile solid-state memory device within the memory 314 includes non-volatile computer-readable storage media. In some embodiments, the memory 314 or the non-volatile computer-readable storage media of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: ● An operating system 316 that includes processes for handling various basic system services and for performing hardware-related tasks; ● A network communication module 318 that is used to connect the server system 112 to other computing devices via at least one network interface 304 (e.g., via wired and / or wireless connections); ● A codec module 320 that is used to perform various functions related to encoding and / or decoding data (such as video data). In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, at least one of the following: o A decoding module 322 that is configured to perform various functions related to decoding encoded data, such as those functions described previously with respect to the decoder component 122; and o An encoding module 340 that is configured to perform various functions related to encoding data, such as those functions described previously with respect to the encoder component 106; and ● A picture memory 352 that is configured to store pictures and picture data, e.g., for use by the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.
[0065] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions described previously with respect to the parser 254), a transform module 326 (e.g., configured to perform various functions described previously with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions described previously with respect to the motion compensation prediction unit 260 and / or the intra prediction unit 262), and a filter module 330 (e.g., configured to perform various functions described previously with respect to the loop filter 256).
[0066] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions described previously with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions described previously with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 a subset of the modules shown. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.
[0067] Each of the above modules stored in the memory 314 corresponds to a set of instructions for performing the functions described herein. The above modules (e.g., instruction sets) need not be implemented as separate software programs, processes, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the above modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.
[0068] Although Figure 3 FIG. illustrates a server system 112 according to some embodiments,Figure 3 It is more a functional description of various features that may exist in at least one server system than a structural schematic diagram of the embodiments described herein. In practice, as will be recognized by those of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in may be implemented on a single server, and a single item may be implemented by at least one server. The actual number of servers used to implement the server system 112, and how the features are distributed among them will vary depending on the implementation and, optionally, in part on the amount of data traffic processed by the server system during peak usage and during average usage periods.
[0069] Figure 4 is a flowchart of an example process 400 for applying a band-offset specific mode 440 in in-loop filtering according to some embodiments. A GOP includes a sequence of image frames, the sequence of image frames further including a current image frame. The current image frame includes a color image, i.e., a non-monochrome image frame, having a plurality of chrominance samples (e.g., chrominance sample 402 and luminance sample 404) located at the same position with respect to each other. After reconstructing the plurality of chrominance samples of the current image frame, in-loop filtering is applied to adjust a subset of the chrominance samples so as to improve the image quality of the current image frame. In some embodiments, the reconstructed samples of a first color component and its adjacent reconstructed samples are combined to derive an offset value for a second color component, and the reconstructed samples of the second color component are located at the same position as the reconstructed samples of the first color component and are adjusted by the offset value. Additionally, in some embodiments, the reconstructed samples of the first color component are used in the band-offset specific mode 440 to derive an offset value for the second color component (e.g., luminance sample 404, chrominance sample 402). The samples of the second color component are adjusted by the offset value. The first color component is optionally the same as or different from the second color component.
[0070] A first syntax element 408 is used to define the band-offset specific mode 440, which indicates whether a first sample offset 406 of a first chrominance sample 410 is determined based on the luminance sample 404 independently of any associated luminance gradient of the luminance sample 404. In the band-offset specific mode 440, the values of the luminance samples 404 (e.g., a first luminance sample 404C and at least one adjacent luminance sample 404X) may be quantized and classified to derive the sample offset 406, which is applied to adjust the first luminance sample 404C itself. Alternatively, the first luminance sample 404C and the adjacent luminance sample 404X may be used to derive the sample offset 406, and the first chrominance sample 402C is located at the same position as the first luminance sample 404AC and may be adjusted by the sample offset 406.
[0071] More specifically, decoder 122 receives video bitstream 116 including a current image frame and a first syntax element 408 for the band offset dedicated mode 440. The first syntax element 408 indicates whether to determine a first sample offset 406 of a first chrominance sample 410 based on the value of at least one luma sample 404, regardless of any associated luma gradient of the at least one luma sample 404. The first syntax element 408 has a first predefined value (e.g., "1") indicating that the band offset dedicated mode 440 is enabled. Based on the band offset dedicated mode 440, decoder 122 generates at least one quantization value 404Q based on at least one luma sample 404, where the at least one luma sample includes a first luma sample 404C co-located with the first chrominance sample 410 in the current image frame. The first chrominance sample 410 is classified based on the at least one quantization value 404Q to determine the first sample offset 406 of the first chrominance sample 410. Decoder 122 reconstructs the current image frame at least by adjusting the first chrominance sample 410 based on the first sample offset 406.
[0072] In some embodiments, the first chrominance sample 410 is one of the following: a first luma sample 404C, a first blue-difference chrominance (Cb) sample 402Cr, and a first blue-difference chrominance (Cb) sample 402Cb. The first luma sample 404C, the first Cb sample 402Cb, and the first Cr sample 402Cr are co-located with each other. In some embodiments, video bitstream 116 includes a first syntax element for the cross-component sample offset (CCSO) mode 442. The CCSO mode 442 indicates whether to determine a first sample offset 406 of the first chrominance sample 410 based on at least one luma sample 404. The first syntax element has a first predefined value indicating that the CCSO mode 442 is enabled.
[0073] Video bitstream 116 includes an alternative image frame and an associated first syntax element 408 for the band offset dedicated mode 440. Decoder 122 determines that the associated first syntax element 408 has a second predefined value (e.g., "0") indicating that the band offset dedicated mode 440 is disabled for the alternative image frame. Based on band offset classification and sideband offset classification, a second sample offset of a second chrominance sample co-located with a second luma sample in the alternative image frame. For the alternative image frame, a set of luma samples and the associated differences of the luma samples are quantized and applied to band offset classification and sideband offset classification, respectively. The band offset classification result and the sideband offset classification result are jointly used to determine the second sample offset.
[0074] In some embodiments, a first syntax element 408 is signaled at the frame level for a first color component corresponding to a first chroma sample 410 (e.g., a luma sample 404, a chroma sample 402) individually. Alternatively, in some embodiments, the first syntax element 408 is a flag signaled at the frame level for multiple color components of a current image frame, and the multiple color components include the first color component corresponding to the first chroma sample 410. Signaling the flag jointly controls the band offset classification of the luma sample 404 and the chroma sample 402. The first chroma sample 410 is one of a first luma sample 404C, a first Cr sample 402Cr, and a first Cb sample 402Cb.
[0075] Alternatively, in some embodiments, a current image frame includes multiple color components, which further includes the first color component corresponding to the first chroma sample 410, and the first syntax element 408 is signaled at the frame level and includes multiple classification syntax flags, each of the multiple classification syntax flags indicating whether a band offset dedicated mode is enabled for a corresponding color component. The first classification syntax flag can be signaled for the luma sample 404. A second first syntax element can be signaled for the Cb sample 402Cb, and a third classification syntax flag can be signaled for the Cr sample 402Cr. In an example, each sample of the current image frame includes a luma component 404, a blue-difference chroma (Cb) component 402Cb, and a red-difference chroma (Cr) component 402Cr. The first chroma sample 410 corresponds to one of the luma component 404, the Cr component 402Cr, and the Cb component 402Cb. The first syntax element 408 is signaled at the frame level and includes a first classification syntax flag and a second classification syntax flag. The first classification syntax flag indicates whether the band offset dedicated mode 440 is enabled for the luma component 404, and the second classification syntax flag indicates whether the band offset dedicated mode 440 is jointly enabled for the Cb component 402Cb and the Cr component 402Cr.
[0076] In some embodiments, when determining that the first syntax element 408 has a first predefined value (e.g., "1") indicating the enabling of the band-offset dedicated mode 440, the decoder 122 identifies, in the video bitstream 116, a band count syntax element 422 indicating a band count limit 432 associated with the band-offset dedicated mode 440. Further, in some embodiments, the band count limit 432 is signaled in binary logarithm (which is a positive integer), thereby reducing signaling overhead. For example, the band count limit 432 is 64 bands and is represented as 5 in binary logarithm. In another example, the band count limit 432 is 128 bands and is represented as 7 in binary logarithm. In another example, the band count limit 432 is 256 bands and is represented as 8 in binary logarithm. In some embodiments, the current image frame has a bit depth, and the band count limit 432 is determined by a bitwise left shift operation associated with the bit depth. For example, the band count limit 432 is equal to 1 << bit depth.
[0077] In some embodiments, when determining that the first syntax element 408 has a second predefined value (e.g., "0") indicating the disabling of the band-offset dedicated mode for different image frames, the decoder 122 identifies, in the video bitstream 116, a band count syntax element 422 indicating a band count limit 432 associated with both the band-offset classification and the sideband-offset classification for different image frames.
[0078] In some embodiments, when determining that the first syntax element 408 has a first predefined value (e.g., "1") indicating the enabling of the band-offset dedicated mode 440 for the current image frame, the decoder 122 identifies a predefined fixed number 434 of bands associated with the band-offset dedicated mode 440. For example, the predefined fixed number 434 of bands is one of the following: 128, 256, and 64.
[0079] In some embodiments, when determining that the first syntax element 408 has a first predefined value (e.g., "1") indicating that the band-offset specific mode 440 is enabled for the current picture frame, the decoder 122 identifies in the video bitstream 116 an additional syntax element 424 indicating the position of the first luma sample, which is applied to the band-offset classification of the first chroma sample 410. Further, in some embodiments, the additional syntax element 424 has a single bit 424-1 that selects one of a third set of two luma samples as the first luma sample 404C, where the third set of two luma samples includes at least the co-located luma sample 404CC sharing the upper left corner with the first chroma sample 410. Alternatively, in some embodiments, the additional syntax element 424 has two bits 424-2 that select one of a first subset of four luma samples 404-1 as the first luma sample 404C, where the four luma samples 404-1 include the current or co-located luma sample 404CC sharing the upper left corner with the first chroma sample 410, the right-adjacent luma sample 404R, the bottom-adjacent luma sample 404B, and the bottom-right adjacent luma sample 404BR. Alternatively, in some embodiments, the additional syntax element 424 has three bits 424-3 that select one of a second subset of nine luma samples 404-2 as the first luma sample 404C, where the nine luma samples 404-2 include the current or co-located luma sample 404CC sharing the upper left corner with the first chroma sample 410 and eight surrounding adjacent luma samples 404X of the co-located luma sample 404CC.
[0080] In some embodiments associated with band-offset classification, the CCSO mode 442 corresponds to the band-offset classifier 412B. Based on the band-offset classifier 412B, the decoder 122 determines that a set of luma samples 404 includes the first luma sample 404C and at least one adjacent luma sample 404X. The set of luma samples 404 is provided to the quantizer 430 and used to generate at least one quantization value 404Q, which is further applied by the band-offset classifier 412B to classify the first chroma sample 410. For example, the filter type has a cross shape and includes four taps. The set of adjacent luma samples includes the first luma sample 404C, the north luma sample 404N, the south luma sample 404S, the west luma sample 404W, and the east luma sample 404E, and these luma samples 404 are further quantized to quantization values 404QC, 404QN, 404QS, 404QW, and 404QE, respectively.
[0081] For example, classifier 412 classifies the first chrominance sample 410 based on quantization value 404Q to determine a first sample offset 406 of the first chrominance sample 410. In the example, quantization value 404Q includes quantization values 404QC, 404QN, 404QS, 404QW, and 404QE. Lookup table 414 maps multiple combinations of quantization values 404QC, 404QN, 404QS, 404QW, and 404QE to different sample offset options SO (e.g., SO1 to SQ16). Based on lookup table 414, quantization value 404Q corresponds to one of the combinations in lookup table 414, and the corresponding sample offset option SO is identified to correspond to the combination of quantization difference 404Q, and is thus selected for the first sample offset 406. In other words, in some embodiments, decoder 122 classifies the first chrominance sample 410 by identifying a combination of at least one quantization value 404Q in lookup table 414 that associates multiple quantization combinations with multiple offset value options SO (e.g., SO1 to SO16) and determining a first sample offset 406 corresponding to the combination of at least one quantization value 404Q in lookup table 414.
[0082] In some embodiments, a scalar quantizer 430 including multiple quantization intervals 418 (QI) and multiple quantization levels 420 (QL) quantizes the value (unassociated difference or gradient) of at least one luma sample 404A into multiple integer values in quantization range 416, and each of the at least one quantization value 404Q includes a corresponding integer in quantization range 416. For each integer value in quantization range 416, quantization interval 418 is defined as the range of values assigned to the corresponding integer value. Quantization level 420 corresponds to the corresponding integer value to which a difference range associated with quantization interval 418 is assigned.
[0083] In some embodiments, in the band-offset dedicated mode 440, only the first luma sample 404 is quantized, e.g., by quantizer 430, and classified to determine a first sample offset 406 of the first chrominance sample 410. The first luma sample 404 is determined to be associated with one of multiple bands (e.g., band 608). Each of the multiple bands corresponds to a respective sample offset value. The first sample offset 406 is determined to be equal to the respective sample offset value corresponding to one of the multiple bands.
[0084] The first chroma sample 410 is adjusted based on the first sample offset 406 of the first chroma sample 410, thereby achieving the reconstruction of the current image frame. In some embodiments, the first chroma sample 410 includes the first chroma sample 402C in the current image frame that is located at the same position as the first luma sample 404C, and the first chroma sample 402C is adjusted based on the first sample offset 406. Alternatively, in some embodiments, the first chroma sample 410 is the first luma sample 404C, and the first luma sample 404C is adjusted based on the first sample offset 406.
[0085] Figure 5 FIG. 4 is a flowchart of an example process 500 of applying cross-component sample offset in in-loop filtering according to some embodiments. The decoder 122 receives a video bitstream 116 including a current image frame and a first syntax element 408. The first syntax element 408 indicates whether to determine the first sample offset 406 of the first chroma sample 410 based on the value of at least one luma sample 404, independently of any associated luma gradient of the at least one luma sample 404. The first syntax element 408 has a first predefined value indicating the enabled offset-only mode 440. Based on the offset-only mode 440, the decoder 122 generates at least one quantization value 404Q based on at least one luma sample 404, and the at least one luma sample includes the first luma sample 404C in the current image frame that is co-located with the first chroma sample 410. The first chroma sample 410 is classified based on the at least one quantization value 404Q to determine the first sample offset 406 of the first chroma sample 410. The decoder 122 reconstructs the current image frame at least by adjusting the first chroma sample 410 based on the first sample offset 406.
[0086] In some embodiments, the video bitstream 116 further includes an alternative syntax element 502 that indicates whether to apply the band offset and the sideband offset to determine the chrominance sample offset 506 for at least one color component 510 (e.g., including the first chrominance sample 410). Additionally, in some embodiments, the alternative syntax element 502 includes a first bit 502B that indicates whether to apply the band offset and a second bit 502E that indicates whether to apply the sideband offset. In some embodiments, at least one color component 510 includes two or more of the luminance component 404, the Cr component 402Cr, and the Cb component 404Cb, and the alternative syntax element 502 is signaled jointly for the at least one color component 510. In some embodiments, at least one color component 510 includes only one of the luminance component 404, the Cr component 402Cr, and the Cb component 402Cb, and the alternative syntax element 502 is signaled individually for that only one of the at least one color component 510. Each remaining color component 510 may have a corresponding alternative syntax element. For example, three alternative syntax elements 502 each having two bits are used for components 404, 402Cr, and 402Cb, respectively.
[0087] In some cases (e.g., associated with a disabled CCSO mode), when it is determined that the first bit 502B has a first value (e.g., "0") that disables the band offset and the second bit 502E has a first value (e.g., "0") that disables the sideband offset, the video decoder 122 may abort the cross-component offset filtering for at least one color component 510. Alternatively, in some cases (e.g., associated with the band-offset-only mode 440), when it is determined that the first bit 502B has a second value (e.g., "1") that enables the band offset and the second bit 502E has a first value (e.g., "0") that disables the sideband offset, the video encoder 122 may determine the chrominance sample offset 506 based on the band offset of at least one color component 510 in the cross-component offset filtering. Alternatively, in some cases (e.g., associated with the sideband-offset-only mode), when it is determined that the first bit 502B has a first value (e.g., "0") that disables the band offset and the second bit 502E has a second value (e.g., "1") that enables the sideband offset, the video encoder may determine the chrominance sample offset 506 based on the sideband offset of at least one color component 510 in the cross-component offset filtering. Alternatively, in some cases (e.g., associated with the hybrid-offset mode), when it is determined that the first bit 502B has a second value (e.g., "1") that enables the band offset and the second bit 502E has a second value (e.g., "1") that enables the sideband offset, the video encoder 122 may determine the chrominance sample offset 506 based on a combination of the band offset and the sideband offset of at least one color component 510 in the cross-component offset filtering.
[0088] In some embodiments, when the second bit 502E has a second value enabling sideband offset, a sideband offset classifier 412E is used in loop filtering. Based on the sideband offset classifier 412E, the decoder 122 determines that at least one luminance sample 404 includes a first luminance sample 404C and at least one neighboring luminance sample 404X, and further determines at least one difference between at least one neighboring luminance sample 404X and the first luminance sample 404C. At least one quantization value 404Q is generated based on at least one difference and applied by the sideband offset classifier 412E to classify the first chrominance sample 410. For example, the filter type has a cross shape and includes four taps. At least one neighboring luminance sample includes a north luminance sample 404N, a south luminance sample 404S, a west luminance sample 404W, and an east luminance sample 404E. The decoder 122 determines at least one difference between at least one neighboring luminance sample and the first luminance sample. For example, at least one difference includes one or more of the following: a north difference, a south difference, a west difference, and an east difference. Each of the differences is the difference between a corresponding one of the neighboring luminance samples 404X and the first luminance sample 404C. At least one difference is quantized to generate at least one quantization value 404QX. For example, at least one quantization value 404QX includes one or more of the following: a north quantization value 404QN, a south quantization value 404QS, a west quantization value 404QW, and an east quantization value 404QE. Each of the differences is provided to a quantizer 430 and quantized to generate a corresponding one of the quantization values 404QN, 404QS, 404QW, and 404QE. The quantization value 404QC is equal to 0.
[0089] In addition, in some embodiments, a look-up table 414 maps a plurality of combinations of the quantization values 404QN, 404QS, 404QW, and 404QE to different sample offset options SO (e.g., SO1 to SQ16). Based on the look-up table 414, the quantization value 404Q corresponds to one of the combinations in the look-up table 414, and the corresponding sample offset option SO is identified to correspond to the combination of the quantized differences 404Q and is thus selected for the first sample offset 406.
[0090] Figure 6Histogram 600 associated with multiple luminance samples 404 of a current image frame according to some embodiments. The decoder 122 receives a video bitstream 116 including the current image frame and a first syntax element 408 for the band-offset dedicated mode 440. The first syntax element 408 indicates whether a first sample offset 406 of a first chrominance sample 410 is determined based on the value of at least one luminance sample 404, regardless of any associated luminance gradient of the at least one luminance sample 404. In some embodiments, a histogram 600 is identified for multiple luminance samples 404 of the current image frame. Based on the histogram 620, the decoder 122 identifies a luminance sample band range 602 of the current image frame, and the luminance sample band range 602 is defined by a luminance value lower threshold 604 and a luminance value upper threshold 606. The luminance sample band range 602 is divided into at least one band 608 of luminance values. Sample counts are plotted on the histogram 600 for the at least one band 608. In some embodiments, the bands 608 may be uniform, having luminance values of equal width. Alternatively, in some embodiments not shown, the bands 608 may be non-uniform, having luminance values of corresponding widths.
[0091] In some embodiments, the luminance samples of the current image frame do not have luminance values below the luminance value lower threshold 604 or above the luminance value upper threshold 606. In an example, the luminance value lower threshold 604 is equal to the minimum luminance value of the multiple luminance samples 404, and the luminance value upper threshold 606 is equal to the maximum luminance value of the multiple luminance samples 404.
[0092] Alternatively, in some embodiments, based on the luminance sample band range 612, a first subset 610L of the luminance samples of the current image frame has luminance values below the luminance value lower threshold 614, and a second subset 610U of the luminance samples has luminance values above the luminance value upper threshold 616. The total number of luminance samples below the luminance value lower threshold 614 is less than a first count threshold, and the total number of luminance samples above the luminance value upper threshold is less than a second count threshold (e.g., which is equal to a value different from the first count threshold).
[0093] Figure 7 Is a flowchart illustrating an example method 700 for encoding video according to some embodiments. The method 700 may be implemented in a computing system having a control circuit and a memory storing instructions executed by the control circuit (e.g., Figure 1executed at the server system 112, the source device 102, or the electronic device 120) herein. In some embodiments, method 700 is applied in conjunction with at least one video codec, including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. The cross-component offset filtering method is implemented based on an edge-preserving loop filter that uses reconstructed samples (e.g., luma sample 404) to determine the sample offset 406 of the luma sample 404 or the chroma sample 402. The luma sample 404 can be classified based on a band offset (e.g., depending on the luma sample 404) or a sideband offset (e.g., depending on the gradient or difference between the first luma sample 404C and the adjacent luma sample 404X). In various embodiments of the present application, the band offset is used as an independent option independent of the sideband offset to simplify the design and increase the gain from certain types of content (e.g., screen content).
[0094] In some embodiments, a flag (e.g., the first syntax element 408) is signaled at the frame level (or other high-level syntax) to indicate whether to use the band-offset dedicated mode 440. If the band-offset dedicated mode 440 is used, only the sample values of the current luma sample or the co-located luma sample (e.g., the first luma sample 404C) are used for offset classification. Otherwise, if the only-band mode 440 is not used, both the band offset and the sideband offset are considered in the classification. For example, a band-offset dedicated flag can be signaled at the frame level for each individual color component (e.g., for the luma sample 404, for the Cr sample). In an example, the band-offset dedicated flag can be signaled at the frame level and jointly shared for all color components (e.g., 402 and 404). In an example, the band-offset dedicated flag can be signaled at the frame level for the luma and chroma components and includes a first flag only for the luma sample 404, and a second flag is shared for the two chroma components 402.
[0095] In some embodiments, the band-offset dedicated flag (e.g., the first syntax element 408) is signaled first. If the flag indicates the use of the band-offset dedicated mode 440, an additional syntax element (e.g., the band number syntax element 422) indicating the maximum number of bands (e.g., the band number limit 422) of the band-offset dedicated mode 440 is further signaled. In an example, the maximum number of bands for the band-offset dedicated mode 440 can be signaled in log2 form to reduce signaling overhead. In another example, the upper limit of the maximum number of bands (e.g., the band number limit 422) can be equal to (1 << bit depth).
[0096] In some embodiments, a band offset dedicated flag is first signaled (e.g., the first syntax element 408). If the flag indicates that the band offset dedicated mode 440 is used, a predefined fixed number 434 of bands for the band offset dedicated mode 440 is used in the band offset classification. In some embodiments, the band offset dedicated mode 440 is true. A predefined number 434 of bands are used in the band offset dedicated mode 440, while a maximum number (e.g., band number limit 422) of bands are used for the band offset plus sideband offset mode 440. Examples of the predefined number 434 of bands include, but are not limited to, 128, 256, and 64.
[0097] In some embodiments, a band offset dedicated flag is first signaled (e.g., a first syntax element 408). In some cases, the flag indicates that a band offset dedicated mode 440 is used, and an additional syntax element 424 ( Figure 4 ) to indicate a subset of luma samples for band offset classification. For example, the additional syntax element 424-2 includes 2 bits configured to identify a subset of four luma samples 404-1, from which the first luma sample 404C is selected to be co-located with the first chroma sample 410. The four luma samples 404-1 ( Figure 4 ) includes a current or co-located luma sample 404CC that shares an upper left corner with the first chroma sample 410, a right luma sample 404R, a bottom luma sample 404B, and a bottom right luma sample 404RB. In another example, the additional syntax element 424 includes 3 bits configured to identify a subset of nine luma samples 404-2, from which the first luma sample 404C is selected to be co-located with the first chroma sample 410. The nine luma samples 404-2 ( Figure 4 ) includes a current or co-located luma sample 404CC and its eight surrounding neighboring luma samples 404X that share an upper left corner with the first chroma sample 410. In yet another example, the additional syntax element 424 includes a single bit 424-1 configured to identify a subset of two luma samples, from which the first luma sample 404C is selected to be co-located with the first chroma sample 410. The subset of two luma samples includes at least the current or co-located luma sample 404CC and another neighboring luma sample 404X.
[0098] In some embodiments, for at least one color component 510 ( Figure 5) Signaling the syntax element 502 to indicate the use of band offset and sideband offset. In the example, the syntax element 502 is signaled jointly for multiple color components. For example, jointly for the luminance samples 404 and the Cb and Cr samples 402; jointly for the Cb and Cr samples 402. Alternatively, in some embodiments, the syntax element is signaled separately for different color components (e.g., separately for the luminance samples 404, the Cb samples, and the Cr samples).
[0099] In some embodiments, the syntax element 502 includes a two-bit index. The first bit 502B indicates whether the band offset is enabled, and the second bit 502E indicates whether the sideband offset is enabled. The syntax element 502 corresponds to four types of cases. First, in some embodiments, when the first bit 502B signals the use of the band offset with a value indicating that the band offset is not enabled (e.g., "0"), and when the second bit 502E signals the use of the sideband offset with a value indicating that the sideband offset is not enabled (e.g., "0"), the cross-component offset filtering is disabled. Second, in some embodiments, when the first bit 502B signals the use of the band offset with a value indicating that the band offset is enabled (e.g., "1"), and when the second bit 502E signals the use of the sideband offset with a value indicating that the sideband offset is not enabled (e.g., "0"), the band-offset-only is applied to the cross-component offset filtering. Third, in some embodiments, when the first bit 502B signals the use of the band offset with a value indicating that the band offset is not enabled (e.g., "0"), and when the second bit 502E signals the use of the sideband offset with a value indicating that the sideband offset is enabled (e.g., "1"), only the sideband offset is applied to the cross-component offset filtering. Fourth, in some embodiments, when the first bit 502B signals the use of the band offset with a value indicating that the band offset is enabled (e.g., "1"), and when the second bit 502E signals the use of the sideband offset with a value indicating that the sideband offset is enabled (e.g., "1"), the combination of the band and sideband offsets is applied to the cross-component offset filtering.
[0100] In some embodiments, if only the band mode 440 is enabled for cross-component sample offset, the histogram 600 ( Figure 6 ) of the current image frame or image sequence with luminance value intensity thresholds (e.g., the lower luminance value thresholds 604 and 614, the upper luminance value thresholds 606 and 616) can be used to narrow the band. For example, referring to Figure 6, the minimum allowable brightness value is equal to 0, and the maximum allowable brightness value is equal to 1 << bit depth. The allowable values of the brightness sample 404 are between 0 and 1 << bit depth, for example, 0 to 255. Starting from the minimum allowable brightness value (e.g., 0), the first brightness value with a non-zero count in the histogram 600 corresponds to the start of the lowest band and is set as the brightness value lower threshold 604. The last brightness value with a non-zero count in the histogram 600 corresponds to the end of the highest band and is set as the brightness value upper threshold 606. The brightness sample band range 602 is set between the brightness value lower threshold 604 and the brightness value upper threshold 606. The brightness sample band range 602 is divided into multiple brightness value bands 608. In the example, the bands 608 correspond to brightness values from 100 to 200 and from 220 to 255, and are configured to cover brightness values from 100 to 255. In another example, the brightness sample band range 602 (e.g., 100 to 255) is divided into 32 equal bands.
[0101] In another example, the brightness value lower threshold 614 is determined based on the sample count. The total sample count below the brightness value lower threshold 614 is equal to the first count threshold (e.g., 10 to 15), and the total sample count above the brightness value upper threshold 616 is equal to the second count threshold (e.g., 10 to 15). The brightness value range 612 between the brightness value lower threshold 614 and the brightness value upper threshold 616 is divided into multiple bands 608.
[0102] Although Figure 7 Multiple logical stages are shown in a specific order, but stages that do not depend on the order can be reordered, and other stages can be combined or split. A certain reordering or other grouping not specifically mentioned will be obvious to those of ordinary skill in the art, so the order and grouping presented herein are not exhaustive. In addition, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.
[0103] Now turning to some example embodiments.
[0104] (A1)In some implementations, a method 700 for decoding video data is implemented. Method 700 includes receiving (operation 702) a video bitstream including a current image frame, where the video bitstream includes (operation 704) a first syntax element that indicates whether a first sample offset of a first chroma sample is determined based on values of at least one luma sample, regardless of any associated luma gradient of the at least one luma sample; when the first syntax element has a first predefined value, determining (operation 706) that a band-offset specific mode is enabled; when the band-offset specific mode is enabled, generating (operation 708) at least one quantization value based on at least one luma sample including a first luma sample co-located with the first chroma sample in the current image frame; classifying (operation 710) the first chroma sample of the current image frame based on the at least one quantization value to determine the first sample offset of the first chroma sample; and reconstructing (operation 712) the current image frame at least by adjusting the first chroma sample based on the first sample offset of the first chroma sample.
[0105] (A2)In some embodiments of A1, where the video bitstream includes an alternative image frame and an associated first syntax element for the band-offset specific mode. Method 700 further includes: determining that the associated first syntax element is associated with the alternative image frame and that a flag of the first syntax element has a second predefined value indicating disabling of the band-offset specific mode; and generating a second sample offset of a second chroma sample co-located with a second luma sample in the alternative image frame based on band classification and sideband-offset classification.
[0106] (A3)In some embodiments of A1 or A2, the first syntax element is signaled separately at a frame level for a first color component corresponding to the first chroma sample.
[0107] (A4)In some embodiments of any one of A1 - A3, the current image frame includes at least two color components, the at least two color components further include a first color component corresponding to the first chroma sample, and the first syntax element is signaled at a frame level and includes at least two classification syntax flags, each flag indicating whether the band-offset specific mode is enabled for a corresponding color component.
[0108] (A5)In some embodiments of any one of A1 - A4, the first syntax element is a flag that signals at a frame level at least two color components of the current image frame, and the at least two color components include a first color component corresponding to the first chroma sample.
[0109] (A6)In some embodiments of any one of A1 - A5, each sample of the current image frame includes a luminance component, a blue - difference chrominance (Cb) component, and a red - difference chrominance (Cr) component, and the first chrominance sample corresponds to one of the luminance component, the Cr component, and the Cb component. The first syntax element is signaled at the frame level and includes a first classification syntax flag and a second classification syntax flag. The first classification syntax flag indicates whether the band - offset dedicated mode is enabled for the luminance component, and the second classification syntax flag indicates whether the band - offset dedicated mode is jointly enabled for the Cb component and the Cr component.
[0110] (A7)In some embodiments of any one of A1 - A6, method 700 further includes, when determining that the first syntax element has the first predefined value indicating enabling the band - offset dedicated mode, identifying a band - count syntax element from the video bitstream, where the band - count syntax element indicates a band - count limit associated with the band - offset dedicated mode.
[0111] (A8)In some embodiments of A7, the band - count limit is signaled in the form of the binary logarithm of a positive integer.
[0112] (A9)In some embodiments of A7, the current image frame has a bit - depth, and the band - count limit is determined by a bit - by - bit left - shift operation associated with the bit - depth.
[0113] (A10)In some embodiments of any one of A1 - A9, method 700 further includes, when determining that the first syntax element has a second predefined value indicating disabling the band - offset dedicated mode for different image frames, identifying a band - count syntax element from the video bitstream, where the band - count syntax element indicates a band - count limit associated with the band - offset classification and the side - band offset classification of the different image frames.
[0114] (A11)In some embodiments of any one of A1 - A10, method 700 further includes, when determining that the first syntax element has the first predefined value indicating enabling the band - offset dedicated mode for the current image frame, identifying a predefined fixed number of bands associated with the band - offset dedicated mode.
[0115] (A12)The predefined fixed number of bands is one of 128, 256, and 64.
[0116] (A13)In some embodiments of any one of A1 - A12, method 700 further includes, when determining that the flag of the first syntax element has the first predefined value indicating that the current picture frame enables the band offset dedicated mode, identifying an additional syntax element from the video bitstream that indicates the position of the first luminance sample, and this additional syntax element is applied to the band offset classification of the first chrominance sample.
[0117] (A14)In some embodiments of A13, the additional syntax element has two bits for selecting one of a first subset of four luminance samples as the first luminance sample, and the four luminance samples include the co - located luminance sample sharing the upper left corner with the first chrominance sample, the right - adjacent luminance sample, the bottom - adjacent luminance sample, and the bottom - right - adjacent luminance sample.
[0118] (A15)In some embodiments of A13, the additional syntax element has three bits for selecting one of a second subset of nine luminance samples as the first luminance sample, and the nine luminance samples include the co - located luminance sample sharing the upper left corner with the first chrominance sample and the eight surrounding adjacent luminance samples of the co - located luminance sample.
[0119] (A16)In some embodiments of A13, the additional syntax element has a single bit that selects one of two positions corresponding to at least one co - located luminance sample sharing the upper left corner with the first chrominance sample as the first luminance sample.
[0120] (A17)In some embodiments of any one of A1 - A16, the video bitstream further includes an alternative syntax element that indicates whether to apply band offset and sideband offset to determine the chrominance sample offset of at least one color component.
[0121] (A18)In some embodiments of A17, the alternative syntax element includes a first bit indicating whether to apply the band offset and a second bit indicating whether to apply the sideband offset.
[0122] (A19) In some embodiments of A18, method 700 further includes one of the following operations: (1) when it is determined that the first bit has a first value that disables the band offset and the second bit has a first value that disables the sideband offset, aborting the cross-component offset filtering for at least one color component; (2) when it is determined that the first bit has a second value that enables the band offset and the second bit has a first value that disables the sideband offset, determining the chrominance sample offset based on the band offset of the at least one color component in the cross-component offset filtering; (3) when it is determined that the first bit has a first value that disables the band offset and the second bit has a second value that enables the sideband offset, determining the chrominance sample offset based on the sideband offset of the at least one color component in the cross-component offset filtering; and (4) when it is determined that the first bit has a second value that enables the band offset and the second bit has a second value that enables the sideband offset, determining the chrominance sample offset based on a combination of the band offset and the sideband offset of at least one color component in the cross-component offset filtering.
[0123] (A20) In some embodiments of A17 or A18, the at least one color component includes two or more of a luminance component, a Cr component, and a Cb component, and the alternative syntax element is signaled jointly for the at least one color component.
[0124] (A21) In some embodiments of A17 or A18, the at least one color component includes only one of a luminance component, a Cr component, and a Cb component, and the alternative syntax element is signaled separately for the at least one color component.
[0125] (A22) In some embodiments of any one of A1 - A21, method 700 further includes identifying a histogram of the luminance samples of the current image frame; based on the histogram, identifying a luminance sample band range of the current image frame, the luminance sample band range being defined by a lower luminance value threshold and an upper luminance value threshold; and dividing the luminance sample band range into at least one luminance value band.
[0126] (A23) In some embodiments of A22, the luminance samples of the current image frame do not have a luminance value lower than the lower luminance value threshold or higher than the upper luminance value threshold.
[0127] (A24) In some embodiments of A23, the total number of luminance samples lower than the lower luminance value threshold is less than a count threshold, and the total number of luminance samples higher than the upper luminance value threshold is less than the count threshold.
[0128] (A25)In some embodiments of any one of A1 - A24, the video bitstream further includes a first syntax element for the cross - component sample offset (CCSO) mode, and this first syntax element indicates whether the first sample offset of the first chrominance samples of the current picture frame is determined based on at least one luma sample.
[0129] (A26)In some embodiments, a computing system includes a control circuit; and a memory storing at least one program configured to be executed by the control circuit. The at least one program further includes instructions for the following operations: receiving video data including a current picture frame and a first syntax element; encoding the current picture frame; based on the first syntax element, determining that the offset - only mode is enabled to determine a first sample offset of a first chrominance sample based on the value of at least one luma sample, regardless of any associated luma gradient of the at least one luma sample; transmitting the encoded current picture frame through the video bitstream; and signaling the first syntax element through the video bitstream to indicate that the offset - only mode is applied to reconstruct the first chrominance sample co - located with the first luma sample based on the first sample offset.
[0130] (A27)In some embodiments, a non - volatile computer - readable storage medium stores at least one program for execution by a control circuit of a computing system. The at least one program includes instructions for the following operations: obtaining a source video sequence including a current picture frame; and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes: the current picture frame; and a first syntax element indicating whether a first sample offset of a first chrominance sample is determined based on the value of at least one luma sample, regardless of any associated luma gradient in the at least one luma sample. When the offset - only mode is enabled, the first sample offset of the first chrominance sample is determined based on at least one luma sample. The at least one luma sample includes a first luma sample co - located with the first chrominance sample in the current picture frame.
[0131] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, the memory storing one or more sets of instructions configured to be executed by the control circuit, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A26 above).
[0132] In yet another aspect, some embodiments include a non - volatile computer - readable storage medium storing one or more sets of instructions for execution by a control circuit of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A26 above).
[0133] The methods described herein can be used alone or in any combination in any order. In addition, each method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., at least one processor, or at least one integrated circuit). For example, at least one processor executes a program stored in a non-volatile computer-readable medium. Hereinafter, the term block may be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU.
[0134] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.
[0135] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of at least one of the associated listed items. It should also be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of at least one other feature, integer, step, operation, element, component, and / or combination thereof.
[0136] As used herein, depending on the context, the term "if" may be interpreted to mean "when the stated precondition is true" or "once the stated precondition is true" or "in response to determining that the stated precondition is true" or "in accordance with determining that the stated precondition is true" or "in response to detecting that the stated precondition is true". Similarly, depending on the context, the phrase "if it is determined that [the stated precondition is true]" or "if [the stated precondition is true]" or "when [the stated precondition is true]" may be interpreted to mean "upon determining that the stated precondition is true" or "in response to determining that the stated precondition is true" or "in accordance with determining that the stated precondition is true" or "upon detecting that the stated precondition is true" or "in response to detecting that the stated precondition is true".
[0137] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the operating principles and the practical application, thereby enabling others skilled in the art to understand.
Claims
1. A method for decoding video data, characterized in that, Comprising: Receiving a video bitstream including a current image frame, wherein the video bitstream includes a first syntax element indicating whether to determine a first sample offset of a first chrominance sample based on values of at least one luma sample, and the first syntax element is independent of any associated luma gradient of the at least one luma sample; When the first syntax element has a first predefined value, determining to enable the band offset dedicated mode; When the band offset dedicated mode is enabled, generating at least one quantization value based on the at least one luma sample, the at least one luma sample including a first luma sample in the current image frame that is co-located with the first chrominance sample; Classifying the first chrominance sample of the current image frame based on the at least one quantization value to determine the first sample offset of the first chrominance sample; And Reconstructing the current image frame at least by adjusting the first chrominance sample based on the first sample offset of the first chrominance sample.
2. The method according to claim 1, wherein The video bitstream includes an alternative image frame and an associated first syntax element, and the method further includes: Determining that the associated first syntax element is associated with the alternative image frame and has a second predefined value indicating disabling the band offset dedicated mode; and Generating a second sample offset of a second chrominance sample that is co-located with a second luma sample in the alternative image frame based on band offset classification and sideband offset classification.
3. The method according to claim 1, wherein Signaling the first syntax element separately at the frame level for a first color component corresponding to the first chrominance sample.
4. The method according to claim 1, characterized in that The current image frame includes at least two color components, the at least two color components further include a first color component corresponding to the first chrominance sample, and the first syntax element is signaled at the frame level and includes at least two classification syntax flags, each of the at least two classification syntax flags indicating whether the band offset dedicated mode is enabled for a corresponding color component.
5. The method according to claim 1, wherein The first syntax element is a flag signaled at the frame level for at least two color components of the current image frame, and the at least two color components include a first color component corresponding to the first chrominance sample.
6. The method according to claim 1, wherein: Each sample of the current image frame includes a luma component, a blue-difference chrominance (Cb) component, and a red-difference chrominance (Cr) component, and the first chrominance sample corresponds to one of the luma component, the Cr component, and the Cb component; And The first syntax element is signaled at the frame level and includes a first syntax flag and a second syntax flag, the first syntax flag indicating whether the band offset dedicated mode is enabled for the luma component, and the second syntax flag indicating whether the band offset dedicated mode is enabled jointly for the Cb component and the Cr component.
7. The method according to claim 1, characterized in that Further comprising: When determining that the first syntax element has the first predefined value indicating enabling the band offset dedicated mode, identifying a band number syntax element from the video bitstream, the band number syntax element indicating a band number limit associated with the band offset dedicated mode.
8. The method according to claim 7, wherein The number of bands is signaled in the form of a binary logarithm, which is a positive integer.
9. The method according to claim 7, wherein The current image frame has a bit depth, and the number of bands is determined by a bitwise left shift operation associated with the bit depth.
10. The method according to claim 1, wherein Further comprising: When it is determined that the first syntax element has a second predefined value indicating the disabling of the band offset specific mode for different image frames, identifying a band number syntax element from the video bitstream, the band number syntax element indicating a band number limit associated with both the band offset classification and the sideband offset classification of the different image frames.
11. The method according to claim 1, characterized in that Further comprising: When it is determined that the first syntax element has the first predefined value indicating the enabling of the band offset specific mode for the current image frame, identifying a predefined fixed number of bands associated with the band offset specific mode.
12. The method according to claim 11, wherein The predefined fixed number of bands is one of the following: 128, 256, and 64.
13. The method according to claim 1, characterized in that, Further comprising: When it is determined that the flag of the first syntax element has the first predefined value indicating the enabling of the band offset specific mode for the current image frame, identifying an additional syntax element from the video bitstream indicating the position of the first luma sample, the additional syntax element being applied to the band offset classification of the first chroma sample.
14. The method according to claim 13, characterized in that, The additional syntax element has two bits, and the two bits select one of a first subset of four luma samples as the first luma sample, the four luma samples including the co-located luma sample sharing the upper left corner with the first chroma sample, the right adjacent luma sample, the bottom adjacent luma sample, and the bottom right adjacent luma sample.
15. The method according to claim 13, wherein The additional syntax element has three bits, and the three bits select one of a second subset of nine luma samples as the first luma sample, the nine luma samples including the co-located luma sample sharing the upper left corner with the first chroma sample and the eight surrounding adjacent luma samples of the co-located luma sample.
16. The method according to claim 13, wherein The additional syntax element has a single bit, and the single bit selects one of two positions corresponding to at least one co-located luma sample sharing the upper left corner with the first chroma sample as the first luma sample.
17. The method according to claim 1, characterized in that, The video bitstream further includes an alternative syntax element indicating whether band offset and sideband offset are applied to determine the chroma sample offset of at least one color component.
18. The method according to claim 17, wherein The alternative syntax element includes a first bit indicating whether the band offset is applied and a second bit indicating whether the sideband offset is applied.
19. The method according to claim 18, wherein Further comprising one of the following operations: When it is determined that the first bit has a first value disabling the band offset and the second bit has the first value disabling the sideband offset, aborting the cross-component offset filtering for the at least one color component; When it is determined that the first bit has a second value enabling the band offset and the second bit has the first value disabling the sideband offset, determining the chroma sample offset based on the band offset of the at least one color component in the cross-component offset filtering; When it is determined that the first bit has the first value that disables the band offset and the second bit has the second value that enables the sideband offset, determine the chrominance sample offset based on the sideband offset of at least one color component in the cross-component offset filtering; and When it is determined that the first bit has the second value that enables the band offset and the second bit has the second value that enables the sideband offset, determine the chrominance sample offset based on a combination of the band offset and the sideband offset of at least one color component in the cross-component offset filtering.
20. The method according to claim 17, wherein The at least one color component includes at least two of a luminance component, a Cr component, and a Cb component, and signals the alternative syntax element jointly for the at least one color component.
21. The method according to claim 17, wherein The at least one color component includes only one of a luminance component, a Cr component, and a Cb component, and signals the alternative syntax element separately for the at least one color component.
22. The method according to claim 1, characterized in that, Further includes: Identify a histogram of luminance samples of the current image frame; Based on the histogram, identify a luminance sample band range of the current image frame, the luminance sample band range being defined by a lower luminance value threshold and an upper luminance value threshold; and Divide the luminance sample band range into at least one luminance value band.
23. The method according to claim 22, wherein The luminance samples of the current image frame do not have a luminance value lower than the lower luminance value threshold or higher than the upper luminance value threshold.
24. The method according to claim 23, wherein The total number of luminance samples below the lower luminance value threshold is less than a count threshold, and the total number of luminance samples above the upper luminance value threshold is less than the count threshold.
25. The method according to claim 1, characterized in that, The video bitstream further includes a first syntax element for a cross-component sample offset (CCSO) mode, and the first syntax element for the cross-component sample offset mode indicates whether the first sample offset of the first chrominance sample of the current image frame is determined based on at least one luminance sample.
26. A computing system, characterized in that, Includes: A control circuit; and A memory that stores at least one program configured to be executed by the control circuit, and the at least one program further includes instructions for the following operations: Receive video data including a current image frame and a first syntax element; Encode the current image frame; Based on the first syntax element, determine that the band offset specific mode is enabled and determine a first sample offset of a first color sample based on the value of the at least one luminance sample, and the first syntax element is independent of any associated luminance gradient of the at least one luminance sample; Transmit the encoded current image frame via the video bitstream; and Signal the first syntax element via the video bitstream to indicate that the band offset specific mode is applied to reconstruct the first color sample co-located with the first luminance sample based on the first sample offset.
27. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores at least one program for execution by a control circuit of a computing system, and the at least one program includes instructions for the following operations: Obtain a source video sequence including a current image frame; and Perform the conversion between the source video sequence and the video bitstream, where the video bitstream includes: The current picture frame; and A first syntax element for the band offset occupancy mode, the first syntax element indicating a first sample offset of a first color sample determined based on the value of the at least one luma sample, the first syntax element being independent of any associated luma gradient of the at least one luma sample; wherein, when the only band offset mode is enabled, the first sample offset of the first color sample is determined based on the at least one luma sample, the at least one luma sample including a first luma sample in the current picture frame that is co-located with the first color sample.