Decoder-side quantization shift offset prediction
By determining the quantized shift offset assumption on the decoder side and calculating the cost value, selecting the lowest cost assumption to reconstruct the video block, the problem of lack of effective prediction on the decoder side in the prior art is solved, and the video decoding efficiency and quality are improved.
Patent Information
- Application Number
- CN202480005373.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-05-30
- Publication Date
- 2025-07-22
AI Technical Summary
The existing video encoding technology lacks effective quantized shift offset prediction methods on the decoder side, resulting in a lack of video decoding efficiency and quality.
By determining the quantized shift offset settings of multiple hypotheses on the decoder side, the associated cost value is calculated, and the assumption of the lowest cost value is selected to reconstruct the transformation coefficients, and then reconstruct the video blocks to realize the quantized shift offset prediction.
It improves the efficiency and quality of video decoding, reduces errors in data quantization, and improves image quality.
Smart Images

Figure CN120359747A_ABST
Abstract
Description
[0001] Incorporation by reference
[0002] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 602,329, filed on Nov. 22, 2023, entitled “DecoderSide Quantization Shifting Offset Prediction”, which is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure describes aspects generally related to video coding. Background Art
[0004] The background art description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work described in this background art section is not otherwise qualified as prior art at the time of filing, aspects of the work of the presently named inventors and descriptions that are not otherwise qualified as prior art are neither expressly nor implicitly admitted as prior art against the present disclosure.
[0005] Image / video compression can facilitate the transmission of image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial redundancy and temporal redundancy. In an example, a video codec can use a technique called intra prediction, which can compress an image based on spatial redundancy. For example, intra prediction can use reference data from the current picture in reconstruction for sample prediction. In another example, a video codec can use a technique called inter prediction, which can compress an image based on temporal redundancy. For example, inter prediction can utilize motion compensation to predict samples in the current picture based on a previously reconstructed picture. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention
[0006] Aspects of the present disclosure include bitstreams, methods, and apparatuses for video encoding / decoding. In some examples, an apparatus for video encoding / decoding includes processing circuitry.
[0007] Some aspects of the present disclosure provide a method for processing visual media data. The method includes performing a conversion between a visual media file and a bitstream of visual media data according to formatting rules. The bitstream includes decoding information for one or more pictures. The formatting rules specify: determining a plurality of hypotheses for decoder-side quantization shift offset prediction, where a hypothesis among the plurality of hypotheses corresponds to a potential quantization shift offset setting in the transform domain of the current block. The formatting rules further specify: calculating cost values respectively associated with the plurality of hypotheses, calculating the cost value associated with a hypothesis among the plurality of hypotheses based on the reconstructed pixels in the current block obtained based on the hypothesis and one or more adjacent reconstructed blocks. The formatting rules further specify: when the lowest cost value is less than a threshold, selecting a particular hypothesis having the lowest cost value, determining one or more quantization shift offsets of transform coefficients in the transform domain of the current block based on the particular hypothesis, reconstructing the transform coefficients based on the one or more quantization shift offsets, calculating a residual in the spatial domain of the current block based on the transform coefficients in the transform domain, and reconstructing the current block according to the residual in the spatial domain.
[0008] Some aspects of the present disclosure provide an apparatus for video decoding. The apparatus includes processing circuitry configured to receive a bitstream including decoding information for a current block and determine a plurality of hypotheses for decoder-side quantization shift offset prediction. A hypothesis among the plurality of hypotheses corresponds to a potential quantization shift offset setting in the transform domain of the current block. The processing circuitry is configured to calculate cost values respectively associated with the plurality of hypotheses, select a particular hypothesis from the plurality of hypotheses based on the cost values, determine one or more quantization shift offsets of transform coefficients in the transform domain based on the particular hypothesis, reconstruct the transform coefficients based on the one or more quantization shift offsets, calculate a residual in the spatial domain of the current block based on the transform coefficients in the transform domain, and reconstruct the current block according to the residual in the spatial domain.
[0009] According to aspects of the present disclosure, the plurality of hypotheses include a hypothesis for setting a sign for at least one quantization shift offset of non-zero transform coefficients. In some examples, the plurality of hypotheses include a first hypothesis for setting a positive sign for the quantization shift offset of non-zero transform coefficients in a transform block and a second hypothesis for setting a negative sign for the quantization shift offset of non-zero transform coefficients in the transform block. In some examples, the plurality of hypotheses include a first hypothesis for setting a positive sign for a first quantization shift offset of a DC coefficient in a transform block and a second hypothesis for setting a negative sign for the first quantization shift offset of the DC coefficient in the transform block. In some examples, the plurality of hypotheses include potential sign combinations of quantization shift offsets of a subset of non-zero transform coefficients in a transform block.
[0010] Aspects according to the present disclosure, the multiple hypotheses include hypotheses for setting values for at least one quantization shift offset that is a non-zero transform coefficient. In some examples, the multiple hypotheses include a first hypothesis for setting a first value for the quantization shift offset of non-zero transform coefficients in a transform block, and a second hypothesis for setting a second value for the quantization shift offset of non-zero transform coefficients in the transform block.
[0011] In some examples, the processing circuitry is configured to select a particular hypothesis having the lowest cost value.
[0012] Aspects according to the present disclosure, the processing circuitry is configured to calculate a cost value associated with a hypothesis based on reconstructed pixels in a current block and one or more adjacent reconstructed blocks resulting from the hypothesis. In some examples, the processing circuitry is configured to calculate a cost value that measures the discontinuity between the reconstructed pixels in the current block resulting from the hypothesis and one or more adjacent reconstructed blocks. In some examples, the processing circuitry is configured to calculate a cost value that measures the derivative discontinuity between the reconstructed pixels in the current block resulting from the hypothesis and one or more adjacent reconstructed blocks.
[0013] Aspects according to the present disclosure, the processing circuitry is configured to decode a flag from a bitstream that indicates whether decoder-side quantization shift offset prediction is performed. The flag can be signaled at any suitable level, such as sequence level, picture level, slice level, tile level, and block level. In some examples, the processing circuitry is configured to determine to use a fixed quantization shift offset when the flag indicates that decoder-side quantization shift offset prediction is not used.
[0014] Some aspects of the present disclosure also provide a method for video coding. The method includes determining to use decoder-side quantization shift offset prediction for a current block, and determining multiple hypotheses for decoder-side quantization shift offset prediction. The hypotheses in the multiple hypotheses correspond to potential quantization shift offset settings in the transform domain of the current block. The method further includes calculating cost values respectively associated with the multiple hypotheses, selecting a particular hypothesis having the lowest cost value from the multiple hypotheses when the lowest cost value is less than a threshold, determining one or more quantization shift offsets of transform coefficients in the transform domain based on the particular hypothesis, reconstructing the transform coefficients based on the one or more quantization shift offsets, calculating a residual in the spatial domain of the current block based on the transform coefficients in the transform domain, and reconstructing the current block according to the residual in the spatial domain.
[0015] In some examples, the method includes signaling a flag in a bitstream to indicate not to use decoder-side quantization shift offset prediction when the lowest cost value is above a threshold. In some examples, the plurality of hypotheses include hypotheses that set signs for at least one quantization shift offset of non-zero transform coefficients. In some examples, the plurality of hypotheses include hypotheses that set values for at least one quantization shift offset of non-zero transform coefficients.
[0016] In some examples, the method includes calculating a cost value associated with a hypothesis based on reconstructed pixels in a current block resulting from the hypothesis and one or more neighboring reconstructed blocks.
[0017] Aspects of the present disclosure also provide an apparatus for video coding. The apparatus for video coding includes processing circuitry configured to implement any of the methods described for video coding.
[0018] Aspects of the present disclosure also provide a method for video decoding. The method includes any of the methods implemented by an apparatus for video decoding.
[0019] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to execute any of the methods described for video decoding / encoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0021] Figure 1 is a schematic illustration of an example of a block diagram of a communication system (100).
[0022] Figure 2 is a schematic illustration of an example of a block diagram of a decoder.
[0023] Figure 3 is a schematic illustration of an example of a block diagram of an encoder.
[0024] Figure 4 is a diagram showing hypotheses in some examples.
[0025] Figure 5 is a diagram showing hypotheses in some examples.
[0026] Figure 6 is a diagram showing hypotheses in some examples.
[0027] Figure 7 is a diagram showing hypotheses in some examples.
[0028] Figure 8 A diagram showing the current block in some examples.
[0029] Figure 9 A flowchart outlining a decoding process according to some aspects of the present disclosure.
[0030] Figure 10 A flowchart outlining an encoding process according to some aspects of the present disclosure.
[0031] Figure 11 A schematic illustration of a computer system according to one aspect. Detailed Description
[0032] Figure 1 A block diagram of a video processing system (100) in some examples. The video processing system (100) is an example of the application of the disclosed subject matter's video encoder and video decoder in a streaming environment. The disclosed subject matter can be equivalently applied to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0033] The video processing system (100) includes a capture subsystem (113), which can include a video source (101), such as a digital camera device, that creates, for example, an uncompressed video picture stream (102). In an example, the video picture stream (102) includes samples taken by the digital camera device. The video picture stream (102) is depicted as a thick line to emphasize the higher data volume when compared to the encoded video data (104) (or decoded video bitstream), and the video picture stream (102) can be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) can include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize the lower data volume when compared to the video picture stream (102), and the encoded video data (104) can be stored on a streaming server (105) for future use. One or more streaming client subsystems such as Figure 1The client subsystems (106) and (108) therein can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an outgoing video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or another rendering device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) can be encoded according to certain video coding standards / video compression standards. Examples of such standards include ITU-T Recommendation H.265. In an example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC.
[0034] Note that the electronic device (120) and the electronic device (130) can include other components (not shown). For example, the electronic device (120) can include a video decoder (not shown), and the electronic device (130) can also include a video encoder (not shown).
[0035] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., receiving circuitry). The video decoder (210) can be used in place of Figure 1 the video decoder (110) in the example of
[0036] A receiver (231) may receive, for example, one or more decoded video sequences included in a bitstream for decoding by a video decoder (210). In one aspect, one decoded video sequence is received at a time, where the decoding of each decoded video sequence is independent of the decoding of other decoded video sequences. The decoded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data as well as other data, such as decoded audio data and / or auxiliary data streams that may be forwarded to their respective entities of use (not depicted). The receiver (231) may separate the decoded video sequences from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and an entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be external to the video decoder (210) (not depicted). In still other applications, there may be a buffer memory (not depicted) external to the video decoder (210) to, for example, prevent network jitter, and additionally there may be another buffer memory (215) inside the video decoder (210) to, for example, handle presentation timing. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer memory (215) may not be needed, or the buffer memory (215) may be small. For use on a best-effort packet network such as the Internet, a buffer memory (215) may be required, which may be relatively large and may advantageously have an adaptive size, and may be implemented at least partially in an operating system or a similar element (not depicted) external to the video decoder (210).
[0037] The video decoder (210) may include a parser (220) to reconstruct symbols (221) according to the decoded video sequences. The categories of these symbols include: information for managing the operation of the video decoder (210); and potential information for controlling a presentation device such as a presentation device (212) (e.g., a display screen), which is not part of the electronic device (230) but may be coupled to the electronic device (230), such as Figure 2As shown. The control information for the presentation device may be in the form of a Supplemental Enhancement Information (SEI) message or a fragment of a Video Usability Information (VUI) parameter set (not depicted). The parser (220) may perform parsing / entropy decoding on the received decoded video sequence. The decoding of the decoded video sequence may be performed according to video decoding techniques or video decoding standards and may follow various principles, including variable length decoding, Huffman coding, arithmetic decoding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set for at least one subgroup of pixels in the video decoder from the decoded video sequence based on at least one parameter corresponding to the group. The subgroups may include Groups of Picture (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser (220) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the decoded video sequence.
[0038] The parser (220) may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (215) to create symbols (221).
[0039] Depending on the type of the decoded video picture or a part of the decoded video picture (such as: inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (221) may involve multiple different units. The subgroup control information parsed by the parser (220) from the decoded video sequence may control which units are involved and how to control these units. For the sake of brevity, such subgroup control information flows between the parser (220) and the following multiple units are not described.
[0040] In addition to the functional blocks already mentioned, the video decoder (210) may be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated into each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.
[0041] The first unit is a scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives the quantized transform coefficients as symbols (221) and control information from the parser (220), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) may output a block including sample values that can be input into the aggregator (255).
[0042] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture but may use prediction information from a previously reconstructed part of the current picture. Such predictive information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses the surrounding reconstructed information obtained from the current picture buffer (258) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds, on a per-sample basis, the prediction information already generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).
[0043] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-coded and possibly motion-compensated block. In this case, the motion compensation prediction unit (253) may access the reference picture memory (257) to obtain samples for prediction. After motion-compensating the obtained samples according to the symbols (221) belonging to the block, these samples may be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or a residual signal) to generate output sample information. The address within the reference picture memory (257) from which the motion compensation prediction unit (253) obtains the prediction samples may be controlled by a motion vector, and the motion vector is in the form of symbols (221) for use by the motion compensation prediction unit (253), and the symbols (221) may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values obtained from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.
[0044] The output samples of the aggregator (255) can undergo various loop filtering techniques in the loop filter unit (256). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the decoded video sequence (also referred to as the decoded video bitstream), and these parameters can be obtained by the loop filter unit (256) as symbols (221) from the parser (220). Video compression can also respond to meta-information obtained during the decoding of the (in decoding order) previous part of the decoded picture or decoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0045] The output of the loop filter unit (256) can be a sample stream, which can be output to the rendering device (212) and stored in the reference picture memory (257) for use in future inter-picture prediction.
[0046] Once fully reconstructed, some decoded pictures can be used as reference pictures for future prediction. For example, once the decoded picture corresponding to the current picture is fully reconstructed and the decoded picture is identified as a reference picture (by, for example, the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting to reconstruct the subsequent decoded pictures.
[0047] The video decoder (210) can perform decoding operations according to a predetermined video compression technique or standard such as ITU-T Recommendation H.265. In the sense that the decoded video sequence follows both the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the decoded video sequence can conform to the syntax specified by the video compression technique or standard being used. Specifically, a profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under that profile. For compliance, it is also required that the complexity of the decoded video sequence be within the bounds defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further restricted by the hypothetical reference decoder (HRD) specification and the metadata of the HRD buffer management signaled in the decoded video sequence.
[0048] In one aspect, a receiver (231) may receive additional (redundant) data along with encoded video. The additional data may be included as part of a decoded video sequence. A video decoder (210) may use the additional data to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.
[0049] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) may be used in place of Figure 1 the video encoder (103) in the example of
[0050] The video encoder (303) may receive video samples from a video source (301) that may capture video images to be encoded by the video encoder (303) (which is not Figure 3 part of the electronic device (320) in the example of
[0051] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 YCrCb, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared video. In a video conferencing system, the video source (301) may be a camera device that captures local image information as a video sequence. Video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be organized as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The following description focuses on the samples.
[0052] According to one aspect, the video encoder (303) can decode and compress pictures of a source video sequence into a decoded video sequence (343) in real time or under any other time constraints required. Implementing an appropriate decoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units described below and is functionally coupled to the other functional units. The coupling is not depicted for the sake of brevity. The parameters set by the controller (350) can include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization techniques, etc.), picture size, group of pictures (group of pictures, GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions that belong to the video encoder (303) optimized for a certain system design.
[0053] In some aspects, the video encoder (303) is configured to operate in a codec loop. As an oversimplified description, in an example, the codec loop may include a source decoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be decoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to what a (remote) decoder would also create. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces a bit-accurate result that is independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.
[0054] The operation of the "local" decoder (333) can be combined with the above Figure 2 The operation of the "remote" decoder described in detail is identical to that of the video decoder (210). However, reference is also briefly made to Figure 2 , when symbols are available and the entropy decoder (345) and the parser (220) can losslessly encode / decode the symbols into a decoded video sequence, the entropy decoding portion of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).
[0055] In one aspect, decoder techniques other than parsing / entropy decoding present in the decoder exist in the corresponding encoder in the same or substantially the same functional form. Thus, the disclosed subject matter focuses on decoder operations. Since encoder techniques are contrary to the fully described decoder techniques, the description of encoder techniques can be simplified. In certain aspects, a more detailed description is provided below.
[0056] During operation, in some examples, a source decoder (330) may perform motion-compensated predictive decoding that predicts an input picture by referring to one or more previously decoded pictures designated as "reference pictures" from a video sequence. In this way, a decoding engine (332) decodes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a prediction reference for the input picture.
[0057] A local video decoder (333) may decode decoded video data of a picture that can be designated as a reference picture based on symbols created by the source decoder (330). The operation of the decoding engine (332) may advantageously be a lossy process. When the decoded video data can be decoded at a video decoder ( Figure 3 not shown), the reconstructed video sequence may generally be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that can be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in a reference picture memory (334). In this way, a video encoder (303) may locally store a copy of the reconstructed reference picture that has common content (in the absence of transmission errors) with the reconstructed reference picture to be obtained by a remote video decoder.
[0058] A predictor (335) may perform a prediction search for the decoding engine (332). That is, for a new picture to be decoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can be used as a suitable prediction reference for the new picture. The predictor (335) may operate on a per-pixel block basis of sample blocks to find an appropriate prediction reference. In some cases, as determined by the search result obtained by the predictor (335), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (334).
[0059] A controller (350) may manage the decoding operations of the source decoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0060] The outputs of all the above functional units can be entropy - coded in an entropy coder (345). The entropy coder (345) converts the symbols into a coded video sequence by applying lossless compression to the symbols generated by the various functional units according to techniques such as Huffman coding, variable - length coding, arithmetic coding, etc.
[0061] The transmitter (340) can buffer the coded video sequence created by the entropy coder (345) in preparation for transmission via a communication channel (360), which can be a hardware / software link to a storage device storing the coded video data. The transmitter (340) can merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (source not shown).
[0062] The controller (350) can manage the operation of the video encoder (303). During coding, the controller (350) can assign a specific coded picture type to each coded picture, which may affect the coding techniques that can be applied to the corresponding picture. For example, pictures can generally be assigned to one of the following picture types:
[0063] An intra - picture (I - picture) can be coded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra - pictures, including, for example, instant decoder refresh (“IDR”) pictures.
[0064] A predictive picture (P - picture) can be coded and decoded using intra - prediction or inter - prediction that uses motion vectors and reference indices to predict the sample values of each block.
[0065] A bi - predictive picture (B - picture) can be coded and decoded using intra - prediction or inter - prediction that uses two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0066] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and decoded block by block. These blocks can be predictively decoded with reference to other (already decoded) blocks, which are determined by the decoding assignments applied to the corresponding pictures of the blocks. For example, blocks of an I picture can be non-predictively decoded, or blocks of an I picture can be predictively decoded (spatial prediction or intra prediction) with reference to already decoded blocks of the same picture. Pixel blocks of a P picture can be predictively decoded with reference to a previously decoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture can be predictively decoded with reference to one or two previously decoded reference pictures via spatial prediction or via temporal prediction.
[0067] The video encoder (303) can perform encoding operations according to a predetermined video decoding technology or standard such as ITU-T Recommendation H.265. In the operation of the video encoder (303), the video encoder (303) can perform various compression operations, including predictive decoding operations that utilize the temporal redundancy and spatial redundancy in the input video sequence. Thus, the decoded video data can conform to the syntax specified by the video decoding technology or video decoding standard used.
[0068] In one aspect, the transmitter (340) can transmit additional data along with the encoded video. The source decoder (330) can include such data as part of the decoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures, and slices, SEI messages, VUI parameter set fragments, etc.
[0069] Video can be captured as multiple source pictures (video pictures) in a time series. Intra picture prediction (commonly abbreviated as intra prediction) utilizes the spatial correlation within a given picture, while inter picture prediction utilizes the (temporal or other) correlation between pictures. In an example, a particular picture in encoding / decoding, which is called the current picture, is segmented into blocks. When a block in the current picture is similar to a reference block in a previously decoded and still buffered reference picture in the video, the block in the current picture can be decoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension that identifies the reference picture.
[0070] In some aspects, bidirectional prediction techniques can be used for inter-picture prediction. According to bidirectional prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture can be decoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.
[0071] In addition, merge mode techniques can be used for inter-picture prediction to improve decoding efficiency.
[0072] According to some aspects of the present disclosure, prediction such as inter-picture prediction and intra-picture prediction is performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU can be split into one 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. According to temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Generally, each PU includes one luminance prediction block (PB) and two chrominance PBs. In one aspect, prediction operations in decoding (encoding / decoding) are performed on a prediction block basis. Using a luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0073] Note that any suitable technique can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In one aspect, one or more integrated circuits can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In another aspect, one or more processors executing software instructions can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210).
[0074] Aspects of the present disclosure provide techniques for decoder-side quantization shift offset prediction. These techniques are used in some examples to predict quantization shift offsets at the decoder side for image and video coding, and can achieve image quality improvement.
[0075] In a video codec, techniques such as transformation, quantization, etc. are used to reduce redundancy in a video signal. For example, transformation techniques can reduce redundancy in a video signal by decorrelating, and quantization techniques can reduce the data represented by transform coefficients by reducing precision, ideally by removing only imperceptible details and thus reducing irrelevance in the data.
[0076] In some examples, transformation decorrelates a signal by transforming the signal from the spatial domain to the transform domain (usually the frequency domain) using a suitable transform basis. For example, the transformation is applied to a prediction residual (regardless of whether it is from inter-picture prediction or intra-picture prediction), i.e., the difference between the prediction and the original input video signal. In the transform domain, the basic information is usually concentrated in a few coefficients. At the decoder, an inverse transform needs to be applied to reconstruct the residual samples.
[0077] Generally, quantization is used to reduce the precision of an input value or a set of input values in order to reduce the amount of data required to represent these values. In some examples, quantization is typically applied to individual transformed residual samples (e.g., transform coefficients), resulting in integer coefficient levels. The transform processing is applied at the encoder. At the decoder, the corresponding processing is called inverse quantization or simply scaling, which restores the original value range without regaining precision.
[0078] In some related video and image codecs, a predefined quantization shift offset ρ* can be applied to the transform coefficient values. In some examples, the quantization shift offset ρ* is used to control the quantization dead zone. For example, at the decoder side, the reconstructed transform coefficient value y i is shifted using the predefined quantization shift offset ρ* according to, for example, Equation (1):
[0079]
[0080] In accordance with some aspects of the present disclosure, the quantization shift offset ρ* of each non-zero coefficient or some subset of all non-zero coefficients can be determined (predicted) at the decoder side using the techniques described in the present disclosure. In some examples, the sign and / or value of the quantization shift offset of one or more non-zero transform coefficients can be determined at the decoder side. For example, the encoder / decoder can determine a plurality of hypotheses for decoder-side quantization shift offset prediction, where the hypotheses in the plurality of hypotheses correspond to potential quantization shift offset settings in the transform domain of the current block. The encoder / decoder can calculate cost values respectively associated with the plurality of hypotheses, select a particular hypothesis from the plurality of hypotheses according to the cost values, determine the quantization shift offset of one or more transform coefficients in the transform domain based on the particular hypothesis, and reconstruct the transform coefficients based on the one or more quantization shift offsets.
[0081] In some examples, a two-step process is applied to obtain the quantization shift offset. For example, in the first step, several hypotheses of quantization shift offset settings are respectively used for reconstruction. In the second step, the final quantization shift offset is determined by selecting one of the hypotheses in the first step based on the decoding information available at the decoder side.
[0082] The present disclosure describes various techniques that can be used in the first and second steps of the two-step process.
[0083] In some examples, in the first step of the two-step process, the sign of the quantization shift offset is determined for the entire block (e.g., transform block, transform sub-block, etc.) or for each non-zero transform coefficient in the block or a subset of the non-zero transform coefficients in the block.
[0084] In an example, the sign of the quantization shift offset is determined for the entire block. Figure 4 A diagram of hypotheses in some examples is shown. Figure 4 A 4×4 block (410) with a positive sign is shown as "hypothesis 1" and a 4×4 block (420) with a negative sign is shown as "hypothesis 2". In Figure 4 , when using "hypothesis 1", a positive sign is used for the quantization shift offset of the reconstructed transform coefficients; otherwise, when using "hypothesis 2", a negative sign is used for the quantization shift offset of the reconstructed transform coefficients at the decoder side. Note that the absolute values of the quantization shift offsets can be the same or can be different.
[0085] In another example, the sign of the quantization shift offset is determined only for the DC coefficient in the block. The other transform coefficients in the block can utilize a fixed quantization shift offset or no quantization is applied.
[0086] Figure 5 A diagram of hypotheses in some examples is shown. Figure 5A 4×4 block (510) with a positive sign at the DC coefficient is shown as "hypothesis 1", and a 4×4 block (520) with a negative sign at the DC coefficient is shown as "hypothesis 2". When using "hypothesis 1", the positive sign is used to quantify the shift offset for reconstructing the DC coefficient; otherwise, when using "hypothesis 2", a negative sign is applied to the quantified shift offset at the decoder side to reconstruct the DC coefficient.
[0087] In another example, the sign of the quantization shift offset is determined for a subset, e.g., n transform coefficients in a block. The other transform coefficients in the block can utilize a fixed quantization shift offset or no quantization shift is applied. The number of hypotheses can be up to 2 n .
[0088] Figure 6 A diagram showing hypotheses in some examples. In Figure 6 the example, hypotheses are formed for the signs of the transform coefficients in the upper left 2×2 sub-block of a 4×4 block. Figure 6 A 4×4 block (601) (with signs "++++" in the 2×2 sub-block) is shown as "hypothesis 1", a 4×4 block (602) (with signs "+++-" in the 2×2 sub-block) is shown as "hypothesis 2", a 4×4 block (603) (with signs "++-+" in the 2×2 sub-block) is shown as "hypothesis 3", and a 4×4 block (616) (with signs "----" in the 2×2 sub-block) is shown as "hypothesis 16". Note that in Figure 6 other hypotheses are not shown, such as "hypothesis 4" with signs "+-++" in the 2×2 sub-block, "hypothesis 5" with signs "-+++" in the 2×2 sub-block, "hypothesis 6" with signs "--++" in the 2×2 sub-block, "hypothesis 7" with signs "-+-+" in the 2×2 sub-block, "hypothesis 8" with signs "-++-" in the 2×2 sub-block, "hypothesis 9" with signs "+--+" in the 2×2 sub-block, "hypothesis 10" with signs "++--" in the 2×2 sub-block, "hypothesis 11" with signs "+-+-" in the 2×2 sub-block, "hypothesis 12" with signs "---+" in the 2×2 sub-block, "hypothesis 13" with signs "--+-" in the 2×2 sub-block, "hypothesis 14" with signs "-+--" in the 2×2 sub-block, "hypothesis 15" with signs "+---" in the 2×2 sub-block. Note that when n = 4, the number of hypotheses can be up to 2 4 = 16.
[0089] In some examples, the value of the quantization shift offset is determined for the entire block or for each non-zero transform coefficient in the block or for a subset of non-zero coefficients in the block.
[0090] In the example, the quantization shift offset value is determined for the entire block.
[0091] Figure 7 Shows a hypothetical diagram in some examples. Figure 7 A 4×4 block (710) with a quantization shift offset of the entire block equal to a is shown as "Hypothesis 1", and a 4×4 block (720) with a quantization shift offset of the entire block equal to b is shown as "Hypothesis 2". In Figure 7 the example, when using "Hypothesis 1", the quantization shift offset ρ* is equal to a; otherwise, when using "Hypothesis 2", the quantization shift offset ρ* is equal to b.
[0092] According to aspects of the present disclosure, for example based on a discontinuity metric, a cost J is calculated for each hypothesis according to the pixels within the current reconstruction block and adjacent reconstruction blocks. The hypothesis providing the lowest cost J is used as the final quantization shift offset setting.
[0093] Figure 8 Shows a diagram of the current block (810) in some examples. The current block (810) can be the current reconstruction block according to the hypothesis. Figure 8 Also shown in gray are the adjacent reconstruction pixels associated with the current block (810). Note that the adjacent reconstruction pixels are reconstructed according to the decoding information of the adjacent blocks. In Figure 8 it, p(x, y) represents the reconstructed pixel value at position (x, y). When x or y is negative, p(x, y) indicates the pixel value of the adjacent reconstruction block.
[0094] In the example, the cost of hypothesis i is calculated as a discontinuity metric according to Equation (2):
[0095]
[0096] where M and N indicate the number of pixels in the columns and rows of the block.
[0097] In another example, the cost of hypothesis i is calculated as a discontinuity metric according to Equation (3):
[0098]
[0099] In another example, the cost of hypothesis i is calculated as a derivative discontinuity metric according to Equation (4):
[0100]
[0101] According to some aspects of the present disclosure, a flag can be signaled in the bitstream at the sequence / picture / slice / tile / block level to control whether to use decoder quantization shift offset prediction.
[0102] In some examples, the same block-level quantization shift offset prediction method is applied at the encoder side, and the flag is signaled in the bitstream to control the use of the quantization shift offset prediction method at the decoder side.
[0103] In an example, when the costs of all hypotheses are greater than a threshold T, a flag indicating that the decoder-side quantization offset prediction is not used is signaled in the bitstream.
[0104] In some examples, when the sequence / picture / slice / tile / block does not use quantization shift offset prediction, a fixed offset value is used for reconstruction or quantization shift is not applied.
[0105] In an example, Figure 7 , two hypotheses are considered at the encoder side. "Hypothesis 1" indicates that the quantization shift offset ρ* of the entire block is equal to a, and "Hypothesis 2" indicates that the quantization shift offset ρ* of the entire block is equal to b. The costs J1 and J2 of the two hypotheses are calculated using Equation (2). If both J1 and J2 are greater than the threshold T, a control flag with a "false" value is signaled in the bitstream at the block level. Otherwise, a control flag with a "true" value is signaled in the bitstream at the block level. At the decoder side, when the control flag has a "true" value, the quantization shift offset prediction method is applied. Otherwise, when the control flag has a "false" value, a fixed offset value is used for reconstruction or quantization shift is not applied.
[0106] Figure 9 FIG. shows a flowchart outlining a process (900) according to aspects of the present disclosure. The process (900) may be used in a video decoder. In various aspects, the process (900) is performed by processing circuitry, such as processing circuitry that performs the functions of video decoder (110), processing circuitry that performs the functions of video decoder (210), etc. In some aspects, the process (900) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry performs the process (900). The process begins at (S901) and proceeds to (S910).
[0107] At (S910), a bitstream including decoding information of a current block is received.
[0108] At (S920), a plurality of hypotheses for decoder-side quantization shift offset prediction are determined. The hypotheses among the plurality of hypotheses correspond to potential quantization shift offset settings in the transform domain of the current block.
[0109] At (S930), cost values respectively associated with the plurality of hypotheses are calculated.
[0110] At (S940), a specific hypothesis is selected from the plurality of hypotheses according to the cost values.
[0111] At (S950), one or more quantization shift offsets of transform coefficients in a transform domain are determined based on specific assumptions.
[0112] At (S960), the transform coefficients are reconstructed based on one or more quantization shift offsets. For example, the quantized coefficients can be decoded from a bitstream, dequantization can be performed to obtain dequantized coefficients, and the dequantized coefficients can be combined with the quantization shift offsets to obtain the transform coefficients.
[0113] At (S970), a residual in a spatial domain of a current block is calculated based on the transform coefficients in the transform domain. For example, an inverse transform can be performed according to the transform coefficients to obtain the residual in the spatial domain.
[0114] At (S980), the current block is reconstructed according to the residual in the spatial domain. For example, a predictor of the current block can be combined with the residual to obtain the reconstructed current block.
[0115] According to aspects of the present disclosure, the multiple assumptions include assumptions for setting signs of at least one quantization shift offset for non-zero transform coefficients. In some examples, the multiple assumptions include a first assumption for setting a positive sign for a quantization shift offset of non-zero transform coefficients in a transform block, and a second assumption including setting a negative sign for the quantization shift offset of non-zero transform coefficients in the transform block. In an example, the multiple assumptions include a first assumption for setting a positive sign for a first quantization shift offset of a DC coefficient in a transform block, and a second assumption including setting a negative sign for the first quantization shift offset of the DC coefficient in the transform block.
[0116] In some examples, the multiple assumptions include potential sign combinations of quantization shift offsets of a subset of non-zero transform coefficients in a transform block. For example, in Figure 6 an example, 16 assumptions correspond to 16 sign combinations of quantization shift offsets of 2×2 non-zero transform coefficients at the upper left corner of a transform block.
[0117] According to aspects of the present disclosure, the multiple assumptions include assumptions for setting values for at least one quantization shift offset for non-zero transform coefficients. In some examples, the multiple assumptions include a first assumption for setting a first value for a quantization shift offset of non-zero transform coefficients in a transform block, and a second assumption including setting a second value for the quantization shift offset of non-zero transform coefficients in the transform block.
[0118] In some examples, a specific assumption with the lowest cost value is selected.
[0119] In aspects according to the present disclosure, a cost value associated with a hypothesis is calculated based on reconstructed pixels in a current block obtained due to the hypothesis and one or more neighboring reconstructed blocks. In some examples, the cost value is calculated, for example, using Equation (2) and Equation (3), and the cost value measures the discontinuity between the reconstructed pixels in the current block obtained due to the hypothesis and one or more neighboring reconstructed blocks. In some examples, the cost value is calculated, for example, using Equation (4), and the cost value measures the derivative discontinuity between the reconstructed pixels in the current block obtained due to the hypothesis and one or more neighboring reconstructed blocks.
[0120] In some examples, a flag is decoded from the bitstream, and the flag indicates whether decoder-side quantization shift offset prediction is performed. The flag is signaled at an appropriate level, such as sequence level, picture level, slice level, tile level, block level, etc. In some examples, when the flag indicates that decoder-side quantization shift offset prediction is not used, a fixed quantization shift offset is determined.
[0121] Then, the process proceeds to (S999) and terminates.
[0122] The process (900) can be adjusted appropriately. The steps in the process (900) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.
[0123] Figure 10 A flowchart outlining a process (1000) according to aspects of the present disclosure is shown. The process (1000) can be used in a video encoder. In various aspects, the process (1000) is performed by processing circuitry, such as processing circuitry that performs the functions of video encoder (103), processing circuitry that performs the functions of video encoder (303), etc. In some examples, the process (1000) is implemented as software instructions, and thus when the processing circuitry executes the software instructions, the processing circuitry performs the process (1000). The process begins at (S1001) and proceeds to (S1010).
[0124] At (S1010), it is determined to use decoder-side quantization shift offset prediction for the current block. In an example, the predictor of the current block is determined, and the residual block of the current block in the spatial domain is calculated. Then, a transform is applied to the residual block to obtain transform coefficients. Additionally, a quantization shift offset is determined, and quantization is applied to obtain quantized coefficients, and the quantized coefficients can be appropriately encoded into the bitstream. In some examples, the quantization shift offset is determined according to decoder-side quantization shift offset prediction.
[0125] At (S1020), a plurality of hypotheses for decoder-side quantization shift offset prediction are determined. The hypotheses among the plurality of hypotheses correspond to potential quantization shift offset settings in the transform domain of the current block.
[0126] At (S1030), cost values associated with multiple hypotheses are calculated.
[0127] At (S1040), a specific hypothesis is selected from the multiple hypotheses based on the cost values.
[0128] At (S1050), one or more quantization shift offsets of transform coefficients in the transform domain are determined based on the specific hypothesis.
[0129] At (S1060), the transform coefficients are reconstructed based on the one or more quantization shift offsets.
[0130] At (S1070), a residual in the spatial domain of the current block is calculated based on the transform coefficients in the transform domain.
[0131] At (S1080), the current block is reconstructed according to the residual in the spatial domain. The current block is encoded accordingly.
[0132] According to aspects of the present disclosure, the multiple hypotheses include hypotheses that set signs for at least one quantization shift offset of non - zero transform coefficients. In some examples, the multiple hypotheses include a first hypothesis that sets a positive sign for the quantization shift offset of non - zero transform coefficients in a transform block, and a second hypothesis that includes setting a negative sign for the quantization shift offset of non - zero transform coefficients in the transform block. In an example, the multiple hypotheses include a first hypothesis that sets a positive sign for a first quantization shift offset of a DC coefficient in a transform block, and a second hypothesis that includes setting a negative sign for the first quantization shift offset of the DC coefficient in the transform block.
[0133] In some examples, the multiple hypotheses include potential sign combinations of quantization shift offsets of a subset of non - zero transform coefficients in a transform block. For example, in Figure 6 an example, 16 hypotheses correspond to 16 sign combinations of quantization shift offsets of 2×2 non - zero transform coefficients at the upper - left corner of a transform block.
[0134] According to aspects of the present disclosure, the multiple hypotheses include hypotheses that set values for at least one quantization shift offset of non - zero transform coefficients. In some examples, the multiple hypotheses include a first hypothesis that sets a first value for the quantization shift offset of non - zero transform coefficients in a transform block, and a second hypothesis that includes setting a second value for the quantization shift offset of non - zero transform coefficients in the transform block.
[0135] In accordance with aspects of the present disclosure, a cost value associated with a hypothesis is calculated based on reconstructed pixels in a current block obtained due to the hypothesis and one or more neighboring reconstructed blocks. In some examples, the cost value is calculated, for example, using Equations (2) and (3), and the cost value measures the discontinuity between the reconstructed pixels in the current block obtained due to the hypothesis and one or more neighboring reconstructed blocks. In some examples, the cost value is calculated, for example, using Equation (4), and the cost value measures the derivative discontinuity between the reconstructed pixels in the current block obtained due to the hypothesis and one or more neighboring reconstructed blocks.
[0136] In some examples, a particular hypothesis is determined to be the hypothesis having the lowest cost value among the cost values. In some examples, when all cost values are higher than a threshold, it is determined not to use decoder-side quantization shift offset prediction, and a flag is signaled in the bitstream to indicate that decoder-side quantization shift offset prediction is not used. In some examples, the flag can be signaled at an appropriate level, such as sequence level, picture level, slice level, tile level, block level, etc. In some examples, when decoder-side quantization shift offset prediction is not used, a fixed quantization shift offset is used.
[0137] Then, the processing proceeds to (S1099) and terminates.
[0138] The processing (1000) can be adjusted appropriately. Steps in the processing (1000) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.
[0139] In accordance with aspects of the present disclosure, a method of processing visual media data is provided. In this method, a conversion between a visual media file and a bitstream of visual media data is performed in accordance with formatting rules. For example, the bitstream can be a bitstream decoded / encoded by any one of the decoding and / or encoding methods described herein. The formatting rules can specify one or more constraints of the bitstream and / or one or more processes to be performed by a decoder and / or an encoder.
[0140] In an example, the bitstream includes decoding information for one or more pictures. The format rules specify: determining a plurality of hypotheses for predicting a quantization shift offset on the decoder side, where a hypothesis in the plurality of hypotheses corresponds to a potential quantization shift offset setting in the transform domain of a current block. The format rules also specify: calculating cost values respectively associated with the plurality of hypotheses, calculating the cost value associated with a hypothesis based on reconstructed pixels in the current block obtained based on the hypothesis and one or more adjacent reconstructed blocks. Further, the format rules specify: when the lowest cost value is less than a threshold, selecting a specific hypothesis having the lowest cost value, determining one or more quantization shift offsets of transform coefficients in the transform domain of the current block based on the specific hypothesis, reconstructing the transform coefficients based on the one or more quantization shift offsets, calculating a residual in the spatial domain of the current block based on the transform coefficients in the transform domain, and reconstructing the current block based on the residual in the spatial domain.
[0141] The above techniques can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 11 FIG. shows a computer system (1100) suitable for implementing certain aspects of the disclosed subject matter.
[0142] The computer software can be decoded using any suitable machine code or computer language, and the machine code or computer language can be processed through mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. or executed through interpretation, microcode execution, etc.
[0143] The instructions can be executed on various types of computers or their components (including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.).
[0144] Figure 11 The components shown in FIG. for the computer system (1100) are exemplary in nature and are not intended to imply any limitation on the scope of use or functionality of the computer software for implementing aspects of the present disclosure. The configuration of the components should also not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary aspects of the computer system (1100).
[0145] A computer system (1100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs by one or more human users through, for example, tactile inputs (such as keystrokes, swipes, data glove movements), audio inputs (such as speech, clapping), visual inputs (such as gestures), and olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as speech, music, ambient sounds), images (such as scanned images, photographic images obtained from a still-image camera device), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).
[0146] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1101), mouse (1102), touchpad (1103), touch screen (1110), data glove (not shown), joystick (1105), microphone (1106), scanner (1107), camera device (1108).
[0147] The computer system (1100) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through the touch screen (1110), data glove (not shown), or joystick (1105), but there may also be tactile feedback devices that do not serve as input devices), audio output devices (such as speakers (1109), headphones (not depicted)), visual output devices (such as screens (1110), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which may output two-dimensional visual output or more than three-dimensional output through means such as stereoscopic graphics output; virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)), and printers (not depicted).
[0148] The computer system (1100) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1120) with media such as CD / DVD (1121), thumb drives (1122), removable hard disk drives or solid-state drives (1123), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0149] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not encompass a transmission medium, a carrier wave, or other transient signals.
[0150] The computer system (1100) may also include an interface (1154) to one or more communication networks (1155). The network may be, for example, wireless, wired, optical. The network may also be local, wide area, urban, vehicular, and industrial, real-time, delay-tolerant, etc. Examples of networks include: local area networks, such as Ethernet, wireless LAN; cellular networks, including GSM, 3G, 4G, 5G, LTE, etc.; television cable or wireless wide area digital networks, including cable television, satellite television, and terrestrial broadcast television; vehicular and industrial networks, including CANBus, etc. Some networks typically require an external network interface adapter attached to certain common data ports or peripheral buses (1149) (such as, for example, the USB port of the computer system (1100)); other networks are typically integrated into the core of the computer system (1100) by attaching to a system bus as described below (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smart phone computer system). The computer system (1100) may communicate with other entities by using any of these networks. Such communication may be unidirectional receive-only (e.g., broadcast television), unidirectional send-only (e.g., CANbus to certain CANbus devices), or bidirectional, such as to other computer systems using local digital networks or wide area digital networks. Specific protocols and protocol stacks may be used on each of these networks and network interfaces as described above.
[0151] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1140) of the computer system (1100).
[0152] The kernel (1140) may include one or more central processing units (CPUs) (1141), a graphics processing unit (GPU) (1142), a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) (1143), a hardware accelerator for specific tasks (1144), a graphics adapter (1150), etc. These devices, together with a read-only memory (ROM) (1145), a random access memory (1146), an internal mass storage device such as an internal non-user-accessible hard disk drive, SSD, etc. (1147), may be connected via a system bus (1148). In some computer systems, the system bus (1148) may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1149) to the system bus (1148) of the kernel. In an example, a screen (1110) may be connected to the graphics adapter (1150). The architecture of the peripheral bus includes PCI, USB, etc.
[0153] The CPU (1141), GPU (1142), FPGA (1143), and accelerator (1144) may execute certain instructions that, when combined, may constitute the above-mentioned computer code. The computer code may be stored in the ROM (1145) or the RAM (1146). Transient data may also be stored in the RAM (1146), while permanent data may be stored in, for example, the internal mass storage device (1147). Fast storage and retrieval of any storage device in the storage devices may be achieved by using a cache memory that may be closely associated with one or more CPUs (1141), GPUs (1142), mass storage devices (1147), ROM (1145), RAM (1146), etc.
[0154] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the medium and the computer code may be of the type well-known and available to those skilled in the field of computer software.
[0155] By way of example and not limitation, a computer system having an architecture (1100), and in particular a kernel (1140), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software included in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage devices introduced above, as well as certain storage devices of the kernel (1140) having a non-transitory nature, such as a mass storage device (1147) internal to the kernel or a ROM (1145). Software implementing aspects of the present disclosure can be stored in such devices and executed by the kernel (1140). Depending on specific requirements, the computer-readable media can include one or more memory devices or chips. The software can cause the kernel (1140), and in particular the processors therein (including the CPU, GPU, FPGA, etc.), to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in the RAM (1146) and modifying such data structures in accordance with the processes defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of being implemented logically hardwired or otherwise in circuitry (such as an accelerator (1144)), which can operate in place of or in conjunction with the software to perform specific processes or specific portions of specific processes described herein. In appropriate cases, references to software can include logic, and vice versa, references to logic can include software. In appropriate cases, references to computer-readable media can include circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry containing logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0156] The use of "at least one of" or "one of" in the present disclosure is intended to include any one or combination of the recited elements. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include A or B or (A and B). Where applicable, the use of "one of" does not exclude any combination of the recited elements, such as when the elements are not mutually exclusive.
[0157] Although the present disclosure has described several examples of aspects, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be understood that those skilled in the art can envision numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the disclosure and thus fall within the spirit and scope of the disclosure.
Claims
1. A method for processing visual media data, the method comprising: Performing a conversion between a visual media file and a bitstream of visual media data according to format rules, wherein: The bitstream includes decoding information of one or more pictures; and The format rules specify: Determining a plurality of hypotheses for decoder-side quantization shift offset prediction, where a hypothesis among the plurality of hypotheses corresponds to a potential quantization shift offset setting in the transform domain of a current block; Calculating cost values respectively associated with the plurality of hypotheses, the cost value associated with a hypothesis among the plurality of hypotheses being calculated based on reconstructed pixels in the current block obtained based on the hypothesis and one or more adjacent reconstructed blocks; When the lowest cost value is less than a threshold, selecting a specific hypothesis having the lowest cost value; Determining one or more quantization shift offsets of transform coefficients in the transform domain of the current block based on the specific hypothesis; Reconstructing the transform coefficients based on the one or more quantization shift offsets; Calculating a residual in the spatial domain of the current block based on the transform coefficients in the transform domain; and Reconstructing the current block according to the residual in the spatial domain.
2. An apparatus for video decoding, comprising processing circuitry configured to: Receive a bitstream including decoding information of a current block; Determine a plurality of hypotheses for decoder-side quantization shift offset prediction, where a hypothesis among the plurality of hypotheses corresponds to a potential quantization shift offset setting in the transform domain of the current block; Calculate cost values respectively associated with the plurality of hypotheses; Selecting a specific hypothesis from the plurality of hypotheses according to the cost values; Determining one or more quantization shift offsets of transform coefficients in the transform domain based on the specific hypothesis; Reconstructing the transform coefficients based on the one or more quantization shift offsets; Calculating a residual in the spatial domain of the current block based on the transform coefficients in the transform domain; And Reconstructing the current block according to the residual in the spatial domain.
3. The apparatus according to claim 2, wherein The plurality of hypotheses include: A hypothesis of setting a sign for at least one quantization shift offset of non-zero transform coefficients.
4. The apparatus according to claim 3, wherein The plurality of hypotheses include: A first hypothesis of setting a positive sign for the quantization shift offset of non-zero transform coefficients in a transform block; and A second hypothesis of setting a negative sign for the quantization shift offset of the non-zero transform coefficients in the transform block.
5. The device according to claim 3, wherein The plurality of hypotheses include: A first hypothesis of setting a positive sign for a first quantization shift offset of a DC coefficient in a transform block; and A second hypothesis of setting a negative sign for the first quantization shift offset of the DC coefficient in the transform block.
6. The apparatus according to claim 3, wherein The plurality of hypotheses include potential sign combinations of quantization shift offsets of a subset of non-zero transform coefficients in a transform block.
7. The apparatus according to claim 2, wherein, The plurality of hypotheses include: A hypothesis of setting a value for at least one quantization shift offset of non-zero transform coefficients.
8. The apparatus according to claim 2, wherein, The plurality of hypotheses include: A first hypothesis of setting a first value for the quantization shift offset of non-zero transform coefficients in a transform block; and A second hypothesis of setting a second value for the quantization shift offset of the non-zero transform coefficients in the transform block.
9. The device according to any one of claims 2 to 8, wherein The processing circuitry is configured to: Select the specific hypothesis having the lowest cost value.
10. The apparatus according to any one of claims 2 to 9, wherein, The processing circuitry is configured to: calculate a cost value associated with the hypothesis based on the reconstructed pixels in the current block obtained due to the hypothesis and one or more neighboring reconstructed blocks.
11. The apparatus according to claim 10, wherein, The processing circuitry is configured to: calculate the cost value that measures the discontinuity between the reconstructed pixels in the current block obtained due to the hypothesis and the one or more neighboring reconstructed blocks.
12. The apparatus according to claim 10, wherein, The processing circuitry is configured to: calculate the cost value that measures the derivative discontinuity between the reconstructed pixels in the current block obtained due to the hypothesis and the one or more neighboring reconstructed blocks.
13. The device according to any one of claims 2 to 12, wherein, The processing circuitry is configured to: decode a flag from the bitstream, the flag indicating whether to perform decoder-side quantization shift offset prediction.
14. The apparatus according to claim 13, wherein signal the flag at least at one of a sequence level, a picture level, a slice level, a tile level, and a block level, and the processing circuitry is configured to: determine to use a fixed quantization shift offset when the flag indicates not to use decoder-side quantization shift offset prediction.
15. A method for video coding, comprising: determining to use decoder-side quantization shift offset prediction for a current block; determining a plurality of hypotheses for the decoder-side quantization shift offset prediction, where a hypothesis in the plurality of hypotheses corresponds to a potential quantization shift offset setting in a transform domain of the current block; calculating cost values respectively associated with the plurality of hypotheses; when the lowest cost value is less than a threshold, selecting a specific hypothesis having the lowest cost value from the plurality of hypotheses; determining one or more quantization shift offsets of transform coefficients in the transform domain based on the specific hypothesis; reconstructing the transform coefficients based on the one or more quantization shift offsets; calculating a residual in a spatial domain of the current block based on the transform coefficients in the transform domain; and reconstructing the current block according to the residual in the spatial domain.