System and Method for Transform Coefficient Sign Prediction and Coding

By predicting coefficient signs using two distinct techniques for different sets of transform coefficients, the method addresses the inefficiencies in existing video coding techniques, reducing computational and bandwidth costs while enhancing coding efficiency.

JP2025516422APending Publication Date: 2025-05-30TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024527539
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-28
Filing Date
2023-03-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing video coding techniques face challenges in efficiently coding coefficient signs of transform coefficients, which requires significant computational resources and bandwidth, especially when predicting signs using different techniques for different coefficients.

Method used

The method involves predicting the coefficient signs of transform coefficients using two different techniques for distinct sets of coefficients, allowing for reduced computational cost, processing time, and bandwidth usage during video coding.

Benefits of technology

This approach effectively reduces the computational burden and bandwidth requirements for video coding by selectively applying different sign prediction techniques, thereby improving coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516422000001_ABST
    Figure 2025516422000001_ABST
Patent Text Reader

Abstract

The various embodiments described herein include a method and system for coding video. In one aspect, the method includes obtaining video data including a plurality of blocks including a first block, and determining a plurality of transform coefficients associated with the first block. The method further includes predicting, using a first technique, a respective coefficient sign for a first set of the plurality of transform coefficients, and predicting, using a second technique different from the first technique, a respective coefficient sign for a second set of the plurality of transform coefficients. The method also includes reconstructing the first block based on the plurality of transform coefficients and the respective coefficient signs predicted for the first and second sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed embodiments generally relate to video coding including, but not limited to, systems and methods for coefficient sign coding for conversion coefficients.

Background Art

[0002] [Related Application] This application claims priority to U.S. Provisional Patent Application No. 63 / 341,057, titled "Harmonization Between Sign Prediction and Cross-Component Sign Coding," filed on May 12, 2022, the entire disclosure of which is incorporated herein by reference. [Background Art] Digital video is supported by a variety of electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. The electronic devices transmit and receive digital video data over a communication network or otherwise communicate and / or store digital video data in a storage device. Since the bandwidth capacity of the communication network is limited and the memory resources of the storage device are limited, video coding may be used to compress the video data according to one or more video coding standards before communicating or storing the video data.

[0003] Multiple video codec standards have been developed. For example, video coding standards include AOMedia Video 1 (AV1), Versatile Video Coding (VVC), Joint Exploration test Model (JEM), High-Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Expert Group (MPEG) coding. Video coding generally utilizes prediction methods (such as inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in video data. The purpose of video coding is to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality.

[0004] HEVC, also known as H.265, is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC), also known as H.266, is a video compression standard intended as a successor to HEVC. ITU-T and ISO / IEC announced the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AV1 is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verified version 1.0.0 was released along with Errata 1 of the specification. SUMMARY OF THE INVENTION

[0005] As described above, video coding techniques include intra coding. In intra coding, sample values are represented without reference to samples from previously reconstructed reference pictures or other data. In some cases, a picture is spatially re - divided into blocks of samples. When all blocks of samples are coded in an intra mode, that picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and a video session or as a still image. Samples of an intra block can be subjected to a transform, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can reduce / minimize the sample values in the pre - transform domain. In some cases, the smaller the DC value after transformation and the smaller the AC coefficients, the fewer bits are required at a given quantization step size to represent the block after entropy coding. In the case of entropy coding of transform coefficients, the coefficient signs can be coded separately from the magnitude / level (the absolute value of the coefficient value) using a bypass mode. This means that each coefficient sign requires 1 bit to code, which is costly.

[0006] According to some embodiments, a method of video coding is provided. The method includes (i) obtaining video data including a plurality of blocks including a first block, (ii) determining a plurality of transform coefficients associated with the first block, (iii) predicting, using a first technique, each coefficient sign of a first set of the plurality of transform coefficients, and (iv) predicting, using a second technique different from the first technique, each coefficient sign of a second set of the plurality of transform coefficients, and (v) reconstructing the first block based on the plurality of transform coefficients and the predicted respective coefficient signs of the first and second sets.

[0007] According to some embodiments, another method of video coding is provided. The method includes: (i) obtaining video data including a plurality of blocks including a first block including a plurality of elements; and (ii) coding an indicator specifying, for each element of the plurality of elements, whether a sign value of a transform coefficient of the element is the same as a sign value of a transform coefficient of a respective second element.

[0008] According to some embodiments, a computing system such as a streaming system, a server system, a personal computer system, or another electronic device is provided. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.

[0009] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions for execution by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.

[0010] Accordingly, devices and systems comprising a method for coding video are disclosed. Such methods, devices, and systems may complement or replace conventional methods, devices, and systems for video coding.

[0011] The features and advantages described in this specification are not necessarily all. By considering the drawings, specification, and claims provided in this disclosure, some additional features and advantages will become apparent to those skilled in the art. Furthermore, note that the language used in this specification is mainly selected for readability and educational purposes, and is not necessarily selected to describe or limit the subject matter described in this specification.

Brief Description of the Drawings

[0012] To better understand the present disclosure, a more specific description can be obtained by referring to the features of various embodiments, some of which are shown in the accompanying drawings. However, the accompanying drawings merely show the relevant features of the present disclosure and should not necessarily be regarded as limiting. The description allows other valid features to be recognized by those skilled in the art upon reading the present disclosure.

Figure 1

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6

[0013] The present disclosure describes, among other things, predicting coefficient signs during coding of transform blocks. For example, a first set of transform coefficients is selected for sign prediction using a first technique, and a second set of transform coefficients is selected for sign prediction using a second technique. Selecting different techniques for predicting the signs of different transform coefficients reduces the computational cost, processing time, and / or the bandwidth for transmitting the bitstream. For example, the computational cost associated with predicting a sign using other color coefficient signs as context may be less compared to predicting a sign using a hypothesis with adjacent blocks as context. However, other color coefficient signs are not always available for all transform coefficients, and predicting the signs of other transform coefficients using adjacent blocks as context may use less bandwidth than not predicting any signs at all.

[0014] Exemplary Systems and Devices FIG. 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system for use with video-enabled applications such as, for example, video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0015] Source device 102 includes video source 104 (e.g., a camera component or media storage) and encoder component 106. In some embodiments, video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). Encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from video source 104 can have a high data volume compared to the encoded video bitstream 108 generated by encoder component 106. Since the encoded video bitstream 108 has a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth for transmission and less storage space for storage compared to the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., is configured to transmit uncompressed video data to one or more networks 110).

[0016] One or more networks 110 represent any number of networks that carry information between source device 102, server system 112, and / or electronic device 120, and include, for example, wireline (wired) and / or wireless communication networks. One or more networks 110 can exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0017] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is or includes a streaming server (e.g., configured to store and / or deliver video content such as an encoded video stream from source device 102). The server system 112 includes an encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the encoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the encoder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the encoder component 114 is configured to decode an encoded video bitstream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video bitstream 108.

[0018] In some embodiments, the server system 112 functions as a Media-Aware Network Element (MANE). For example, the server system 112 can be configured to prune the encoded video bitstream 108 to potentially adjust different bitstreams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0019] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an outgoing video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., communicatively coupled to an external display device and / or include media storage). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0020] The source device and / or the plurality of electronic devices 120 may be referred to as a "terminal device" or "user device". In some embodiments, one or more of the source device 102 and / or the electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smartphone, a tablet, or a laptop), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0021] In an exemplary operation of communication system 100, source device 102 transmits an encoded video bitstream 108 to server system 112. For example, source device 102 may code a stream of pictures captured by the source device. Server system 112 receives encoded video bitstream 108 and may decode and / or encode encoded video bitstream 108 using coder component 114. For example, server system 112 may apply an encoding that is more suitable for network transmission and / or storage to the video data. Server system 112 may transmit encoded video data 116 (e.g., one or more coded video bitstreams) to one or more of electronic devices 120. Each electronic device 120 may decode encoded video data 116, restore the video pictures, and optionally display them.

[0022] In some embodiments, the transmission described above is unidirectional data transmission. Unidirectional data transmission may be used in media serving applications and the like. In some embodiments, the transmission described above is bidirectional data transmission. Bidirectional data transmission may be used in video conferencing applications and the like. In some embodiments, encoded video bitstream 108 and / or encoded video data 116 are encoded and / or decoded according to any of the video coding / compression standards described herein, such as HEVC, VVC, and / or AV1.

[0023] Figure 2A is a block diagram showing exemplary elements of an encoder component 106 according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 can provide the source video sequence in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCb, or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures that give the appearance of motion when viewed sequentially. The pictures themselves can be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. One of ordinary skill in the art can readily understand the relationship between pixels and samples. The following description focuses on samples.

[0024] The encoder component 106 is configured to code and / or compress pictures of a source video sequence into a coded video sequence 216 in real time or under other timing constraints required by the application. Implementing an appropriate coding speed is one function of the controller 204. In some embodiments, the controller 204 controls other functional units and is functionally coupled to other functional units as described below. The parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or lambda value of rate distortion optimization techniques), picture size, picture group (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller 204, but they may be related to the encoder component 106 that is optimized for a particular system design.

[0025] In some embodiments, the encoder component 106 is configured to operate in a coding loop. In a simplified example, the coding loop includes a source coder 202 (which is responsible for creating symbols such as a symbol stream based on, for example, an input picture to be coded and one or more reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols and creates sample data in a manner similar to a (remote) decoder when the compression between the symbols and the coded video bitstream is reversible. The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream results in bit-accurate results independent of the location of the decoder (local or remote), the content of the reference picture memory 208 is also bit-accurate between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the same sample values as the reference picture samples that the decoder interprets when using prediction during decoding. This principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example, due to channel errors) is known to those skilled in the art.

[0026] The operation of the decoder 210 may be the same as the operation of a remote decoder such as the decoder component 122, which will be described in detail below in connection with FIG. 2B. However, referring briefly to FIG. 2B, since symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder 214 and the parser 254 can be lossless, the entropy decoding part of the decoder component 122, which includes the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.

[0027] The observations that can be made at this point are that any decoder technology other than parsing / entropy decoding present in the decoder must necessarily also be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operation. The description of encoder technology can be omitted since it is the reverse of the decoder technology described comprehensively. Only in certain areas is a more detailed description required and is provided below.

[0028] As part of its operation, source coder 202 can perform motion compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from a video sequence designated as a reference frame. In this way, coding engine 212 codes the difference between a pixel block of the input frame and a pixel block of a reference frame that can be selected as a prediction reference for the input frame. Controller 204 can manage the coding operation of source coder 202, including, for example, setting parameters and subgroup parameters used to code video data.

[0029] Decoder 210 decodes the coded video data of a frame that can be specified as a reference frame based on the symbols created by source coder 202. The operation of coding engine 212 can advantageously be a lossy process. When the coded video data is decoded in a video decoder (not shown in FIG. 2A), the reconstructed video sequence can be a replica of the source video sequence with some errors. Decoder 210 can replicate the decoding process that can be performed by a remote video decoder for the reference frame and store the reconstructed reference frame in reference picture memory 208. In this way, encoder component 106 locally stores a copy of the reconstructed reference frame that has the same content as the reconstructed reference frame obtained by the remote video decoder (without transmission errors).

[0030] Predictor 206 can perform a prediction search for coding engine 212. That is, for a new frame to be coded, predictor 206 can search reference picture memory 208 to obtain sample data (as candidate reference pixel blocks) that can function as an appropriate prediction reference for the new picture, or certain metadata such as reference picture motion vectors, block shapes, etc. Predictor 206 can operate on a sample block-by-pixel block basis to find an appropriate prediction reference. In some cases, the input picture can have prediction references drawn from a plurality of reference pictures stored in reference picture memory 208 as determined by the search results obtained by predictor 206.

[0031] The outputs of all the aforementioned functional units are entropy encoded by the entropy encoding unit 214. The entropy coder 214 converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).

[0032] In some embodiments, the output of the entropy coder 214 is coupled to a transmitter. The transmitter may be configured to buffer the coded video sequences created by the entropy coder 214 and prepare them for transmission via the communication channel 218, where the communication channel 210 may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to merge the encoded video data from the source coder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter may transmit additional data along with the encoded video. The source coder 202 may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, and the like.

[0033] The controller 204 can manage the operation of the encoder component 106. During coding, the controller 204 can assign to each coded picture a specific coded picture type that may affect the coding technique applied to that picture. For example, a picture can be assigned as an intra picture (I picture), a predicted picture (P picture), or a bi - directionally predicted picture (B picture). An intra picture can be coded and decoded without using other frames within the sequence as a source of prediction. Some video codecs allow for different types of intra pictures, such as Independent Decoder Refresh (IDR) Pictures. These variations of I pictures, as well as their respective uses and characteristics, will not be repeated here as they are well known to those skilled in the art. A predicted picture can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block. A bi - directionally predicted picture can be coded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use more than two reference pictures and related metadata for the reconstruction of a single block.

[0034] The source picture is generally spatially re - divided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block - by - block. The blocks can be predictively coded by referring to other (already coded) blocks as determined by the coding assignment applied to each picture of the block. For example, blocks of an I - picture may be non - predictively coded, or may be predictively coded by referring to already - coded blocks of the same picture (spatial prediction or intra - prediction). Pixel blocks of a P - picture can be coded non - predictively, via spatial prediction or via temporal prediction, by referring to one previously coded reference picture. Blocks of a B - picture can be coded non - predictively, via spatial prediction or via temporal prediction, by referring to one or two previously coded reference pictures.

[0035] Video can be captured as a plurality of source pictures (video pictures) in a time series. Intra - picture prediction (often abbreviated as intra - prediction) utilizes the spatial correlation within a given picture, and inter - picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been previously coded and is still buffered in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension identifying the reference picture when multiple reference pictures are being used.

[0036] The encoder component 106 may perform a coding operation in accordance with any of the video coding techniques or standards described herein. In that operation, the encoder component 106 may perform various compression operations, including a predictive coding operation that exploits temporal redundancy and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0037] FIG. 2B is a block diagram showing exemplary elements of a decoder component 122 according to some embodiments. The decoder component 122 of FIG. 2B is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter configured to transmit data to the display 124 (e.g., via a wired or wireless connection) and is coupled to a loop filter section 256.

[0038] In some embodiments, decoder component 122 includes a receiver configured to be coupled to channel 218 and receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, and channel 210 may be a hardware / software link to a storage device that stores the encoded video data. The receiver may receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams that may be transferred to their respective using entities (not shown). The receiver can separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by decoder component 122 to decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.

[0039] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (which may also be referred to as an entropy decoder), scaler / inverse transform unit 258, intra picture prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, decoder component 122 is implemented at least partially in software.

[0040] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to handle network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 within decoder component 122 (configured, e.g., to handle playout timing), a separate buffer memory is provided external to decoder component 122 (e.g., to handle network jitter). When receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isosynchronous network, buffer memory 252 may not be necessary or may be small. For use on a best-effort packet network such as the Internet, buffer memory 252 is required, is relatively large, advantageously of an adaptive size, and may be implemented at least in part by an operating system or similar element (not shown) external to decoder component 122.

[0041] Parser 254 is configured to reconstruct symbol 270 from the coded video sequence. The symbol may include, for example, information used to manage the operation of decoder component 122 and / or information for controlling a rendering device such as display 124. The control information for the rendering device may be in the form of, for example, a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). Parser 254 analyzes (entropy decodes) the encoded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 254 can extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups can include, for example, picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. Parser 254 can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.

[0042] The reconstruction of symbol 270 can include multiple different parts depending on the type of the encoded video picture or a part thereof (such as inter-picture and intra-picture, inter-block and intra-block, etc.), and other factors. Which parts are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by parser 254. Such a flow of subgroup control information between parser 254 and the following multiple parts is not illustrated for clarity of the figure.

[0043] In addition to the function blocks already described, decoder component 122 can conceptually be subdivided into multiple functional parts as described below. In an actual implementation operating under commercial constraints, many of these parts can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the conceptual re-division into the following functional parts is maintained.

[0044] Scaler / inverse transform unit 258 receives, as symbol 270, the quantized transform coefficients, as well as control information (such as which transform to use, block size, quantization coefficient, and / or quantization scaling matrix, etc.) from parser 254. Scaler / inverse transform unit 258 can output a block including sample values that can be input to aggregator 268.

[0045] In some cases, the output samples of the scaler / inverse transform unit 258 are related to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 may generate a block of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information taken from the current (partially reconstructed) picture in the current picture memory 264. The aggregator 268 can add, for each sample, the prediction information generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258.

[0046] In other cases, the output samples of the scaler / inverse transform unit 258 are related to inter-coded, and potentially motion-compensated, blocks. In such cases, the motion compensation prediction unit 260 can access the reference picture memory 266 to fetch the samples used for prediction. According to the symbol 270 related to the block, after motion-compensating the fetched samples, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (in this case, called the residual samples or residual signal) to generate the output sample information. The address in the reference picture memory 266 where the motion compensation prediction unit 260 fetches the prediction samples can be controlled by the motion vector. The motion vector can be available to the motion compensation prediction unit 260, for example, in the form of a symbol 270 that can have X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory 266 when an exact sub-sample motion vector is used, a motion vector prediction mechanism, etc.

[0047] The output samples of aggregator 268 can follow various loop filtering techniques in loop filter section 256. Video compression techniques can include in-loop filtering techniques that are included in the coded video bitstream and are controlled by parameters made available to loop filter section 256 as symbols 270 from parser 254, but can also respond to meta information obtained during the decoding of previous parts (in decoding order) of a coded picture or coded video sequence, and can also respond to previously reconstructed and loop-filtered sample values.

[0048] The output of loop filter section 256 can be a sample stream that is output to a rendering device such as display 124 and can be stored in reference picture memory 266 for use in future inter-picture prediction.

[0049] Once a coded picture is fully reconstructed, it can be used as a reference picture for future prediction. When a coded picture is fully reconstructed and is identified as a reference picture (e.g., by parser 254), the current reference picture can become part of reference picture memory 266, and the new current picture memory can be reallocated before starting the reconstruction of subsequent coded pictures.

[0050] The decoder component 122 can perform a decoding operation according to a predetermined video compression technique that can be documented in a standard, such as any of the standards described herein. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that it conforms to the syntax of a video compression technique or standard as specified in a video compression technique document or standard, particularly a profile document therein. Also, in order to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within a range as defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level can, in some cases, be further restricted through the virtual reference decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.

[0051] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., a CPU, a GPU, and / or a DPU). In some embodiments, the control circuit includes one or more field programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuits).

[0052] The network interface 304 can be configured to interface with one or more communication networks (e.g., wireless, wireline, and / or optical networks). The communication network can be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of communication networks include Ethernet®, wireless LAN, GSM®, cellular networks including 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, and local area networks including vehicle and industrial such as CANBus. Such communication can be unidirectional, receive only (e.g., broadcast TV), transmit only unidirectional (e.g., CANbus to a specific CANbus device), or bidirectional (e.g., to other computer systems using local or wide area digital networks). Such communication can include communication to one or more cloud computing networks.

[0053] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 can include one or more of a keyboard, mouse, trackpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. The output device 308 can include one or more of an audio output device (e.g., speaker), visual output device (e.g., display or monitor), etc.

[0054] Memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices), and / or non-volatile memory (such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Memory 314 optionally includes one or more storage devices located remotely from control circuit 302. Memory 314, or alternatively, the non-volatile solid-state memory device within memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, memory 314, or the non-transitory computer-readable storage medium of memory 314, stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: · An operating system 316 that processes various basic system services and includes procedures for performing hardware-dependent tasks. · A network communication module 318 used to connect server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections). · A coding module 320 for performing various functions related to encoding and / or decoding data such as video data. In some embodiments, coding module 320 is an instance of coder component 114. Coding module 320 includes, without limitation: ○ A decoding module 322 for performing various functions related to decoding encoded data, such as those previously described with respect to decoder component 122; and ○ An encoding module 340 for performing various functions related to encoding data, such as those previously described with respect to encoder component 106. · A picture memory 352 for storing pictures and picture data for use, for example, with an encoding module 320. In some embodiments, the picture memory 352 includes one or more of a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0055] In some embodiments, the decoding module 322 includes a parsing module 324 (configured to perform various functions previously described, for example, with respect to the parser 254), a conversion module 326 (configured to perform various functions previously described, for example, with respect to the scaler / inverse transform unit 258), a prediction module 328 (configured to perform various functions previously described, for example, with respect to the motion compensation prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (configured to perform various functions previously described, for example, with respect to the loop filter unit 256).

[0056] In some embodiments, the encoding module 340 includes a code module 342 (configured to perform various functions previously described, for example, with respect to the source coder 202, the coding engine 212, and / or the entropy coder 214), and a prediction module 344 (configured to perform various functions previously described, for example, with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes a subset of the modules shown in FIG. 3. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.

[0057] Each of the identified modules stored in the memory 314 corresponds to a set of instructions for performing the functions described herein. The modules identified above (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise reconfigured in various embodiments. For example, the encoding module 320 optionally does not include separate decoding and encoding modules, but rather uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the modules and data structures identified above. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0058] In some embodiments, the server system 112 includes web pages and applications implemented using, for example, a web or Hypertext Transfer Protocol (HTTP) server, a File Transfer Protocol (FTP) server, and Common Gateway Interface (CGI) scripts, PHP Hyper-text Preprocessor (PHP), Active Server Pages (ASP), Hyper Text Markup Language (HTML), Extensible Markup Language (XML), Java (registered trademark), JavaScript (registered trademark), Asynchronous JavaScript and XML (AJAX), XHP, Javelin, Wireless Universal Resource File (WURFL), and the like.

[0059] FIG. 3 shows a server system 112 according to some embodiments. FIG. 3 is not a structural schematic diagram of the embodiments described herein, but is intended to provide a functional description of various features that may exist in one or more server systems. In practice, as will be recognized by those skilled in the art, the separately shown items may be combined and some items may be separated. For example, some items separately shown in FIG. 3 may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the functions are allocated among them vary from implementation to implementation and, optionally, depend in part on the amount of data traffic that the server system processes during peak and average usage periods.

[0060] Examples of Coding Approaches FIGS. 4A-4D show exemplary coding tree structures according to some embodiments. As shown in the first coding tree structure (400) of FIG. 4A, some coding approaches (e.g., VP9) use a 4-way partition tree starting from the 64×64 level to the 4×4 level with some additional restrictions for 8×8 blocks. In FIG. 4A, the partition designated as R can be called recursive in that the same partition tree is repeated at a lower scale until the lowest 4×4 level is reached.

[0061] As shown in the second coding tree structure (402) of FIG. 4B, some coding techniques (e.g., AV1) extend the partition tree to a 10-way structure and increase the maximum size (e.g., called a superblock in VP9 / AV1 terminology) to start from 128×128. The second coding tree structure includes 4:1 / 1:4 rectangular partitions not present in the first coding tree structure. The partition type with three sub-partitions in the second row of FIG. 4B is called a T-type partition. The rectangular partitions in this tree structure cannot be further subdivided. In addition to the coding block size, the coding tree depth can be defined to indicate the split depth from the root node. For example, the coding tree depth for the root node, e.g., for 128×128, is set to 0, and after the tree block is split one more time, the coding tree depth is increased by only 1.

[0062] As an example, instead of implementing a fixed transform unit size like VP9, AV1 allows the luma coding block to be divided into transform units of multiple sizes that can be represented by a recursive partitioning that goes down a maximum of 2 levels. To incorporate the extended coding block partitions of AV1, square, 2:1 / 1:2, and 4:1 / 1:4 transform sizes from 4×4 to 64×64 are supported. For chroma blocks, only the largest possible transform units are allowed.

[0063] As an example, the CTU can be divided into CUs by using a quadtree structure shown as a coding tree to adapt to various local characteristics in HEVC and the like. In some embodiments, the decision as to whether to code a picture area using inter-picture (temporal) prediction or to code a picture area using intra-picture (spatial) prediction is made at the CU level. Each CU can be further divided into one, two, or four PUs according to the PU partition type. Within one PU, the same prediction process is applied and the relevant information is sent to the decoder for each PU. After obtaining the residual block by applying the prediction process based on the PU partition type, the CU can be partitioned into TUs according to another quadtree structure such as the coding tree for the CU. One of the important features of the HEVC structure is that it has a plurality of partition concepts including CUs, PUs, and TUs. In HEVC, a CU or TU can only be square in shape, while a PU can be square or rectangular in shape in the case of an inter-predicted block. In HEVC, one coding block can be further divided into four square sub-blocks, and the transformation is performed on each sub-block (TU). Each TU can be further recursively divided into smaller TUs (using quadtree partitioning), which is called a Residual Quad-Tree (RQT). At the picture boundary in HEVC and the like, an implicit quadtree partitioning can be adopted, whereby the block maintains quadtree partitioning until its size conforms to the picture boundary.

[0064] In VVC and the like, a quadtree with a nested multi-type tree using binary and ternary segmentation structures can replace the concept of multiple partition unit types, for example, removing the separation of the CU, PU, and TU concepts and supporting more flexibility in the CU partition shape, except when required for a CU having a size too large for the maximum transform length. In the coding tree structure, a CU can have a shape of either a square or a rectangle. The ACTUA is first divided by a quadtree (also called a four-ary tree) structure. A quadtree leaf node can be further divided by a multi-type tree structure. As shown in the third coding tree structure (404) of FIG. 4C, the multi-type tree structure includes four split types. For example, the multi-type tree structure includes a vertical binary split (SPLIT_BT_VER), a horizontal binary split (SPLIT_BT_HOR), a vertical ternary split (SPLIT_TT_VER), and a horizontal ternary split (SPLIT_TT_HOR). A multi-type tree leaf node is called a CU, and this segmentation is used for prediction and transform processing without further subdivision as long as the CU is not too large for the maximum transform length. This means that in most cases, the CU, PU, and TU have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is smaller than the width or height of the color component of the CU. An example of the block partition of one CTU (406) is shown in FIG. 4D, and FIG. 4D shows an exemplary quadtree having a nested multi-type tree coding block structure.

[0065] In VVC and the like, the maximum supported luma transform size can be 64×64, and the maximum supported chroma transform size can be 32×32. When the width or height of a CB is larger than the maximum transform width or height, the CB is automatically split in the horizontal direction and / or the vertical direction to meet the transform size limit in that direction.

[0066] The coding tree method supports the luma and chroma capabilities to have a separate block tree structure, such as in VTM7. In some cases, for P slices and B slices, the luma CTB and chroma CTB in one CTU share the same coding tree structure. However, in the case of I slices, luma and chroma can have separate block tree structures. When the separate block tree mode is applied, the luma CTB is partitioned into CUs by one coding tree structure, and the chroma CTB is partitioned into chroma CUs by another coding tree structure. This means that a CU in an I slice can contain or consist of a coding block of the luma component or coding blocks of two chroma components, and a CU in a P slice or B slice can always contain or consist of coding blocks of all three color components as long as the video is not monochrome.

[0067] To support extended coding block partitioning, in AV1 etc., multiple transform sizes (e.g., in the range from 4 points to 64 points for each dimension) and transform shapes (e.g., squares or rectangles with width / height ratios of 2:1 / 1:2 and 4:1 / 1:4) can be utilized.

[0068] The 2D transform process can involve the use of a hybrid transform kernel (e.g., composed of different 1D transforms for each dimension of the coded residual block). The primary 1D transform can include at least one of a) 4-point, 8-point, 16-point, 32-point, 64-point discrete cosine transform DCT-2, b) 4-point, 8-point, 16-point asymmetric discrete sine transforms (DST-4, DST-7) and their inverse versions, or c) 4-point, 8-point, 16-point, 32-point identity transforms. The basis functions of DCT-2 and asymmetric DST as used in AV1 are listed in Table 1.

[0069]

Table 1

[0070] [Table 2] For the JPEG2025516422000004.jpg 128×168 chroma components, the transform type selection is performed in an implicit manner. For the intra prediction residual, the transform type is selected according to the intra prediction mode, as specified in Table 3 for example. For the inter prediction residual, the transform type can be selected according to the transform type selection of the collocated luma block. Therefore, for the chroma components, transform type signaling in the bitstream is not required.

[0071] [Table 3] For the purpose of replacing and extending the above-mentioned 1D DST (by introducing 32 points and 64 points), the line graph transform (LGT) has been introduced.

[0072] A graph is a general mathematical structure that includes or consists of a set of vertices and edges used to model affinity relationships between objects of interest. A weighted graph (a set of weights is assigned to the edges and potentially to the vertices) provides a sparse representation for robust modeling of signals / data. LGT can improve coding efficiency by providing better adaptation to various block statistics. Separable LGT is designed and optimized by learning a line graph from the data to model the statistics for each row and column underlying the blocks in the residual signal, and the associated generalized graph Laplacian (GGL) matrix is used to derive the LGT. FIG. 5A shows an exemplary LGT characterized by self-loop weights v c1 and v c2 and edge weight w c .

[0073] Given a weighted graph G(W, V), the GGL matrix can be defined as follows.

[0074]

Number

[0075]

Number

[0076] The LGT can be implemented as a matrix multiplication. The 4p LGT core is derived by setting v c in L c1 = 2w c , which means it is DST-4. The 8p LGT core can be derived by setting vc1 = 1.5 w c in Lc. The 16p, 32p, and 64p LGT cores may be derived by setting v c1 = w c in Lc, which means it is DST-7.

[0077] In an example of residual coding in AV1, for each transform unit, coefficient coding starts by signaling a skip flag. When the skip flag is 0, the transform kernel system and the end-of-block (eob) position follow. Next, each coefficient value is mapped to a plurality of level maps and signs. After the EOB position is coded, the lower-level map and the middle-level map are coded in reverse scan order, where the former indicates whether the magnitude of the coefficient is between 0 and 2, and the latter indicates whether the range is between 3 and 14. The next step is to code the sign of the coefficient and the residual value of the coefficient greater than 14 by an exponential Golomb code in forward scan order.

[0078] Regarding the use of context modeling, the lower-level map coding incorporates the transform size and direction, as well as up to five adjacent coefficient information. On the other hand, the middle-level map coding follows a similar approach to the lower-level map coding, except that the number of adjacent coefficients is reduced to two. The exponential Golomb (Exp-Golomb) code for the residual level as well as the signs of the AC coefficients are coded without a context model, and the sign of the DC coefficient is coded using the DC sign of its adjacent transform unit.

[0079] In an example of residual coding for transform skip, such as in VVC, a CU coded in transform skip mode (TSM) may use a modified transform coding process. The modifications can be summarized as (a) all sub-blocks and positions within the sub-blocks are scanned in forward scan order, (b) the last significant coefficient position is not signaled, (c) the syntax element coded_sub_block_flag is not coded for the last sub-block, (d) changes are made to the context modeling for the syntax sig_coeff_flag, abs_level_gt1, and par_level_flag, and (e) the sign_flag is context-coded based on the left and upper adjacent values.

[0080] During the development of AV2, a new mode called forward skip coding (FSC) was introduced and the transform coding process for IDTX (two-dimensional transform skip) was modified. The modifications introduced by FSC are similar in function to the above-mentioned changes introduced in the VVC transform skip mode and can be summarized as follows: (a) All coded blocks and positions within the coded block are scanned in forward scan order; (b) The EOB syntax is skipped; (c) The reduced context is used for the coefficient levels; (d) The sign flag is context-coded based on the left, bottom, and bottom-left.

[0081] In the case of an intra block, when the FSC mode is selected, the transform type is not signaled for the transform block. Rather, the transform type signaling is bundled with the FSC mode at the coded block level. Inter blocks do not signal the FSC mode, but when the transform type is IDTX and the screen content flag is active, the FSC method is implicitly selected.

[0082] For the entropy coding of transform coefficients, the coefficient sign can be coded separately from the magnitude / level (the absolute value of the coefficient value) using the bypass mode. Separate coding means that 1 bit is required to code each coefficient sign, which is costly. To improve the entropy coding efficiency of the coefficient sign, sign prediction techniques can be used. For example, instead of signaling the sign value, a flag indicating whether the predicted sign is the same as the actual sign can be entropy-coded using the context. The context value can depend on the level of the coefficient (the absolute value of the coefficient value) because larger level values result in more accurate predicted sign values.

[0083] In one example, a group of transform coefficients for which the associated signs need to be predicted is identified. Next, a set of hypotheses for the predicted sign values of these coefficients is generated. For example, for three coefficients, the number of hypotheses can be up to 8 (2^3). To predict the sign value, there is a cost value associated with each hypothesis, and the hypothesis with the minimum cost is used to specify the predicted sign value of the coefficients covered by the hypothesis.

[0084] FIG. 5A shows an example of pixel positions of a transform block 500 and adjacent rows 502 and adjacent columns 504. In some embodiments, the cost of each hypothesis is calculated as follows. A reconstruction block (hypothesis reconstruction) associated with a given hypothesis is generated following a reconstruction process (e.g., inverse quantization, inverse transform), and boundary samples of the reconstructed block, e.g., p 0,y and p x,0 are derived. For each reconstructed pixel p 0,y at the left boundary of the reconstruction block, a simple linear prediction using the two previously reconstructed adjacent pixels to the left is performed to obtain its prediction pred 0,y =(2 p-1,y - p -2,y ). The absolute difference between this prediction and the reconstructed pixel p 0,y is added to the cost of the hypothesis. Similar processing is done for the pixels in the top row of the reconstructed block, and the absolute differences between each prediction pred x,0 =(2 p x,-1 - p x,-2 ) and the reconstructed pixel p x,0 are summed. Thus, the calculation of the cost of each coefficient sign prediction hypothesis is given by Equation 3 below.

[0085]

Equation

[0086] FIG. 5B shows an exemplary scan order and exemplary low-frequency pixels in transform block 550 according to some embodiments. Transform block 550 includes 64 pixels (corresponding to 8 columns and 8 rows). In the example of FIG. 5B, box 552 shows the low-frequency pixels of transform block 550 (e.g., pixels P 0,0 ~P 2,2 ). In other examples, box 552 has a rectangular shape and may include more or fewer rows and / or columns (e.g., box 552 may include P 3,n columns and / or P n,3 rows). FIG. 5B also shows an exemplary scan order 554 (e.g., proceeding in a zigzag pattern from P 0,0 to P 7,7 ). In some embodiments, other scan orders are used.

[0087] FIG. 6 is a flowchart showing a method 600 for coding video according to some embodiments. Method 600 can be executed in a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions for execution by the control circuit. In some embodiments, method 600 is performed by executing instructions stored in the memory of the computing system (e.g., memory 314).

[0088] When entropy coding the sign values of the first color component (e.g., Cr), different context designs can be used. For example, the context can be derived using either sign prediction based on adjacent blocks or inter-component sign coding. In some embodiments, different sign prediction techniques are used for different transform coefficients to improve (maximize) efficiency and reduce computational cost and bandwidth usage.

[0089] The system obtains (602) video data including a plurality of blocks including a first block. The system determines (604) a plurality of transform coefficients associated with the first block. The system predicts (606) the respective coefficient signs for a first set of the plurality of transform coefficients using a first technique. The system predicts (608) the respective coefficient signs for a second set of the plurality of transform coefficients using a second technique, where the second technique is different from the first technique. The system reconstructs (610) the first block based on the plurality of transform coefficients and the respective coefficient signs predicted for the first and second sets. In some embodiments, the system forgoes (612) predicting the respective coefficient signs for a third set of the plurality of transform coefficients. Method 600 is optionally applied to luminance blocks and / or chrominance blocks.

[0090] As used herein, "block" may be interpreted to mean a prediction block, a coding block, a coding unit (CU), etc. The term "block" may also be used to refer to a transform block. In various situations, the term "block size" may refer to the width or height of a block, or the maximum value of the width and height, or the minimum value of the width and height, or the area size (width * height), or the aspect ratio of the block (width:height, or height:width).

[0091] In some embodiments, when determining which signs of the current transform coefficient block are coded by the cross-component sign coding technique and which signs of the current transform coefficient block are coded by the sign prediction technique, one of the following methods is used. In a first exemplary method, the transform coefficients for which signs are coded by the sign prediction technique are first determined, and then the cross-component sign coding technique is applied to the remaining transform coefficient signs. In a second exemplary method, the transform coefficients for which signs are coded by the cross-component sign coding technique are first determined, and then the sign prediction technique is applied to the remaining transform coefficient signs.

[0092] In some embodiments, for transform coefficients that have non-zero Cr transform coefficients but do not have zero co-located Cb transform coefficients, the sign prediction technique is applied. Otherwise, the sign prediction technique cannot be applied.

[0093] In some embodiments, for transform coefficients located beyond the first N transform coefficients along the scan order, the signs are coded by the cross-component sign coding technique. Otherwise, the cross-component sign coding technique cannot be applied.

[0094] In some embodiments, instead of explicitly coding the sign value, an indicator is signaled that specifies whether the sign value of the current conversion coefficient of the first color component is the same as the sign value of another conversion coefficient at a specified position of the second color component. In some embodiments, the first color component is the Cr (or Cb) color component and the second color component is the Cb (or Cr) color component. In some embodiments, the specified position of the second color component refers to a co-located conversion coefficient position of the conversion coefficient of the first color component. In some embodiments, the method is applied only when the conversion coefficients of both the first color component and the second color component are non-zero. In some embodiments, when the conversion coefficient of the second color component is 0, the sign value of the conversion coefficient at the same position is coded without using context, for example, bypass coded or literal coded.

[0095] In some embodiments, during analysis (entropy decoding), values are converted to signs according to the techniques and methods described herein. For example, obtaining video data comprising a plurality of blocks including a first block, determining a plurality of conversion coefficients associated with the first block, using a first technique to predict the coefficient sign of each of a first set of the plurality of conversion coefficients, using a second technique to predict the coefficient sign of each of a second set of the plurality of conversion coefficients, the second technique being different from the first technique, and reconstructing the first block based on the plurality of conversion coefficients and the predicted coefficient signs for the first set and the second set.

[0096] FIG. 6 shows several logical stages in a particular order, but stages that are not order-dependent may be reordered, other stages may be combined, or separated. Some order changes or other groupings not specifically recited will be apparent to those of ordinary skill in the art, and thus the orderings and groupings presented herein are not exhaustive. Further, it should be recognized that the various stages may be implemented in hardware, firmware, software, or any combination thereof.

[0097] Next, some exemplary embodiments are referred to.

[0098] (A1) In one aspect, some embodiments include a method of video decoding (e.g., method 600). In some embodiments, the method is performed in a computing system (e.g., server system 112) having a memory and a control circuit. In some embodiments, the method is executed in a coding module (e.g., coding module 320). In some embodiments, the method is executed in an entropy coder (e.g., entropy coder 214). In some embodiments, the method is performed in a parser (e.g., parser 254). The method includes (i) obtaining video data including a plurality of blocks including a first block from a video bitstream, (ii) determining a plurality of transform coefficients associated with the first block, wherein coefficient signs associated with the plurality of transform coefficients are not explicitly signaled in the video bitstream, (iii) predicting each coefficient sign for a first set of the plurality of transform coefficients using a first technique, (iv) predicting each coefficient sign for a second set of the plurality of transform coefficients using a second technique, wherein the second technique is different from the first technique, and (v) reconstructing the first block based on the plurality of transform coefficients and the predicted respective coefficient signs for the first and second sets. For example, the plurality of blocks are transform blocks. In some embodiments, the first block includes a luma block. In some embodiments, the first block includes a chroma block. In some embodiments, each element of the first block corresponds to a reconstructed pixel.

[0099] (A2)In some embodiments of A1, the first technique includes predicting the coefficient sign of a color coefficient based on one or more adjacent components (e.g., an adjacent coefficient, a block, or a reconstructed sample). For example, a first sign is selected, the cost of the first sign is determined based on one or more adjacent components, then a different sign is selected, the cost of the second sign is determined based on one or more adjacent components, and then the sign with the lowest associated cost is used as the predicted sign. In some situations, the first technique uses more computational resources than the second technique (e.g., because the second technique includes testing multiple hypotheses while the first technique does not).

[0100] (A3)In some embodiments of A1 or A2, the second technique includes predicting the coefficient sign of a color coefficient based on the sign values of different color coefficients. For example, predicting the sign of the Cb component based on the sign of the Cr component (e.g., at the same block location). In some embodiments, each block is associated with a plurality (e.g., 3) of color components. In some situations, the color component coefficient signs are strongly correlated. Thus, the probability of sign prediction based on color context information can be 90 / 10 instead of 50 / 50, which is more efficient for entropy coding. In some situations, the second technique cannot be used for all transform coefficients. For example, if the coefficients of different color coefficients are 0, the coefficient sign of the color coefficient cannot be predicted based on different color coefficients.

[0101] (A4)In some embodiments of any of A1 - A3, the method further includes refraining from predicting each coefficient sign for a third set of a plurality of transform coefficients. For example, the sign values of the third set of elements are assigned using bypass coding or literal coding.

[0102] (A5)In some embodiments of any one of A1 to A4, the method further includes selecting a first set for use with a first technique based on one or more frequency-based references. For example, the elements of the first set of elements have associated frequencies that are less than a preset threshold frequency.

[0103] (A6)In some embodiments of any one of A1 to A5, the method further includes selecting a first set for use with a first technique based on the scanning order of the first block. In some embodiments, the first set of elements consists of a sequence of the first elements of the scanning order of the first block (e.g., the first 5 elements along the scanning order).

[0104] (A7)In some embodiments of any one of A1 to A6, the method further includes selecting a first set for use with a first technique based on a predetermined area of the first block. For example, the predefined area is an upper left N×N area or N×M of the first block, such as a 1×1, 2×2, 2×3, 3×2, 4×3, or 3×4 area.

[0105] (A8)In some embodiments of any one of A1 to A7, the method further includes selecting a second set for use with a second technique in accordance with the determination that each transformation coefficient of the second set is non-zero.

[0106] (A9)In some embodiments of any one of A1 to A8, (i) the second set includes a set of color elements of a first color, and (ii) the method further includes selecting a second set for use with a second technique in accordance with the determination that the corresponding color element of the second color has a non-zero transformation coefficient for each element of the second set. In some embodiments, the second set of elements is selected for use with the second technique in accordance with the determination that each element having a non-zero transformation coefficient and the corresponding second color element have non-zero transformation coefficients.

[0107] (B1)In another aspect, some embodiments include a method of video coding. In some embodiments, the method is performed in a computing system (e.g., server system 112) having a memory and a control circuit. In some embodiments, the method is executed in a coding module (e.g., coding module 320). In some embodiments, the method is executed in an entropy coder (e.g., entropy coder 214). In some embodiments, the method is performed in a parser (e.g., parser 254). The method includes (i) obtaining video data including a plurality of blocks including a first block including a plurality of elements, and (ii) coding an indicator for each element of the plurality of elements specifying whether the sign value of the transform coefficient of the element is the same as the sign value of the transform coefficient of a respective second element. The probability associated with the indicator may be 90 / 10 rather than a 50 / 50 probability associated with the coefficient sign, where 90 / 10 is more efficient for entropy coding.

[0108] (B2)In some embodiments of B1, the plurality of elements correspond to a first color and the second element corresponds to a second color different from the first color. For example, the first color is red (or blue) and the second color is blue (or red).

[0109] (B3)In some embodiments of B1 or B2, for each element of the plurality of elements, the respective second element is co-located with that element. For example, for each element, the second element is an element at the same location of a different color.

[0110] (B4)In some embodiments of any of B1 - B3, the method further includes, for each element of the plurality of elements, selecting a plurality of elements for coding according to a determination that the element and the respective second element each have a non-zero transform coefficient.

[0111] In some embodiments of any of B1 - B4, the method further includes coding the sign of an element for each element of the second plurality of elements of the first block using bypass coding or literal coding. For example, the sign of each element of the second plurality of elements is coded without using context information.

[0112] The methods described herein may be used separately or combined in any order. Each of the methods may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In some embodiments, the processing circuit executes a program stored on a non - transitory computer - readable medium.

[0113] In another aspect, some embodiments include a computing system (e.g., server system 112) including a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, the memory storing one or more sets of instructions configured to be executed by the control circuit, and the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 - A9 and B1 - B5 above).

[0114] In yet another aspect, some embodiments include a non - transitory computer - readable storage medium storing one or more sets of instructions for execution by a control circuit of a computing system, and the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 - A9 and B1 - B5 above).

[0115] Although terms such as "first," "second," etc. may be used herein to describe various elements, it should be understood that these elements are not to be limited by these terms. These terms are only used to distinguish one element from another.

[0116] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. When used in the description of embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Also, the term "and / or" as used herein refers to any and all possible combinations of one or more of the associated listed items and is to be understood to be inclusive. As used herein, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0117] As used herein, the term "when" can be interpreted, depending on the context, as "when" the precedent of the stated condition is true, or "upon", or "in response to a determination", or "in accordance with a determination", or "in response to a detection". Similarly, the phrases "when determined [that the precedent of the stated condition is true]" or "when [the precedent of the stated condition is true]" or "when [the precedent of the stated condition is true]" can be interpreted, depending on the context, as "when it is determined" or "in response to a determination" or "in accordance with a determination" or "when detected" or "in response to a detection" that the precedent of the stated condition is true.

[0118] The foregoing description has been presented for purposes of illustration and is described with reference to particular embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. Embodiments were chosen and described in order to best explain the principles of operation and the practical application and thereby enable others skilled in the art.

Claims

1. A method of video decoding executed in a computing system having a memory and one or more processors, comprising: obtaining video data including a plurality of blocks including a first block from a video bitstream; determining a plurality of transform coefficients associated with the first block, wherein coefficient signs associated with the plurality of transform coefficients are not explicitly signaled in the video bitstream; predicting respective coefficient signs for a first set of the plurality of transform coefficients using a first technique; predicting respective coefficient signs for a second set of the plurality of transform coefficients using a second technique, wherein the second technique is different from the first technique; reconstructing the first block based on the plurality of transform coefficients and respective coefficient signs predicted for the first set and the second set. A method comprising the steps above.

2. The method according to claim 1, wherein the first technique includes predicting a coefficient sign of a color coefficient based on one or more adjacent components.

3. The method according to claim 1, wherein the second technique includes predicting a coefficient sign of a color coefficient based on sign values of different color coefficients.

4. The method according to claim 1, further comprising refraining from predicting respective coefficient signs for a third set of the plurality of transform coefficients.

5. The method according to claim 1, further comprising selecting the first set for use with the first technique based on one or more frequency-based criteria.

6. The method according to claim 1, further comprising selecting the first set for use with the first technique based on a scan order of the first block.

7. The method according to claim 1, further comprising selecting the first set for use with the first technique based on a predetermined region of the first block.

8. The method according to claim 1, further comprising selecting the second set for use with the second technique according to a determination that each transform coefficient of the second set is non-zero.

9. The second set includes a set of color elements of a first color. The method further includes selecting the second set for use with the second technique in accordance with a determination that, for each element of the second set, a corresponding color element of the second color has a non-zero conversion coefficient. The method according to claim 1. **Claim 10** A computing system, comprising: a control circuit; a memory; one or more sets of instructions stored in the memory and configured to be executed by the control circuit, the one or more sets of instructions, when executed by the control circuit, causing the control circuit to perform the steps of the method according to any one of claims 1 to 9. A computing system. **Claim 11** A non-transitory computer-readable storage medium storing one or more sets of instructions configured to be executed by a computing device having a control circuit and a memory, the one or more sets of instructions, when executed by the control circuit, causing the control circuit to perform the steps of the method according to any one of claims 1 to 9. A non-transitory computer-readable storage medium. **Claim 12** A computer program configured to be executed by a computing device having a control circuit, the computer program, when executed by the control circuit, causing the control circuit to perform the steps of the method according to any one of claims 1 to 9.