Extended directional prediction for residual blocks

By applying intra prediction modes in different directions in video encoding to generate corrected residual blocks, the problem of high redundancy of residual signals in the prior art is solved, and the encoding efficiency and quality are improved.

CN120239968APending Publication Date: 2025-07-01TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380080569.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-04
Filing Date
2023-10-30
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing video encoding technology has problems of high redundancy and low encoding efficiency in residual signal compression, especially in intra prediction, where direction information is not effectively used for optimization.

Method used

By adopting the intra prediction mode in video encoding, residual blocks are generated in the first direction, and short-distance intra prediction is used in the second direction to generate a corrected residual block, thereby reducing redundancy in the residual domain.

Benefits of technology

The encoding efficiency of video encoding is improved, the transmission bandwidth requirement is reduced, and the redundancy is further reduced by choosing the most effective direction, improving the encoding quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120239968A_ABST
    Figure CN120239968A_ABST
Patent Text Reader

Abstract

Various embodiments described herein include methods and systems for encoding and decoding a video. In one aspect, a method of video decoding includes receiving video data including a plurality of blocks from a video bitstream, the plurality of blocks including a first block and a plurality of residual coefficients. The first block is encoded by applying a first intra prediction. The plurality of residual coefficients are generated by applying a second intra prediction in a first direction to a residual block of the first block. The residual block is generated by applying a first intra prediction mode to a first block in a second direction. The method further includes generating a modified residual block of the first block based on the plurality of residual coefficients, and reconstructing the first block using the modified residual block.
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This application claims priority to U.S. Patent Application No. 18 / 480,966, titled "Extended Directional Predictions for Residual Blocks," filed on October 4, 2023. Technical Field

[0002] The disclosed embodiments generally relate to image and video encoding and compression, including but not limited to systems and methods for predicting residual information. Background Art

[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices transmit and receive digital video data over a communication network or otherwise convey digital video data and / or store digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited memory resources of the storage device, video encoding can be used to compress video data according to one or more video coding standards before transmitting or storing the video data. Video encoding and decoding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services. Video encoding typically uses prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that exploit the redundancy inherent in video data. Video encoding aims to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation of video quality. A variety of video codec standards have been developed.

[0004] For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard designed to be a successor to HEVC. The ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the confirmed version 1.0.0 with Specification Errata 1 was released. Summary of the Invention

[0005] A comprehensive video codec typically includes multiple components, such as intra / inter prediction, transform coding, quantization, residual coding, and in-loop filtering. To further reduce the residual signal, various residual prediction techniques have been developed. These techniques predict the residual signal and encode the high-order residual into the bitstream. This disclosure describes methods and systems for enhancing video (image) compression, including advanced residual prediction techniques.

[0006] According to some embodiments, a method of video encoding is provided. The method includes (i) receiving video data including a plurality of blocks, the plurality of blocks including a first block, wherein the first block is to be encoded in a first intra prediction mode; (ii) generating a residual block of the first block by applying the first intra prediction mode to the first block along a first direction; (iii) generating a corrected residual block of the first block by applying a second intra prediction mode to the residual block along a second direction; and (iv) signaling the corrected residual block via a video bitstream.

[0007] According to some embodiments, a method of video decoding is provided. The method includes (i) receiving from a video bitstream video data including a plurality of blocks, the plurality of blocks including a first block and a plurality of residual coefficients of the first block; (ii) generating a corrected residual block of the first block from the plurality of residual coefficients; (iii) generating a reconstructed residual block by applying a first intra prediction to the corrected residual block along a first direction; and (iv) reconstructing the first block by applying a second intra prediction to the reconstructed residual block along a second direction.

[0008] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic device. The computing system includes a control circuit and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.

[0009] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions for execution by a computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.

[0010] Accordingly, apparatuses and systems having methods for encoding and decoding video are disclosed. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, and systems for video encoding and / or decoding. The features and advantages described in the specification are not necessarily all-inclusive, and in particular, given the figures, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Additionally, it should be noted that the language used in the specification is primarily selected for readability and guidance purposes and is not necessarily selected to depict or limit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To understand the present disclosure in more detail, a more specific description may be made with reference to the features of various embodiments, some of which are illustrated in the figures. However, the figures only illustrate the relevant features of the present disclosure and are therefore not necessarily considered restrictive, as those skilled in the art will understand, upon reading this disclosure, that the specification may permit other effective features.

[0012] Figure 1 is a block diagram illustrating an example communication system according to some embodiments.

[0013] Figure 2A is a block diagram illustrating example elements of an encoder component according to some embodiments.

[0014] Figure 2B is a block diagram illustrating example elements of a decoder component according to some embodiments.

[0015] Figure 3 is a block diagram illustrating an example server system according to some embodiments.

[0016] Figure 4A illustrates the calculation of a prediction block according to some embodiments.

[0017] Figure 4B illustrates the calculation of a residual block according to some embodiments.

[0018] Figure 4C illustrates the calculation of a reconstruction block according to some embodiments.

[0019] Figure 4D illustrates the calculation of a modified residual block and a difference block according to some embodiments.

[0020] Figure 4E illustrates the calculation of a residual block and a reconstruction block according to some embodiments.

[0021] Figure 4F and Figure 4G illustrates example line-by-line prediction according to some embodiments.

[0022] Figure 5A and Figure 5B illustrates an example residual block in accordance with some embodiments.

[0023] In accordance with conventional practice, the various features illustrated in the drawings need not be drawn to scale and, throughout the specification and drawings, like reference numerals may be used to denote like features. DETAILED DESCRIPTION

[0024] The present disclosure describes systems and methods for predicting residual information. The systems and methods described herein can improve the performance of lossless coding by reducing redundancy in the residual domain. In some embodiments, short-range intra prediction (e.g., a line-by-line residual domain prediction mode) is implemented for vertical intra prediction mode and / or horizontal intra prediction mode. For the luminance plane and the chrominance plane, a flag indicating the use of short-range intra prediction can be signaled separately, while the U plane and the V plane can share a flag.

[0025] In some embodiments, a residual block is generated (e.g., by applying an intra prediction mode to a current block along a first direction), and then a modified residual block is generated (e.g., by applying short-range intra prediction to the residual block). Generating and using the modified residual block can reduce redundancy in the residual domain. The reduction in redundancy can reduce the number of bits required to signal the residual (e.g., improve coding efficiency and reduce transmission bandwidth). Additionally, being able to specify a first direction and a second direction allows for further reduction of redundancy by selecting the most efficient direction. Example Systems and Devices

[0026] Figure 1 is a block diagram illustrating a communication system 100 in accordance with some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, used with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0027] The source device 102 includes a video source 104 (e.g., a camera component or a media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from the video source 104 can be of high data volume compared to the encoded video bitstreams 108 generated by the encoder component 106. Since the encoded video bitstreams 108 are of lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less bandwidth to transmit and less storage space to store compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video data to the network 110).

[0028] One or more networks 110 represent any number of networks for communicating information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. One or more networks 110 include the server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content such as the encoded video stream from the source device 102). The server system 112 includes an encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the encoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the encoder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the encoder component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video bitstream 108.

[0029] In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 can be configured to trim the encoded video bitstream (108) to tailor potentially different bitstreams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0030] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., communicatively coupled to an external display device and / or including a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0031] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are examples of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0032] In an example operation of the communication system 100, the source device 102 transmits the encoded video stream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video stream 108 and may decode and / or encode the encoded video stream 108 using the encoder component 114. For example, the server system 112 may apply an encoding that is more optimized for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video streams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 to recover and optionally display the video pictures.

[0033] Figure 2AFIG. is a block diagram showing example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that produce motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. Those of ordinary skill in the art can readily understand the relationship between pixels and samples. The following description focuses on samples.

[0034] The encoder component 106 is configured to encode and / or compress pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. The parameters set by the controller 204 may include parameters related to rate control (e.g., picture skipping, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as they may belong to the encoder component 106 optimized for a specific system design.

[0035] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols such as a symbol stream based on an input picture and reference pictures to be encoded) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data (when the compression between the symbols and the encoded video bitstream is lossless). This reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream results in a bit-exact result regardless of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. Thus, when prediction is used during decoding, the prediction part of the encoder interprets the same sample values as the decoder interprets as reference picture samples. The principle of reference picture synchronization (and if synchronization cannot be maintained, e.g., drift due to channel errors) is known to those of ordinary skill in the art.

[0036] The operation of the decoder 210 can be the same as the operation of a remote decoder (such as the decoder component 122), which is described in detail below in conjunction with Figure 2B which. However, briefly referring to Figure 2B , since the symbols are available and the entropy encoder 214 and the parser 254 encoding / decoding the symbols into the encoded video sequence can be lossless, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be entirely implemented in the local decoder 210.

[0037] In addition to the parsing / entropy decoding present in the decoder, the decoder techniques described herein may also exist in the corresponding encoder in substantially the same functional form. To this end, the disclosed subject matter focuses on decoder operations. The description of encoder techniques may be abbreviated since they may be the inverse process of decoder techniques.

[0038] As part of its operation, the source encoder 202 may perform motion-compensated predictive coding, which predictively encodes an input frame by referring to one or more previously encoded frames designated as reference frames from a video sequence. In this way, the coding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of one or more reference frames that can be selected as one or more prediction references for the input frame. The controller 204 may manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroups of parameters for encoding video data.

[0039] The decoder 210 may decode the encoded video data of a frame that can be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the coding engine 212 may advantageously be a lossy process. When in a video decoder ( Figure 2AWhen decoding the encoded video data at the (not shown in the figure), the reconstructed video sequence can be a replica of the source video sequence with some errors. The decoder 210 replicates the decoding process that can be performed on the reference frame by the remote video decoder, and can cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frame, which has the same content as the reconstructed reference frame that will be obtained by the remote video decoder (without transmission errors).

[0040] The predictor 206 can perform a prediction search on the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search in the reference picture memory 208 for sample data (as candidate reference pixel blocks) that can be used as an appropriate prediction reference for the new picture or some metadata such as reference picture motion vectors, block shapes, etc. The predictor 206 can operate on a per-pixel-block sample basis to find an appropriate prediction reference. As determined by the search results obtained by the predictor 206, the input picture can have a prediction reference extracted from a plurality of reference pictures stored in the reference picture memory 208.

[0041] The outputs of all the above functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those of ordinary skill in the art (such as Huffman coding, variable length coding, and / or arithmetic coding).

[0042] In some embodiments, the output of the entropy encoder 214 is coupled to the transmitter. The transmitter can be configured to buffer one or more encoded video sequences created by the entropy encoder 214 to prepare them for transmission via the communication channel 218, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter can be configured to merge the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (not shown in the source). In an embodiment, the transmitter can transmit additional data along with the encoded video. The video encoder 202 can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0043] The controller 204 can manage the operation of the encoder component 106. During encoding, the controller 204 can assign a specific encoded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture can be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). Intra pictures can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example independent decoder refresh (IDR) pictures. Those of ordinary skill in the art are aware of these variations of I pictures and their respective applications and characteristics, and thus they will not be repeated herein. Predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multi-predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0044] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each), and encoded on a block-by-block basis. Blocks can be encoded predictively with reference to other (already encoded) blocks determined by the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture can be encoded non-predictively, or they can be encoded predictively with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be encoded non-predictively with reference to one previously encoded reference picture, either via spatial prediction or via temporal prediction. Blocks of a B picture can be encoded non-predictively with reference to one or two previously encoded reference pictures, either via spatial prediction or via temporal prediction.

[0045] Video can be captured as multiple source pictures (video pictures) in a time sequence. Intra picture prediction (commonly abbreviated as intra prediction) exploits the spatial correlation within a given picture, and inter picture prediction exploits the (temporal or other) correlation between picture pictures. In an example, a specific picture under encoding / decoding, referred to as the current picture, is divided into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. In the case of using multiple reference pictures, the motion vector points to the reference block in the reference image and can have a third dimension that identifies the reference picture.

[0046] The encoder component 106 may perform encoding operations according to a predetermined video encoding technique or standard (such as any described herein). In its operation, the encoder component 106 may perform various compression operations, including predictive encoding operations that utilize temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.

[0047] Figure 2B is a block diagram illustrating example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to a loop filter unit 256 and configured to transmit data to the display 124 (e.g., via a wired or wireless connection).

[0048] In some embodiments, the decoder component 122 includes a receiver coupled to the channel 218 and configured to receive data from the channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from the channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data with other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective entities using an entity (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of one or more encoded video sequences. The additional data may be used by the decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0049] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaling / inverse transform unit 258, an intra picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, the decoder component 122 is implemented at least partially in software.

[0050] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to counter network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 inside decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided outside decoder component 122 (e.g., to counter network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, buffer memory 252 may not be needed, or this buffer memory 252 may be small. For use on a best-effort packet network such as the Internet, buffer memory 252 may be required, which can be relatively large and can advantageously have an adaptive size, and can be implemented at least in part in an operating system or similar element (not shown) outside decoder component 122.

[0051] Parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. The symbols can include, for example, information for managing the operation of decoder component 122, and / or information for controlling a rendering device such as display 124. The control information for one or more rendering devices can be in the form of, for example, Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). Parser 254 can parse (entropy decode) the encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technique or standard and can follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 254 can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the set. Subgroups can include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. Parser 254 can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0052] The reconstruction of symbols 270 can involve multiple different units, depending on the type of the encoded video picture or its part (such as: inter-frame and intra-frame pictures, inter-frame and intra-frame blocks) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information, which is parsed by parser 254 from the encoded video sequence. For clarity, this subgroup control information flow between parser 254 and the multiple units is not depicted below.

[0053] The decoder component 122 can be conceptually subdivided into multiple functional units. In some implementations, many of these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, the conceptual subdivision of the functional units is maintained here.

[0054] The scaling / inverse transform unit 258 receives the quantized transform coefficients and control information (such as the transform to be used, block size, quantization factor, and / or quantization scaling matrix) as one or more symbols 270 from the parser 254. The scaling / inverse transform unit 258 can output a block including sample values that can be input to the aggregator 268.

[0055] In some cases, the output samples of the scaling / inverse transform unit 258 belong to in - frame - coded blocks; that is: blocks that do not use prediction information from previously reconstructed pictures, but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the in - frame picture prediction unit 262. The in - frame picture prediction unit 262 can use the surrounding reconstructed information obtained from the current (partially reconstructed) picture in the current picture memory 264 to generate a block having the same size and shape as the block being reconstructed. The aggregator 268 can add, on a per - sample basis, the prediction information already generated by the in - frame picture prediction unit 262 to the output sample information provided by the scaling / inverse transform unit 258.

[0056] In other cases, the output samples of the scaling / inverse transform unit 258 belong to inter - frame - coded and possibly motion - compensated blocks. In this case, the motion - compensation prediction unit 260 can access the reference picture memory 266 to obtain samples for prediction. After motion - compensating the obtained samples according to the symbols 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaling / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The address from which the motion - compensation prediction unit 260 obtains the prediction samples from within the reference picture memory 266 can be controlled by a motion vector. The motion vector can be in the form of a symbol 270 for the motion - compensation prediction unit 260, which can have, for example, X, Y, and reference - picture components. Motion compensation can also include interpolation of sample values obtained from the reference picture memory 266 when using sub - sampled accurate motion vectors, motion - vector prediction mechanisms, etc.

[0057] The output samples of aggregator 268 can be processed by various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and are available to loop filter unit 256 as symbols 270 from parser 254, but can also respond to meta-information obtained during the decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to sample values of previously reconstructed and loop-filtered samples. The output of loop filter unit 256 can be a sample stream that can be output to a rendering device such as display 124 and stored in reference picture memory 266 for future inter-picture prediction use.

[0058] Once some encoded pictures are reconstructed, they can be used as reference pictures for future prediction. Once an encoded picture is reconstructed and the encoded picture has been identified as a reference picture (e.g., by parser 254), the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.

[0059] Decoder component 122 can perform decoding operations according to predetermined video compression techniques that can be recorded in a standard (such as any of the standards described herein). The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that it adheres to the syntax of the video compression technique or standard as specified in the video compression technique documentation or standard and particularly in the profile documentation therein. Also, in order to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the assumptions of the hypothetical reference decoder (HRD) buffer management and metadata signaled in the encoded video sequence can further limit the limits set by the level.

[0060] Figure 3 is a block diagram illustrating server system 112 according to some embodiments. Server system 112 includes control circuit 302, one or more network interfaces 304, memory 314, user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, control circuit 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes one or more field programmable gate arrays (FPGA), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuit).

[0061] The network interface 304 may be configured to interface with one or more communication networks (e.g., wireless, wired, and / or optical networks). The communication network may be local, wide area, metropolitan area, vehicular and industrial, real-time, delay-tolerant, etc. Examples of communication networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANBus, etc. Such communication may be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., to other computer systems using local digital networks or wide area digital networks). Such communication may include communication to one or more cloud computing networks.

[0062] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 may include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. The output device 308 may include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., display screens or monitors), etc.

[0063] The memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices remote from the control circuit 302. Alternatively, the memory 314 or the non-volatile solid-state storage device within the memory 314 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: · An operating system 316, which includes programs for handling various basic system services and for performing hardware-related tasks; · A network communication module 318, which is used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); · A codec module 320 for performing various functions regarding encoding and / or decoding data (e.g., video data). In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following: ο A decoding module 322, which is configured to perform various functions regarding decoding encoded data, such as those previously described with respect to decoder component 122; and ο An encoding module 340, which is configured to perform various functions regarding encoding data, such as those previously described with respect to encoder component 106; and ● A picture memory 352, which is configured to store pictures and picture data, for example for use by the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0064] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to parser 254), a transformation module 326 (e.g., configured to perform various functions previously described with respect to the scaling / inverse transformation unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to the loop filter unit 256).

[0065] In some embodiments, the encoding module 340 includes an encoding module 342 (e.g., configured to perform various functions previously described with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 a subset of the modules shown. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.

[0066] Each of the above-identified modules stored in the memory 314 corresponds to a set of instructions for performing the functions described herein. The above-identified modules (e.g., instruction sets) need not be implemented as separate software programs, programs, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the above-identified modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0067] Although Figure 3 illustrates a server system 112 according to some embodiments, but Figure 3Rather, it is intended to be a functional description of various features that may exist in one or more server systems, rather than a structural diagram of the embodiments described herein. In practice, and as would be recognized by one of ordinary skill in the art, items shown separately may be combined and some items may be separated. For example, Figure 3 Some of the items shown separately in Figure 3 may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement server system 112, and how the features are distributed among them, will vary from one implementation to another and, optionally, will depend in part on the volume of data traffic processed by the server system during peak usage periods as well as during average usage periods. Example Coding and Decoding Processes and Techniques

[0068] As described above, some codecs (e.g., AV1) operate on pixel blocks. Each pixel block may be processed in a predictive transform coding scheme where prediction is obtained using intra-frame reference pixels, inter-frame motion compensation, or some combination of both. The residual from the prediction may undergo a transform (e.g., a 2-D unitary transform) to further remove spatial correlation, and the transform coefficients are quantized. Then, both the predicted syntax elements and the quantized transform coefficient indices may be entropy coded using arithmetic coding.

[0069] Figure 4A Illustrated is the calculation of a predicted block according to some embodiments. In Figure 4A the example of Figure 4A , intra-frame prediction is performed on the current block 402 to generate a predicted block 404. The current block 402 includes a set of samples (e.g., a pixel block) S 11 through S 44 , and the predicted block 404 includes a set of predictions P 11 through P 44 . Figure 4B Illustrated is the calculation of a residual block according to some embodiments. As Figure 4B shown, the predicted block 404 is subtracted from the current block 402 to generate a residual block 406 that includes a set of residuals R 11 through R 44 . For example, the corresponding difference between each sample and the corresponding prediction is calculated. Figure 4C Illustrated is the calculation of a reconstructed block according to some embodiments. As Figure 4C shown, the residual block 406 undergoes one or more transforms and quantization to generate a set of residual coefficients. The set of residual coefficients may be transmitted from an encoder component to a decoder component. The set of residual coefficients undergoes inverse quantization and inverse transform to generate a reconstructed residual block 408. The reconstructed residual block 408 is combined with the predicted block 404 (e.g., the reconstructed residuals of the reconstructed residual block 408 are added to the predictions of the predicted block 404) to generate a reconstructed block 410 corresponding to the current block 402.

[0070] To reduce redundancy in the residual signal, various residual prediction techniques have been developed. These techniques predict the residual signal and encode the corrected residual. Residual differential pulse code modulation (RDPCM) requires sample-based differential pulse code modulation along the horizontal or vertical axis. By doing so, each residual row in the horizontal mode (or each residual column in the case of vertical orientation) can be reconstructed at the decoder by summing the scaled differential pulse code modulation residual levels along the corresponding row (or column). RDPCM can be of an explicit type or an implicit type. The explicit type requires supplementary signaling of the direction and its application is limited to inter-frame prediction blocks. On the other hand, the implicit type does not require direction signaling and can only be applied to intra-frame prediction blocks, where the prediction direction is associated with the intra-frame prediction mode. Block-based differential pulse code modulation (BDPCM) performs sample-based differential pulse code modulation on the reconstructed samples rather than the residual samples. The indication of the use of the second mode occurs during the prediction mode reconstruction process. This signaling involves two syntax elements, each for both luminance and chrominance. For example, the initial syntax element flag indicates its utilization, while the second syntax element flag specifies the horizontal or vertical direction.

[0071] To reduce redundancy in the residual domain, a line-by-line residual prediction mode can be used to generate corrected residual blocks. The line-by-line residual domain prediction can be performed in the horizontal or vertical direction. For the horizontal prediction case, the prediction can be defined as shown in Equation 1 below, and for the vertical prediction case, the prediction can be defined as shown in Equation 2 below. where x and y represent the row index and column index respectively, and r(*,*) represents the pixel value in the original residual block (e.g., generated by the intra-frame prediction process).

[0072] In some embodiments, a bi-directional prediction mode is used to generate corrected residual blocks. The bi-directional prediction mode can be performed in the horizontal or vertical direction. For the horizontal prediction case for encoding, the prediction can be defined as shown in Equation 3 below, and for the vertical prediction case for encoding, the prediction can be defined as shown in Equation 4 below. where r(*,*) represents the pixel value in the original residual block (e.g., generated by the intra-frame prediction process).

[0073] For the horizontal prediction case for decoding, the prediction can be defined as shown in Equation 5 below, and for the vertical prediction case for decoding, the prediction can be defined as shown in Equation 6 below. where r ′ (*,*) represents the pixel values in the modified residual block. For example, the residual block values are determined by determined.

[0074] In some embodiments, intra prediction is performed on an encoded block or each sub-block in an encoded block, and a residual block is generated by subtracting a predicted block from the reconstructed samples of adjacent blocks. Then, short-range intra prediction is applied to the residual block to generate a modified residual block. For example, the modified residual signal can be calculated as shown in Equation 7 below. where represents the modified residual signal.

[0075] The reconstruction process can be performed by summing samples along a determined direction, as shown in Equation 8 and Equation 9 below.

[0076] Thus, a decoder according to an embodiment of the present disclosure can receive video data including a plurality of blocks and a plurality of residual coefficients from a video bitstream, the plurality of blocks including a first block. The first block is encoded in an intra prediction mode. In addition, a plurality of residual coefficients are generated by applying short-range intra prediction to the residual block of the first block. In addition, a residual block is generated by applying an intra prediction mode to the first block. Then, the decoder can generate a modified residual block of the first block according to the plurality of residual coefficients and use the modified residual block to reconstruct the first block.

[0077] In some embodiments, a flag is signaled to indicate whether to use a progressive residual prediction mode. In some embodiments, the flag is signaled separately for the luminance component and the chrominance component. In some embodiments, if the flag indicates the use of a progressive residual prediction mode, another flag is signaled to indicate whether the direction of the progressive residual prediction mode is vertical or horizontal. In some embodiments, when using the progressive residual prediction mode, the angular increment and / or the multi-reference line (mrl) index are inferred to be zero. In some embodiments, the transform block size is fixed at the minimum transform size (e.g., 4×4), and the progressive residual prediction mode is implemented on 4×4 residual blocks.

[0078] In some embodiments, forward skip coding (FSC) mode is enabled together with the row-by-row residual prediction mode (e.g., in a lossless coding scheme). FSC can be a simpler and more efficient residual coding method for coefficients obtained after a 2-D identity transform (IDTX). FSC-encoded blocks have less TU-level signaling because transform type signaling and block end index signaling are avoided for IDTX blocks, where the former reduces the symbol count in the TX_SET_INTRA set by 1 each. Finally, since FSC disables the signaling of multiple reference line (MRL) indices, filter intra-frame modes, and angle delta syntax when the transform type is IDTX, this can simplify the reconstruction process of FSC blocks, so it is designed as a more economical intra-frame block coding mode alternative.

[0079] Figure 4D FIG. 4 illustrates the calculation of the modified residual block and the difference block according to some embodiments. Figure 4D As shown, short-range intra prediction is applied to the residual block 406 to generate a residual Z 11 To Z 44 The modified residual block 420. As discussed in more detail below, short-range intra prediction may include row-by-row prediction and / or bidirectional prediction. Figure 4D Also shown is a diagram of a residual difference D generated by subtracting the modified residual block 420 from the residual block 406. 11 To D 44 According to some embodiments, the residual difference values ​​are used to generate residual coefficients (eg, via one or more transforms and quantizations). For example, the residual difference values ​​are used as a substitute for the residual of the residual block 406.

[0080] Figure 4E The calculation of residual blocks and reconstructed blocks during the decoding process according to some embodiments is illustrated. For example, modified residual coefficients are received and the modified residual blocks are recovered from these coefficients (e.g., using an inverse transform and inverse quantization process). After recovering the modified residual blocks, short-range intra prediction can be applied to recover the residual blocks. After recovering the residual blocks, intra prediction can be applied to recover the reconstructed blocks.

[0081] Figure 4F An example row-by-row prediction for an encoding process according to some embodiments is illustrated. Figure 4F In the example of ij , and the modified prediction block 420 includes sample p′ ij The residual block 406 may be obtained by applying intra prediction to the current block. Figure 4F The residual prediction block 420 in is obtained by column-by-column vertical prediction. Figure 4F In the example, from the residual block 406 sample r ij Subtract the residual prediction block 420 samples p′ij to obtain a modified residual block 422 having samples r′ ij In some embodiments, the modified residual block 422 is used to generate modified residual coefficients, and these modified residual coefficients are signaled in a subsequent bitstream.

[0082] Figure 4G illustrates an example line-by-line prediction for a decoding process according to some embodiments. In Figure 4G the example of, the modified residual block 422 is obtained (e.g., by applying an inverse transform to the residual coefficients received via the bitstream), and the modified residual block 422 includes samples r′ ij . Figure 4G The residual prediction block 420 in having samples p′ ij is obtained via column-by-column vertical prediction. In Figure 4G the example of, the samples p′ of the residual prediction block 420 ij are added to the samples r′ of the modified residual block 422 ij to recover a residual block 406 having samples r ij In some embodiments, the recovered residual block is used to obtain a reconstructed block of a current block.

[0083] Figure 5A and Figure 5B illustrate example residual blocks according to some embodiments. Figure 5A shows a residual block 502 that includes a set of residuals R corresponding to rows 1 to 8 and columns 1 to 8 11 to R 88 . Figure 5B shows the residual block 502 having index row 1 (e.g., corresponding to row 1), index row 2 (e.g., corresponding to row 4), index row 3 (e.g., corresponding to row 2), and index row 4 (e.g., corresponding to row 3). In Figure 5B the residual block 502 also includes index row A (e.g., corresponding to column 1), index row B (e.g., corresponding to column 4), index row C (e.g., corresponding to column 2), and index row D (e.g., corresponding to column 3).

[0084] Additional residual prediction techniques are described below. The disclosed techniques can be used alone or in any combination. These techniques as well as the encoder and decoder methods can be performed using a processing circuit, which can include one or more processors or integrated circuits. As an example, a program stored in a non-transitory computer-readable medium can be executed by one or more processors.

[0085] Figure 6AFIG. is a flowchart of a method 600 for encoding a video according to some embodiments. The method 600 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions for execution by the control circuit. In some embodiments, the method 600 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.

[0086] The system receives (602) video data including a plurality of blocks, the plurality of blocks including a first block (e.g., current block 402) to be encoded in a first intra prediction mode. For example, the system receives video data from a video source (e.g., video source 104). The system generates (604) a residual block (e.g., residual block 406) of the first block by applying the first intra prediction mode in a first direction (e.g., horizontal or vertical) to the first block. In some embodiments, the intra prediction mode is a directional intra prediction. The system generates (606) a modified residual block (e.g., modified residual block 420) of the first block by applying a second intra prediction to the residual block in a second direction. In some embodiments, the second intra prediction is a short-range intra prediction. In some embodiments, the short-range intra prediction is a progressive prediction. In some embodiments, the short-range intra prediction is a bidirectional prediction. The system signals (608) the modified residual block via a video bitstream. In some embodiments, the system signals the difference between the modified residual block and the residual block (e.g., as a residual coefficient).

[0087] Figure 6B FIG. is a flowchart of a method 650 for decoding a video according to some embodiments. The method 650 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions for execution by the control circuit. In some embodiments, the method 650 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.

[0088] The system receives (652) from a video bitstream (e.g., via channel 218) a first block (e.g., current block 402) and a plurality of residual coefficients for the first block (e.g., Figure 4Dvideo data of the residual coefficients shown in []. The system generates (654) a modified residual block (e.g., reconstructed residual block 408) of the first block based on the plurality of residual coefficients. For example, the system applies inverse quantization and one or more inverse transforms to the plurality of residual coefficients to generate a modified residual block (e.g., reconstructed residual block 408). The system generates (656) a reconstructed residual block by applying an intra prediction in a first direction (e.g., short-range intra prediction) to the modified residual block. The system decodes (658) the first block by applying an intra prediction in a second direction to the reconstructed residual block. For example, the system generates a reconstructed block (e.g., reconstructed block 410) for the first block.

[0089] Although Figure 6A and Figure 6B a plurality of logical stages are illustrated in a particular order, stages that are not order-dependent can be reordered, and other stages can be combined or split. A particular reordering or other grouping not specifically mentioned will be apparent to those of ordinary skill in the art, and thus the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that the stages can be implemented in hardware, firmware, software, or any combination thereof.

[0090] In some embodiments, intra prediction is performed on an encoded block or each sub-block in an encoded block to generate a residual block by subtracting a predicted block from reconstructed samples of adjacent blocks. In some embodiments, short-range intra prediction is applied to the residual block to generate a modified residual block. In some embodiments, intra prediction is performed on an M×N block regardless of the size of the encoded block. Example values of M and N include, but are not limited to, 1, 2, 4, 8, 16, 32, and 64. In some embodiments, short-range intra prediction is applied to the residual block for lossless coding mode.

[0091] In some embodiments, row-by-row prediction is used as short-range intra prediction. For example, the residuals in a particular row or column are predicted using their neighboring previous row / column, and the difference between these residuals and the predicted residuals is used as input to subsequent transform, quantization, or entropy coding processes. In some embodiments, row-by-row prediction is applied to the samples / residuals in all rows or columns except the samples / residuals in the first row / column. In some embodiments, for the first row / column, prediction is performed using residual samples generated from multiple reconstructed rows / columns of adjacent blocks. In some embodiments, row-by-row prediction is performed on these residuals in the horizontal direction. For example, the predicted residual of the first row is set to zero, and the predicted residuals of subsequent rows are predicted from their neighboring previous rows. In some embodiments, column-by-column prediction is performed on these residuals in the vertical direction. For example, the predicted residual of the first column is set to zero, and the predicted residuals of subsequent columns are predicted by their neighboring previous columns.

[0092] In some embodiments, the short-range intra prediction is bidirectional prediction. In some embodiments, bidirectional prediction is applied to each M×N residual block. For example, the prediction residuals of the residuals in the first index row and the second index row of the M×N residual block are set to zero, and the weighted average of the residuals in the first index row and the second index row is used to predict the residuals in the third index row and the fourth index row. In some embodiments, the difference between the residual and the prediction residual is used as the input for subsequent transformation, quantization, or entropy coding processes. Example values of M and N include but are not limited to 1, 2, 4, 8, 16, 32, and 64. In some embodiments, the first row and the second row are not adjacent rows. Examples of the first index row and the second index row include but are not limited to the first row and the fourth row along a given direction, respectively. Examples of the third index row and the fourth index row include but are not limited to the second row and the third row along a given direction, respectively. In some embodiments, the weighting factors for weighted averaging the residuals in the first index row and the second index row depend on the distance between the residual and its predicted value. For example, when predicting the residuals in the second row / column, the weighting factors for the residuals in the first index row and the second index row are {2 / 3, 1 / 3} or {3 / 4, 1 / 4}. As another example, when predicting the residuals in the third row / column, the weighting factors for the residuals in the first index row and the second index row are {1 / 3, 2 / 3} or {1 / 4, 3 / 4}.

[0093] In some embodiments, bidirectional prediction is performed in the horizontal direction. For example, the prediction residuals in the first index row and the second index row are set to zero. In this example, the residuals in the third index row and the fourth index row are predicted by weighted averaging the residuals in the first row and the fourth row. In some embodiments, bidirectional prediction is performed in the vertical direction. For example, the prediction residuals in the first index row and the second index row are set to zero. In this example, the residuals in the third index row and the fourth index row are predicted by weighted averaging the residuals in the first row and the fourth row.

[0094] In some embodiments, for short-range intra prediction, the weighted average of the residuals from multiple adjacent rows is used to predict the residuals in subsequent rows. In some embodiments, the weighting factors for weighted averaging the residuals in multiple adjacent rows depend on the distance between the residuals in the current row and the residuals in the adjacent rows. In some embodiments, the residuals in two adjacent rows are used to predict the residuals in the current row, the weighting factor of the residual in the nearest adjacent row is set to a first value, and the weighting factor of the residual in the other row is set to a second value. Examples of the first value and the second value include but are not limited to 2 / 3 and 1 / 3, respectively.

[0095] In some embodiments, multiple short - range prediction methods are sequentially applied to the residual block. For example, first, the row - by - row prediction method is applied to the residual block to generate a modified residual block, and then the bidirectional prediction method is applied to the modified residual block to generate the final residual block.

[0096] In some embodiments, a flag is signaled in the bitstream to indicate which short - range intra - prediction method is applied to the residual block. In some embodiments, two separate flags are used to signal the use of short - range prediction for the luma residual block plane and the chroma residual block plane. In some embodiments, the direction of short - range prediction for the residual block is signaled in the bitstream. In some embodiments, the direction of short - range residual block prediction is inferred from the intra - prediction mode. In some embodiments, the context for entropy - coding the flag for short - range prediction for the chroma residual block depends on the corresponding luma flag. In some embodiments, whether one or more short - range intra - predictions in short - range intra - prediction are applied is signaled in the high - level syntax (including but not limited to sequence flags, GOP flags, picture flags, sub - picture flags, slice flags, or tile - level flags).

[0097] In some embodiments, a first direction is adopted in the intra - prediction on the encoded block or each sub - block in the encoded block, so that a residual block is generated by subtracting the predicted block from the reconstructed samples of the adjacent blocks. Then, a second direction is adopted in the short - range residual prediction to predict the residuals, and the difference between these residuals and the predicted residuals is used as the input for subsequent processing, which includes but is not limited to transformation, quantization, entropy - coding, and in - loop filtering. At the decoder, the difference between the residuals and the predicted residuals is parsed, and then these differences are added to the predicted residuals to derive the reconstructed residual samples.

[0098] In some embodiments, the direction of the first direction is different from the direction of the second direction. In some embodiments, the direction of the second direction is the same as the direction of the first direction. In some embodiments, the value of the first direction is used as the context for entropy - coding the second direction. In some embodiments, the high - level syntax (including but not limited to sequence level, frame level, slice level, super - block level) is signaled in the bitstream to indicate whether the second direction is the same as the first direction.

[0099] In some embodiments, row - by - row prediction or bidirectional prediction is used as short - range prediction. For example, row - by - row prediction is adopted in the residual block as a short - range intra - prediction, and the previous row adjacent to it is used to predict the residuals in a specific row. As another example, bidirectional prediction is adopted in the residual block as a short - range intra - prediction, and the predicted residuals of the residuals in the first index row and the second index row of the residual block are set to zero, while the weighted average of the residuals in the first index row and the second index row is used to predict the residuals in the third index row and the fourth index row.

[0100] In some embodiments, the angle of a second direction for short - range residual prediction is implicitly determined based on the angle of a first direction for intra - prediction. In some embodiments, if the prediction angle for intra - prediction is closer to the horizontal direction than to the vertical direction, short - range prediction is performed on the residual in the horizontal direction. In some embodiments, if the prediction angle for intra - prediction is closer to the vertical direction than to the horizontal direction, short - range prediction is performed on the residual in the vertical direction. In some embodiments, if the prediction angle for intra - prediction is closer to the diagonal direction than to the horizontal direction or the vertical direction, short - range prediction is performed on the residual in the diagonal direction. In some embodiments, short - range prediction is applied to N nominal angles for intra - prediction. For example, N is equal to 4, 6, 8, or 10.

[0101] In some embodiments, one or more syntax elements are signaled in the bitstream to indicate the direction / angle of short - range prediction. In some embodiments, a syntax element indicating the direction of short - range prediction is signaled in the bitstream. In some embodiments, a first syntax element is used to indicate the nominal / principal direction, and a second syntax element is used to indicate the angular increment relative to the nominal direction. In some embodiments, the supported values of the increment angle are predefined in a lookup table, and the index of the increment angle in the lookup table is signaled in the bitstream. As an example, a first syntax is signaled to indicate whether the direction for short - range residual prediction is vertical or horizontal, and then a second syntax is signaled to indicate the angular increment relative to the specified principal direction.

[0102] In some embodiments, a first syntax element is used to indicate the nominal direction; a second syntax element is used to indicate whether the angular increment is zero. If the angular increment is not zero, a third syntax element and a fourth syntax element are further used. The third syntax element is used to indicate the positive or negative value of the angular increment. The fourth syntax element is used to indicate the absolute angular increment.

[0103] In some embodiments, only the angular increment syntax is signaled to derive the second direction (e.g., the prediction direction used for residual prediction), and then the direction used for residual prediction is derived by adding the angular increment value to the nominal prediction direction (or prediction direction) of the intra - prediction mode.

[0104] Some embodiments include transform coding techniques for short - range residual prediction. In some embodiments, to process a residual block, a transform coding mode is signaled at a first processing unit level and a transform process is performed at a second processing unit level, where the transform coding mode refers to any parameter or operation involved in the transform process and a transform method can be applied in the forward transform process at the encoder or the inverse transform process at the encoder and / or decoder. For example, the scope of the first processing unit level and the second processing unit level can include sequence level, frame level, super - block level, coded - block level, prediction - block level, or transform - block level.

[0105] In some embodiments, a residual block or a modified residual block generated from intra - prediction is used as an input to the transform process. In some embodiments, intra - prediction is performed on a coded block or each sub - block in a coded block, so as to generate a residual block by subtracting the prediction block from the reconstructed samples of adjacent blocks. In some embodiments, a short - range intra - prediction method is applied to the residual block, so as to generate a modified residual block. In some embodiments, the modified residual block is used as an input for subsequent transform coding. In some embodiments, different transform kernels can be applied to the modified residual block. For example, no transform is applied to the modified residual block (or the transform kernel is an identity transform). As an example, no transform (or an identity transform) is applied in one direction (e.g., horizontal or vertical), and a lossless transform (e.g., Hadamard transform) is applied in the other direction.

[0106] In some embodiments, the first processing unit level is the same as the second processing unit level. In some embodiments, the type of the transform coding mode is signaled at the coded - block level, the modified residual block size is the same as the coded - block size, and the transform - coded block size is also the same as the coded - block size. For example, a syntax element for the transform - block size can be inferred based on the modified residual block size or the coded - block size, and a syntax element for the transform - block size does not need to be signaled.

[0107] In some embodiments, the first processing unit level is different from the second processing unit level. In some embodiments, the type of the transform coding kernel is signaled at the coding block level, but the transform block size used to perform the transform process is smaller than the coding block size. In some embodiments, the transform block size is fixed, and the same transform coding kernel is applied to the transform coding blocks within an encoded block. For example, there is no need to signal a syntax element for the transform block size in the bitstream. In one example, the identity transform is signaled at the coding block level, and the transform size is fixed at M×N regardless of the coding block size. In another example, the Hadamard transform (or a different lossless transform) is signaled at the coding block level, and the transform size is fixed at M×N regardless of the coding block size. In some embodiments, M and N are selected to correspond to the minimum allowable transform size (e.g., both M and N are equal to 4).

[0108] In some embodiments, the transform block size and the transform coding mode are determined by a given cost metric at the encoder. For example, the best transform coding type and the best transform size are signaled at the coding block level. In one example, the cost metric is the rate-distortion cost used in rate-distortion optimization.

[0109] In some embodiments, whether the first processing unit level is different from the second processing unit level is signaled at a high-level syntax (including but not limited to sequence, GOP, frame, or slice level).

[0110] In some embodiments, the transform scheme includes applying a transform skip (or identity transform) in one direction (e.g., horizontal or vertical direction) and applying the Hadamard transform in the other direction. In some embodiments, this transform scheme is only applied to the lossless coding mode. In some embodiments, this transform scheme is only applied to a specific M×N transform block size (e.g., 4×4 transform block size). In some embodiments, it is signaled separately for each direction whether to select the transform skip (or identity transform) or the Hadamard transform.

[0111] Now turning to some example embodiments.

[0112] (A1)In one aspect, some embodiments include a method of video encoding (e.g., method 600). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). In some embodiments, the method is performed at a source coding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving video data including a plurality of blocks, the plurality of blocks including a first block (e.g., a chrominance or luminance block), wherein the first block is to be encoded in a first intra prediction mode (e.g., a first type of intra prediction); (ii) generating a residual block of the first block by applying the first intra prediction mode to the first block along a first direction; (iii) generating a corrected residual block of the first block by applying a second intra prediction mode (e.g., a second type of intra prediction) to the residual block along a second direction; and (iv) signaling the corrected residual block via a video bitstream. For example, in intra prediction on an encoded block (or each sub-block in the encoded block), the first direction is adopted, so that a residual block is generated by subtracting a predicted block from reconstructed samples of adjacent blocks. Then, the second direction is adopted in short-range residual prediction to predict the residual, and the difference between these residuals and the predicted residual is used as an input for subsequent processing such as transformation, quantization, entropy encoding, and / or in-loop filtering. In some embodiments, generating the corrected residual block of the first block includes applying multiple short-range intra predictions.

[0113] (A2)In some embodiments of A1, the method further includes transmitting the encoded first block via the video bitstream.

[0114] (A3)In some embodiments of A1 or A2, the method further includes determining the difference between the residual block and the corrected residual block. For example, the residuals in a particular row or column are predicted using their neighboring previous rows. The difference between these residuals and the predicted residual is used as an input for subsequent transformation, quantization, or entropy encoding processes.

[0115] (A4)In some embodiments of any one of A1 to A3, the plurality of residual coefficients correspond to the difference between: the residuals of the residual block, and the corrected residuals obtained by applying the second intra prediction to the residual block.

[0116] (A5)In some embodiments of any one of A1 to A4, the first direction and the second direction are different directions. In some embodiments, the first direction and the second direction are the same direction.

[0117] (A6)In some embodiments of any one of A1 to A5, the first direction is used as a context for entropy encoding the second direction.

[0118] In some embodiments of any one of A1 to A6, the method further includes signaling a syntax element indicating whether the second direction is the same as the first direction.

[0119] In some embodiments of any one of A1 to A7, the second intra prediction is short - range intra prediction. In some embodiments, the short - range intra prediction includes a progressive prediction in which the residuals in a particular row or column are predicted using a neighboring previous row or column. In some embodiments, the short - range intra prediction includes a bidirectional prediction in which the residuals in a third index row and a fourth index row are predicted using a weighted average of the residuals in a first index row and a second index row.

[0120] In some embodiments of any one of A1 to A8, the angle of the second direction is determined based on the angle of the first direction.

[0121] In some embodiments of any one of A1 to A9, the second direction is selected from the group consisting of N nominal angles for intra prediction, where N is a positive integer.

[0122] In some embodiments of any one of A1 to A10, the method further includes signaling one or more syntax elements indicating the direction and / or angle of the second intra prediction. In some embodiments, the one or more syntax elements include a first syntax indicating a nominal direction and a second syntax indicating an angular increment applied to the nominal direction. In some embodiments, the one or more syntax elements include a first syntax element indicating a nominal direction, a second syntax element indicating whether the angular increment is zero, a third syntax element indicating whether the angular increment is positive or negative, and a fourth syntax element indicating an absolute angular increment.

[0123] (B1)In another aspect, some embodiments include a method of video decoding (e.g., method 650). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). In some embodiments, the method is performed at a parser (e.g., parser 254), a motion prediction component (e.g., motion compensation prediction unit 260), and / or an intra prediction component (e.g., intra prediction unit 262). The method includes: (i) receiving video data (e.g., an encoded video sequence) including a plurality of blocks from a video bitstream, the plurality of blocks including the first block and a plurality of residual coefficients for the first block; (ii) generating a modified residual block for the first block from the plurality of residual coefficients; (iii) generating a reconstructed residual block by applying a first intra prediction to the modified residual block along a first direction; and (iv) reconstructing the first block by applying a second intra prediction to the reconstructed residual block along a second direction. For example, a residual block represents the difference between a reconstructed sample and a corresponding predicted value. In some embodiments, the first intra prediction and the second intra prediction are applied as part of a lossless coding scheme. In some embodiments, the plurality of residual coefficients are generated by applying the first intra prediction to the residual block of the first block in the first direction, and the residual block is generated by applying the second prediction mode to the first block in the second direction. In some embodiments, the modified residual block is generated by applying an inverse transform and / or inverse quantization to the plurality of residual coefficients.

[0124] (B2)In some embodiments of B1, the plurality of residual coefficients correspond to the difference between: (i) the residual of the residual block of the first block, and (ii) a predicted residual obtained by applying the first intra prediction to the residual block. For example, the differences between these residuals and the predicted residual are parsed and then added to the predicted residual to derive reconstructed residual samples.

[0125] (B3)In some embodiments of B1 or B2, the first direction and the second direction are different directions. For example, the first direction is vertical and the second direction is horizontal, or vice versa. In some embodiments, the first direction and the second direction are the same direction.

[0126] (B4)In some embodiments of B1 to B3, the second direction is used as a context for entropy coding of the first direction. For example, the direction of the intra prediction applied to the current block is used as a context for the short-range intra prediction of the residual block applied to the current block.

[0127] (B5)In some embodiments of B1 to B4, the video bitstream further includes a syntax element indicating whether the second direction is the same as the first direction. In some embodiments, the syntax element is a high-level syntax (HLS) element. In some embodiments, the HLS is signaled at a level higher than the block level. For example, the HLS may correspond to a sequence level, a frame level, a slice level, or a tile level. As another example, the HLS may be signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), a slice header, a picture header, a tile header, and / or a CTU header.

[0128] (B6)In some embodiments of B1 to B5, the first intra prediction is short-range intra prediction. For example, progressive prediction or bi-directional prediction may be used as short-range prediction.

[0129] (B7)In some embodiments of B6, the short-range intra prediction includes progressive prediction, in which the residuals in a particular row or column are predicted using neighboring previous rows or columns. For example, progressive prediction is adopted as a short-range intra prediction. The residuals in a particular row or column are predicted using its neighboring previous row. In some embodiments, the residuals in a subsequent row are predicted using a weighted average of the residuals from multiple neighboring rows and / or multiple previous rows. For example, the weighting factor for weighting the residuals in multiple neighboring rows may depend on the distance between the residuals in the current row and the neighboring rows. In some embodiments, the residuals in two neighboring rows are used to predict the residuals in the current row, and the weighting factor for the residuals in the nearest neighboring row is set to a first value, and the weighting factor for the residuals in the other row is set to a second value. Examples of the first value and the second value include, but are not limited to, 2 / 3 and 1 / 3, respectively.

[0130] (B8)In some embodiments of B6, the short-range intra prediction includes bi-directional prediction, in which the residuals in a third index row and a fourth index row are predicted using a weighted average of the residuals in a first index row and a second index row. For example, bi-directional prediction is applied to each M×N residual block. In some embodiments, the residuals in a particular index row are predicted using a weighted average of the residuals from multiple neighboring rows and / or multiple previous rows. For example, the weighting factor for weighting the residuals in multiple neighboring rows may depend on the distance between the residuals in the current row and the neighboring rows.

[0131] (B9)In some embodiments of any one of B1 to B8, the angle of the first direction is determined based on the angle of the second direction. For example, the angle of the first direction for short - range residual prediction can be implicitly determined based on the angle of the second direction for intra - prediction. As an example, if the prediction angle for intra - prediction is closer to the horizontal direction than the vertical direction, short - range prediction of the residual is performed in the horizontal direction. As another example, if the prediction angle for intra - prediction is closer to the vertical direction than the horizontal direction, short - range prediction of the residual is performed in the vertical direction. As another example, if the prediction angle for intra - prediction is closer to the diagonal direction than the horizontal or vertical direction, short - range prediction of the residual is performed in the diagonal direction.

[0132] (B10)In some embodiments of any one of B1 to B9, the first direction is selected from the group consisting of N nominal angles for intra - prediction, where N is a positive integer. For example, short - range prediction can be applied to the N (e.g., 8) nominal angles for intra - prediction.

[0133] (B11)In some embodiments of any one of B1 to B10, the video bitstream includes one or more syntax elements indicating the direction and / or angle of the first intra - prediction. In some embodiments, the one or more syntax elements are high - level syntax (HLS) elements. As an example, a syntax element indicating the direction of short - range prediction is signaled in the bitstream. In some embodiments, an angle - increment syntax is signaled to derive the second direction (e.g., the prediction direction used for residual prediction), and then the direction used for residual prediction is derived by adding the angle - increment value to the nominal prediction direction (or prediction direction) of the intra - prediction mode.

[0134] (B12)In some embodiments of B11, the one or more syntax elements include a first syntax indicating the nominal direction and a second syntax indicating the angle increment applied to the nominal direction. For example, the first syntax element is used to indicate the nominal / principal direction, and the second syntax element is used to indicate the angle increment relative to the nominal direction. In some embodiments, the supported values of the increment angle are predefined in a look - up table, and the index of the increment angle in the look - up table is signaled in the bitstream. In some embodiments, the first syntax is signaled to indicate whether the direction for short - range residual prediction is vertical or horizontal, and then the second syntax is signaled to indicate the angle increment relative to the specified principal direction.

[0135] (B13) In some embodiments of B11, the one or more syntax elements include a first syntax element indicating a nominal direction, a second syntax element indicating whether the angular increment is zero, a third syntax element indicating whether the angular increment is positive or negative, and a fourth syntax element indicating an absolute angular increment. For example, the first syntax element is used to indicate the nominal direction, and the second syntax element is used to indicate whether the angular increment is zero. If the angular increment is not zero, the third and fourth syntax elements are further used. The third syntax element is used to indicate the positive or negative value of the angular increment. The fourth syntax element is used to indicate the absolute angular increment.

[0136] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including a control circuit (e.g., control circuit 302) and a memory coupled to the control circuit (e.g., memory 314), the memory storing one or more sets of instructions configured to be executed by the control circuit, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A13 and B1 to B8 above). In another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions, the instructions being executed by a control circuit of a computing system, the instructions including instructions for performing any of the methods described herein (e.g., A1 - A11, B1 - B13 above).

[0137] It will be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should also be understood that when used in this specification, the terms "comprises" and / or "comprising" indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0138] As used herein, the term "if" can be interpreted to mean "when" or "while" or "in response to determining" or "in accordance with determining" or "in response to detecting" that the stated prerequisite is true, depending on the context. Similarly, depending on the context, the phrase "if it is determined that [the stated prerequisite is true]" or "if [the stated prerequisite is true]" or "when [the stated prerequisite is true]" can be interpreted to mean "upon determining" or "in response to determining" or "in accordance with determining" or "upon detecting" or "in response to detecting" that the stated prerequisite is true.

[0139] For purposes of explanation, the above description has been presented with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teaching. The embodiments were chosen and described in order to best explain the operating principles and the practical application, thereby enabling others skilled in the art to understand.

Claims

1. A video decoding method executed in a computing system having a memory and one or more processors, characterized in that, The method includes: Receiving video data including a plurality of blocks from a video bitstream, the plurality of blocks including a first block and a plurality of residual coefficients of the first block; Generating a modified residual block of the first block from the plurality of residual coefficients; Generating a reconstructed residual block by applying a first intra prediction to the modified residual block along a first direction; and Reconstructing the first block by applying a second intra prediction to the reconstructed residual block along a second direction.

2. The method according to claim 1, characterized in that, The plurality of residual coefficients correspond to the difference between: (i) the residual of the residual block of the first block, and (ii) the predicted residual obtained by applying the first intra prediction to the residual block.

3. The method according to claim 1, wherein The first direction and the second direction are different directions.

4. The method according to claim 1, wherein The second direction is used as a context for entropy coding the first direction.

5. The method according to claim 1, characterized in that The video bitstream further includes a syntax element indicating whether the second direction is the same as the first direction.

6. The method according to claim 1, wherein The first intra prediction is a short-range intra prediction.

7. The method according to claim 6, wherein The short-range intra prediction includes a line-by-line prediction, in which the residuals in a specific row or column are predicted using adjacent previous rows or columns.

8. The method according to claim 6, characterized in that The short-range intra prediction includes a bidirectional prediction, in which the residuals in a third index row and a fourth index row are predicted using a weighted average of the residuals in a first index row and a second index row.

9. The method according to claim 1, characterized in that, The angle of the first direction is determined based on the angle of the second direction.

10. The method according to claim 1, characterized in that The first direction is selected from a group consisting of N nominal angles for intra prediction, where N is a positive integer.

11. The method according to claim 1, wherein The video bitstream includes one or more syntax elements indicating the direction and / or angle of the first intra prediction.

12. The method according to claim 11, wherein The one or more syntax elements include: a first syntax indicating a nominal direction, and a second syntax indicating an angle increment applied to the nominal direction.

13. The method according to claim 11, wherein The one or more syntax elements include: a first syntax element indicating a nominal direction, a second syntax element indicating whether the angle increment is zero, a third syntax element indicating whether the angle increment is positive or negative, and a fourth syntax element indicating an absolute angle increment.

14. A computing system, characterized in that, Comprising: A control circuit; A memory; And One or more sets of instructions stored in the memory and configured to be executed by the control circuit, the one or more sets of instructions including instructions for: Receiving video data including a plurality of blocks from a video bitstream, the plurality of blocks including a first block and a plurality of residual coefficients of the first block; Generating a modified residual block of the first block from the plurality of residual coefficients; Generating a reconstructed residual block by applying a first intra prediction to the modified residual block along a first direction; And Reconstructing the first block by applying a second intra prediction to the reconstructed residual block along a second direction.

15. The computing system according to claim 14, wherein The plurality of residual coefficients correspond to the difference between: (i) the residual of the residual block of the first block, and (ii) the predicted residual obtained by applying the first intra prediction to the residual block.

16. The computing system according to claim 14, wherein The first direction and the second direction are different directions.

17. The computing system according to claim 14, wherein The second direction is used as a context for entropy coding the first direction.

18. A non-volatile computer-readable storage medium, characterized in that, The non - volatile computer - readable storage medium stores one or more sets of instructions configured to be executed by a computing device having a control circuit and a memory, the one or more sets of instructions including instructions for: Receiving video data including a plurality of blocks from a video bitstream, the plurality of blocks including a first block and a plurality of residual coefficients of the first block; Generating a modified residual block of the first block from the plurality of residual coefficients; Generating a reconstructed residual block by applying a first intra - prediction to the modified residual block along a first direction; And Reconstructing the first block by applying a second intra - prediction to the reconstructed residual block along a second direction.

19. The non-volatile computer-readable storage medium according to claim 18, wherein The plurality of residual coefficients corresponds to the difference between (i) the residual of the residual block of the first block and (ii) the predicted residual obtained by applying the first intra - prediction to the residual block.

20. The non-volatile computer-readable storage medium according to claim 18, wherein The first direction and the second direction are different directions.