Short range prediction for residual blocks

By using short-distance intra prediction technology to correct the residual blocks in the video encoding and decoding technology, the problem of residual signal redundancy in the prior art is solved, the encoding efficiency is improved and the transmission bandwidth is reduced.

CN120226345APending Publication Date: 2025-06-27TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380080684.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-04
Filing Date
2023-10-30
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the existing video encoding and decoding technology, the redundancy of residual signals cannot be effectively reduced, resulting in inefficiency in encoding efficiency and transmission bandwidth.

Method used

Short-distance intra prediction technology is used to correct the generated residual blocks, generate corrected residual blocks, and signal the corrected residual blocks through video bit streams to reduce redundancy.

Benefits of technology

By reducing redundancy in the residual domain, the performance of lossless encoding is improved, the number of encoded bits is reduced, the encoding and decoding efficiency is improved, and the transmission bandwidth is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226345A_ABST
    Figure CN120226345A_ABST
Patent Text Reader

Abstract

Various embodiments described herein include methods and systems for video coding and decoding. In one aspect, a method of video decoding includes receiving video data including a first block and at least two residual coefficients from a video bitstream. The first block is encoded in an intra prediction mode, and the at least two residual coefficients are generated by applying short range intra prediction to a residual block of the first block. The residual block is generated by applying the intra prediction mode to the first block. The method includes generating a modified residual block of the first block according to the at least two residual coefficients, and generating a reconstructed residual block of the first block, where the reconstructed residual block is generated using the intra prediction block and the modified residual block. The method also includes reconstructing the first block using the modified residual block.
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This application claims priority to U.S. Patent Application No. 18 / 480,957, filed on October 4, 2023, entitled "Short-Distance Prediction for Residual Blocks". Technical Field

[0002] The disclosed embodiments generally relate to the encoding, decoding, and compression of images and videos, including but not limited to systems and methods for predicting residual information. Background Art

[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablets, digital cameras, digital video recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. These electronic devices transmit and receive digital video data over a communication network or otherwise convey digital video data, and / or store the digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the memory resources of the storage device, video encoding can be used to compress the video data according to at least one video encoding standard before transmitting or storing the video data. Video encoding can be performed by hardware and / or software on an electronic / client device or server that provides cloud services.

[0004] Video encoding typically uses prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), which utilize the redundancy inherent in video data. Video encoding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4), respectively. Versatile Video Coding (VVC / H.266) is a video compression standard designed to succeed HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (version 1) and 2022 (version 2), respectively. Alliance for Open Media Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verified version 1.0.0 with Specification Errata 1 was released. Summary of the Invention

[0005] A comprehensive video codec typically includes at least two components, such as intra / inter prediction, transform coding, quantization, residual coding, and loop filtering. To further reduce the residual signal, various residual prediction techniques have been developed. These techniques predict the residual signal and encode the high-order residual into the bitstream. This disclosure describes methods and systems for enhancing video (image) compression, including advanced residual prediction techniques.

[0006] According to some embodiments, a method of video encoding is provided. The method includes: (i) receiving video data including at least two blocks, the at least two blocks including a first block, wherein the first block is to be encoded in an intra prediction mode; (ii) generating a residual block of the first block by applying the intra prediction mode to the first block; (iii) generating a modified residual block of the first block by applying short-range intra prediction to the residual block; and (iv) signaling the modified residual block through a video bitstream.

[0007] According to some embodiments, a method of video decoding is provided. The method includes: (i) receiving from a video bitstream video data including at least two blocks, the at least two blocks including a first block and at least two residual coefficients of the first block; (ii) generating a modified residual block of the first block based on the at least two residual coefficients; (iii) generating a reconstructed residual block of the first block, wherein the reconstructed residual block is generated using an intra prediction block and the modified residual block; and (iv) reconstructing the first block using the reconstructed residual block.

[0008] According to some embodiments, a computing system, such as a streaming system, a server system, a personal computer system, or other electronic device, is provided. The computing system includes a control circuit and a memory storing at least one set of instructions. The at least one set of instructions includes instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).

[0009] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores at least one set of instructions executable by a computing device. The at least one set of instructions includes instructions for performing any of the methods described herein.

[0010] Accordingly, at least two devices and systems employing methods for encoding and decoding video are disclosed. These methods, devices, and systems may supplement or replace conventional methods, devices, and systems for video encoding / decoding. The features and advantages described in the specification are not necessarily all inclusive, and in particular, given the figures, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Additionally, it should be noted that the language used in the specification is primarily chosen for readability and guidance purposes and is not necessarily chosen to depict or limit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To understand the present disclosure in more detail, a more specific description may be made with reference to the features of various embodiments, some of which are illustrated in the drawings. However, the drawings only illustrate the relevant features of the present disclosure and thus should not be considered limiting, as those skilled in the art will understand, after reading this disclosure, that the specification may permit other effective features.

[0012] Figure 1 A block diagram of an example communication system according to some embodiments is shown.

[0013] Figure 2A A block diagram of example elements of an encoder assembly according to some embodiments is shown.

[0014] Figure 2B A block diagram of example elements of a decoder assembly according to some embodiments is shown.

[0015] Figure 3 A block diagram of an example server system according to some embodiments is shown.

[0016] Figure 4A Calculation of a prediction block according to some embodiments is shown.

[0017] Figure 4B Calculation of a residual block according to some embodiments is shown.

[0018] Figure 4C Calculation of a reconstruction block according to some embodiments is shown.

[0019] Figure 4D Calculation of a residual prediction block and a corrected residual block according to some embodiments is shown.

[0020] Figure 4E Calculation of a residual block and a reconstruction block according to some embodiments is shown.

[0021] Figure 4F and 4G Example line-by-line prediction according to some embodiments is shown.

[0022] Figure 5A and 5B illustrates an example residual block according to some embodiments.

[0023] Figure 6A illustrates a flowchart of a method for encoding a video according to some embodiments.

[0024] Figure 6B illustrates a flowchart of a method for decoding a video according to some embodiments.

[0025] By convention, the various features shown in the drawings need not be drawn to scale, and throughout the specification and drawings, the same reference numerals may be used to denote the same features. Detailed Description

[0026] The present disclosure describes systems and methods for predicting residual information. The systems and methods described herein can improve the performance of lossless coding by reducing redundancy in the residual domain. In some embodiments, a per-line residual domain prediction mode is implemented for vertical and / or horizontal intra prediction modes. For the luma plane and the chroma plane, flags indicating the use of the proposed method can be signaled separately, while the U plane and the V plane can share a flag.

[0027] In some embodiments, a residual block is generated (e.g., by applying an intra prediction mode to a current block), and then a modified residual block is generated (e.g., by applying a short-range intra prediction to the residual block). Generating and using the modified residual block can reduce redundancy in the residual domain. Reducing redundancy reduces the number of bits required to signal the residual (e.g., improves coding / decoding efficiency and reduces transmission bandwidth). Example Systems and Devices

[0028] Figure 1 illustrates a block diagram of a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and at least two electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via at least one network. In some embodiments, the communication system 100 is a streaming system, such as for use with video-enabled applications such as video conferencing applications, digital television applications, and media storage and / or distribution applications.

[0029] The source device 102 includes a video source 104 (e.g., a camera assembly or a media memory) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates at least one encoded video bitstream from the video stream. The video stream from the video source 104 may have a high data volume compared to the encoded video bitstream 108 generated by the encoder component 106. Since the encoded video bitstream 108 has a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video to at least one network 110).

[0030] The at least one network 110 represents any number of networks that convey information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wired (wired) and / or wireless communication networks. The at least one network 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0031] The at least one network 110 includes a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from the source device 102). The server system 112 includes a decoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the decoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the decoder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the decoder component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate at least two video formats and / or encodings from the encoded video bitstream 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to trim the encoded video bitstream 108 in order to tailor potentially different bitstreams for at least one of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0032] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be rendered on a display or other type of rendering device. In some embodiments, at least one of the electronic devices 120 does not include a display component (e.g., is communicatively coupled to an external display device and / or includes a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0033] The source device and / or at least two electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or at least one of the electronic devices 120 is an instance of a server system, a personal computer, a portable device (e.g., a smartphone, a tablet, or a laptop), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0034] In an example operation of the communication system 100, the source device 102 transmits the encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may use the decoder component 114 to decode and / or encode the encoded video bitstream 108. For example, the server system 112 may apply an encoding that is more suitable for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., at least one encoded video bitstream) to at least one of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.

[0035] Figure 2AA block diagram showing example elements of an encoder component 106 in accordance with some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide a source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCB or RGB), and any suitable sampling structure (e.g., YCrCB 4:2:0 or Y CrCB 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously acquired / prepared video. In some embodiments, the video source 104 is a camera that acquires local image information as a video sequence. The video data may be provided as at least two separate pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial array of pixels, where each pixel may include at least one sample depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can readily understand the relationship between pixels and samples. The following focuses on describing samples.

[0036] The encoder component 106 is configured to encode and / or compress pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. Implementing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls and is functionally coupled to other functional units as described below. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204, as these functions may involve the encoder component 106 optimized for a particular system design.

[0037] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified embodiment, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact corresponding between the local encoder and the remote encoder. Thus, the reference picture samples "interpreted" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "interpret" when using prediction during decoding. This principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is known to those of ordinary skill in the art.

[0038] The operation of the decoder 210 can be the same as that of the "remote" decoder of the decoder component 122 described in detail below, for example. However, briefly referring to Figure 2B When the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210. Figure 2B

[0039] Except for parsing / entropy decoding, the decoder techniques described herein can exist in the corresponding encoder in a substantially identical functional form. For this reason, this application focuses on decoder operations. The description of encoder techniques can be simplified because encoder techniques can be inverse to decoder techniques.

[0040] As part of the operation, the source encoder 202 may perform motion-compensated predictive coding. Referring to at least one previously encoded frame designated as a "reference frame" in the video sequence, the motion-compensated predictive coding performs predictive coding on the input frame. In this way, the coding engine 212 encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, and the reference frame can be selected as the prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.

[0041] The decoder 210 decodes the encoded video data of the frame that can be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the coding engine 212 can be a lossy process. When the encoded video data is in the video decoder (Figure 2A When decoded at a location not shown (in the figure), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process, which can be performed by a remote video decoder on a reference frame, and can cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frame, which has the same content (without transmission errors) as the reconstructed reference frame to be obtained by the remote video decoder.

[0042] The predictor 206 can perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search the reference picture memory 208 for sample data (as a candidate reference pixel block) or some metadata, such as a reference picture motion vector, block shape, etc., that can serve as an appropriate prediction reference for the new picture. The predictor 206 can operate on a per-pixel block basis to find a suitable prediction reference. As determined by the search results obtained by the predictor 206, the input picture can have prediction references taken from at least two reference pictures stored in the reference picture memory 208.

[0043] The outputs of all the above functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by various functional units according to techniques known to those of ordinary skill in the art (such as Huffman coding, variable length coding, and / or arithmetic coding), thereby converting the symbols into an encoded video sequence.

[0044] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer the encoded video sequence created by the entropy encoder 214 to prepare for transmission via the communication channel 218, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter can be configured to merge the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown). In some embodiments, the transmitter can transmit additional data when transmitting the encoded video. The source encoder 202 can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data, such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0045] The controller 204 can manage the operation of the encoder component 106. During encoding, the controller 204 can assign a certain type of encoded picture to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, a picture can be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art understand the variants of I pictures and their corresponding applications and features, so they will not be repeated here. Predictive pictures can be encoded and decoded using intra prediction or inter prediction, which use at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures can be encoded and decoded using intra prediction or inter prediction, which use at most two motion vectors and reference indices to predict the sample values of each block. Similarly, at least two predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0046] Source pictures can generally be spatially subdivided into at least two sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (encoded) blocks, which are determined according to the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-prediction-encoded, or the blocks can be prediction-encoded with reference to already-encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be non-prediction-encoded by spatial prediction or by temporal prediction with reference to one previously encoded reference picture. Blocks of a B picture can be non-prediction-encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.

[0047] Video can be acquired as at least two source pictures (video pictures) in a time sequence. Intra picture prediction (often simplified to intra prediction) exploits the spatial correlation in a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In an embodiment, the particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using at least two reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0048] The encoder component 106 may perform encoding operations according to a predetermined video encoding technique or standard (such as any of the techniques or standards described herein). In its operation, the encoder component 106 may perform various compression operations, including predictive encoding operations that utilize the temporal and spatial redundancies in the input video sequence. Accordingly, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.

[0049] Figure 2B A block diagram showing example elements of a decoder component 122 according to some embodiments is shown. Figure 2B The decoder component 122 in is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to a loop filter 256 and is configured to transmit data to the display 124 (e.g., via a wired or wireless connection).

[0050] In some embodiments, the decoder component 122 includes a receiver that is coupled to the channel 218 and is configured to receive data from the channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive at least one encoded video sequence to be decoded by the decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from the channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data with other data (such as encoded audio data and / or auxiliary data streams), which may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequence from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of at least one encoded video sequence. The decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0051] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, the decoder component 122 is implemented at least partially in software.

[0052] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to counter network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 inside decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided outside decoder component 122 (e.g., to counter network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, buffer memory 252 may not be needed, or buffer memory 252 can be small. For use on a best-effort packet network such as the Internet, buffer memory 252 may be required, which can be relatively large and advantageously can have an adaptive size and can be implemented at least in part in an operating system or similar element (not depicted) outside decoder component 122.

[0053] Parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. The symbols can include, for example, information for managing the operation of decoder component 122, and / or information for controlling a rendering device such as display 124. The control information for at least one rendering device can be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not depicted). Parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technique or standard and can follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 254 can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the set. The subgroups can include group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. Parser 254 can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0054] The reconstruction of symbols 270 can involve at least two different units, depending on the type of the encoded video picture or a portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information that is parsed by parser 254 from the encoded video sequence. For clarity, this subgroup control information flow between parser 254 and the following at least two units is not depicted.

[0055] The decoder component 122 can be conceptually subdivided into at least two functional units, which in some embodiments interact closely with each other and can be at least partially integrated with each other. However, for clarity, the conceptual subdivision of the functional units is retained herein.

[0056] The scaler / inverse transform unit 258 receives the quantized transform coefficients and control information (such as the transform to be used, block size, quantization factor, and / or quantization scaling matrix) as at least one symbol 270 from the parser 254. The scaler / inverse transform unit 258 can output a block including sample values, which can be input into the aggregator 268.

[0057] In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is, blocks that do not use prediction information from a previous reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by the intra prediction unit 262. The intra prediction unit 262 can use the surrounding already reconstructed information obtained from the current (partially reconstructed) picture in the current picture memory 264 to generate a block having the same size and shape as the block being reconstructed. The aggregator 268 can add the prediction information already generated by the intra prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 on a per-sample basis.

[0058] In other cases, the output samples of the scaler / inverse transform unit 258 belong to inter-coded and possibly motion-compensated blocks. In this case, the motion compensation prediction unit 260 can access the reference picture memory 266 to obtain samples for prediction. After motion-compensating the obtained samples according to the symbol 270 belonging to the block, the aggregator 268 can add these samples to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The address from which the motion compensation prediction unit 260 obtains the prediction samples from within the reference picture memory 266 can be controlled by a motion vector. The motion vector can be provided to the motion compensation prediction unit 260 in the form of a symbol 270, which can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values obtained from the reference picture memory 266 when using sub-sampled accurate motion vectors, a motion vector prediction mechanism, etc.

[0059] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filtering techniques, which are controlled by parameters included in the encoded video bitstream and provided to loop filter unit 256 as symbols 270 from parser 254, but can also respond to meta-information obtained during the decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, and to sample values of previously reconstructed and loop-filtered samples. The output of loop filter unit 256 can be a sample stream, which can be output to a rendering device such as display 124 or stored in reference picture memory 266 for future inter-frame prediction use.

[0060] Once certain encoded pictures are reconstructed, they can be used as reference pictures for future prediction. Once an encoded picture is reconstructed and the encoded picture has been identified as a reference picture (e.g., by parser 254), the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct the next encoded picture.

[0061] Decoder component 122 can perform decoding operations according to a predetermined video compression technique, which can be documented in a standard, such as any of the standards described herein. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that it follows the syntax of the video compression technique or standard specified in the video compression technique document or standard, particularly the syntax specified in the profile document thereof. Moreover, in order to conform to certain video compression techniques or standards, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further restricted by HRD specifications and metadata for hypothetical reference decoder (HRD) buffer management signaled in the encoded video sequence.

[0062] Figure 3 A block diagram of server system 112 according to some embodiments is shown. Server system 112 includes control circuit 302, at least one network interface 304, memory 314, user interface 306, and at least one communication bus 312 for interconnecting these components. In some embodiments, control circuit 302 includes at least one processor (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes at least one field programmable gate array (FPGA), hardware accelerator, and / or integrated circuit (e.g., application specific integrated circuit).

[0063] At least one network interface 304 may be configured to interface with at least one communication network (e.g., wireless, wired, and / or optical networks). The communication network may be local, wide area, metropolitan area, vehicular and industrial, real-time, delay tolerant, etc. Examples of communication networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicular and industrial networks including CANBus, etc. Such communication may be unidirectional only receive (e.g., broadcast television), unidirectional only send (e.g., CAN bus to certain CAN bus devices), or bidirectional (e.g., to other computer systems using local digital networks or wide area digital networks). Such communication may include communication to at least one cloud computing network.

[0064] The user interface 306 includes at least one output device 308 and / or at least one input device 310. The input device 310 may include at least one of the following: keyboard, mouse, touchpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. At least one output device 308 may include at least one of the following: audio output device (e.g., speaker), visual output device (e.g., display or monitor), etc.

[0065] The memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as at least one disk storage device, optical disk storage device, flash device, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes at least one storage device remote from the control circuit 302. Alternatively, the memory 314 or at least one non-volatile solid-state storage device within the memory 314 includes a non-volatile computer-readable storage medium. In some embodiments, the memory 314 or the non-volatile computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: ● An operating system 316, which includes programs for handling various basic system services and for performing hardware-related tasks; ● A network communication module 318, which is used to connect the server system 112 to other computing devices via at least one network interface 304 (e.g., via wired and / or wireless connections); ● A codec module 320, which is used to perform various functions related to encoding and / or decoding data (e.g., video data). In some embodiments, the codec module 320 is an instance of the decoder component 114. The codec module 320 includes but is not limited to at least one of the following: o A decoding module 322, which is configured to perform various functions related to decoding encoded data, such as those functions described previously with respect to the decoder component 122; and o An encoding module 340, which is configured to perform various functions related to encoding data, such as those functions described previously with respect to the encoder component 106; and ● A picture memory 352 for storing pictures and picture data, for example for use by the codec module 320. In some embodiments, the picture memory 352 includes at least one of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0066] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions described previously with respect to the parser 254), a transform module 326 (e.g., configured to perform various functions described previously with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions described previously with respect to the motion compensation prediction unit 260 and / or the intra prediction unit 262), and a filter module 330 (e.g., configured to perform various functions described previously with respect to the loop filter 256).

[0067] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions described previously with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions described previously with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 a subset of the modules shown. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.

[0068] Each of the above modules stored in the memory 314 corresponds to a set of instructions for performing the functions described herein. The above modules (e.g., instruction sets) need not be implemented as separate software programs, routines, or modules, so various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but instead uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the above modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0069] Although Figure 3 a server system 112 according to some embodiments is shown, however Figure 3It is more intended as a functional description of various features that may exist in at least one server system rather than a structural schematic of the embodiments described herein. In practice, as recognized by those of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 Some of the items shown separately in Figure 3 may be implemented on a single server, and a single item may be implemented by at least one server. The actual number of servers used to implement the server system 112 and how the features are distributed among them will vary depending on the implementation and, optionally, partly on the data traffic processed by the server system during peak usage periods and average usage periods. Example Encoding Processes and Techniques

[0070] As described above, some codecs (e.g., AV1) operate on pixel blocks. Each pixel block can be processed in a predictive transform coding scheme, where intra-frame reference pixels, inter-frame motion compensation, or some combination of both are used to obtain a prediction. The residuals from the prediction can undergo a transform (e.g., 2-D unitary transform) to further remove spatial correlation, and the transform coefficients are quantized. Then, arithmetic coding can be used to entropy code the prediction syntax elements and the quantized transform coefficient indices.

[0071] Figure 4A An example calculation of a prediction block is shown. In Figure 4A the example of Figure 4A , intra-frame prediction is performed on the current block 402 to generate a prediction block 404. The current block 402 includes a set of samples (e.g., a pixel block) S 11 to S 44 , and the prediction block 404 includes a set of predictions P 11 to P 44 . Figure 4B A calculation of a residual block according to some embodiments is shown. As Figure 4B shown, the prediction block 404 is subtracted from the current block 402 to generate a residual block 406 that includes a set of residuals R 11 to R 44 . For example, the corresponding difference between each sample and the corresponding prediction is calculated. Figure 4C A calculation of a reconstructed block according to some embodiments is shown. As Figure 4C shown, the residual block 406 undergoes at least one transform and quantization to generate a set of residual coefficients. The set of residual coefficients can be transmitted from the encoder component to the decoder component. The set of residual coefficients undergoes inverse quantization and inverse transform to generate a reconstructed residual block 408. The reconstructed residual block 408 is combined with the prediction block 404 (e.g., the reconstructed residuals of the reconstructed residual block 408 are added to the predictions of the prediction block 404) to generate a reconstructed block 410 corresponding to the current block 402.

[0072] To reduce redundancy in the residual signal, various residual prediction techniques have been developed. These techniques predict the residual signal and encode the corrected residual. Residual differential pulse code modulation (RDPCM) requires sample-based differential pulse code modulation along the horizontal or vertical axis. By doing so, each residual row in the horizontal mode (or column in the vertical direction) can be reconstructed at the decoder by summing the scaled differential pulse code modulation residuals levels along the corresponding row (or column). RDPCM can be of an explicit type or an implicit type. The explicit type requires supplementary signaling of the direction, and its application is limited to inter-prediction blocks. On the other hand, the implicit type does not require direction signaling and can only be applied to intra-prediction blocks, where the prediction direction is associated with the intra-prediction mode. Block-based differential pulse code modulation (BDPCM) performs sample-based differential pulse code modulation on the reconstructed samples rather than the residual samples. The indication of using the second mode occurs during the prediction mode reconstruction process. This signaling involves two syntax elements, each for luminance and chrominance. For example, the initial syntax element flag indicates its utilization, while the second syntax element flag specifies the horizontal or vertical direction.

[0073] To reduce redundancy in the residual domain, a per-line residual prediction mode can be used to generate corrected residual blocks. The per-line residual domain prediction can be performed in the horizontal or vertical direction. For the horizontal prediction case, the prediction can be defined as shown in Equation 1 below, and for the vertical prediction case, the prediction can be defined as shown in Equation 2 below. Equation 1 - Horizontal per-line prediction Equation 2 - Vertical per-line prediction where x and y represent the row index and column index respectively, and r(*,*) represents the pixel value in the original residual block (e.g., generated by the intra-prediction process).

[0074] In some embodiments, a bi-prediction mode is used to generate corrected residual blocks. The bi-prediction mode can be performed in the horizontal or vertical direction. For the horizontal prediction case for encoding, the prediction can be defined as shown in Equation 3 below, and for the vertical prediction case for encoding, the prediction can be defined as shown in Equation 4 below. Equation 3 - Horizontal encoded bi-prediction Equation 4 - Vertical encoded bi-prediction where r(*,*) represents the pixel value in the original residual block (e.g., generated by the intra-prediction process).

[0075] For the horizontal prediction case for decoding, the prediction can be defined as shown in Equation 5 below, and for the vertical prediction case for decoding, the prediction can be defined as shown in Equation 6 below. Equation 5 - Horizontal Decoding Bi - direction Prediction Equation 6 - Vertical Decoding Bi - direction Prediction Where r′(*,*) represents the pixel value in the modified residual block. For example, the residual block value is determined by Determined.

[0076] In some embodiments, intra - prediction is performed on an encoded block or each sub - block in an encoded block, and a residual block is generated by subtracting the predicted block from the reconstructed samples of adjacent blocks. Then, short - range intra - prediction is applied to the residual block to obtain a modified residual block. For example, the modified residual signal can be calculated as shown in Equation 7 below. Equation 7 - Modified Residual Signal Where Represents the modified residual signal.

[0077] The reconstruction process can be performed by summing the samples along the determined direction, as shown in Equations 8 and 9 below. Equation 8 - Horizontal Reconstruction Equation 9 - Vertical Reconstruction

[0078] Therefore, a decoder according to an embodiment of the present disclosure can receive video data including at least two blocks from a video bitstream, where the at least two blocks include a first block and at least two residual coefficients. The first block is encoded in an intra - prediction mode. In addition, the at least two residual coefficients are generated by applying short - range intra - prediction to the residual block of the first block. In addition, the residual block is generated by applying the intra - prediction mode to the first block. Then, the decoder can generate a modified residual block of the first block according to the at least two residual coefficients and reconstruct the first block using the modified residual block.

[0079] In some embodiments, a flag is signaled to indicate whether a line-by-line residual prediction mode is used. In some embodiments, the flag is signaled separately for the luminance and chrominance components. In some embodiments, if the flag indicates that the line-by-line residual prediction mode is used, another flag is signaled to indicate whether the direction of the line-by-line residual prediction mode is vertical or horizontal. In some embodiments, when the line-by-line residual prediction mode is used, the angular increment and / or the multi-reference line (MRL) index are inferred to be zero. In some embodiments, the transform block size is fixed at the minimum transform size (e.g., 4×4), and the line-by-line residual prediction mode is implemented on 4×4 residual blocks.

[0080] In some embodiments, the line-by-line residual prediction mode enables the forward skip coding (FSC) mode (e.g., in a lossless coding scheme). For the coefficients obtained after a 2-D identity transform (IDTX), FSC may be a simpler and more efficient way of residual coding. The FSC-coded blocks have less TU-level signaling because for IDTX blocks, the transform type signaling and the block end index signaling are avoided, where the former reduces the symbol count in TX_SET_INTRA by 1 respectively. Finally, FSC is designed as a cheaper coding mode alternative for intra blocks because it disables the signaling of the multi-reference line (MRL) index, the filter intra mode, and the angular increment syntax when the transform type is IDTX, which can simplify the reconstruction process of FSC blocks.

[0081] Figure 4D The calculation of a residual prediction block and a modified residual block according to some embodiments is shown. As Figure 4D shown, short-range intra prediction is applied to the residual block 406 to generate a residual prediction block 420 including residuals Z 11 through Z 44 . As discussed in more detail below, the short-range intra prediction may include line-by-line prediction and / or bidirectional prediction. Figure 4D The generation of a modified residual block 422 is also shown, which includes residual differences D 11 through D 44 obtained by subtracting the residual prediction block 420 from the residual block 406. According to some embodiments, the residual differences are used to generate residual coefficients (e.g., through at least one transform and quantization). For example, the residual differences are used as an alternative to the residuals of the residual block 406.

[0082] Figure 4E The calculation of a residual block and a reconstructed block during the decoding process according to some embodiments is shown. For example, the modified residual coefficients are received, and the modified residual block is recovered from the coefficients (e.g., using an inverse transform and an inverse quantization process). After the modified residual block is recovered, short-range intra prediction may be applied to recover the residual block. After the residual block is recovered, intra prediction may be applied to recover the reconstructed block.

[0083] Figure 4F Shows an example line-by-line prediction for an encoding process. In Figure 4F the example, the residual block 406 includes the sample r ij , and the modified prediction block 420 includes the sample p′ ij . The residual block 406 can be obtained by applying intra prediction to the current block. Figure 4F The residual prediction block 420 in Figure 4F is obtained by line-by-line vertical prediction. In ij the example, the sample p′ ij of the residual prediction block 420 is subtracted from the sample r ij of the residual block 406 to obtain a modified residual block 422 with the sample r′

[0084] Figure 4G Shows an example line-by-line prediction for a decoding process. In Figure 4G the example, the modified residual block 422 is obtained (e.g., by applying an inverse transform to the residual coefficients received via the bitstream), and the modified residual block 422 includes the sample r′ ij . Figure 4G The residual prediction block 420 with the sample p′ ij in Figure 4G is obtained by line-by-line vertical prediction. In ij the example, the sample p′ ij of the residual prediction block 420 is added to the sample r′ ij of the modified residual block 422 to recover the residual block 406 with the sample r

[0085] Figure 5A and 5B show example residual blocks according to some embodiments. Figure 5A Shows the residual block 502, which includes a set of residuals R 11 to R 88 corresponding to rows 1 to 8 and columns 1 to 8. Figure 5B Shows the residual block 502 with index row 1 (e.g., corresponding to row 1), index row 2 (e.g., corresponding to row 4), index row 3 (e.g., corresponding to row 2), and index row 4 (e.g., corresponding to row 3). In Figure 5B the residual block 502 also includes index row A (e.g., corresponding to column 1), index row B (e.g., corresponding to column 4), index row C (e.g., corresponding to column 2), and index row D (e.g., corresponding to column 3).

[0086] Other residual prediction techniques are described below. The disclosed techniques can be used alone or in any order of combination. These techniques, as well as the encoder and decoder methods, can be performed using a processing circuit, which can include at least one processor or integrated circuit. As an example, a program stored in a non-volatile computer-readable storage medium can be executed by at least one processor.

[0087] Figure 6A A flowchart of a method 600 for video encoding according to some embodiments is shown. Method 600 can be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120), which includes a control circuit and a memory storing instructions executed by the control circuit. In some embodiments, method 600 is performed by executing instructions stored in the memory of the computing system (e.g., memory 314).

[0088] The system receives (602) video data including at least two blocks, the at least two blocks including a first block (e.g., current block 402) to be encoded in an intra prediction mode. For example, the system receives video data from a video source (e.g., video source 104). The system generates (604) a residual block (e.g., residual block 406) of the first block by applying an intra prediction mode to the first block. In some embodiments, the intra prediction mode is selected from the group including: at least one directional intra prediction, at least one smooth intra prediction, at least one cross-color intra prediction (e.g., predicting chrominance from luminance), at least one recursive intra prediction, intra block copy prediction, and palette prediction. The system generates (606) a corrected residual block (e.g., corrected residual block 420) of the first block by applying short-range intra prediction to the residual block. In some embodiments, the short-range intra prediction is a line-by-line prediction. In some embodiments, the short-range intra prediction is a bi-directional prediction. The system signals (608) the corrected residual block via a video bitstream. In some embodiments, the system signals the difference between the corrected residual block and the residual block (e.g., as a residual coefficient).

[0089] Figure 6B A flowchart of a method 650 for decoding video according to some embodiments is shown. Method 650 can be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120), which includes a control circuit and a memory storing instructions executed by the control circuit. In some embodiments, method 650 is performed by executing instructions stored in the memory of the computing system (e.g., memory 314).

[0090] The system receives (652) video data from a video bitstream, the video data including a first block (e.g., current block 402) and at least two residual coefficients of the first block (e.g., the residual coefficients shown in Figure 4D ). The system generates (654) a modified residual block (e.g., reconstructed residual block 408) of the first block based on the at least two residual coefficients. For example, the system applies inverse quantization and at least one inverse transform to generate the modified residual block (e.g., reconstructed residual block 408). The system generates (656) a residual block of the first block, where the residual block is generated using an intra-predicted block and the modified residual block. The system decodes (658) the first block using the residual block. For example, the system generates a reconstructed block (e.g., reconstructed block 410) of the first block.

[0091] Although Figure 6A and 6B show at least two logical stages in a particular order, stages that are not order-dependent can be reordered, and other stages can be combined or decomposed. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, so the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.

[0092] In some embodiments, intra-prediction is performed on an encoded block or each sub-block in an encoded block, and a residual block is generated by subtracting a predicted block from reconstructed samples of adjacent blocks. In some embodiments, short-range intra-prediction is applied to the residual block, thereby generating a modified residual block. In some embodiments, intra-prediction is performed on M×N blocks regardless of the size of the encoded block. Example values of M and N include, but are not limited to, 1, 2, 4, 8, 16, 32, and 64. In some embodiments, for lossless coding mode, short-range intra-prediction is applied to the residual block.

[0093] In some embodiments, row-by-row prediction is used as short-range intra-prediction. For example, the residuals in a particular row or column are predicted using its adjacent previous row, and the difference between the residual and the predicted residual is used as input for subsequent transform, quantization, or entropy coding processes. In some embodiments, row-by-row prediction is applied to the samples / residuals in all rows or columns except the samples / residuals in the first row / column. In some embodiments, for the first row / column, prediction is performed using residual samples generated from at least two reconstructed rows / columns of adjacent blocks. In some embodiments, row-by-row prediction of the residuals is performed in the horizontal direction. For example, the predicted residual of the first row is set to zero, and the predicted residuals of subsequent rows are predicted based on their adjacent previous rows. In some embodiments, row-by-row prediction of the residuals is performed in the vertical direction. For example, the predicted residual of the first column is set to zero, and the predicted residuals of subsequent columns are predicted by their adjacent previous columns.

[0094] In some embodiments, short - range intra - prediction is bi - directional prediction. In some embodiments, bi - directional prediction is applied to each M×N residual block. For example, the prediction residuals of the residuals in the first index row and the second index row of the M×N residual block are set to zero, and the weighted average of the residuals in the first index row and the second index row is used to predict the residuals in the third index row and the fourth index row. In some embodiments, the difference between the residual and the prediction residual is used as the input for subsequent transformation, quantization, or entropy - coding processes. Example values of M and N include, but are not limited to, 1, 2, 4, 8, 16, 32, and 64. In some embodiments, the first row and the second row are not adjacent rows. Examples of the first index row and the second index row include, but are not limited to, the first row and the fourth row respectively along a given direction. Examples of the third index row and the fourth index row include, but are not limited to, the second row and the third row respectively along a given direction. In some embodiments, the weighting factors used for weighted - averaging the residuals in the first index row and the second index row depend on the distance between the residual and its predicted value. For example, when predicting the residuals in the second row / column, the weighting factors of the residuals in the first index row and the second index row are {2 / 3, 1 / 3} or {3 / 4, 1 / 4}. As another example, when predicting the residuals in the third row / column, the weighting factors of the residuals in the first index row and the second index row are {1 / 3, 2 / 3} or {1 / 4, 3 / 4}.

[0095] In some embodiments, bi - directional prediction is performed in the horizontal direction. For example, the prediction residuals in the first index row and the second index row are set to zero. In this example, the residuals in the third index row and the fourth index row are predicted by weighted - averaging the residuals in the first row and the fourth row. In some embodiments, bi - directional prediction is performed in the vertical direction. For example, the prediction residuals in the first index row and the second index row are set to zero. In this example, the residuals in the third index row and the fourth index row are predicted by weighted - averaging the residuals in the first row and the fourth row.

[0096] In some embodiments, for short - range intra - prediction, the weighted average of the residuals from at least two adjacent rows is used to predict the residuals in the subsequent row. In some embodiments, the weighting factors used for weighted - averaging the residuals in at least two adjacent rows depend on the distance between the residual in the current row and the residuals in the adjacent rows. In some embodiments, two adjacent rows' residuals are used to predict the residual in the current row, the weighting factor of the residual in the nearest adjacent row is set to a first value, and the weighting factor of the residual in the other row is set to a second value. Examples of the first value and the second value include, but are not limited to, 2 / 3 and 1 / 3 respectively.

[0097] In some embodiments, at least two short - distance prediction methods are sequentially applied to a residual block. For example, first, a line - by - line prediction method is applied to the residual block to generate a modified residual block, and then a bidirectional prediction method is applied to the modified residual block to generate a final residual block.

[0098] In some embodiments, a flag is signaled in the bitstream to indicate which short - distance intra - prediction method is applied to the residual block. In some embodiments, two separate flags are used to signal the application of short - distance prediction to the luminance and chrominance residual block planes. In some embodiments, the direction of short - distance prediction of the residual block is signaled in the bitstream. In some embodiments, the direction of short - distance residual block prediction is inferred from the intra - prediction mode. In some embodiments, the context of the flag for entropy - coding the short - distance prediction of the chrominance residual block depends on the corresponding luminance flag. In some embodiments, whether at least one of the short - distance intra - predictions is applied is signaled in a high - level syntax, which includes but is not limited to sequence flags, GOP flags, picture flags, sub - picture flags, slice flags, or tile - level flags.

[0099] In some embodiments, a first direction is adopted in the intra - prediction of an encoded block or each sub - block in the encoded block, and a residual block is generated by subtracting the predicted block from the reconstructed samples of adjacent blocks. Then, a second direction is adopted in the short - distance residual prediction to predict the residual, and the difference between the residual and the predicted residual is used as the input for subsequent processing, which includes but is not limited to transformation, quantization, entropy coding, and loop filtering. At the decoder, the difference between the residual and the predicted residual is parsed and then added to the predicted residual to derive the reconstructed residual samples.

[0100] In some embodiments, the directions of the first direction and the second direction are different. In some embodiments, the direction of the second direction is the same as the direction of the first direction. In some embodiments, the value of the first direction is used as the context for entropy - coding the second direction. In some embodiments, a high - level syntax (including but not limited to sequence level, frame level, slice level, super - block level) is signaled in the bitstream to indicate whether the second direction is the same as the first direction.

[0101] In some embodiments, line - by - line prediction or bidirectional prediction is used as short - distance prediction. For example, line - by - line prediction is used as a short - distance intra - prediction in a residual block, and the residuals in a specific line are predicted using its adjacent previous line. As another example, bidirectional prediction is used as a short - distance intra - prediction in a residual block, and the predicted residuals of the residuals in the first index row and the second index row of the residual block are set to zero, while the weighted average of the residuals in the first index row and the second index row is used to predict the residuals in the third index row and the fourth index row.

[0102] In some embodiments, the angle for the second direction for short - range residual prediction is implicitly determined based on the angle of the first direction for intra - prediction. In some embodiments, if the prediction angle of intra - prediction is closer to the horizontal direction than to the vertical direction, short - range prediction is performed on the residual in the horizontal direction. In some embodiments, if the prediction angle of intra - prediction is closer to the vertical direction than to the horizontal direction, short - range prediction is performed on the residual in the vertical direction. In some embodiments, if the prediction angle of intra - prediction is closer to the diagonal direction than to the horizontal or vertical direction, short - range prediction is performed on the residual in the diagonal direction. In some embodiments, short - range prediction is applied to N nominal angles for intra - prediction. For example, N is equal to 4, 6, 8, or 10.

[0103] In some embodiments, at least one syntax element is signaled in the bitstream to indicate the direction / angle of short - range prediction. In some embodiments, a syntax element indicating the direction of short - range prediction is signaled in the bitstream. In some embodiments, a first syntax element is used to indicate the nominal / master direction, and a second syntax element is used to indicate the angle increment relative to the nominal direction. In some embodiments, the supported values of the increment angle are predefined in a lookup table, and the index of the increment angle in the lookup table is signaled in the bitstream. As an example, a first syntax is signaled to indicate whether the direction for short - range residual prediction is vertical or horizontal, and then a second syntax is signaled to indicate the angle increment to a specified master direction.

[0104] In some embodiments, a first syntax element is used to indicate the nominal direction; a second syntax element is used to indicate whether the angle increment is zero. If the angle increment is non - zero, third and fourth syntax elements are further used. The third syntax element is used to indicate the positive or negative value of the angle increment. The fourth syntax element is used to indicate the absolute angle increment.

[0105] In some embodiments, only the angle increment syntax is signaled to derive the second direction (e.g., the prediction direction used for residual prediction), and then the direction used for residual prediction is derived by adding the angle increment value to the nominal prediction direction (or prediction direction) of the intra - prediction mode.

[0106] Some embodiments include transform coding techniques for short - range residual prediction. In some embodiments, to process a residual block, a selection of a transform coding mode is signaled at a first processing unit level, and a transform process is performed at a second processing unit level, where the transform coding mode refers to any parameter or operation involved in the transform process, and the transform method can be applied in the forward transform process at the encoder or the inverse transform process at the encoder and / or decoder. For example, the scope of the first processing unit level and the second processing unit level can include sequence level, frame level, super - block level, coding block level, prediction block level, or transform block level.

[0107] In some embodiments, a residual block or a modified residual block generated from intra - prediction is used as an input to the transform process. In some embodiments, intra - prediction is performed on a coding block or each sub - block in the coding block, and a residual block is generated by subtracting the predicted block from the reconstructed samples of adjacent blocks. In some embodiments, a short - range intra - prediction method is applied to the residual block to obtain a modified residual block. In some embodiments, the modified residual block is used as an input for subsequent transform coding. In some embodiments, different transform kernels can be applied to the modified residual block. For example, no transform is applied to the modified residual block (or the transform kernel is an identity transform). As an example, no transform (or an identity transform is applied) in one direction (e.g., horizontal or vertical), and a lossless transform (e.g., Hadamard transform) is applied in the other direction.

[0108] In some embodiments, the first processing unit level and the second processing unit level are the same. In some embodiments, the type of the transform coding mode is signaled at the coding block level, the size of the modified residual block is the same as the coding block size, and the transform coding block size is also the same as the coding block size. For example, the syntax element for the transform block size can be inferred from the modified residual block size or the coding block size, and there is no need to signal the syntax element for the transform block size.

[0109] In some embodiments, the first processing unit level and the second processing unit level are different. In some embodiments, the type of the transform coding kernel is signaled at the coding block level, but the transform block size used to perform the transform process is smaller than the coding block size. In some embodiments, the transform block size is fixed, and the same transform coding kernel is applied to the transform coding blocks within one coding block. For example, there is no need to signal the syntax element for the transform block size in the bitstream. In one example, the identity transform is signaled at the coding block level, and the transform size is fixed at M×N regardless of the coding block size. In another example, the Hadamard transform (or a different lossless transform) is signaled at the coding block level, and the transform size is fixed at M×N regardless of the coding block size. In some embodiments, M and N are selected to correspond to the minimum allowable transform size (e.g., both M and N are equal to 4).

[0110] In some embodiments, the transform block size and the transform coding mode are determined by a given cost metric at the encoder. For example, the best transform coding type and the best transform size are signaled at the coding block level. In one example, the cost metric is the rate-distortion cost used in rate-distortion optimization.

[0111] In some embodiments, it is signaled at a high-level syntax (including but not limited to sequence, GOP, frame, or slice level) whether the first processing unit level and the second processing unit level are different.

[0112] In some embodiments, the transform scheme includes applying a transform skip (or identity transform) in one direction (e.g., horizontal or vertical direction) and applying the Hadamard transform in the other direction. In some embodiments, the transform scheme is only applied to the lossless coding mode. In some embodiments, the transform scheme is only applied to a specific M×N transform block size (e.g., 4×4 transform block size). In some embodiments, it is signaled separately for each direction whether to select the transform skip (or identity transform) or the Hadamard transform. Now turn to some example embodiments.

[0113] (A1) In one aspect, some embodiments include a method of video encoding (e.g., method 600). In some embodiments, the method is performed at a computing system (e.g., server system 112) including a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). In some embodiments, the method is performed at a source coding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving video data including at least two blocks, the at least two blocks including a first block (e.g., current block 402), wherein the first block is to be encoded in an intra prediction mode; (ii) generating a residual block of the first block by applying the intra prediction mode to the first block; (iii) generating a modified residual block of the first block by applying short-range intra prediction to the residual block; and (iv) signaling the modified residual block via a video bitstream. In some embodiments, generating the modified residual block of the first block includes applying at least two short-range intra predictions. For example, at least two short-range prediction methods may be sequentially applied to the residual block. As an example, a horizontal prediction is applied to the residual block to generate a modified residual block, and then a bidirectional prediction is applied to the modified residual block to generate a final residual block. In some embodiments, the short-range intra prediction mode is applied as part of a lossless coding scheme.

[0114] (A2) In some embodiments of A1, the method further includes: transmitting the encoded first block via the video bitstream.

[0115] (A3) In some embodiments of A1 or A2, the method further includes: determining the difference between the residual block and the modified residual block. For example, the residuals in a particular row or column are predicted using its adjacent previous row. The difference between the residual and the predicted residual is used as an input for subsequent transform, quantization, or entropy coding processes.

[0116] (A4) In some embodiments of any one of A1 to A3, the intra prediction mode is applied to an M×N portion of the residual block, where M and N are positive integers.

[0117] (A5) In some embodiments of any one of A1 to A4, the short-range intra prediction includes horizontal prediction, in which the residuals in a particular row or column are predicted using an adjacent previous row or column.

[0118] (A6) In some embodiments of any one of A1 to A5, the short-range intra prediction includes bidirectional prediction, in which the residuals in a third index row and a fourth index row of the residual block are predicted using a weighted average of the residuals in a first index row and a second index row of the residual block.

[0119] (A7) In some embodiments of any one of A1 to A6, the method further includes: signaling a syntax element in the video bitstream, where the syntax element indicates which type of short-range intra prediction (e.g., progressive prediction or bi-directional prediction) is to be applied.

[0120] (A8) In some embodiments of any one of A1 to A7, the method further includes: signaling a first flag to indicate the use of a short-range intra prediction mode for the chrominance plane, and a second flag to indicate the use of a short-range intra prediction mode for the luminance plane.

[0121] (B1) In another aspect, some embodiments include a method of video decoding (e.g., method 650). In some embodiments, the method is performed at a computing system (e.g., server system 112) including a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). In some embodiments, the method is performed at a parser (e.g., parser 254), a motion prediction component (e.g., motion compensation prediction unit 260), and / or an intra prediction component (e.g., intra prediction unit 262). The method includes: (i) receiving video data (e.g., an encoded video sequence) including at least two blocks from a video bitstream, the at least two blocks including a first block and at least two residual coefficients of the first block; (ii) generating a modified residual block of the first block based on the at least two residual coefficients; (iii) generating a residual block of the first block, where the residual block is generated using an intra prediction block and the modified residual block; and (iv) reconstructing the first block using the residual block. For example, the modified residual block is generated by applying inverse quantization and / or inverse transform to the at least two residual coefficients. As an example, the residual block represents the difference between the reconstructed samples and the corresponding predicted values.

[0122] (B2) In some embodiments of B1, an intra prediction mode is applied to an M×N portion of the first block, where M and N are positive integers. In some embodiments, the intra prediction mode is applied at each M×N unit of the first block. For example, intra prediction is performed on the M×N block regardless of the size of the residual block. Example values of M and N include 1, 2, 4, 8, 16, 32, and 64.

[0123] (B3) In some embodiments of B1 or B2, short - range intra - prediction includes line - by - line prediction. In line - by - line prediction, the residuals in a particular row or column are predicted using the adjacent previous row or column. For example, line - by - line prediction is used as a short - range intra - prediction. The residuals in a particular row or column are predicted using its adjacent previous row. In some embodiments, the residuals in a subsequent row are predicted using a weighted average of the residuals from at least two adjacent rows and / or at least two previous rows. For example, the weighting factor used to weight - average the residuals in at least two adjacent rows may depend on the distance between the residuals in the current row and the residuals in the adjacent rows. In some embodiments, the residuals in two adjacent rows are used to predict the residuals in the current row, with the weighting factor for the residuals in the nearest adjacent row set to a first value and the weighting factor for the residuals in the other row set to a second value. Examples of the first value and the second value include, but are not limited to, 2 / 3 and 1 / 3, respectively.

[0124] (B4) In some embodiments of B3, line - by - line prediction is applied to each row or column of the residual block except the first row or the first column. For example, the first row is the left - most row. As another example, the first column is the top - most column. For example, this prediction technique is applied to the samples / residuals in all rows or columns except the samples / residuals in the first row / column. In some embodiments, for the first row / column, prediction is performed using the residual samples generated from at least two reconstructed rows / columns of the adjacent block. In some embodiments, line - by - line prediction is applied to each row or column of the M×N portion of the residual block except the first row or the first column.

[0125] (B5) In some embodiments of B3 or B4, line - by - line prediction is performed in the horizontal direction, where the residuals of the first column of the residual block are set to zero, and the residuals of the other columns of the residual block are predicted based on the previous columns. For example, the predicted residuals of the first column are set to zero, while the predicted residuals of the subsequent columns are predicted based on their adjacent previous rows. In some embodiments, the residuals of the first column of the M×N portion of the residual block are set to zero, while the residuals of the other columns of the M×N portion are predicted based on the previous columns.

[0126] (B6) In some embodiments of B3 or B4, line - by - line prediction is performed in the vertical direction, where the residuals of the first row of the residual block are set to zero, and the residuals of the other rows of the residual block are predicted based on the previous rows. For example, the predicted residuals of the first column are set to zero, while the predicted residuals of the subsequent columns are predicted based on their adjacent previous rows. In some embodiments, the residuals of the first row of the M×N portion of the residual block are set to zero, while the residuals of the other rows of the M×N portion are predicted based on the previous rows.

[0127] (B7)In some embodiments of any one of B1 to B6, short - distance intra - frame prediction includes bi - directional prediction, in which the weighted average of the residuals in the first index row and the second index row of the residual block is used to predict the residuals in the third index row and the fourth index row of the residual block. For example, bi - directional prediction is applied to each M×N residual block. In some embodiments, the weighted average of the residuals from at least two adjacent rows and / or at least two previous rows is used to predict the residuals in a specific index row. For example, the weighting factor for weighted - averaging the residuals in at least two adjacent rows may depend on the distance between the residual in the current row and the residuals in the adjacent rows.

[0128] (B8)In some embodiments of B7, bi - directional prediction includes setting the predicted residuals of the residuals in the first index row and the second index row to zero. In some embodiments, the predicted residuals of the residuals in the first index row and the second index row of the M×N residual block are set to zero.

[0129] (B9)In some embodiments of B7 or B8, the first index row and the second index row are not adjacent rows. Examples of the first index row and the second index row include, but are not limited to, the first row (line) and the fourth row (line) along a given direction, respectively. Examples of the third index row and the fourth index row include, but are not limited to, the second row and the third row along a given direction, respectively.

[0130] (B10)In some embodiments of any one of B7 to B9, the weights of the weighted average of the residuals are determined based on the distance between each residual and the corresponding predicted value. For example, the weighting factor for weighted - averaging the residuals in the first index row and the second index row may depend on the distance between the residual and its predicted value. As an example, when predicting the residuals in the second row / column, the weighting factors of the residuals in the first index row and the second index row are {2 / 3, 1 / 3} or {3 / 4, 1 / 4}. As another example, when predicting the residuals in the third row / column, the weighting factors of the residuals in the first index row and the second index row are {1 / 3, 2 / 3} or {1 / 4, 3 / 4}.

[0131] (B11)In some embodiments of any one of B7 to B10, bi - directional prediction is performed in the horizontal direction, where the predicted residuals in the third index row and / or the fourth index row are predicted by the weighted average of the residuals in the first column and the fourth column of the residual block. For example, the predicted residuals in the first index row and the second index row are set to zero, and the residuals in the third index row and the fourth index row are predicted by weighted - averaging the residuals in the first column and the fourth column.

[0132] In some embodiments of any one of B7 to B10, bidirectional prediction is performed in the vertical direction, wherein the predicted residual in the third index row and / or the fourth index row is predicted by a weighted average of the residuals in the first row and the fourth row of the residual block. For example, the predicted residuals in the first index row and the second index row are set to zero, and the residuals in the third index row and the fourth index row are predicted by weighted averaging the residuals in the first row and the fourth row.

[0133] In some embodiments of any one of B1 to B12: (i) the video bitstream further includes a syntax element; and (ii) a first type of short-range intra prediction is applied to the residual block according to the syntax element having a first value, and (iii) a second type of short-range intra prediction is applied to the residual block according to the syntax element having a second value. For example, a flag is signaled in the bitstream to indicate which short-range intra prediction method is applied to the residual block. In some embodiments, the direction of short-range prediction of the residual block is signaled in the bitstream. In some embodiments, the direction of short-range residual block prediction is inferred by the intra prediction mode. In some embodiments, the syntax element is a High-Level Syntax (HLS) element. In some embodiments, the HLS element is signaled at a level higher than the block level. For example, HLS may correspond to the sequence level, frame level, slice level, or tile level. As another example, the HLS element may be signaled in a Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), slice header, picture header, tile header, and / or Coding Tree Unit (CTU) header.

[0134] In some embodiments of any one of B1 to B13, the video bitstream includes: a first flag to signal the use of a short-range intra prediction mode for the chrominance plane, and a second flag to signal the use of a short-range intra prediction mode for the luminance plane. For example, two separate flags are used to signal the use of short-range prediction for the luminance and chrominance residual block planes. In some embodiments, the context of the flag for entropy coding the short-range prediction of the chrominance residual block depends on the corresponding luminance flag.

[0135] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more sets of instructions configured to be executed by the control circuitry, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A8, B1 to B14 above). In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions executed by the control circuitry of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A8, B1 to B14 above).

[0136] It should be understood that although the terms “first,” “second,” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of at least one of the associated listed items. It should be further understood that when the terms “comprises” and / or “comprising” are used in this specification, it indicates the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of at least one other feature, integer, step, operation, element, component, and / or group thereof.

[0137] As used herein, depending on the context, the term “if” may be interpreted to mean “when” or “upon” or “in response to determining” or “in accordance with determining” or “in response to detecting” that the stated precondition is true. Similarly, depending on the context, the phrases “if it is determined [that the stated precondition is true]” or “if [the stated precondition is true]” or “when [the stated precondition is true]” may be interpreted to mean “upon determining” or “in response to determining” or “in accordance with determining” or “upon detecting” or “in response to detecting” that the stated precondition is true.

[0138] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the operating principles and the practical application, thereby enabling others skilled in the art to understand.

Claims

1. A method for video decoding performed at a computing system including a memory and one or more processors, characterized in that, The method includes: Receiving video data including at least two blocks from a video bitstream, the at least two blocks including a first block and at least two residual coefficients of the first block; Generating a modified residual block of the first block according to the at least two residual coefficients; Generating a reconstructed residual block of the first block, wherein the reconstructed residual block is generated using an intra-prediction block and the modified residual block; and Reconstructing the first block using the reconstructed residual block.

2. The method according to claim 1, characterized in that, Applying an intra-prediction mode to an M×N portion of the first block, where M and N are positive integers.

3. The method according to claim 1, wherein The short-range intra-prediction includes progressive prediction, in which the residuals in a specific row or column are predicted using adjacent previous rows or previous columns.

4. The method according to claim 3, wherein Applying the progressive prediction to each row or column of the residual block except the first row or the first column.

5. The method according to claim 3, characterized in that Performing the progressive prediction in a horizontal direction, where the residuals of the first column of the residual block are set to zero, and the residuals of the other columns of the residual block are predicted according to the previous columns.

6. The method according to claim 3, wherein Performing the progressive prediction in a vertical direction, where the residuals of the first row of the residual block are set to zero, and the residuals of the other rows of the residual block are predicted according to the previous rows.

7. The method according to claim 1, wherein The short-range intra-prediction includes bidirectional prediction, in which the residuals in a third index row and a fourth index row of the residual block are predicted using a weighted average of the residuals in a first index row and a second index row of the residual block.

8. The method according to claim 7, wherein The bidirectional prediction includes setting the predicted residuals of the residuals in the first index row and the second index row to zero.

9. The method according to claim 7, wherein The first index row and the second index row are not adjacent rows.

10. The method according to claim 7, wherein The weights of the weighted average of the residuals are determined based on the distance between each residual and the corresponding predicted value.

11. The method according to claim 7, characterized in that, Performing the bidirectional prediction in a horizontal direction, and predicting the predicted residuals in the third index row and / or the fourth index row through a weighted average of the residuals in the first column and the fourth column of the residual block.

12. The method according to claim 7, characterized in that, Performing the bidirectional prediction in a vertical direction, and predicting the predicted residuals in the third index row and / or the fourth index row through a weighted average of the residuals in the first row and the fourth row of the residual block.

13. The method according to claim 1, wherein: The video bitstream further includes syntax elements; Applying a first type of short-range intra-prediction to the residual block according to a syntax element having a first value; And Applying a second type of short-range intra-prediction to the residual block according to a syntax element having a second value.

14. The method according to claim 1, wherein The video bitstream includes a first flag for signaling the use of a short-range intra-prediction mode for a chrominance plane, and a second flag for using the short-range intra-prediction mode for a luminance plane.

15. A computing system, characterized in that, Including: A control circuit, A memory, and At least one set of instructions stored in the memory and configured to be executed by the control circuit, the at least one set of instructions including instructions for the following operations: Receiving video data including at least two blocks from a video bitstream, the at least two blocks including a first block and at least two residual coefficients of the first block; Generate a corrected residual block of the first block according to the at least two residual coefficients; Generate a reconstructed residual block of the first block, wherein the reconstructed residual block is generated using an intra prediction block and the corrected residual block; and Reconstruct the first block using the reconstructed residual block.

16. The computing system according to claim 15, wherein Apply an intra prediction mode to an M×N portion of the first block, where M and N are positive integers.

17. The computing system according to claim 15, wherein The short-range intra prediction includes a line-by-line prediction, in which the residuals in a specific row or column are predicted using adjacent preceding rows or preceding columns.

18. The computing system according to claim 15, wherein The short-range intra prediction includes a bidirectional prediction, in which the residuals in a third index row and a fourth index row of the residual block are predicted using a weighted average of the residuals in a first index row and a second index row of the residual block.

19. A non-volatile computer-readable storage medium, characterized in that, Store at least one set of instructions configured to be executed by a computing device including a control circuit and a memory, the at least one set of instructions including instructions for: Receiving video data including at least two blocks from a video bitstream, the at least two blocks including a first block and at least two residual coefficients of the first block; Generating a corrected residual block of the first block according to the at least two residual coefficients; Generating a reconstructed residual block of the first block, wherein the reconstructed residual block is generated using an intra prediction block and the corrected residual block; and Reconstructing the first block using the reconstructed residual block.

20. The computing system according to claim 19, wherein The short-range intra prediction includes a line-by-line prediction, in which the residuals in a specific row or column are predicted using adjacent preceding rows or preceding columns.