Flexible transform scheme for residual block
By using short-distance intra prediction technology in video encoding to generate optimized residual blocks and perform the transformation process at the second processing unit level, the problem of high redundancy of residual signals is solved, and encoding efficiency and transmission performance are improved.
Patent Information
- Application Number
- CN202380080771.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-04
- Filing Date
- 2023-10-30
- Publication Date
- 2025-07-04
AI Technical Summary
The existing video encoding technology has problems such as high redundancy and low encoding efficiency in residual signal processing, especially in the intra prediction mode, which is difficult to effectively reduce transmission bandwidth requirements.
The short-distance intra prediction technology is adopted to reduce redundant signaling by generating the optimized residual block in the intra prediction mode of the video block and performing the transformation process at the second processing unit level without signaling this level.
By reducing redundancy in the residual domain, encoding efficiency is improved, transmission bandwidth requirements are reduced, and the compression performance of video encoding is improved.
Smart Images

Figure CN120266476A_ABST
Abstract
Description
Related Applications
[0001] This application claims priority to U.S. Patent Application No. 18 / 480,973, entitled "Flexible Transform Scheme for Residual Blocks," filed on October 4, 2023. Technical Field
[0002] The disclosed embodiments generally relate to image and video coding and compression, including but not limited to systems and methods for predicting residual information. Background Art
[0003] Various electronic devices support digital video, such as digital televisions, laptop or desktop computers, tablets, digital cameras, digital video recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. These electronic devices transmit and receive digital video data over a communication network or otherwise convey digital video data, and / or store digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited memory resources of the storage device, video coding can be used to compress video data according to at least one video coding standard before transmitting or storing the video data. Video coding can be performed by hardware and / or software on an electronic device / client device or a server providing cloud services.
[0004] Video coding typically uses prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), which utilize the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was published by ITU-T and ISO / IEC in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard designed to be a successor to HEVC. The VVC / H.266 standard was published by ITU-T and ISO / IEC in 2020 (version 1) and 2022 (version 2). Alliance for Open Media Video 1 (AOMedia Video 1, AV1) is an open video coding format designed to replace HEVC. On January 8, 2019, the verified version 1.0.0 of the specification with errata 1 was released. Summary of the Invention
[0005] A comprehensive video codec typically includes at least two components, such as intra / inter prediction, transform coding, quantization, residual coding, and loop filtering. To further reduce the residual signal, various residual prediction techniques have been developed. These techniques predict the residual signal and encode the high-order residual into the bitstream. This disclosure describes methods and systems for enhancing video (image) compression, including improved residual prediction techniques.
[0006] According to some embodiments, a method of video encoding is provided. The method includes: (i) receiving video data that includes at least two blocks, the at least two blocks including a first block (e.g., a chrominance or luminance block); (ii) selecting a transform coding mode for the first block, the transform coding mode using a first processing unit level; and (iii) performing a transform process on the first block using the selected transform coding mode, wherein the transform process is performed on the transformed block at a second processing unit level, where the second processing unit level is not signaled.
[0007] According to some embodiments, a method of video decoding is provided. The method includes: (i) receiving video data from a video bitstream, the video data including at least two blocks and a syntax element, the at least two blocks including a first block, where the syntax element is signaled at a first processing unit level; (ii) selecting a transform coding mode based on the syntax element; (iii) performing a transform process on the first block using the selected transform coding mode, wherein the transform process is performed on the transformed block at a second processing unit level, where the second processing unit level is not signaled.
[0008] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic device. The computing system includes a control circuit and a memory storing at least one set of instructions. The at least one set of instructions includes instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder). According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores at least one set of instructions for execution by a computing system. The at least one set of instructions includes instructions for performing any of the methods described herein.
[0009] Accordingly, devices and systems, as well as video encoding and decoding methods, are disclosed. Such methods, devices, and systems can supplement or replace related methods, devices, and systems for video encoding / decoding. The features and advantages described in the specification are not necessarily fully enumerated, and in particular, given the accompanying drawings, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Additionally, it should be noted that the language used in the specification is mainly chosen for readability and guidance purposes and is not necessarily chosen to depict or limit the subject matter described herein. Description of the Drawings
[0010] To understand the present disclosure in more detail, a more specific description can be obtained by referring to the features of various embodiments, some of which are shown in the accompanying drawings. However, the accompanying drawings only show the relevant features of the present disclosure and are therefore not necessarily considered restrictive, as those skilled in the art will understand that there may be other valid features in this specification after reading the present disclosure.
[0011] Figure 1 is a schematic block diagram of an example communication system of some embodiments.
[0012] Figure 2A is a schematic block diagram of an example element of an encoder component of some embodiments.
[0013] Figure 2B is a schematic block diagram of an example element of a decoder component of some embodiments.
[0014] Figure 3 is a schematic block diagram of an example server system of some embodiments.
[0015] Figure 4A is a schematic diagram of the calculation of a prediction block of some embodiments.
[0016] Figure 4B is a schematic diagram of the calculation of a residual block of some embodiments.
[0017] Figure 4C is a schematic diagram of the calculation of a reconstruction block of some embodiments.
[0018] Figure 4D is a schematic diagram of the calculation of an optimized residual block and difference block of some embodiments.
[0019] Figure 4E is a schematic diagram of the calculation of a residual block and a reconstruction block of some embodiments.
[0020] Figure 4F and Figure 4G is a schematic diagram of an example line-by-line prediction of some embodiments.
[0021] Figure 5A and Figure 5B are schematic diagrams of example residual blocks of some embodiments.
[0022] Figure 6A is a flowchart of an example method for encoding a video of some embodiments.
[0023] Figure 6B is a flowchart of an example method for decoding a video of some embodiments.
[0024] According to conventional practice, the various features illustrated in the drawings need not be drawn to scale, and throughout the specification and drawings, like reference numerals may be used to represent like features. Detailed Description
[0025] The present disclosure describes systems and methods for predicting residual information. The systems and methods described herein can improve the performance of lossless coding by reducing redundancy in the residual domain. In some embodiments, short-range intra prediction (e.g., line-by-line residual domain prediction mode) is implemented for vertical intra prediction mode and / or horizontal intra prediction mode. For the luminance plane and the chrominance plane, flags can be signaled separately to indicate the use of short-range intra prediction, while the U plane and the V plane can share one flag.
[0026] In some embodiments, a residual block is generated (e.g., by using an intra prediction mode for a current block in a first direction), and then an optimized residual block is generated (e.g., by using short-range intra prediction for the residual block). Generating and using the optimized residual block can reduce redundancy in the residual domain. Reducing redundancy reduces the number of bits required to signal the residual (e.g., improves coding efficiency and reduces transmission bandwidth).
[0027] In some embodiments, a transform coding mode is selected for a first block, the transform coding mode uses a first processing unit level, and then a transform process (e.g., a transform or an inverse transform) is performed on the first block using the selected transform coding mode, wherein the transform process is performed on the transformed block at a second processing unit level. In some embodiments, the second processing unit level is not signaled. The advantage of not signaling the transform block size is that it reduces the transmission bandwidth. Additionally, this signaling scheme allows short-range intra prediction to be used on a set transform block size (e.g., the minimum transform block size). Example Systems and Devices
[0028] Figure 1FIG. is a schematic block diagram of a communication system 100 of some embodiments. The communication system 100 includes a source device 102 and at least two electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m), which are communicatively coupled to each other via at least one network. In some embodiments, the communication system 100 is a streaming system, such as for use with video-enabled applications, such as video conferencing applications, digital television applications, and media storage and / or distribution applications.
[0029] The source device 102 includes a video source 104 (e.g., a camera component or a media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., for creating an uncompressed video sample stream). The encoder component 106 generates at least one encoded video bitstream using the video stream. Compared with the encoded video bitstream 108 generated by the encoder component 106, the video stream from the video source 104 may have a higher data volume. Since the encoded video bitstream 108 has a lower data volume (less data) compared with the video stream from the video source, the encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store. In some embodiments, the source device 102, such as for transmitting uncompressed video to at least one network 110, may not include the encoder component 106.
[0030] The at least one network 110 represents any number of networks that transfer information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wired (cable) communication networks and / or wireless communication networks. The at least one network 110 may exchange data in a circuit-switched channel and / or a packet-switched channel. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0031] The at least one network 110 includes a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is, or includes, a streaming server (e.g., for storing and / or distributing video content, such as an encoded video stream from the source device 102). The server system 112 includes a coder component 114 (e.g., for encoding and / or decoding video data). In some embodiments, the coder component 114 includes an encoder component and / or a decoder component. In various embodiments, the coder component 114 is exemplified as hardware, software, or a combination thereof. In some embodiments, the coder component 114 is used to decode the encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is used to generate at least two video formats and / or encodings based on the encoded video bitstream 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 can be used to trim the encoded video bitstream 108 to customize potentially different bitstreams for at least one electronic device 120. In some embodiments, the MANE is provided separately from the server system 112.
[0032] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is used to decode the encoded video data 116 to generate an outgoing video stream that can be presented on a display or other type of presentation device. In some embodiments, at least one electronic device 120 does not include a display component (e.g., is communicatively coupled to an external display device and / or includes a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is used to access the server system 112 to obtain the encoded video data 116.
[0033] The source device and / or the at least two electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or the at least one electronic device 120 are examples of server systems, personal computers, portable devices (e.g., smartphones, tablets, or laptops), wearable devices, video conferencing devices, and / or other types of electronic devices.
[0034] In an example operation of communication system 100, source device 102 transmits encoded video bitstream 108 to server system 112. For example, source device 102 may encode a stream of images captured by the source device. Server system 112 receives encoded video bitstream 108 and may use encoding component 114 to decode and / or encode the encoded video bitstream 108. For example, server system 112 may apply an encoding that is more suitable for network transmission and / or storage to the video data. Server system 112 may transmit encoded video data 116 (e.g., at least one encoded video bitstream) to at least one electronic device 120. Each electronic device 120 may decode the encoded video data 116 and optionally display video images.
[0035] Figure 2A FIG. is a schematic block diagram of example elements of encoder component 106 of some embodiments. Encoder component 106 receives a source video sequence from video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component for receiving the source video sequence. In some embodiments, encoder component 106 receives a video sequence from a remote video source (e.g., a video source of a component located in a different device than encoder component 106). Video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 YCrCb or RGB), and any suitable sampling structure (e.g., YCrCb4:2:0 or YCrCb 4:4:4). In some embodiments, video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as at least two separate images that, when viewed in sequence, produce a motion effect. These images themselves may be organized as a spatial array of pixels, where each pixel may include at least one sample determined by the sampling structure, color space, etc. in use. Those of ordinary skill in the art can readily understand the relationship between pixels and samples. The following description focuses on samples.
[0036] The encoder component 106 is configured to encode and / or compress the images of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls the other functional units as described below and is functionally coupled to these other functional units. The parameters set by the controller 204 may include rate control related parameters (e.g., picture skipping, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204, as these functions may involve optimizing the encoder component 106 for a particular system design.
[0037] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and at least one reference image) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference image memory 208. Since the decoding of the symbol stream results in a bit-level exact result independent of the decoder location (local or remote), the content in the reference image memory 208 is also bit-level exact between the local encoder and the remote encoder. Thus, when prediction is used during decoding, the prediction part of the encoder interprets the same sample values that the decoder would interpret as reference image samples. The principle of reference image synchronization (and if synchronization cannot be maintained, e.g., due to drift caused by channel errors) is known to those of ordinary skill in the art.
[0038] The operation of the decoder 210 may be the same as the operation of a remote decoder (such as the decoder component 122), which will be described in detail below in connection with Figure 2B Briefly referring to Figure 2B , however, since the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode these symbols into the encoded video sequence losslessly, the entropy decoding part of the decoder component 122 (including the buffer memory 252 and the parser 254) may not be fully implemented in the local decoder 210.
[0039] Except for parsing / entropy decoding, the decoder techniques described herein may exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder techniques may be simplified, as encoder techniques may be the inverse operation of decoder techniques.
[0040] As part of its operation, the source encoder 202 may perform motion-compensated predictive coding that predictively encodes an input frame with reference to at least one previously encoded frame designated as a reference frame from a video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of at least one reference frame, which at least one reference frame may be selected as at least one prediction reference for the input frame. The controller 204 may manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroups of parameters for encoding video data.
[0041] Based on the symbols created by the source encoder 202, the decoder 210 decodes the encoded video data of frames that may be designated as reference frames. Advantageously, the operation of the encoding engine 212 may be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 2A (not shown), the reconstructed video sequence may be a copy of the source video sequence with some errors. The decoder 210 duplicates the decoding process that may be performed by a remote video decoder on a reference frame and may cause the reconstructed reference frame to be stored in the reference image memory 208. In this way, the encoder component 106 locally stores copies of the reconstructed reference frames that have the same content as the reconstructed reference frames that would be obtained by a remote video decoder (in the absence of transmission errors).
[0042] The predictor 206 may perform a prediction search for the encoding engine 212. That is, for a new image to be encoded, the predictor 206 may search the reference image memory 208 for sample data (as candidate reference pixel blocks) or some metadata, such as reference image motion vectors, block shapes, etc., that may serve as an appropriate prediction reference for the new image. The predictor 206 may operate on a per-pixel-block basis of sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 206, it may be determined that the input image may have prediction references taken from at least two reference images stored in the reference image memory 208.
[0043] The outputs of all the above functional units may be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by the various functional units according to techniques well known to those of ordinary skill in the art (such as Huffman coding, variable length coding, arithmetic coding, etc.), thereby converting the symbols into an encoded video sequence.
[0044] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter can be used to buffer at least one encoded video sequence created by the entropy encoder 214 to prepare for transmission via a communication channel 218, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter can be used to combine the encoded video data from the source encoder 202 with other data to be transmitted (e.g., encoded audio data and / or auxiliary data streams (not shown in the source)). In some embodiments, the transmitter can transmit additional data and the encoded video. The source encoder 202 can include such data as part of the encoded video sequence. The additional data can include temporal enhancement layers / spatial enhancement layers / SNR enhancement layers, other forms of redundant data (such as redundant pictures and slices), supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set segments, etc.
[0045] The controller 204 can manage the operation of the video coding component 106. During encoding, the controller 204 can assign a certain encoded picture type to each encoded picture, but this may affect the coding techniques available for the corresponding picture. For example, pictures can typically be assigned as intra pictures (I pictures), predictive pictures (P pictures), and bi-predictive pictures (B pictures). Intra pictures can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are familiar with the variants of I pictures and their corresponding applications and characteristics, and thus will not be elaborated here. Predictive pictures can be encoded and decoded using intra prediction or inter prediction, which use at most one motion vector and reference index to predict the sample values of each block. Bi-predictive pictures can be encoded and decoded using intra prediction or inter prediction, which use at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, at least two predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0046] The source image can typically be spatially subdivided into at least two sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (already encoded) blocks, and the other blocks are determined according to the encoding assignment of the corresponding image used by the block. For example, blocks of an I image can be non-prediction-encoded, or the blocks can be prediction-encoded with reference to already encoded blocks of the same image (spatial prediction or intra-frame prediction). Pixel blocks of a P image can be prediction-encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference image. Blocks of a B image can be non-predictively encoded, or encoded by spatial prediction, or encoded by temporal prediction with reference to one or two previously encoded reference images.
[0047] The captured video can be at least two source images (video images) in a time series. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes the spatial correlation in a given image, while inter-frame image prediction utilizes the (temporal or other) correlation between images. In an embodiment, a particular image being encoded / decoded is segmented into blocks, and the particular image being encoded / decoded is referred to as the current image. When a block in the current image is similar to a reference block in a reference image that has been previously encoded and is still buffered in the video, the block in the current image can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference image, and in the case of using at least two reference images, the motion vector can have a third dimension identifying the reference image.
[0048] The encoder component 106 can perform encoding operations according to a predetermined video coding technique or standard (such as any video coding technique or standard described herein). In its operation, the encoder component 106 can perform various compression operations, including predictive coding operations that utilize the temporal redundancy and spatial redundancy in the input video sequence. Thus, the encoded video data can comply with the syntax specified by the video coding technique or standard used.
[0049] Figure 2B is a schematic block diagram of example elements of a decoder component 122 in some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to the loop filter 256 and is used to transmit data (e.g., via a wired or wireless connection) to the display 124.
[0050] In some embodiments, decoder component 122 includes a receiver that is coupled to channel 218 and is operative to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be operative to receive at least one encoded video sequence to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data and other data (e.g., encoded audio data and / or auxiliary data streams), which may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data and the encoded video. The additional data may be included as part of at least one encoded video sequence. The additional data may be used by decoder component 122 to decode the data and / or, more precisely, reconstruct the original video data. The additional data may be in the form of, for example, a temporal enhancement layer, a spatial enhancement layer, or an SNR enhancement layer, redundant strips, redundant pictures, forward error correction codes, etc.
[0051] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaler / inverse transform unit 258, intra prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference image memory 266, and current image memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, decoder component 122 is implemented at least in part by software.
[0052] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to counter network jitter). In some embodiments, the buffer memory 252 is independent of the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 inside the decoder component 122 (e.g., this buffer memory 252 is used to handle playback timing), a separate buffer memory is provided outside the decoder component 122 (e.g., to counter network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory 252 may not be needed, or the buffer memory 252 can be small. For use in a best-effort packet network such as the Internet, the buffer memory 252 may be needed, which can be relatively large and can advantageously have an adaptive size, and can be implemented at least partially in an operating system or a similar element (not depicted) outside the decoder component 122.
[0053] The parser 254 is used to reconstruct symbols 270 from the encoded video sequence. These symbols can include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The form of the control information for at least one rendering device can be, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technology or standard, and can follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 can extract a set of subgroup parameters of at least one subgroup of pixels used by the video decoder from the encoded video sequence based on at least one parameter corresponding to the set. Subgroups can include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The parser 254 can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0054] The reconstruction of the symbols 270 can involve at least two different units, depending on the type of the encoded video image or its part (such as: inter-frame image and intra-frame image, inter-frame block and intra-frame block) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information, which is parsed by the parser 254 from the encoded video sequence. For clarity, this subgroup control information flow between the parser 254 and at least two units is not described below.
[0055] The decoder component 122 can be conceptually subdivided into a number of functional units, and in some embodiments, these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, the conceptual subdivision of the functional units is retained herein.
[0056] The scaler / inverse transform unit 258 receives quantized transform coefficients in the form of at least one symbol 270 and control information (such as which transform to use, block size, quantization factor, and / or quantization scaling matrix) from the parser 254. The scaler / inverse transform unit 258 can output a block including sample values, and these sample values can be input into the aggregator 268.
[0057] In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is, blocks that do not use predictive information from a previously reconstructed image, but can use predictive information from a previously reconstructed part of the current image. Such predictive information can be provided by the intra prediction unit 262. The intra prediction unit 262 can use the surrounding reconstructed information obtained from the current (partially reconstructed) image in the current image memory 264 to generate a block having the same size and shape as the block being reconstructed. The aggregator 268 can add the predictive information already generated by the intra prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 on a sample-by-sample basis.
[0058] In other cases, the output samples of the scaler / inverse transform unit 258 belong to inter-coded blocks that may be motion compensated. In such cases, the motion compensation prediction unit 260 can access the reference image memory 266 to obtain samples for prediction. After motion compensating the obtained samples according to the symbol 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The motion compensation prediction unit 260 obtains prediction samples from an address within the reference image memory 266, and this address can be controlled by a motion vector. These motion vectors can be provided to the motion compensation prediction unit 260 in the form of symbols 270, and these symbols 270 can have, for example, an X component, a Y component, and a reference image component. Motion compensation can also include interpolation of sample values obtained from the reference image memory 266 when using sub-sample level accurate motion vectors, a motion vector prediction mechanism, and the like.
[0059] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream, which are provided to loop filter unit 256 as symbols 270 in parser 254. However, these in-loop filter techniques can also respond to meta-information obtained during the decoding of a previous (in decoding order) portion of the encoded image or encoded video sequence, as well as in response to previously reconstructed and loop-filtered sample values. The output of loop filter unit 256 can be a sample stream that can be output to a rendering device, such as display 124, and stored in reference image memory 266 for use in subsequent inter-frame prediction.
[0060] Once some encoded images are reconstructed, they can serve as reference images in subsequent predictions. Once an encoded image is reconstructed and the encoded image has been determined to be a reference image (e.g., by parser 254), the current reference image can become part of reference image memory 266, and a new current image memory can be reallocated before starting to reconstruct subsequent encoded images.
[0061] Decoder component 122 can perform decoding operations according to a predetermined video compression technique, which can be documented in a standard (such as any of the standards described herein). The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, as it follows the syntax of the video compression technique or standard, such as that specified in the video compression technique document or standard and specifically in its profile. Moreover, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within the bounds defined by the level of the video compression technique or standard. In some cases, the level limits the maximum image size, maximum frame rate, maximum reconstructed sample rate (e.g., measured in megasamples per second), maximum reference image size, and so on. In some cases, the limits set by the level can be further restricted by the hypothetical reference decoder (HRD) specification and the metadata signaled in the encoded video sequence for HRD buffer management.
[0062] Figure 3 is a schematic block diagram of server system 112 of some embodiments. The server system 112 includes control circuit 302, at least one network interface 304, memory 314, user interface 306, and at least one communication bus 312 for interconnecting these components. In some embodiments, the control circuit 302 includes at least one processor (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes at least one field programmable gate array (FPGA), hardware accelerator, and / or at least one integrated circuit (e.g., application specific integrated circuit).
[0063] At least one network interface 304 can be used to connect to at least one communication network (e.g., wireless network, wired network, and / or optical network). The communication network can be a local area network, wide area network, metropolitan area network, vehicular network, industrial network, real-time network, delay-tolerant network, etc. Examples of communication networks include local area networks (such as Ethernet, WLAN), cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), wired or wireless wide-area digital television networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicular and industrial networks including CANBus, etc. Such communication can be one-way receive-only (e.g., broadcast television), one-way transmit-only (e.g., CANbus to certain CANbus devices), or two-way (e.g., to other computer systems using local digital networks or wide-area digital networks). Such communication can include communication to at least one cloud computing network.
[0064] The user interface 306 includes at least one output device 308 and / or at least one input device 310. The at least one input device 310 can include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. The at least one output device 308 can include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., display screens or monitors), etc.
[0065] The memory 314 can include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as at least one disk storage device, optical disk storage device, flash memory device, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes at least one storage device located remotely from the control circuit 302. The memory 314 or at least one non-volatile solid-state memory device within the memory 314 includes a non-volatile computer-readable storage medium. In some embodiments, the memory 314 or the non-volatile computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: · An operating system 316, which includes processes for handling various basic system services and for performing hardware-related tasks; · A network communication module 318, which is used to connect the server system 112 to other computing devices via at least one network interface 304 (e.g., via wired and / or wireless connections); ● A codec module 320 for performing various functions regarding encoding and / or decoding of data such as video data. In some embodiments, the codec module 320 is an instance of the encoding component 114. The codec module 320 includes, but is not limited to, one or more of the following: ο A decoding module 322 for performing various functions regarding decoding of encoded data, such as those functions previously described with respect to the decoder component 122; and ο An encoding module 340 for performing various functions regarding encoding of data, such as those functions previously described with respect to the encoder component 106; and · An image memory 352 for storing images and image data, e.g., for use with the codec module 320. In some embodiments, the image memory 352 includes one or at least two of the following: a reference image memory 208, a buffer memory 252, a current image memory 264, and a reference image memory 266.
[0066] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., for performing various functions previously described with respect to the parser 254), a transformation module 326 (e.g., for performing various functions previously described with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., for performing various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra prediction unit 262), and a filter module 330 (e.g., for performing various functions previously described with respect to the loop filter 256).
[0067] In some embodiments, the encoding module 340 includes a code module 342 (e.g., for performing various functions previously described with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., for performing various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 a subset of the modules shown. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.
[0068] Each of the above-mentioned modules stored in the memory 314 corresponds to a set of instructions for performing the functions described herein. The above-mentioned modules (e.g., sets of instructions) need not be implemented as separate software programs, processes, or modules, and thus various subsets of these modules may be combined or otherwise reorganized in various embodiments. For example, optionally, the codec module 320 does not include separate decoding and encoding modules, but uses the same set of modules to perform two sets of functions. In some embodiments, the memory 314 stores a subset of the above-mentioned modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.
[0069] Although Figure 3 FIG. illustrates a server system 112 according to some embodiments, but Figure 3 is more intended as a functional description of the various features that may exist in at least one server system than as a structural diagram of the embodiments described herein. In practice, and as will be appreciated by those of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in may be implemented on a single server, and a single item may be implemented by at least one server. The actual number of servers used to implement the server system 112, and how the features are distributed among them, will vary from one implementation to another and will optionally depend in part on the amount of data traffic processed by the server system during peak usage periods as well as during average usage periods. Example Encoding and Decoding Processes and Techniques
[0070] As mentioned above, some codecs (e.g., AV1) operate on pixel blocks. Each pixel block may be processed in a predictive transform coding scheme, where prediction is obtained using in-frame reference pixels, inter-frame motion compensation, or some combination of both. The prediction residuals may undergo a transform (e.g., 2-D unitary transform) to further remove spatial correlation, and the transform coefficients are quantized. Then, arithmetic coding may be used to entropy code the predicted syntax elements and the quantized transform coefficient indices.
[0071] Figure 4A Schematic for the calculation of a prediction block for some embodiments. In Figure 4A the example of, intra-frame prediction is performed on the current block 402 to generate a prediction block 404. The current block 402 includes a set of samples (e.g., a pixel block) S 11 to S 44 , and the prediction block 404 includes a set of predictions P 11 to P 44 . Figure 4BSchematic diagram of the calculation of the residual block for some embodiments. As Figure 4B shown, the predicted block 404 is subtracted from the current block 402 to generate a residual block 406 including a set of residuals R 11 to R 44 . For example, the difference between each sample and the corresponding predicted sample is calculated. Figure 4C Schematic diagram of the calculation of the reconstruction block for some embodiments. As Figure 4C shown, the residual block 406 undergoes at least one transformation and quantization to generate a set of residual coefficients. The set of residual coefficients can be transmitted from the encoder component to the decoder component. The set of residual coefficients undergoes inverse quantization and inverse transformation to generate a reconstructed residual block 408. The reconstructed residual block 408 is combined with the predicted block 404 (e.g., adding the reconstructed residuals of the reconstructed residual block 408 to the prediction of the predicted block 404) to generate a reconstructed block 410 corresponding to the current block 402.
[0072] To reduce the redundancy in the residual signal, various residual prediction techniques have been developed. These techniques predict the residual signal and encode the optimized residual. Residual Difference Pulse Code Modulation (RDPCM) includes using sample-based differential pulse code modulation along the horizontal or vertical axis. By doing so, each residual row in the horizontal mode (or each residual column in the vertical direction) can be reconstructed at the decoder by summing the scaled differential pulse code modulation residual levels along the corresponding row (or column). RDPCM can be of explicit type or implicit type. The explicit type requires supplementary signaling of the direction and its application is limited to inter-frame prediction blocks. On the other hand, the implicit type does not require direction signaling and can only be applied to intra-frame prediction blocks, where the prediction direction is determined by the intra-frame prediction mode. Block-based Differential Pulse Code Modulation (BDPCM) performs sample-based differential pulse code modulation on the reconstructed samples rather than the residual samples. The indication of the use of the second mode occurs during the prediction mode reconstruction process. This signaling involves two syntax elements, each for luminance and chrominance. For example, the initial syntax element flag indicates whether it is used, while the second syntax element flag specifies the horizontal or vertical direction.
[0073] To reduce the redundancy in the residual domain, a line-by-line residual prediction mode can be used to generate an optimized residual block. This line-by-line residual domain prediction can be performed in the horizontal or vertical direction. For the case of horizontal prediction, the prediction can be defined as shown in Equation 1 below, and for the case of vertical prediction, the prediction can be defined as shown in the following Equation 2. Wherein, x and y respectively represent the row index and the column index, and r(*,*) represents the pixel values in the original residual block (e.g., generated by the intra prediction process).
[0074] In some embodiments, a bi - directional prediction mode is used to generate the optimized residual block. The bi - directional prediction mode can be performed in the horizontal or vertical direction. For the case of horizontal prediction for encoding, the prediction can be defined as shown in Equation 3 below, and for the case of vertical prediction for encoding, the prediction can be defined as shown in the following Equation 4. Where r(*,*) represents the pixel values in the original residual block (e.g., generated by the intra prediction process).
[0075] For the case of horizontal prediction for decoding, the prediction can be defined as shown in Equation 5 below, and for the case of vertical prediction for decoding, the prediction can be defined as shown in Equation 6 below. Where r′(*,*) represents the pixel values in the optimized residual block. For example, the residual block values are determined by Determined.
[0076] In some embodiments, intra prediction is performed on an encoded block or each sub - block in an encoded block, and a residual block is generated by subtracting the predicted block from the reconstructed samples of neighboring blocks. Then, short - range intra prediction is used for the residual block, thereby generating an optimized residual block. For example, the optimized residual signal can be calculated as shown in Equation 7 below. Where Represents the optimized residual signal.
[0077] The reconstruction process can be performed by summing the samples along the determined direction, as shown in Equation 8 and the following Equation 9.
[0078] Thus, a decoder according to an embodiment of the present disclosure can receive video data from a video bitstream, where the video data includes at least two blocks (including a first block) and at least two residual coefficients. The first block is encoded by an intra prediction mode. In addition, at least two residual coefficients are generated by using short - range intra prediction on a residual block of the first block. Further, a residual block is generated by using the intra prediction mode on the first block. Then, the decoder can generate an optimized residual block of the first block according to the at least two residual coefficients and use the optimized residual block to reconstruct the first block.
[0079] In some embodiments, a flag is signaled to indicate whether a line - by - line residual prediction mode is used. In some embodiments, the flag is signaled separately for a luminance component and a chrominance component. In some embodiments, if the flag indicates the use of the line - by - line residual prediction mode, another flag is signaled to indicate whether the direction of the line - by - line residual prediction mode is a vertical direction or a horizontal direction. In some embodiments, when the line - by - line residual prediction mode is used, an angle increment and / or a multi - reference line (mrl) index is inferred to be zero. In some embodiments, the transform block size is fixed at a minimum transform size (e.g., 4×4), and the line - by - line residual prediction mode is performed on 4×4 residual blocks.
[0080] In some embodiments, a forward skip coding (FSC) mode is enabled together with the line - by - line residual prediction mode (e.g., in a lossless coding scheme). For coefficients obtained after a 2 - D identity transform (IDTX), FSC can be a simpler and more efficient residual coding method. A block encoded by FSC has less TU - level signaling because the transform type signaling and the block - end index signaling of the IDTX block are avoided, where the former reduces the symbol count in the TX_SET_INTRA set by 1 each. Finally, since FSC disables the signaling of the multi - reference line (MRL) index, the filter intra mode, and the angle increment syntax when the transform type is IDTX, this can simplify the reconstruction process of the FSC block, so it is designed as a more economical alternative to the intra - block coding mode.
[0081] Figure 4D is a schematic diagram of the calculation of an optimized residual block and a differential block in some embodiments. As Figure 4D shown, short - range intra prediction is used on the residual block 406 to generate an optimized residual block 420 including residuals Z 11 to Z 44 . As discussed in more detail below, the short - range intra prediction can include line - by - line prediction and / or bidirectional prediction. Figure 4D Also shown is the generation of a differential block including residuals D 11 to D 44The differential block 422. According to some embodiments, the difference of the residuals is used to generate residual coefficients (e.g., via at least one transformation and quantization). For example, the difference of the residuals is used as a substitute for the residual of the residual block 406.
[0082] Figure 4E is a schematic diagram of the calculation of the residual block and the reconstructed block during the decoding process in some embodiments. For example, the optimized residual coefficients are received, and the optimized residual block is recovered according to these coefficients (e.g., using an inverse transformation and an inverse quantization process). After recovering the optimized residual block, short-range intra prediction can be applied to recover the residual block. After recovering the residual block, intra prediction can be applied to recover the reconstructed block.
[0083] Figure 4F is a schematic diagram of an example line-by-line prediction for the encoding process in some embodiments. In Figure 4F the example, the residual block 406 includes samples r ij , and the optimized prediction block 420 includes samples p′ ij . The residual block 406 can be obtained by using intra prediction on the current block. Figure 4F The residual prediction block 420 in is obtained via line-by-line vertical prediction. In Figure 4F the example, the samples p′ ij of the residual prediction block 420 are subtracted from the samples r ij of the residual block 406 to obtain the optimized residual block 422 including samples r′ ij . In some embodiments, the optimized residual block 422 is used to generate optimized residual coefficients, and these optimized residual coefficients are signaled in the subsequent bitstream.
[0084] Figure 4G is a schematic diagram of an example line-by-line prediction for the decoding process in some embodiments. In Figure 4G the example, the optimized residual block 422 including samples r′ ij is obtained (e.g., by inverse-transforming the residual coefficients received via the bitstream). Figure 4G The residual prediction block 420 including samples p′ j in is obtained via line-by-line vertical prediction. In Figure 4G the example, the samples p′ ij of the residual prediction block 420 are added to the samples r′ ij of the optimized residual block 422 to recover the residual block 406 including samples r ij . In some embodiments, the recovered residual block is used to obtain the reconstructed block of the current block.
[0085] Figure 5A and Figure 5B are schematic diagrams of example residual blocks in some embodiments. Figure 5AResidual block 502 is shown, which includes a set of residuals R corresponding to rows 1 to 8 and columns 1 to 8 11 to R 88 . Figure 5B Residual block 502 with index line 1 (e.g., corresponding to row 1), index line 2 (e.g., corresponding to row 4), index line 3 (e.g., corresponding to row 2), and index line 4 (e.g., corresponding to row 3) is shown. In Figure 5B , residual block 502 further includes index line A (e.g., corresponding to column 1), index line B (e.g., corresponding to column 4), index line C (e.g., corresponding to column 2), and index line D (e.g., corresponding to column 3).
[0086] Additional residual prediction techniques are described below. The disclosed techniques can be used alone or in any order of combination. These techniques, as well as the encoder and decoder methods, can be performed using a processing circuit, which may include at least one processor or integrated circuit. As an example, a program stored in a non - volatile computer - readable medium can be executed by at least one processor.
[0087] Figure 6A is a flowchart of a video encoding method 600 of some embodiments. The method 600 can be performed by a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory that stores instructions executed by the control circuit. In some embodiments, method 600 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.
[0088] The system receives (602) video data. The video data includes at least two blocks, and the at least two blocks include a first block (e.g., current block 402) to be encoded in a first intra-frame prediction mode. For example, the system receives video data from a video source (e.g., video source 104). The system generates (604) a residual block (e.g., residual block 406) of the first block by using the first intra-frame prediction mode in a first direction (e.g., horizontal direction or vertical direction). In some embodiments, the intra-frame prediction mode is directional intra-frame prediction. The system generates (606) an optimized residual block (e.g., optimized residual block 420) of the first block by using a second intra-frame prediction on the residual block in a second direction. In some embodiments, the second intra-frame prediction is short-range intra-frame prediction. In some embodiments, the short-range intra-frame prediction is line-by-line prediction. In some embodiments, the short-range intra-frame prediction is bidirectional prediction. The system signals (608) the optimized residual block via a video bitstream. In some embodiments, the system signals the difference between the optimized residual block and the residual block (e.g., as a residual coefficient).
[0089] Figure 6B is a flowchart of a method 650 for decoding video in some embodiments. The method 650 may be executed by a computing system (e.g., server system 112, source device 102, or electronic device 120). The computing system has a control circuit and a memory that stores instructions for execution by the control circuit. In some embodiments, the method 650 is executed by executing instructions stored in the memory (e.g., memory 314) of the computing system.
[0090] The system receives (652) video data from a video bitstream (e.g., via channel 218), the video data including a first block (e.g., current block 402) and at least two residual coefficients (e.g., Figure 4DThe residual coefficients shown). Among them, the first block is encoded using the first intra prediction (e.g., directional intra prediction). At least two residual coefficients are generated by using the second intra prediction (e.g., short - range intra prediction) for the residual block (e.g., residual block 406) of the first block in a first direction (e.g., horizontal direction or vertical direction), and the residual block is generated by using the first intra prediction mode for the first block in a second direction (e.g., vertical direction or horizontal direction). The system generates (654) an optimized residual block (e.g., reconstructed residual block 408) of the first block according to at least two residual coefficients. For example, the system performs inverse quantization and at least one inverse transform on at least two residual coefficients to generate the optimized residual block (e.g., reconstructed residual block 408). The system decodes (656) the first block by using the optimized residual block. For example, the system generates a reconstructed block (e.g., reconstructed block 410) for the first block.
[0091] Although Figure 6A and Figure 6B show at least two logical stages with a specific order, the stages that do not depend on the order can be reordered, and other stages can be combined or split. A certain reordering or other grouping method not specifically mentioned is obvious to those of ordinary skill in the art, so the ordering and grouping methods presented herein are not exhaustive. In addition, it should be recognized that the stages can be implemented in hardware, firmware, software, or any combination thereof.
[0092] In some embodiments, intra prediction is performed on the encoded block or each sub - block in the encoded block, so as to generate a residual block by subtracting the predicted block from the reconstructed samples of adjacent blocks. In some embodiments, short - range intra prediction is used for the residual block, so as to generate an optimized residual block. In some embodiments, intra prediction is performed on M×N blocks regardless of the size of the encoded block. Example values of M and N include, but are not limited to, 1, 2, 4, 8, 16, 32, and 64. In some embodiments, short - range intra prediction is performed on the residual block for lossless coding mode.
[0093] In some embodiments, line-by-line prediction is used as short-distance intra prediction. For example, the residuals in a specific row or column are predicted using their neighboring previous lines, and the difference between these residuals and the predicted residuals is used as the input for subsequent transformation, quantization, or entropy coding processes. In some embodiments, line-by-line prediction is performed on the samples / residuals in all rows or columns except for the samples / residuals in the first row / column. In some embodiments, for the first row / column, residual samples generated from at least two reconstructed rows / columns of neighboring blocks are used for prediction. In some embodiments, line-by-line prediction is performed on these residuals in the horizontal direction. For example, the predicted residual of the first row is set to zero, and the predicted residuals of subsequent rows are predicted using their neighboring previous rows. In some embodiments, line-by-line prediction is performed on these residuals in the vertical direction. For example, the predicted residual of the first column is set to zero, and the predicted residuals of subsequent columns are predicted using their previous neighboring columns.
[0094] In some embodiments, the short-distance intra prediction is bidirectional prediction. In some embodiments, bidirectional prediction is performed on each M×N residual block. For example, the predicted residuals of the residuals in the first index line and the second index line of the M×N residual block are set to zero, and the weighted average of the residuals in the first index line and the second index line is used to predict the residuals in the third index line and the fourth index line. In some embodiments, the difference between the residual and the predicted residual is used as the input for subsequent transformation, quantization, or entropy coding processes. Example values of M and N include, but are not limited to, 1, 2, 4, 8, 16, 32, and 64. In some embodiments, the first line and the second line are not neighboring lines. Examples of the first index line and the second index line include, but are not limited to, the first line and the fourth line respectively along a given direction. Examples of the third index line and the fourth index line include, but are not limited to, the second line and the third line respectively along a given direction. In some embodiments, the weighting factors used for weighted averaging of the residuals in the first index line and the second index line depend on the distance between the residual and its predicted value. For example, when predicting the residuals in the second row / column, the weighting factors of the residuals in the first index line and the second index line are {2 / 3, 1 / 3} or {3 / 4, 1 / 4}. As another example, when predicting the residuals in the third row / column, the weighting factors of the residuals in the first index line and the second index line are {1 / 3, 2 / 3} or {1 / 4, 3 / 4}.
[0095] In some embodiments, bidirectional prediction is performed in the horizontal direction. For example, the prediction residuals in the first index line and the second index line are set to zero. In this example, the residuals in the third index line and the fourth index line are predicted by weighted averaging of the residuals in the first row and the fourth row. In some embodiments, bidirectional prediction is performed in the vertical direction. For example, the prediction residuals in the first index line and the second index line are set to zero. In this example, the residuals in the third index line and the fourth index line are predicted by weighted averaging of the residuals in the first column and the fourth column.
[0096] In some embodiments, for short-range intra prediction, the weighted average of the residuals from at least two neighboring lines is used to predict the residuals in the subsequent line. In some embodiments, the weighting factor for weighted averaging of the residuals in at least two neighboring lines depends on the distance between the residuals in the current line and the residuals in the neighboring lines. In some embodiments, the residuals in two neighboring rows are used to predict the residuals in the current line, the weighting factor for the residuals in the nearest neighboring line is set to a first value, and the weighting factor for the residuals in the other line is set to a second value. Examples of the first value and the second value include, but are not limited to, 2 / 3 and 1 / 3, respectively.
[0097] In some embodiments, multiple short-range prediction methods are used for the residual block order. For example, a line-by-line prediction method is first used for the residual block to generate an optimized residual block, and then a bidirectional prediction method is used for the optimized residual block to generate a final residual block.
[0098] In some embodiments, a flag is signaled in the bitstream to indicate which short-range intra prediction method is used for the residual block. In some embodiments, two separate flags are used to signal whether short-range prediction is used for the luma residual block plane and the chroma residual block plane. In some embodiments, the direction of short-range prediction used for the residual block is signaled in the bitstream. In some embodiments, the direction of short-range residual block prediction is inferred from the intra prediction mode. In some embodiments, the context for entropy coding of the flag for short-range prediction of the chroma residual block depends on the corresponding luma flag. In some embodiments, whether at least one short-range intra prediction is used is signaled in the high-level syntax (including, but not limited to, sequence flags, GOP flags, picture flags, slice flags, or tile-level flags).
[0099] In some embodiments, a first direction is employed in intra prediction of an encoded block or each sub - block in an encoded block, such that a residual block is generated by subtracting a predicted block from the reconstructed samples of neighboring blocks. Then, a second direction is employed in short - range residual prediction to predict these residuals, and the difference between these residuals and the predicted residuals is used as an input for subsequent processing, which includes, but is not limited to, transformation, quantization, entropy coding, and in - loop filtering. At the decoder, the differences between the residuals and the predicted residuals are parsed, and then these differences are added to the predicted residuals to obtain the reconstructed residual samples.
[0100] In some embodiments, the direction of the first direction is different from the direction of the second direction. In some embodiments, the direction of the second direction is the same as the direction of the first direction. In some embodiments, the value of the first direction is used as a context for entropy - coding the second direction. In some embodiments, high - level (including, but not limited to, sequence - level, frame - level, slice - level, super - block - level) syntax is signaled in the bitstream to indicate whether the second direction is the same as the first direction.
[0101] In some embodiments, line - by - line prediction or bi - directional prediction is used as short - range prediction. For example, line - by - line prediction is employed as a type of short - range intra prediction in a residual block, and the residuals in a particular line are predicted using its neighboring previous line. As another example, bi - directional prediction is employed as short - range intra prediction in a residual block, and the predicted residuals of the residuals in the first index line and the second index line of the residual block are set to zero, while the weighted average of the residuals in the first index line and the second index line is used to predict the residuals in the third index line and the fourth index line.
[0102] In some embodiments, the angle of the second direction of short - range residual prediction is implicitly determined based on the angle of the first direction of intra prediction. In some embodiments, when the prediction angle of intra prediction is closer to the horizontal direction compared to the vertical direction, then short - range prediction of the residuals is performed in the horizontal direction. In some embodiments, when the prediction angle of intra prediction is closer to the vertical direction than the horizontal direction, then short - range prediction of the residuals is performed in the vertical direction. In some embodiments, when the prediction angle of intra prediction is closer to the diagonal direction than the horizontal direction or the vertical direction, then short - range prediction of the residuals is performed in the diagonal direction. In some embodiments, short - range prediction is used for N nominal angles of intra prediction. For example, N is equal to 4, 6, 8, or 10.
[0103] In some embodiments, at least one syntax element is signaled in the bitstream to indicate the direction / angle of short - range prediction. In some embodiments, a syntax element indicating the direction of short - range prediction is signaled in the bitstream. In some embodiments, a first syntax element is used to indicate the nominal / master direction, while a second syntax element is used to indicate the angular increment relative to the nominal direction. In some embodiments, the values of the supported increment angles are predefined in a lookup table, and an index of the increment angle in the lookup table is signaled in the bitstream. As an example, a first syntax is signaled to indicate whether the direction of short - range residual prediction is the vertical direction or the horizontal direction, and then a second syntax is signaled to indicate the angular increment relative to the specified master direction.
[0104] In some embodiments, a first syntax element is used to indicate the nominal direction; a second syntax element is used to indicate whether the angular increment is zero. If the angular increment is not zero, a third syntax element and a fourth syntax element are further used. The third syntax element is used to indicate whether the angular increment is positive or negative. The fourth syntax element is used to indicate the absolute value of the angular increment.
[0105] In some embodiments, only the angular increment syntax is signaled to derive a second direction (e.g., the prediction direction used for residual prediction), and then the direction used for residual prediction is derived by adding the angular increment value to the nominal prediction direction (or prediction direction) of the intra - prediction mode.
[0106] Some embodiments include transform coding techniques used for short - range residual prediction. In some embodiments, to process a residual block, a selected transform coding mode is signaled at a first processing unit level, and a transform process is performed at a second processing unit level, where the transform coding mode refers to any parameter or operation involved in the transform process, and a transform method can be applied in the forward transform process of the encoder or the inverse transform process of the encoder and / or decoder. For example, the ranges of the first processing unit level and the second processing unit level can include sequence level, frame level, super - block level, coding - block level, prediction - block level, or transform - block level.
[0107] In some embodiments, the residual block generated by intra prediction or the optimized residual block is used as the input to the transformation process. In some embodiments, intra prediction is performed on the encoded block or each sub-block in the encoded block to generate a residual block by subtracting the predicted block from the reconstructed samples of neighboring blocks. In some embodiments, a short-range intra prediction method is used for the residual block to generate an optimized residual block. In some embodiments, the optimized residual block is used as the input to subsequent transform coding. In some embodiments, different transform kernels can be used for the optimized residual block. For example, no transformation is performed on the optimized residual block (or the transform kernel is an identity transform). As an example, no transformation (or an identity transform is applied) is performed in one direction (e.g., the horizontal direction or the vertical direction), while a lossless transformation (e.g., Hadamard transform) is performed in the other direction.
[0108] In some embodiments, the first processing unit level is the same as the second processing unit level. In some embodiments, the type of transform coding mode is signaled at the encoded block level, the size of the optimized residual block is the same as the size of the encoded block, and the size of the transform coded block is also the same as the size of the encoded block. For example, the syntax element of the transform block size can be inferred based on the size of the optimized residual block or the encoded block size, and the syntax element of the transform block size does not need to be signaled.
[0109] In some embodiments, the first processing unit level is different from the second processing unit level. In some embodiments, the type of transform coding kernel is signaled at the encoded block level, and the size of the transform block used to perform the transformation process is smaller than the size of the encoded block. In some embodiments, the size of the transform block is fixed, and the same transform coding kernel is used for the transform coded blocks within an encoded block. For example, the syntax element of the transform block size does not need to be signaled in the bitstream. In one example, an identity transform is signaled at the encoded block level, and the transform size is fixed to M×N regardless of the size of the encoded block. In another example, a Hadamard transform (or a different lossless transform) is signaled at the encoded block level, and the transform size is fixed to M×N regardless of the size of the encoded block. In some embodiments, M and N are selected to correspond to the smallest allowed transform size (e.g., both M and N are equal to 4).
[0110] In some embodiments, the transform block size and the transform coding mode are determined by a given cost metric at the encoder. For example, the best transform coding type and the best transform size are signaled at the encoded block level. In one example, the cost metric is the rate-distortion cost used in rate-distortion optimization.
[0111] In some embodiments, it is signaled at an advanced syntax (including but not limited to sequence, GOP, frame, or slice level) whether a first processing unit level is different from a second processing unit level.
[0112] In some embodiments, a transform scheme includes applying a transform skip (or identity transform) in one direction (e.g., horizontal or vertical direction) and applying a Hadamard transform in another direction. In some embodiments, the transform scheme is only applied to lossless coding modes. In some embodiments, the transform scheme is only applied to a specific M×N transform block size (e.g., 4×4 transform block size). In some embodiments, it is signaled separately for each direction whether to select a transform skip (or identity transform) or a Hadamard transform.
[0113] Some example embodiments are now described.
[0114] (A1) In one aspect, some embodiments include a video coding method (e.g., method 600). In some embodiments, the method is performed by a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed by a codec module (e.g., codec module 320). In some embodiments, the method is performed by a source coding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving video data including at least two blocks, the at least two blocks including a first block (e.g., a chrominance or luminance block); (ii) selecting a transform coding mode for the first block, the transform coding mode using a first processing unit level; and (iii) performing a transform process on the first block using the selected transform coding mode, wherein the transform process is performed on the transform block at a second processing unit level and the second processing unit level is not signaled.
[0115] (A2) In some embodiments of A1, a transform process is performed on an optimized residual block corresponding to the first block (e.g., optimized residual block 420), and the optimized residual block is generated by performing intra prediction on the residual block of the first block.
[0116] (A3) In some embodiments of A2, a first transform kernel is used for the residual block and a second transform kernel different from the first transform kernel is used for the optimized residual block.
[0117] (A4) In some embodiments of A2 or A3, a residual block is generated by subtracting a predicted block from reconstructed samples of neighboring blocks.
[0118] (A5) In some embodiments of any of A2 to A4, the intra prediction is short-range intra prediction.
[0119] (A6) In some embodiments of A5, the short-range intra prediction includes line-by-line prediction, in which the residuals in a particular row or column are predicted using the neighboring prior rows or columns.
[0120] (A7) In some embodiments of A5, the short-range intra prediction includes bidirectional prediction, in which the residuals in the third and fourth index lines of a residual block are predicted using a weighted average of the residuals in the first and second index lines of the residual block.
[0121] (A8) In some embodiments of any of A1 to A7, the first processing unit level is the same as the second processing unit level.
[0122] (A9) In some embodiments of any of A1 to A7, the first processing unit level is different from the second processing unit level.
[0123] (A10) In some embodiments of any of A1 to A9, the method further includes signaling a second syntax element that indicates whether the first processing unit is the same as the second processing unit.
[0124] (A11) In some embodiments of any of A1 to A10, the transform process includes a first transform performed in a first direction and a second transform performed in a second direction.
[0125] (B1)In another aspect, some embodiments include a video decoding method (e.g., method 650). In some embodiments, the method is performed by a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed by a codec module (e.g., codec module 320). In some embodiments, the method is performed by a parser (e.g., parser 254), a motion prediction component (e.g., motion compensation prediction unit 260), and / or an intra prediction component (e.g., intra prediction unit 262). The method includes: (i) receiving video data (e.g., an encoded video sequence) from a video bitstream, the video data including at least two blocks and syntax elements, the at least two blocks including a first block, wherein the syntax elements are signaled at a first processing unit level; (ii) selecting a transform coding mode based on the syntax elements; and (iii) performing a transform process (e.g., an inverse transform) on the first block using the selected transform coding mode, wherein the transform process is performed on the transformed block at a second processing unit level, and wherein the second processing unit level is not signaled. For example, to process a residual block, the selected transform coding mode is signaled at the first processing unit level, and the transform process is performed at the second processing unit level. As an example, the transform block size is fixed, and the same transform coding kernel is used for the transform coding blocks within an encoded block. Thus, in this example, there is no need to signal a syntax element for the transform block size in the bitstream (which reduces bandwidth consumption).
[0126] The transform coding mode refers to any parameter or operation involved in the transform process, and a transform method can be applied in the forward transform process of the encoder or the inverse transform process of the encoder and / or decoder. The scope of the first processing unit level and the second processing unit level includes but is not limited to sequence level, frame level, superblock level, encoded block level, prediction block level, and transform block level. In some embodiments, the transform block size and the transform coding mode are determined by a given cost metric of the encoder, and the best transform coding type and the best transform size are signaled at the encoded block level. For example, the cost metric is the rate-distortion cost used in rate-distortion optimization. The method of A1 can be applied as part of a lossless or lossy coding scheme.
[0127] (B2)In some embodiments of B1, the transform process includes performing an inverse transform on at least two residual coefficients of the first block to generate an optimized residual block corresponding to the first block.
[0128] (B3) In some embodiments of B2, the method further includes generating a residual block based on the optimized residual block by using short - range intra - prediction on the optimized residual block, where the residual block is used to reconstruct the first block. The short - range intra - prediction defines a method for predicting the residual block along the horizontal or vertical direction, thereby generating the optimized residual block. The difference between this short - range residual prediction and the uni - directional method / bidirectional method lies in the utilization of neighboring samples in the residual domain. Therefore, line - by - line prediction or bidirectional prediction can be used as short - range prediction. For example, the short - range intra - prediction is line - by - line prediction, in which neighboring prior rows or columns are used to predict the residuals in a specific row or column. As another example, the short - range intra - prediction is bidirectional prediction, in which the weighted average of the residuals in the first index line and the second index line of the residual block is used to predict the residuals in the third index line and the fourth index line of the residual block.
[0129] (B4) In some embodiments of B3, a first transform kernel is used for the residual block, and a second transform kernel different from the first transform kernel is used for the optimized residual block. For example, different transform kernels can be used for the optimized residual block. In some embodiments, no transformation is performed on the optimized residual block, or an identity transformation is performed on the optimized residual block. For example, no transformation (or an identity transformation) is performed in one direction (e.g., the horizontal direction or the vertical direction), while a lossless transformation (e.g., Hadamard transformation) is performed in the other direction.
[0130] (B5) In some embodiments of B3 or B4, the residual block is generated by adding the predicted block to the residual block. For example, a second intra - prediction is performed on the encoded block (or each sub - block in the encoded block).
[0131] (B6) In some embodiments of any of B3 to B5, the short - range intra - prediction includes line - by - line prediction, in which neighboring prior rows or columns are used to predict the residuals in a specific row or column. For example, line - by - line prediction is adopted as a short - range intra - prediction. The residuals in a specific row or column are predicted using its neighboring prior lines.
[0132] (B7) In some embodiments among any of B3 to B5, short-range intra prediction includes bidirectional prediction, in which the weighted average of the residuals in the first index line and the second index line of the residual block is used to predict the residuals in the third index line and the fourth index line of the residual block. For example, bidirectional prediction is performed on each M×N residual block. In some embodiments, the weighted average of the residuals from at least two neighboring lines and / or at least two prior lines is used to predict the residuals in a specific index line. As an example, bidirectional prediction is adopted in the residual block as short-range intra prediction. In this example, the predicted residuals of the residuals in the first index line and the second index line of the residual block are set to zero, while the weighted average of the residuals in the first index line and the second index line is used to predict the residuals in the third index line and the fourth index line.
[0133] (B8) In some embodiments among any of B1 to B7, the first processing unit level is the same as the second processing unit level. For example, the type of transform coding mode is signaled at the coding block level, and the optimized residual block size is the same as the coding block size and the transform coding block size is also the same as the coding block size. As an example, the syntax element of the transform block size can be inferred based on the optimized residual block size or the coding block size, without signaling the syntax element of the transform block size.
[0134] (B9) In some embodiments among any of B1 to B7, the first processing unit level is different from the second processing unit level. In some embodiments, the type of transform coding kernel is signaled at the coding block level, and the transform block size for performing the transform process is smaller than the coding block size. As an example, the identity transform is signaled at the coding block level, and regardless of the coding block size, the transform size is fixed to M×N, where M and N are positive integers (e.g., equal to 4). In some embodiments, M×N is the minimum transform size. As another example, the Hadamard transform (or a different lossless transform) is signaled at the coding block level, and regardless of the coding block size, the transform size is fixed to M×N.
[0135] (B10) In some embodiments among any of B1 to B9, the video bitstream further includes a second syntax element that indicates whether the first processing unit is the same as the second processing unit. In some embodiments, the second syntax element is a high-level syntax (HLS) element. In some embodiments, HLS is signaled at a level higher than the block level. For example, HLS can correspond to the sequence level, frame level, slice level, or tile level. As another example, HLS can be signaled in the video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptive parameter set (APS), slice header, picture header, tile header, and / or CTU header.
[0136] (B11)In any of some embodiments among B1 to B10, the transformation process includes a first transformation performed in a first direction and a second transformation performed in a second direction. For example, the transformation process includes performing a transform skip (or identity transform) in one direction (e.g., the horizontal direction or the vertical direction), while performing a Hadamard transform in the other direction. In some embodiments, the transformation process is applied in a lossless coding mode. In some embodiments, the transformation process is performed with a specific M×N transform block size (e.g., 4×4). In some embodiments, it is signaled separately for each direction whether to select a transform skip (or identity transform) or a Hadamard transform.
[0137] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, the memory storing at least one set of instructions for execution by the control circuit, the at least one set of instructions including instructions for performing any of the methods described herein (e.g., A1 to A11 and B1 to B11 above).
[0138] In yet another aspect, some embodiments include a non - volatile computer - readable storage medium storing at least one set of instructions for execution by a control circuit of a computing system, the at least one set of instructions including instructions for performing any of the methods described herein (e.g., A1 to A11 and B1 to B11 above).
[0139] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of at least one of the associated listed items. It should also be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of at least one other feature, integer, step, operation, element, component, and / or combination thereof. As used herein, depending on the context, the term "if" can be construed to mean "when the stated precondition is true" or "once the stated precondition is true" or "in response to determining that the stated precondition is true" or "in accordance with determining that the stated precondition is true" or "in response to detecting that the stated precondition is true". Similarly, depending on the context, the phrases "if it is determined that [the stated precondition is true]" or "if [the stated precondition is true]" or "when [the stated precondition is true]" can be construed to mean "upon determining that the stated precondition is true" or "in response to determining that the stated precondition is true" or "in accordance with determining that the stated precondition is true" or "upon detecting that the stated precondition is true" or "in response to detecting that the stated precondition is true".
[0140] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the operating principles and practical applications, thereby enabling others skilled in the art to implement them.
Claims
1. A method for video decoding, executed by a computing system, the computing system comprising a memory and at least one processor, characterized in that, The method includes: Receiving video data from a video bitstream, the video data including at least two blocks and syntax elements, the at least two blocks including a first block, wherein the syntax elements are signaled at a first processing unit level; Selecting a transform coding mode based on the syntax elements; and Performing a transform process on the first block using the selected transform coding mode, wherein the transform process is performed on the transformed block at a second processing unit level, wherein the second processing unit level is obtained by inference.
2. The method according to claim 1, characterized in that, The transform process includes performing an inverse transform on at least two residual coefficients of the first block to generate an optimized residual block corresponding to the first block.
3. The method according to claim 2, wherein Further includes: Generating a residual block according to the optimized residual block by performing short-range intra prediction on the optimized residual block, wherein the residual block is used to reconstruct the first block.
4. The method according to claim 3, characterized in that, Using a first transform kernel for the residual block and a second transform kernel for the optimized residual block, the second transform kernel being different from the first transform kernel.
5. The method according to claim 3, wherein Generating the residual block by adding a predicted block to the residual block.
6. The method according to claim 3, wherein The short-range intra prediction includes line-by-line prediction, in which prior adjacent rows or columns are used to predict the residuals in a specific row or column.
7. The method according to claim 3, wherein The short-range intra prediction includes bidirectional prediction, in which the weighted average of the residuals in a first index line and a second index line of the residual block is used to predict the residuals in a third index line and a fourth index line of the residual block.
8. The method according to claim 1, characterized in that, The first processing unit level is the same as the second processing unit level.
9. The method according to claim 1, wherein The first processing unit level is different from the second processing unit level.
10. The method according to claim 1, wherein The video bitstream further includes a second syntax element that indicates whether the first processing unit is the same as the second processing unit.
11. The method according to claim 1, wherein The transform process includes an identity transform in a first direction and a Hadamard transform in a second direction.
12. A computing system, characterized in that, Includes: A control circuit; A memory; And At least one set of instructions stored in the memory for execution by the control circuit, the at least one set of instructions including instructions for: Receiving video data from a video bitstream, the video data including at least two blocks and syntax elements, the at least two blocks including a first block, wherein the syntax elements are signaled at a first processing unit level; Selecting a transform coding mode based on the syntax elements; and Performing a transform process on the first block using the selected transform coding mode, wherein the transform process is performed on the transformed block at a second processing unit level, wherein the second processing unit level is obtained by inference.
13. The computing system according to claim 12, wherein The transform process includes performing an inverse transform on at least two residual coefficients of the first block to generate an optimized residual block corresponding to the first block.
14. The computing system according to claim 13, wherein Using a first transform kernel for the residual block and a second transform kernel for the optimized residual block, the second transform kernel being different from the first transform kernel.
15. The computing system according to claim 13, wherein The at least one set of instructions further includes instructions for generating a residual block according to the optimized residual block by performing short-range intra prediction on the optimized residual block, wherein the residual block is used to reconstruct the first block.
16. The computing system according to claim 15, wherein The short-range intra prediction includes line-by-line prediction, in which neighboring previous lines or columns are used to predict residuals in a specific line or column.
17. A non-volatile computer-readable storage medium, characterized in that, The storage medium stores at least one set of instructions for execution by a computing device, the computing device including a control circuit and a memory, the at least one set of instructions including instructions for: Receiving video data from a video bitstream, the video data including at least two blocks and syntax elements, the at least two blocks including a first block, wherein the syntax elements are signaled at a first processing unit level; Selecting a transform coding mode based on the syntax elements; and Performing a transform process on the first block using the selected transform coding mode, wherein the transform process is performed on the transformed block at a second processing unit level, and the second processing unit level is obtained by inference.
18. The non-volatile computer-readable storage medium according to claim 17, wherein The transform process includes inverse-transforming at least two residual coefficients of the first block to generate an optimized residual block corresponding to the first block.
19. The non-volatile computer-readable storage medium according to claim 18, wherein A first transform kernel is used for the residual block, and a second transform kernel different from the first transform kernel is used for the optimized residual block.
20. The non-volatile computer-readable storage medium according to claim 18, wherein The at least one set of instructions further includes instructions for generating a residual block based on the optimized residual block by using short-range intra prediction on the optimized residual block, wherein the residual block is used to reconstruct the first block.