Video decoding methods, computing systems and storage media

By employing an angle prediction mode in video coding and utilizing a weighted sum of top and left reference samples, the problem of insufficient prediction accuracy caused by uneven distribution of reference samples is solved, achieving more efficient video data compression and decoding.

CN119325706BActive Publication Date: 2026-03-10TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing video coding techniques fail to effectively utilize the uneven distribution of reference samples, resulting in insufficient prediction accuracy.

Method used

An angle prediction mode is adopted, which derives the final prediction value by using a weighted sum of top and left reference samples, thereby improving prediction accuracy.

Benefits of technology

It improves the prediction accuracy of video encoding and decoding, reduces the amount of data, and saves transmission bandwidth and storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119325706B_ABST
    Figure CN119325706B_ABST
Patent Text Reader

Abstract

The various implementations described herein include methods and systems for encoding and decoding video. In one aspect, a video decoding method includes receiving video data from a video bitstream, the video data including a first block, wherein the first block is predicted using an angle prediction mode. The method further includes: identifying a reference sample set for a portion of the first block using predicted angles from the angle prediction mode; and deriving a first angle prediction value for this portion using at least a first subset of the reference sample set. The method further includes: deriving a second angle prediction value for this portion using a weighted sum of at least a second subset of the reference sample set and the first angle prediction value; and decoding this portion using the second angle prediction value.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 467,213, filed May 17, 2023, entitled "Angular Blend Mode in Intra Mode Coding," and is a continuation-into-the-file of U.S. Patent Application No. 18 / 474,532, filed September 26, 2023, entitled "Systems and Methods for Angular Intra Mode Coding," and claims priority to that U.S. Patent Application. Technical Field

[0003] The disclosed embodiments generally relate to video encoding and decoding, including but not limited to systems and methods for predicting angle patterns for video encoding / decoding. Background Technology

[0004] Digital video is supported by a variety of electronic devices such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, and video streaming devices. Electronic devices send and receive, or otherwise transmit, digital video data via communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video encoding can be used to compress video data according to one or more video encoding standards before transmission or storage.

[0005] Several video codec standards have been developed. For example, video coding standards include the Open Media Consortium Video 1 (AV1), Universal Video Coding (VVC), Joint Explore Test Model (JEM), High Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Experts Group (MPEG) coding. Video coding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that take advantage of the inherent redundancy in video data. Video coding aims to compress video data into a form using a lower bitrate while avoiding or minimizing video quality degradation.

[0006] HEVC, also known as H.265, is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was published by ITU-T and ISO / IEC in 2013 (1st edition), 2014 (2nd edition), 2015 (3rd edition), and 2016 (4th edition). Versatile Video Coding (VVC), also known as H.266, is a video compression standard designed to succeed HEVC. The VVC / H.266 standard was published by ITU-T and ISO / IEC in 2020 (1st edition) and 2022 (2nd edition). AV1 is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the 1.0.0 version of the specification was released with Errata 1. SUMMARY

[0007] As described in more detail below, the angular prediction mode in some systems does not take into account the uneven distribution of available reference samples (e.g., only top and left reference samples can be available). The systems and methods described herein can improve prediction accuracy by biasing the prediction process that uses the top and / or left reference samples. For example, to predict a sample in the current block using a spatially neighboring reference sample, an angular prediction value is first derived, and then a final prediction value is derived using a weighted sum of the left reference sample and the angular prediction value. In this example, the final prediction value can then be used to encode / decode the sample in the current block.

[0008] According to some embodiments, a video decoding method is provided. The method includes: (i) receiving video data from a video bitstream, the video data including a first block, wherein the first block is predicted in an angular prediction mode; (ii) identifying a reference sample set for a portion of the first block (e.g., a current sample / pixel of the first block), wherein the reference sample set is identified using a prediction angle of the angular prediction mode; (iii) deriving a first angular prediction value for the portion of the first block using at least a first subset of the reference sample set; (iv) deriving a second angular prediction value for the portion of the first block using a weighted sum of at least a second subset of the reference sample set and the first angular prediction value; and (v) decoding the portion of the first block using the second angular prediction value.

[0009] According to some embodiments, a method of video encoding is provided. The method includes: (i) receiving video data, the video data including a first block, wherein the first block is to be encoded in an angular mode; (ii) identifying a reference sample set for a portion of the first block (e.g., a current sample / pixel of the first block), wherein the reference sample set is identified using a prediction angle of the angular prediction mode; (iii) deriving a first angular prediction value for the portion of the first block using at least a first subset of the reference sample set; (iv) deriving a second angular prediction value for the portion of the first block using a weighted sum of at least a second subset of the reference sample set and the first angular prediction value; and (v) encoding the portion of the first block using the second angular prediction value.

[0010] According to some embodiments, a computing system, such as a streaming system, a server system, a personal computer system, or other electronic device, is provided. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder component).

[0011] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets that are executed by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.

[0012] Thus, devices and systems having methods for encoding and decoding video are disclosed. Such methods, devices, and systems can supplement or replace conventional methods, devices, and systems for video encoding / decoding.

[0013] The features and advantages described in the specification are not all-inclusive, and particularly, many additional features and advantages will be apparent to one of ordinary skill in the art in view of the drawings, specification, and claims provided herein. Moreover, it should be noted that the language used in the specification has been principally selected for readability and instructional purposes and can not have been selected to delineate or circumscribe the subject matter described herein. BRIEF DESCRIPTION OF DRAWINGS

[0014] So that the disclosure can be more readily understood, a more particular description of the features of the various embodiments will be rendered by reference to the appended drawings, of which some embodiments are shown. However, it is to be understood that the figures are not necessarily to scale, and that the drawings are merely intended to conceptually illustrate the features of the disclosure. It should be noted that the figures are not necessarily drawn to scale.

[0015] FIG. 1 is a block diagram illustrating an exemplary communication system, in accordance with some embodiments.

[0016] FIG. 2A is a block diagram illustrating exemplary elements of an encoder component, in accordance with some embodiments.

[0017] FIG. 2B is a block diagram illustrating exemplary elements of a decoder component, in accordance with some embodiments.

[0018] FIG. 3 is a block diagram illustrating an exemplary server system, in accordance with some embodiments.

[0019] FIG. 4A through FIG. 4D illustrates an exemplary coding tree structure, in accordance with some embodiments.

[0020] FIG. 5A illustrates examples of directional intra prediction mode angles, in accordance with some embodiments.

[0021] FIG. 5B and FIG. 5C illustrates exemplary reference samples for samples in a current block, in accordance with some embodiments.

[0022] FIG. 5D and FIG. 5E illustrates exemplary reference samples for samples in a current block, in accordance with some embodiments.

[0023] FIG. 6A is a flowchart illustrating an exemplary method of encoding video, in accordance with some embodiments.

[0024] FIG. 6B is a flowchart illustrating an exemplary method of decoding video, in accordance with some embodiments.

[0025] In accordance with common practice the various features illustrated in the drawings can not be drawn to scale, in which case, the dimensions of the various features can be exaggerated or reduced relative to each other for clarity. DETAILED DESCRIPTION

[0026] The present disclosure describes, among other things, biasing angular mode prediction using top and / or left reference samples. For example, a refined angular prediction value for a portion (samples) of a first block can be derived using a weighted sum of (top / left) reference samples and an unrefined angular prediction value. Biasing angular mode prediction in this way can improve prediction accuracy, as only top and left reference samples are available (e.g., other reference samples can be padded).

[0027] Exemplary Systems and Devices

[0028] FIG. 1This is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, used with video-enabled applications such as video conferencing applications, digital television applications, media storage and / or distribution applications.

[0029] Source device 102 includes a video source 104 (e.g., a camera component or media storage) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video streams from the video stream. The video stream from video source 104 can have a high data volume compared to the encoded video stream 108 generated by encoder component 106. Because the encoded video stream 108 has a lower data volume (less data) compared to the video stream from video source 104, it requires less bandwidth for transmission and less storage space for storage compared to the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., configured to transmit uncompressed video to network 110).

[0030] One or more networks 110 represent any number of networks that transmit information between source device 102, server system 112, and / or electronic device 120, including, for example, wired (connected) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet.

[0031] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from source device 102). The server system 112 includes an encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the encoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the encoder component 114 is instantiated as hardware, software, or a combination of hardware and software. In some embodiments, the encoder component 114 is configured to decode the encoded video stream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video stream 108.

[0032] In some embodiments, server system 112 functions as a media-aware network element (MANE). For example, server system 112 may be configured to trim encoded video streams 108 to customize potentially different streams for one or more electronic devices 120. In some embodiments, the MANE is provided separately from server system 112.

[0033] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode encoded video data 116 to generate an output video stream that can be displayed on a display or other type of presentation device. In some embodiments, one or more electronic devices 120 do not include a display component (e.g., communicatively coupled to an external display device and / or includes media storage). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.

[0034] The source device and / or multiple electronic devices 120 are sometimes referred to as “terminal devices” or “user devices”. In some embodiments, the source device 102 and / or one or more electronic devices 120 are instances of server systems, personal computers, portable devices (e.g., smartphones, tablets, or laptops), wearable devices, video conferencing equipment, and / or other types of electronic devices.

[0035] In an exemplary operation of the communication system 100, the source device 102 transmits an encoded video stream 108 to the server system 112. For example, the source device 102 may encode a stream of images captured by the source device. The server system 112 receives the encoded video stream 108 and may decode and / or encode the encoded video stream 108 using an encoder component 114. For example, the server system 112 may apply encoding to video data, which is more optimized for network transmission and / or storage. The server system 112 may send encoded video data 116 (e.g., one or more encoded video streams) to one or more electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display video images.

[0036] FIG. 2AThis is a block diagram illustrating exemplary elements of an encoder assembly 106 according to some embodiments. The encoder assembly 106 receives a source video sequence from a video source 104. In some embodiments, the encoder assembly includes a receiver (e.g., transceiver) assembly configured to receive the source video sequence. In some embodiments, the encoder assembly 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder assembly 106). The video source 104 may provide a source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601YCrCB, or RGB), and any suitable sampling structure (e.g., YCrCb4:2:0, or YCrCb4:4:4). In some embodiments, the video source 104 is a storage device storing previously acquired / prepared video. In some embodiments, the video source 104 is a camera that acquires local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed sequentially. An image itself can be organized into a spatial pixel array, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. Those skilled in the art will readily understand the relationship between pixels and samples. The following focuses on describing samples.

[0037] Encoder component 106 is configured to encode and / or compress images of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. Implementing an appropriate encoding rate is a function of controller 204. In some embodiments, controller 204 controls and is functionally coupled to other functional units described below. Parameters set by controller 204 may include rate control-related parameters (e.g., image skipping, quantizer, and / or the λ value of rate-distortion optimization techniques), image size, group of images (GOP) layout, maximum motion vector search range, etc. Other functions of controller 204 will be readily recognizable to those skilled in the art, as these functions may relate to encoder component 106 optimized for a particular system design.

[0038] In some embodiments, encoder component 106 is configured to operate within an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder 210. Decoder 210 reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (when compression between the symbols and the encoded video stream is lossless). The reconstructed sample stream (sample data) is input to reference image memory 208. Since decoding of the symbol stream produces bit-accurate results independent of the decoder's location (local or remote), the contents of reference image memory 208 are also bit-accurately corresponding between the local encoder and the remote encoder. Thus, the encoder's prediction portion interprets the reference image samples as the same sample values ​​that the decoder will interpret during prediction. This principle of reference image synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is known to those skilled in the art.

[0039] The operation of decoder 210 can be combined with, for example, the following: FIG. 2B The decoder component 122 described in detail is the same as the remote decoder. However, a brief reference is provided. FIG. 2B Since symbols are available and the entropy encoder 214 and parser 254 can encode / decode symbols into an encoded video sequence in a lossless manner, the entropy decoding portion of the decoder component 122, which includes buffer memory 252 and parser 254, may not be fully implemented in the local decoder 210.

[0040] The decoder techniques described in this paper (aside from parsing / entropy decoding) can exist in the corresponding encoders with essentially the same functional form. Therefore, the subject matter presented here focuses on decoder operations. The description of the encoder techniques can be simplified, as the encoder and decoder techniques are inverses of each other.

[0041] As part of its operation, the source encoder 202 can perform motion-compensated predictive coding, which predictively encodes the input frame by referencing one or more previously encoded frames from the video sequence designated as reference frames. In this way, the encoding engine 212 encodes the differences between pixel blocks of the input frame and pixel blocks of the reference frame, which can be selected as the prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding the video data.

[0042] Decoder 210 decodes encoded video data of frames that can be designated as reference frames, based on symbols created by source encoder 202. The operation of encoding engine 212 can advantageously be a lossy process. When encoded video data is processed by video decoder (…FIG. 2A When decoded at (not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. Decoder 210 replicates the decoding process, which can be performed on the reference frame by a remote video decoder, and allows the reconstructed reference frame to be stored in reference image memory 208. In this way, encoder component 106 locally stores a copy of the reconstructed reference frame that shares common content (no transmission errors) with the reconstructed reference frame that will be obtained by the remote video decoder.

[0043] Predictor 206 can perform a prediction search against encoding engine 212. For a new frame to be encoded, predictor 206 can search the reference image memory 208 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new image. Predictor 206 can operate pixel-by-pixel based on sample blocks to find suitable prediction references. In some cases, as determined by the search results obtained by predictor 206, the input image may have prediction references obtained from multiple reference images stored in reference image memory 208.

[0044] The outputs of all the aforementioned functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by the various functional units according to techniques known to those skilled in the art (e.g., Huffman coding, variable-length coding, and / or arithmetic coding), thereby converting the symbols into an encoded video sequence.

[0045] In some embodiments, the output of entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the encoded video sequence created by entropy encoder 214 in preparation for transmission via communication channel 218, which may be a hardware / software link to a storage device capable of storing the encoded video data. The transmitter may be configured to combine encoded video data from source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter may transmit additional data while transmitting the encoded video. Source encoder 202 may include such data as part of the encoded video sequence. Additional data may include temporal / spatial / signal-to-noise ratio (SNR) enhancement layers, other forms of redundant data such as redundant pictures and slices, auxiliary enhancement information (SEI) messages, video availability information (VUI) parameter set fragments, etc.

[0046] Controller 204 manages the operation of encoder component 106. During encoding, controller 204 can assign a specific type of encoded picture to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, a picture can be assigned as an intra-picture (I-picture), a predictive picture (P-picture), or a bidirectional predictive picture (B-picture). Intra-pictures can be encoded and decoded without using any other frames in the sequence as prediction sources. Some video codecs allow different types of intra-pictures, including, for example, Independent Decoder Refresh (IDR) pictures. These variations of I-pictures and their corresponding applications and characteristics are familiar to those skilled in the art and will not be described in detail here. Predictive pictures can be encoded and decoded using intra-prediction or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​for each block. Bidirectional predictive pictures can be encoded and decoded using intra-prediction or inter-prediction, which uses at most two motion vectors and reference indices to predict the sample values ​​for each block. Similarly, multiple predictive images can be used to reconstruct a single block using more than two reference images and associated metadata.

[0047] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and each block is encoded sequentially. These blocks can be predictively coded with reference to other (already coded) blocks, which are determined by the coding assignments of the corresponding images applied to the blocks. For example, blocks of an I-image can be nonpredictively coded, or blocks of an I-image can be predictively coded (spatial prediction or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be nonpredictively coded with reference to a previously coded reference image via spatial prediction or temporal prediction. Blocks of a B-image can be nonpredictively coded with reference to one or two previously coded reference images via spatial prediction or temporal prediction.

[0048] Video can be captured as a series of source images (video images) over time. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In one example, a specific image being encoded / decoded is divided into blocks; this specific image is called the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference image.

[0049] Encoder component 106 can perform encoding operations according to any predetermined video coding technique or standard such as those described herein. In operation, encoder component 106 can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0050] FIG. 2B This is a block diagram illustrating exemplary elements of a decoder component 122 according to some embodiments. FIG. 2B The decoder component 122 is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter 256 and configured to (e.g., via a wired or wireless connection) transmit data to the display 124.

[0051] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured (e.g., via a wired or wireless connection) to receive data from channel 218. The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, decoding of each encoded video sequence is independent of decoding of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link leading to a storage device storing the encoded video data. The receiver may receive encoded video data and other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective user entities (not depicted). The receiver may separate the encoded video sequences from other data. In some embodiments, the receiver receives additional (redundant) data when receiving encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0052] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra-frame image prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference image memory 266, and a current image memory 264. In some embodiments, the decoder component 122 is implemented as one or more integrated circuits and / or other electronic circuits. In some embodiments, the decoder component 122 is at least partially implemented in software.

[0053] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to prevent network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is located between the output of channel 218 and decoder component 122. In some embodiments, a separate buffer memory is located outside decoder component 122 (e.g., to prevent network jitter), in addition to buffer memory 252 located inside decoder component 122 (e.g., configured to handle playback timing). Buffer memory 252 may not be needed or may be small when receiving data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network. Buffer memory 252 may be required for use on packet-switched networks such as the Internet. Buffer memory 252 may be relatively large, advantageously having an adaptive size, and may be implemented at least partially in an operating system or a similar element (not depicted) outside decoder component 122.

[0054] Parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols may include, for example, information for managing the operation of decoder component 122, and / or information for controlling a presentation device such as display 124. Control information for the presentation device may be in the form of, for example, Supplementary Enhancement Information (SEI) messages or fragments of Video Usability Information (VUI) parameter sets (not depicted). Parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 254 may extract a subset of parameters from the encoded video sequence for at least one subset of pixels in a subgroup for use in the video decoder, based on at least one parameter corresponding to a group. Subgroups may include group of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The parser 254 can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0055] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of symbol 270 may involve multiple different units. Which units are involved and how they are involved can be controlled by parser 254 through subgroup control information parsed from the encoded video sequence. For clarity, the flow of such subgroup control information between parser 254 and the various units described below is not depicted.

[0056] The decoder component 122 can be conceptually subdivided into multiple functional units, which in some implementations interact closely with each other and can be at least partially integrated with each other. However, for clarity, it will remain conceptually subdivided into multiple functional units hereinafter.

[0057] The scaler / inverse transform unit 258 receives from the parser 254 the quantization transform coefficients as symbols 270, as well as control information (e.g., which transform to use, block size, quantization factor, and / or quantization scaling matrix). The scaler / inverse transform unit 258 can output a block containing sample values, which can be input into the aggregator 268.

[0058] In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed images, but can use prediction information from previously reconstructed portions of the current image. Such prediction information can be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 can use surrounding reconstructed information extracted from the current (partially reconstructed) image from the current image memory 264 to generate blocks of the same size and shape as the block being reconstructed. The aggregator 268 can add the prediction information generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 based on each sample.

[0059] In other cases, the output samples of the scaler / inverse transform unit 258 belong to blocks of inter-frame coding and potential motion compensation. In this case, the motion compensation prediction unit 260 can access the reference image memory 266 to extract samples for prediction. After motion compensation is performed on the extracted samples according to the symbols 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (in this case, referred to as residual samples or residual signals) to generate output sample information. The extraction of prediction samples from the address in the reference image memory 266 by the motion compensation prediction unit 260 can be controlled by motion vectors. These motion vectors can be provided to the motion compensation prediction unit 260 in the form of symbols 270, which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values ​​extracted from the reference image memory 266 when using subsample precise motion vectors, motion vector prediction mechanisms, etc.

[0060] The output samples of aggregator 268 can be subjected to various loop filtering techniques in loop filter unit 256. The video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video bitstream and available to loop filter unit 256 as symbols 270 from parser 254, and may also respond to metadata obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0061] The output of the loop filter unit 256 can be a sample stream, which can be output to a presentation device (e.g., display 124) and stored in a reference image memory 266 for future inter-frame image prediction.

[0062] Once reconstructed, certain encoded images can be used as reference images for future predictions. Once the encoded images have been reconstructed and the encoded images (via, for example, parser 254) are identified as reference images, the current reference images can become part of the reference image memory 266, and new current image memory can be reallocated before reconstructing subsequent encoded images begins.

[0063] Decoder component 122 can perform decoding operations according to a predetermined video compression technique that can be recorded in a standard (e.g., any of the standards described herein). In the sense that the encoded video sequence follows the syntax of a video compression technique or standard, the encoded video sequence may conform to the syntax specified by the video compression technique or standard used (as specified in the video compression technique documentation or standard, particularly in the configuration file of the video compression technique or standard). Furthermore, in order to conform to some video compression techniques or standards, the complexity of the encoded video sequence may be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, megasamples per second), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the hypothetical reference decoder (HRD) specification and metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.

[0064] FIG. 3 This is a block diagram illustrating a server system 112 according to some embodiments. The server system 112 includes control circuitry 302, one or more network interfaces 304, memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuitry includes one or more field-programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).

[0065] Network interface 304 can be configured to connect to one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). Communication networks can be local area networks, wide area networks, metropolitan area networks, vehicle and industrial networks, real-time networks, latency-tolerant networks, etc. Examples of communication networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital television networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., CANBus connected to certain CANBus devices), or bidirectional (e.g., connecting to other computer systems using a local area network or wide area network). Such communication may include communication to one or more cloud computing networks.

[0066] User interface 306 includes one or more output devices 308 and / or one or more input devices 310. Input devices 310 may include one or more of a keyboard, mouse, touchpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. Output devices 308 may include one or more of an audio output device (e.g., a speaker), a visual output device (e.g., a monitor or display screen), etc.

[0067] Memory 314 may include high-speed random access memory (e.g., DRAM, SRAM, DDR RAM, and / or other solid-state random access memory devices) and / or non-volatile memory (e.g., one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Optionally, memory 314 includes one or more storage devices disposed remotely from control circuitry 302. Memory 314, or alternatively, the non-volatile solid-state memory devices within memory 314 include non-transitory computer-readable storage media. In some embodiments, memory 314 or the non-transitory computer-readable storage media of memory 314 stores programs, modules, instructions, and data structures, or subsets or supersets thereof:

[0068] • Operating system 316, which includes programs for handling various basic system services and for performing hardware-related tasks;

[0069] • Network communication module 318, which is used to connect server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connection);

[0070] • Encoding module 320, which performs various functions related to encoding and / or decoding data (e.g., video data). In some embodiments, encoding module 320 is an instance of encoder component 114. Encoding module 320 includes, but is not limited to, one or more of the following:

[0071] Decoding module 322, which performs various functions related to decoding encoded data, such as those previously described for decoder component 122; and

[0072] The encoding module 340 performs various functions related to encoding data, such as those previously described for encoder component 106; and

[0073] • Image memory 352 is used to store images and image data, for example, for use with encoding module 320. In some embodiments, image memory 352 includes one or more of reference image memory 208, buffer memory 252, current image memory 264, and reference image memory 266.

[0074] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described for the parser 254), a transformation module 326 (e.g., configured to perform various functions previously described for the scaler / inverse transformation unit 258), a prediction module 328 (e.g., configured to perform various functions previously described for the motion compensation prediction unit 260 and / or the intra-frame image prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described for the loop filter 256).

[0075] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions previously described for the source encoder 202 and / or encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described for the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes... FIG. 3 A subset of the modules shown. For example, both decoding module 322 and encoding module 340 use a shared prediction module.

[0076] Each of the modules identified above and stored in memory 314 corresponds to a set of instructions for performing the functions described herein. The modules identified above (e.g., instruction sets) do not need to be implemented as separate software programs, processes, or modules; therefore, subsets of these modules can be combined or otherwise rearranged in various embodiments. For example, optionally, encoding module 320 does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, memory 314 stores a subset of the modules and data structures identified above. In some embodiments, memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0077] Although FIG. 3 A server system 112 according to some embodiments is shown, but FIG. 3 This is intended more as a functional description of various features that may exist in one or more server systems, rather than a structural schematic diagram of the embodiments described herein. In practice, as those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, FIG. 3 Some items shown individually can be implemented on a single server, and a single item can be implemented by one or more servers. The actual number of servers used to implement server system 112, and how features are distributed among the servers, will vary depending on the implementation method, and optionally, in part, on the amount of data traffic processed by the server system during peak usage periods and during average usage periods.

[0078] Exemplary Encoding Processes and Techniques

[0079] The encoding processes and techniques described below can be performed at the devices and systems described above (e.g., source device 102, server system 112, and / or electronic device 120). FIG. 4A through FIG. 4D An exemplary encoding tree structure according to some embodiments is shown. FIG. 4A As shown in the first coding tree structure (400), some coding methods (e.g., VP9) use a 4-way partition tree that decreases from a 64×64 level to a 4×4 level, which imposes some additional constraints on 8×8 blocks. FIG. 4A In this context, a partition designated as R can be called a recursive partition because the same partition tree is repeated in smaller increments until the minimum 4×4 level is reached.

[0080] like FIG. 4B As shown in the second coding tree structure (402), some coding methods (e.g., AV1) expand the partition tree into a 10-way structure and increase the maximum size (e.g., called a superblock in VP9 / AV1 terms) to start from 128×128. The second coding tree structure includes 4:1 / 1:4 rectangular partitions that are not present in the first coding tree structure. FIG. 4B The partition type with 3 sub-partitions in the second row is called a T-partition. In addition to the code block size, the code tree depth can also be defined to indicate the partition depth from the root node.

[0081] As an example, in HEVC, for instance, a CTU can be divided into CUs using a quadtree structure represented as a coding tree to accommodate various local characteristics. In some embodiments, the decision on whether to use inter-frame picture (temporal) prediction or intra-frame picture (spatial) prediction to encode picture regions is made at the CU level. Depending on the PU partitioning type, each CU can be further divided into one, two, or four PUs. Within a PU, the same prediction process is applied, and relevant information is transmitted to the decoder based on the PU. After obtaining residual blocks by applying prediction processing based on the PU partitioning type, the CUs can be divided into TUs according to another quadtree structure (similar to the coding tree used for CUs).

[0082] For example, in VVC, binary and ternary partitioning structures, and quadtrees with nested multi-type trees, can replace the concept of multiple partition unit types. For instance, quadtrees eliminate the separation of CU, PU, ​​and TU concepts (except for CUs that are too large for the maximum transform length), and quadtrees support greater flexibility in CU partition shapes. In the coding tree structure, CUs can be square or rectangular. CTUs are first partitioned using a four-level tree (also called a quadtree) structure. Leaf nodes of the four-level tree can be further partitioned using multi-type tree structures. For example...FIG. 4C As shown in the third coding tree structure (404), the multi-type tree structure includes four partition types. The leaf nodes of the multi-type tree are called CUs, and unless the CU is too large for the maximum transform length, this segment is used for prediction and transform processing without requiring any further partitioning. In most cases, CUs, PUs, and TUs have the same block size in a quadtree with a nested multi-type tree coding block structure. An example of a block partition of a CTU (406) is shown below. FIG. 4D As shown, FIG. 4D An exemplary quadtree with a nested multi-type tree coding block structure is shown.

[0083] A fixed set of intra-frame prediction angles can be used to remove spatial redundancy from the video signal. For directional intra-frame prediction, some methods support eight directional modes, corresponding to angles from 45 degrees to 207 degrees. To utilize more types of spatial redundancy in directional textures, directional intra-frame modes can be extended to a more fine-grained set of angles. For example, the eight angles can be represented as nominal angles. FIG. 5A Eight nominal angles are shown, named V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED. Also... FIG. 5A As shown, for each nominal angle, there can be 7 finer angles, for a total of 56 orientation angles. The predicted angle can be described by adding an angle increment (e.g., a step of -3 to 3 multiplied by 3 degrees) to the nominal intra-frame angle. To implement orientation prediction modes in a general way, the 56 orientation intra-frame prediction modes can be implemented using a unified orientation predictor, which projects each pixel to a reference sub-pixel location and interpolates the reference pixel using a 2-tap bilinear filter. This type of orientation intra-frame prediction is sometimes also called unidirectional intra-frame prediction.

[0084] In some methods, a lookup table is used to map each intra-prediction angle to a horizontal and vertical offset between each pixel in the current block and a reference sample. For each intra-prediction angle, the offset in the lookup table can be an integer value that is the tangent of the angle multiplied by 64. For example, 64 is equal to the tangent of (45°) multiplied by 64, where the tangent of (45°) is equal to 1. For example, for an intra-prediction angle of 45 degrees, the associated offset is 64, and the horizontal offset between each pixel in the current block and the reference pixel increases by 1 pixel as the row number of the pixel increases by 1.

[0085] In some methods, there are also non-directional smooth intra-frame prediction modes, such as DC, PAETH, SMOOTH, SMOOTH-V, and SMOOTH-H. For DC prediction, the average of the left and top neighboring samples can be used as the predictor for the block to be predicted. For PAETH prediction, top, left, and top-left reference samples can be extracted, and the value closest to a given value (e.g., top + left - top-left) is set as the predictor for the pixel to be predicted. FIG. 5B This shows the positions of the top sample 514, left sample 518, and top-left sample 516 of the current pixel (or block or sub-block) in the current block. For SMOOTH mode, SMOOTH-V mode, and SMOOTH-H mode, the block can be predicted using quadratic interpolation in the vertical or horizontal direction, or the average of the two directions.

[0086] To capture attenuation spatial correlation using edge-based references, intra-frame filter modes can be used for luma blocks. In some methods, five intra-frame filter modes are defined, each represented by a set of eight 7-tap filters, reflecting the correlation between pixels in a 4×2 sub-block and their seven neighboring pixels. In this way, the weighting factor of the 7-tap filters depends on location. For example, as... FIG. 5C As shown, an 8×8 block can be divided into eight 4×2 smaller blocks. FIG. 5C In the diagram, small blocks are indicated by B0, B1, B2, B3, B4, B5, B6, and B7. For each small block, seven neighbors (indicated by R0 to R7) are used to predict the pixels within that block. For block B0, all neighbors have been reconstructed. However, for other small blocks, some neighbors may not need to be reconstructed. In this case, the predictions of direct neighbors can be used as a reference. For example, it is not necessary to reconstruct all neighbors of block B7; therefore, instead, the predicted samples of neighbors (e.g., B5 and B6) can be used.

[0087] The directional intra-prediction process may include the following as inputs: a variable plane specifying which plane is being predicted; variables x and y specifying the position of the top-left sample in the CurrFrame[plane] array of the current transform block; variables haveLeft (which is equal to 1 if there is a valid sample to the left of the transform block); variables haveAbove (which is equal to 1 if there is a valid sample above the transform block); a variable mode specifying the type of intra-prediction to be applied; a variable w specifying the width of the region to be predicted; a variable h specifying the height of the region to be predicted; a variable maxX specifying the maximum valid x-coordinate of the current plane; and a variable maxY specifying the maximum valid y-coordinate of the current plane.

[0088] The output of the directional intra-frame prediction process can be the angle prediction value P. A The angle prediction can be based on the predicted angle pAngle, the upper row AboveRow, or the left column LeftCol. For example, if the predicted angle is less than 180 degrees, the angle prediction can be derived from the AboveRow sample, and if the predicted angle is greater than 180 degrees, the angle prediction can be derived from the LeftCol sample.

[0089] In some systems, angular (orientation) prediction models do not account for the uneven distribution of available reference samples (e.g., only top and left reference samples may be available). Therefore, the prediction accuracy of these angular models can be reduced compared to models that more heavily weight the top and left reference samples. The systems and methods described below improve prediction accuracy and precision compared to angular prediction models. The methods and procedures described below can be used individually or in any combination. In the following description, a model is called an orientation model if it generates prediction samples based on a given prediction direction.

[0090] As used in this article, the left-hand reference sample refers to a reference sample whose vertical coordinate value lies within the minimum and maximum vertical coordinate values ​​of the current block (e.g., current block 522). FIG. 5D The left part of Figure 528 is shown. The top reference sample refers to the reference sample whose horizontal coordinate value is within the minimum and maximum horizontal coordinate values ​​of the current block, such as... FIG. 5D The top portion of 524 is shown. The lower left reference sample refers to a reference sample whose vertical coordinate value is greater than the maximum vertical coordinate value of the current block, such as... FIG. 5D The lower left portion (530) is shown in the image. The upper right reference sample refers to a reference sample whose horizontal coordinate value is greater than the maximum horizontal coordinate value of the current block, such as... FIG. 5D The upper right portion of the image is shown in 526. The upper left reference sample refers to a reference sample whose horizontal coordinate value is less than the minimum horizontal coordinate value of the current block and whose vertical coordinate value is less than the minimum vertical coordinate value of the current block, such as... FIG. 5D The upper left part is shown in 532.

[0091] FIG. 5E The prediction angle 538 for sample 536 of the current block 522 is also shown. Prediction angle 538 leads to the top sample 542 in the upper right portion 526 and the bottom sample 540 in the lower left portion 530. FIG. 6A The predicted angle 550 for sample 552 is shown. Predicted angle 550 leads to the top sample 556 and the bottom sample 554. In some embodiments, predicted angles 538 and 550 are based on an angle pattern (e.g., a specific angle pattern indicates the corresponding predicted angle).

[0092] FIG. 6BThis is a flowchart illustrating a method 600 for encoding video according to some embodiments. Method 600 may be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions executed by the control circuitry. In some embodiments, method 600 is executed by executing instructions stored in the memory of the computing system (e.g., memory 314).

[0093] The system receives (602) video data comprising multiple blocks, including a first block, wherein the first block is encoded in an angle pattern (e.g., a directional intra-frame prediction pattern). The system identifies (604) a reference sample set (e.g., top sample 542 and / or bottom sample 540) for a portion of the first block, wherein the reference sample set is identified using the predicted angle of the angle prediction pattern. The system derives (606) a first angle prediction value for this portion of the first block using at least a first subset of the reference sample set. The system derives (608) a second angle prediction value for this portion of the first block using a weighted sum of at least a second subset of the reference sample set and the first angle prediction value. The system encodes (610) this portion of the first block using the second angle prediction value. For example, the system reconstructs the first block based on the second angle prediction value, evaluates the angle prediction pattern based on the reconstructed first block, and selects an angle prediction pattern (e.g., an angle prediction pattern with minimum error) for encoding the first block based on the evaluation.

[0094] FIG. 6A This is a flowchart illustrating a method 650 for decoding video according to some embodiments. Method 650 may be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions executed by the control circuitry. In some embodiments, method 650 is executed by executing instructions stored in the memory of the computing system (e.g., memory 314).

[0095] The system receives (652) video data from a video stream, the video data comprising multiple blocks, including a first block, wherein the first block is predicted using an angle prediction mode (e.g., a directional intra-frame prediction mode). The system identifies (654) a reference sample set (e.g., top sample 542 and / or bottom sample 540) for a portion of the first block, wherein the reference sample set is identified using the predicted angle of the angle prediction mode. The system derives (656) a first angle prediction value for this portion of the first block using at least a first subset of the reference sample set. The system derives (658) a second angle prediction value for this portion of the first block using a weighted sum of at least a second subset of the reference sample set and the first angle prediction value. The system decodes (660) this portion of the first block using the second angle prediction value.

[0096] Although FIG. 6B and FIG. 5D Multiple logical stages are shown in a specific order, but stages independent of the order can be reordered, and other stages can be combined or decomposed. Some reorderings or other groupings not specifically mentioned will be obvious to those skilled in the art, and therefore the orderings and groupings presented herein are not exhaustive. Furthermore, it should be recognized that the stages can be implemented in hardware, firmware, software, or any combination thereof.

[0097] In some embodiments, in order to predict samples in the current block using spatially adjacent reference samples, the angle prediction value P is first derived. A Then use the left-side reference sample and P A The final predicted value P' is derived from the weighted sum. A (For example, a horizontal blending mode corresponding to an angle pattern). In some embodiments, an angle prediction value P is derived along a given prediction angle using reference samples from the top adjacent position and / or the upper right adjacent position of the current block. A In some embodiments, P can be generated. A The reference sample is filtered before the value is generated. For example, in generating P... A Before the values ​​are calculated, one or more smoothing filters (e.g., Gaussian filters or bilateral filters) are used to filter the left reference sample, top reference sample, upper left reference sample, lower left reference sample, and upper right reference sample, where the coefficients of the smoothing filters are all non-negative integers. For example, the disclosed method can identify a set of reference samples for a first block (e.g., a portion of the first block), wherein the set of reference samples is identified using the predicted angles of the angle prediction pattern. The disclosed method can then derive a first angle prediction value for this portion of the first block using at least a first subset of the reference sample set. Subsequently, the disclosed method can derive a second angle prediction value for this portion of the first block using (i) at least a second subset of the reference sample set and (ii) a weighted sum of the derived first angle prediction values.

[0098] In some embodiments, P' is derived using the following Equation 1. A .

[0099] P' A =(W L ·L+(NW L )·P A |+r) / N

[0100] Equation 1 - Refined Angle Prediction Value

[0101] Here, w is derived using the horizontal coordinate value of the current sample. LN is a predetermined value (e.g., an integer that is a power of 2, such as 2, 4, 8, 16, 32, 64, 128, 256, 512, or 1024), and r is the rounding offset (e.g., equal to 0.5·N). L is the left / bottom-left sample and can be obtained by extending the direction from the top reference to the left / bottom-left sample of the predicted value, as shown below. FIG. 5A As shown.

[0102] In some embodiments, w L This is derived as K >> ((x << 1)) >> s, where K is a predetermined value (e.g., 16, 32, or 64), x is the horizontal coordinate of the current sample to be predicted, and s is a scaling factor based on the block size. For example, s can be equal to (log2(W) + log2(H) + 2) >> 2, where W and H are the width and height of the block, respectively.

[0103] In some embodiments, if L is outside the lower left reference sample, L is refined using predetermined values ​​on the left or lower left. In some embodiments, when L is outside the lower left reference sample, predetermined samples located on the left or lower left are used. In some embodiments, L is filled with the nearest available sample located in the lower left region, left region, upper left region, top region, and / or upper right region.

[0104] In some embodiments, for refining the (final) predictor, P is calculated. A The division operation for predicting samples is moved to the last step. For example, during internal computation, parameters can be multiplied by a value so that these parameters have the same divisor.

[0105] In some embodiments, when the angle mode is located FIG. 5D When the system finds the left sample by extending a given angular direction to the right of V_PRED, and uses the refined predictor P' from Equation 1. A (For example, apply a horizontal blending mode).

[0106] In some embodiments, in order to predict samples in the current block using spatially adjacent reference samples, the angle prediction value P is first derived. A Then use the top reference sample and P A The final predicted value P' is derived from the weighted sum. A (For example, the vertical blending mode corresponding to the angle pattern). In some embodiments, the angle prediction value P is derived along a given prediction angle using reference samples from the left-hand adjacent position and / or the lower-left adjacent position of the current block. A In some embodiments, P can be generated. A The reference sample is filtered before the value is generated. For example, in generating P... ABefore the values ​​are calculated, one or more smoothing filters (e.g., Gaussian filters or bilateral filters) are used to filter the left reference sample, top reference sample, top left reference sample, bottom left reference sample, and top right reference sample, where the coefficients of the smoothing filters are all non-negative integers.

[0107] In some embodiments, P' is derived using the following Equation 2. A .

[0108] P' A =(W T ·T+(NW T )·P A +r) / N

[0109] Equation 2 - Refined Angle Prediction Value

[0110] Specifically, w is derived using the vertical coordinate values ​​of the current sample. T N is a predetermined value (e.g., an integer that is a power of 2, such as 2, 4, 8, 16, 32, 64, 128, 256, 512, or 1024), and r is the rounding offset (e.g., equal to 0.5·N). T is the top / top-right sample and can be obtained by extending the direction from the top reference to the left / bottom-left sample of the predicted value, as shown below. FIG. 5A As shown.

[0111] In some embodiments, w T This is derived as K >> ((y << 1)) >> s, where K is a predetermined value (e.g., 16, 32, or 64), y is the vertical coordinate of the current sample to be predicted, and s is a scaling factor based on the block size. For example, s can be equal to (log2(W) + log2(H) + 2) >> 2, where W and H are the width and height of the block, respectively.

[0112] In some embodiments, if T is outside the top / top right reference sample, a predetermined sample located in the top / top right region is assigned to T. In some embodiments, T is populated with the nearest available sample located in the top right region, top region, top left region, left side region, and / or bottom left region.

[0113] In some embodiments, for refining the (final) predictor, P is calculated. A The division operation for predicting samples is moved to the last step. For example, during internal computation, parameters can be multiplied by a value so that these parameters have the same divisor.

[0114] In some embodiments, when the angle mode is located FIG. 5E When H_PRED is below, use the refined predictor P' from Equation 2. A (For example, apply a vertical blending mode).

[0115] In some embodiments, in order to predict samples in the current block using spatially adjacent reference samples, the top (or left) reference sample is first used (e.g., FIG. 5E The top sample (556) and the bottom (or right) reference sample (e.g., FIG. 5E The angle prediction value P is derived by interpolating the bottom sample (554) in the data. A Then use the top (or left) reference sample (e.g., FIG. 5D The top sample 556) and P A The weighted sum is used to derive the refined (final) predicted value P'. A (For example, the angle blending mode corresponding to the angle mode).

[0116] In some embodiments, reference samples from the top, left, bottom, and / or right reference samples of the current block are used to derive the angle prediction value P along a given prediction angle. A In some embodiments, the top (and / or left) reference sample (e.g., top sample 556) is derived using interpolations of multiple top (and / or left) reference samples located at integer positions as input. In some embodiments, the right (or bottom) reference sample (e.g., bottom sample 554) is derived using reference samples from the top or left reference. For example, the derivation of the right (or bottom) reference sample (e.g., bottom sample 554) can be the same as in a smooth mode, a blend mode, or a planar mode. In some embodiments, P can be generated... A The reference sample is filtered before the value is generated. For example, in generating P... A Before the values ​​are calculated, one or more smoothing filters (e.g., Gaussian filters or bilateral filters) are used to filter the left reference sample, top reference sample, top left reference sample, bottom left reference sample, and top right reference sample, where the coefficients of the smoothing filters are all non-negative integers.

[0117] In some embodiments, when the top reference sample is used to generate P A Then, use the following equation 3 to derive P' A For example, P can be generated through bilinear interpolation of T and B. A .

[0118] P' A =(w·T+((NW)·P) A +r) / N

[0119] Equation 3 - Detailed Angle Prediction Value

[0120] Here, w is derived using the vertical coordinate value of the current sample, N is a predetermined value (e.g., an integer that is a power of 2, such as 2, 4, 8, 16, 32, 64, 128, 256, 512, or 1024), and r is the rounding offset (e.g., equal to 0.5·N). T is the top / top right sample and can be obtained by extending from the left sample along a predetermined angle to the top or bottom right, as shown below. ​ As shown.

[0121] In some embodiments, w is derived as K >> ((y << 1)) >> s, where K is a predetermined value (e.g., 16, 32, or 64), y is the vertical coordinate of the current sample to be predicted, and s is a scaling factor based on the block size. For example, s can be equal to (log2(W) + log2(H) + 2) >> 2, where W and H are the width and height of the block, respectively.

[0122] In some embodiments, if T is outside the top / top right reference sample, a predetermined sample located in the top / top right region is assigned to T. In some embodiments, T is populated with the nearest available sample located in the top right region, top region, top left region, left side region, and / or bottom left region.

[0123] In some embodiments, when the left-side reference sample is used to generate P A (For example, when replacing T with a reference sample on the left), use Equation 3 to derive P'. A .

[0124] In some embodiments, for refining the (final) predictor, P is calculated. A The division operation for predicting samples is moved to the last step. For example, during internal computation, parameters can be multiplied by a value so that these parameters have the same divisor.

[0125] In some embodiments, the angle pattern starts along the angular direction from the top sample or the left sample, and when blended to obtain P' A When using the bottom sample B, for example, derive P' using Equation 4 below. A .

[0126] P' A =(w·B+((Nw)·P) A +r) / N

[0127] Equation 4 - Refined Angle Prediction Value

[0128] In some embodiments, whether to use a mixed mode (e.g., whether to export P') A Using P' AThe predicted samples are represented by signals at the sequence header, slice header, superblock, or coded block level. In some embodiments, if the block is in intra-frame mode, a mixed mode is used (e.g., deriving P'). A Using P' A (For prediction), and if the block is in intra-frame / inter-frame mode, the blending mode is not applied. For example, when the block is in intra-frame / inter-frame mode, the traditional intra-frame mode is used, such as smooth mode, smooth-v mode, or smooth-h mode in AV1.

[0129] (A1) In one aspect, some embodiments include a video coding method (e.g., method 600). In some embodiments, the method is performed at a computing system (e.g., server system 112) having memory and control circuitry. In some embodiments, the method is performed at an encoding module (e.g., encoding module 320). In some embodiments, the method is performed at a source encoding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving video data comprising a plurality of blocks, the plurality of blocks including a first block, wherein the first block is encoded in a smoothing mode; (ii) identifying a set of reference samples for the first block; (iii) deriving a first prediction value for the first block, wherein the first prediction value is derived by a linear interpolation function relating to a first reference sample in the set of reference samples and the height and width of the first block; and (iv) encoding the first block based on a first angle prediction value.

[0130] (A2) In some embodiments of A1, the second subset of the reference samples includes left-side samples located to the left of this portion of the first block. In some embodiments, the left-side samples are identified by extending a direction from the top reference through this portion to the left-side samples. In some embodiments, the left-side samples include refined samples refined using one or more predetermined values. In some embodiments, a first angle prediction value is derived using interpolation of the left-side and right-side samples of the second subset of the reference samples, the right-side samples being located to the right of this portion of the first block.

[0131] (A3) In some embodiments of A1, the second subset of the reference samples includes a top sample located above this portion of the first block. In some embodiments, the top sample is identified by extending a direction from the left reference through this portion to the top sample. In some embodiments, the top sample includes a refined sample refined using one or more predetermined values. In some embodiments, a first angle prediction value is derived by interpolation of the top and bottom samples of the second subset of the reference samples, the bottom sample located below this portion of the first block.

[0132] (A4) In some embodiments of any of A1 to A3, the method further includes filtering the reference sample set before using at least a first subset of the reference sample set to derive the first angle prediction value for this part of the first block.

[0133] (A5) In some embodiments of any of A1 to A4, deriving the second angle prediction value includes deriving a weighting factor based on the coordinates of this portion of the first block.

[0134] (A6) In some embodiments of any of A1 to A5, the first angle prediction value is derived without division. In some embodiments, the first angle prediction value is derived using interpolation of two or more reference samples. In some embodiments, the first angle prediction value is derived using a top reference sample, a left reference sample, a bottom reference sample, and a right reference sample.

[0135] (A7) In some embodiments of any of A1 to A6, based on the fact that the predicted angle of the determined angle prediction pattern is within a predetermined angle range, this portion of the first block is encoded using a second angle prediction value.

[0136] (A8) In some embodiments of any of A1 to A7, the second subset of the reference sample includes the bottom sample located below this portion of the first block.

[0137] (A9) In some embodiments of any of A1 to A8, the method further includes transmitting the encoded first block via a video stream. In some embodiments, the method further includes signaling the use of a second angle prediction value (e.g., signaling that an angle blending mode is enabled).

[0138] (B1) In another aspect, some embodiments include a video decoding method (e.g., method 650). In some embodiments, the method is executed at a computing system (e.g., server system 112) having memory and control circuitry. In some embodiments, the method is executed at an encoding module (e.g., encoding module 320). In some embodiments, the method is executed at a parser (e.g., parser 254), a motion prediction component (e.g., motion compensation prediction unit 260), and / or an intra-frame prediction component (e.g., intra-frame picture prediction unit 262). The method includes: (i) receiving video data (e.g., an encoded video sequence) from a video stream, the video data comprising a plurality of blocks, the plurality of blocks including a first block, wherein the first block is predicted in an angle prediction mode; (ii) identifying a reference sample set for a portion of the first block, wherein the reference sample set is identified using a prediction angle of the angle prediction mode; (iii) deriving a first angle prediction value for this portion of the first block using at least a first subset of the reference sample set; (iv) deriving a second angle prediction value for this portion of the first block using a weighted sum of at least a second subset of the reference sample set and the first angle prediction value; and (v) decoding this portion of the first block using the second angle prediction value. For example, the angle prediction mode is an intra-frame prediction mode, and the angle prediction value P is derived along a given prediction angle using reference samples from the top adjacent position and / or the upper right adjacent position of the current block. A .

[0139] (B2) In some embodiments of B1, the second subset of the reference samples includes the left-hand samples located to the left of this portion of the first block. For example, to predict samples in the current block using spatially adjacent reference samples, the angle prediction value P is first derived. A Then use the left-side reference sample and P A The final predicted value P' is derived from the weighted sum. A In some embodiments, P' is derived using Equation 1 above. A .

[0140] (B3) In some embodiments of B2, the left sample is identified by extending a direction from the top reference through this portion to the left sample. For example, L is obtained by extending this direction from the top reference to the left / lower left sample of the predicted value.

[0141] (B4) In some embodiments of B2 or B3, the left-side sample includes a refined sample refined using one or more predetermined values. For example, if L is outside the lower-left reference sample, L can be refined using some predetermined values ​​located in the left and / or lower-left regions. As an example, when L is outside the lower-left reference sample, predetermined samples located in the left or lower-left regions can be used. As another example, the left-side sample is populated using the nearest available sample located in the lower-left region, left region, upper-left region, upper region, and / or upper-right region.

[0142] (B5) In some embodiments of any of B2 to B4, the first angle prediction is derived using interpolation of the left and right samples of a second subset of the reference samples, where the right sample is located to the right of this portion of the first block. For example, P is generated by bilinear interpolation of the top sample T and the bottom sample B. A In some embodiments, P' is derived using Equation 3 above. A .

[0143] (B6) In some embodiments of B1, the second subset of the reference samples includes the top samples located above this portion of the first block. For example, to predict samples in the current block using spatially adjacent reference samples, the angle prediction value P is first derived. A Then use the top reference sample and P A The final predicted value P' is derived from the weighted sum. A In some embodiments, P' is derived using Equation 2 above. A .

[0144] (B7) In some embodiments of B6, the top sample is identified by extending the direction from the left reference through this portion to the top sample. For example, the top sample T is obtained by extending the left sample along the prediction angle to the top / upper right portion.

[0145] (B8) In some embodiments of B6 or B7, the top sample includes a refined sample refined using one or more predetermined values. For example, if T is outside the top reference sample or the upper right reference sample, T can be refined using some predetermined samples located in the top region and / or the upper right region. As an example, T is populated using the nearest available sample located in the upper right region, the top region, the upper left region, the left side region, and / or the lower left region. In some embodiments, if T is outside the top / upper right reference sample, T is assigned a predetermined sample located in the top region or the upper right region.

[0146] (B9) In some embodiments of any of B6 through B8, the first angle prediction is derived using interpolation of the top and bottom samples of a second subset of the reference samples, the bottom samples being located below this portion of the first block. For example, P is generated by bilinear interpolation of the top sample T and the bottom sample B. A In some embodiments, P' is derived using Equation 3 above. A .

[0147] (B10) In some embodiments of any of B1 to B9, the method further includes filtering the reference sample set before deriving the first angle prediction value for this part of the first block using at least a first subset of the reference sample set. For example, when generating P A Before the values ​​are calculated, a smoothing filter (e.g., a Gaussian filter or a bilateral filter) is used to filter the left reference sample, the top reference sample, the top left reference sample, the bottom left reference sample, and / or the top right reference sample, where the coefficients of the smoothing filter are all non-negative integers.

[0148] (B11) In some embodiments of any of B1 to B10, deriving the second angle prediction includes deriving a weighting factor based on the coordinates of this portion of the first block. For example, w L This can be derived as K >> ((x << 1)) >> s, where K is a predetermined value, x is the horizontal coordinate of the current sample to be predicted, and s is a scaling factor based on the block size. For example, s can be equal to (log2(W) + log2(H) + 2) >> 2, where W and H are the width and height of the block, respectively.

[0149] (B12) In some embodiments of any of B1 to B11, the first angle prediction value is derived without performing a division operation. For example, for the final predictor, P is calculated. A The division when predicting samples is moved to the last step, and during internal computation, parameters can be multiplied by a value so that these parameters have the same divisor.

[0150] (B13) In some embodiments of any of B1 to B12, the first angle prediction value is derived using interpolation of two or more reference samples. For example, to predict a sample in the current block using spatially adjacent reference samples, the angle prediction value P is first derived using interpolation of a top (or left) reference sample and a bottom (or right) reference sample. A Then use the top (or left) reference sample and P A The weighted sum is used to derive the refined (final) predicted value P'. AAs an example, a top (or left) reference sample can be derived using interpolations of multiple top (or left) reference samples (e.g., at integer positions) as input. As another example, a right (or bottom) reference sample can be derived using a reference sample from a top or left reference. For example, the derivation of the right (or bottom) reference sample can be the same as in smooth mode, blend mode, and / or planar mode.

[0151] (B14) In some embodiments of B13, a first angle prediction value is derived using a top reference sample, a left reference sample, a bottom reference sample, and a right reference sample. For example, an angle prediction value P is derived along a given prediction angle using reference samples from the top, left, bottom, and right reference samples of the current block. A .

[0152] (B15) In some embodiments of B1 through B14, this portion of the first block is decoded using a second angle prediction value, based on the fact that the predicted angle of the determined angle prediction mode is within a predetermined angle range. For example, a horizontal blending mode may be applied when the predicted angle is to the right of V_PRED (e.g., when the decoder can find the left sample by extending the predicted angle to the left reference sample). As another example, a vertical blending mode may be applied when the predicted angle is below H_PRED.

[0153] (B16) In some embodiments of any of B1 through B15, the second subset of the reference samples includes the bottom samples located below this portion of the first block. In some embodiments, P' is derived using Equation 3. A .

[0154] (B17) In some embodiments of any of B1 through B15, a second angle prediction value is derived in response to a mixing mode represented by a signal in the video bitstream. For example, whether a mixing mode is used to represent a signal at the sequence header, slice header, superblock, or coded block level. In some embodiments, a mixing mode is used if the block is encoded in intra-frame mode. In some embodiments, a mixing mode is not applied if the block is encoded in intra-frame-inter-frame mode. For example, for blocks encoded in intra-frame-inter-frame mode, an intra-frame mode is applied without mixing, such as using smooth mode, smooth-v mode, or smooth-h mode.

[0155] On the other hand, some embodiments include a computing system (e.g., server system 112) that includes control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry. The memory stores one or more instruction sets configured to be executed by the control circuitry. The one or more instruction sets include instructions for performing any of the methods described herein (e.g., A1 to A9 and B1 to B17 above).

[0156] In another aspect, some embodiments include a non-transitory computer-readable storage medium that stores one or more instruction sets executed by control circuitry of a computing system, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 to A9 and B1 to B17 above).

[0157] It should be understood that while the terms “first,” “second,” etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. The terminology used herein is for describing particular embodiments only and is not intended to limit the claims. As used in the description of embodiments and the appended claims, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations of one or more of the associated listed items, and includes any and all possible combinations of one or more of the associated listed items. It should also be further understood that, when used in this specification, the terms “comprising” and / or “including” specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0158] As used herein, the term "if" can be interpreted as meaning "when," "at," "in response to determination," "according to determination," or "in response to detection," depending on the context. Similarly, the phrases "if it is determined [the stated prerequisite is true]," "if [the stated prerequisite is true]," or "when [the stated prerequisite is true]" can be interpreted as meaning "in response to determination," "according to determination," "in response to detection," or "in response to detection," depending on the context.

[0159] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the teachings above. The embodiments were chosen and described in order to best explain the operating principles and practical applications, thereby enabling others skilled in the art to implement them.

Claims

1. A method of video decoding performed on a computing system having a memory and one or more processors, the method comprising: The method comprises: receiving video data from a video bitstream, the video data comprising a plurality of blocks, the plurality of blocks comprising a first block, wherein the first block is predicted in an angular prediction mode; identifying a reference sample set for a portion of the first block, wherein the reference sample set is identified using a prediction angle of the angular prediction mode; deriving a first angular prediction value for the portion of the first block using at least a first subset of the reference sample set, wherein the first angular prediction value is derived using an interpolation of a left reference sample and a right reference sample of the first subset of the reference sample set, the right reference sample being located right of the portion of the first block, the left reference sample being located left of the portion of the first block; deriving a second angular prediction value for the portion of the first block using a weighted sum of (i) at least a second subset of the reference sample set and (ii) the derived first angular prediction value; and decoding the portion of the first block using the second angular prediction value.

2. The method of claim 1, wherein, The second angular prediction value corresponds to a horizontal blending mode, and wherein the horizontal blending mode derives the second angular prediction value using a weighted sum of the first angular prediction value and a left reference sample, the second subset of the reference sample set comprising the left reference sample.

3. The method of claim 2, wherein, The left reference sample is identified by extending a direction through the portion from a top reference sample to the left reference sample.

4. The method of claim 2, wherein, The left reference sample comprises a refined reference sample refined using one or more predetermined values.

5. The method of claim 1, wherein, The second angular prediction value corresponds to a vertical blending mode, and wherein the vertical blending mode derives the second angular prediction value using a weighted sum of the first angular prediction value and a top reference sample, the second subset of the reference sample set comprising the top reference sample, the top reference sample being located above the portion of the first block.

6. The method of claim 5, wherein, The top reference sample is identified by extending a direction through the portion from a left reference sample to the top reference sample.

7. The method of claim 5, wherein, The top reference sample comprises a refined reference sample refined using one or more predetermined values.

8. The method of claim 5, wherein, The first angular prediction value is derived using an interpolation of the top reference sample and a bottom reference sample of the first subset of the reference sample set, the bottom reference sample being located below the portion of the first block.

9. The method of claim 1, wherein, Further comprising: filtering the reference sample set prior to deriving the first angular prediction value for the portion of the first block using at least the first subset of the reference sample set.

10. The method of claim 1, wherein, Deriving the second angular prediction value comprises deriving a weighting factor based on coordinates of the portion of the first block.

11. The method of claim 1, wherein, The first angular prediction value is derived without performing a division operation.

12. The method of claim 1, wherein, The first angular prediction value is derived using an interpolation of two or more reference samples.

13. The method of claim 12, wherein, The first angular prediction value is derived using a top reference sample, a left reference sample, a bottom reference sample, and a right reference sample.

14. The method of claim 1, wherein, According to a determination that the prediction angle of the angular prediction mode lies within a predetermined angle range, the portion of the first block is decoded using the second angle prediction value.

15. The method of claim 1, wherein, The second subset of the reference sample set includes bottom reference samples that are located below the portion of the first block.

16. The method of claim 1, wherein, When a hybrid mode is signaled in the video bitstream, the second angle prediction value is derived using a weighted sum of (i) at least a second subset of the reference sample set and (ii) a derived first angle prediction value.

17. A computing system, comprising: comprising: control circuitry; memory; and one or more sets of instructions stored in the memory and configured for execution by the control circuitry, the one or more sets of instructions including instructions for implementing the video decoding method of any of claims 1 to 16.

18. A non-transitory computer-readable storage medium, comprising: storing one or more sets of instructions configured for execution by a computing device having control circuitry and memory, the one or more sets of instructions including instructions for implementing the video decoding method of any of claims 1 to 16.

19. A method of storing a video bitstream, the method comprising: executing the video decoding method of any of claims 1 to 16 to decode a video bitstream; and storing the video bitstream.

20. A method of transmitting a video bitstream, the method comprising: executing the video decoding method of any of claims 1 to 16 to decode a video bitstream; and transmitting the video bitstream.

Citation Information

Patent Citations

  • Method and apparatus for low-complexity bi-directional intra prediction in video encoding and decoding

    US20200204826A1