Systems and methods for signaling downsampling filters for luma-to-chroma intra-prediction modes
The method optimizes the signaling of downsampling filters in luma-to-chroma intra-prediction modes by retrieving syntax elements and predicting chroma blocks, addressing inconsistencies and redundancy in existing methods, thereby enhancing decoding quality and efficiency.
Patent Information
- Application Number
- JP2025523009
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-05
- Filing Date
- 2023-05-08
- Publication Date
- 2025-11-12
- Estimated Expiration
- 2043-05-08
AI Technical Summary
Current signaling approaches for downsampling filters in luma-to-chroma intra-prediction modes result in differing outcomes between sequential and parallel operation of multiple groups of pictures and redundant signaling overhead for inter-coded frames.
A method for signaling downsampling filters in luma-to-chroma intra-prediction modes that involves retrieving a syntax element associated with a key frame, downsampling luma blocks using the filter type, and predicting chroma blocks based on the downsampled frames, improving decoding quality and efficiency.
Enhances decoding quality and efficiency by optimizing the signaling of downsampling filters, reducing redundancy and ensuring consistent results across different operational scenarios.
Smart Images

Figure 2025536967000001_ABST
Abstract
Description
[Technical Field]
[0001] Related Applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 428,714, entitled "Signaling of Downsampling Filters for Chroma from Luma Intra Prediction Mode," filed November 29, 2022, and is a continuation of U.S. Patent Application No. 18 / 144,042, entitled "Systems and Methods for Signaling of Downsampling Filters for Chroma from Luma Intra Prediction Mode," filed May 5, 2023, all of which are incorporated herein by reference.
[0002] FIELD OF THE INVENTION
[0002] The disclosed embodiments relate generally to video coding, including, but not limited to, systems and methods for signaling downsampling filters for luma-to-chroma intra-prediction modes. [Background technology]
[0003]
[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video gaming consoles, smartphones, video teleconferencing devices, video streaming devices, etc. The electronic devices transmit and receive or sometimes communicate digital video data across communication networks and / or store the digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding may be used to compress the video data according to one or more video coding standards before the video data is communicated or stored.
[0004]
[0004] Multiple video codec standards have been developed. For example, video coding standards include AOMedia Video 1 (AV1), Generic Video Coding (VVC), Joint Exploration and Test Model (JEM), High Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Experts Group (MPEG) coding. Video coding generally uses prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality.
[0005]
[0005] HEVC, also known as H.265, is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Generic Video Coding (VVC), also known as H.266, is a video compression standard intended as the successor to HEVC. ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AV1 is an open video coding format designed as a replacement for HEVC. Validated version 1.0.0, accompanied by Errata 1 of the specification, was released on January 8, 2019. Summary of the Invention [Problem to be solved by the invention]
[0006]
[0006] This disclosure describes various techniques for signaling downsampling filters used in cross-component intra-prediction modes. Luma-to-chroma (CfL) prediction is an efficient video coding tool that models chroma pixels as a linear function of simultaneous reconstructed luma pixels. When there are multiple downsampling filters used in CfL modes, these filters need to be signaled in a high-level syntax. Current signaling approaches may have two problems: first, results may differ between sequential and parallel operation of multiple groups of pictures (GOPs); and second, there may be redundant signaling overhead for inter-coded frames. [Means for solving the problem]
[0007]
[0007] Therefore, there is a need for improved methods and systems for signaling downsampling filters in CfL modes. This disclosure describes various techniques for signaling downsampling filters used in cross-component intra-prediction modes. The disclosed techniques may be used by a decoder of a video bitstream to improve decoding quality and / or efficiency. Video encoders may also implement these techniques during encoding (e.g., to reconstruct coded frames and / or test hypotheses).
[0008] According to some embodiments, a method of video decoding is provided. The method includes receiving a video stream having a sequence of frames, the sequence of frames including one or more key frames, each key frame having a respective downsampling filter type. In accordance with determining that a current frame corresponds to a first key frame of the one or more key frames, the method includes (i) retrieving from the video stream a syntax element associated with a first downsampling filter type associated with the first key frame, (ii) downsampling luma blocks of the current frame and a predefined set of frames immediately following the current frame using the first downsampling filter type to obtain downsampled frames, and (iii) predicting chroma blocks of the current frame based on the downsampled frames.
[0009] According to some embodiments, a computing system, such as a streaming system, a server system, a personal computer system, or other electronic device, is provided. The computing system includes control circuitry and a memory that stores one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.
[0010]
[0010] According to some embodiments, a non-transitory computer-readable storage medium is provided that stores one or more sets of instructions for execution by a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein.
[0011]
[0011] Accordingly, disclosed are methods, as well as devices and systems, for coding video, which may complement or replace conventional methods, devices, and systems for video coding.
[0012]
[0012] The features and advantages described herein are not necessarily all-inclusive, and in particular, some additional features and advantages will become apparent to those skilled in the art in view of the drawings, specification, and claims provided in this disclosure. Moreover, it should be noted that the language used herein has been selected primarily for readability and educational purposes, and has not necessarily been selected to define or limit the subject matter described herein.
[0013]
[0013] So that the present disclosure may be more fully understood, a more particular description may be made by reference to features of various embodiments, some of which are illustrated in the accompanying drawings, which, however, illustrate only pertinent features of the present disclosure and, therefore, should not necessarily be considered limiting, since the description may lead to other useful features, as will be appreciated by those skilled in the art upon reading the present disclosure. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary communication system, according to some embodiments. [Figure 2A]
[0015] FIG. 2 is a block diagram illustrating exemplary elements of an encoder component, according to some embodiments. [Figure 2B]
[0016] 3 is a block diagram illustrating exemplary elements of a decoder component, according to some embodiments. [Figure 3]
[0017] FIG. 1 is a block diagram illustrating an exemplary server system, according to some embodiments. [Figure 4]
[0018] FIG. 10 illustrates nominal angles in directional intra prediction, according to some embodiments. [Figure 5]
[0019] FIG. 10 is a diagram illustrating top, left, and top-left positions for PAETH intra-prediction modes for predicting coding blocks, according to some embodiments. [Figure 6]
[0020] FIG. 1 is a block diagram illustrating a luma-to-chroma (CfL) prediction process, according to some embodiments. [Figure 7]
[0021] FIG. 1 is a block diagram of luma samples inside and outside a picture boundary according to some embodiments. [Figure 8]
[0022] 1A-1C illustrate different chroma downsampling formats according to some embodiments. [Figure 9]
[0023] FIG. 1 illustrates an AVI CfL downsampling filter, according to some embodiments. [Figure 10]
[0024] FIG. 1 illustrates binarization processes and their corresponding codes, according to some embodiments. [Figure 11]
[0025] FIG. 10 illustrates temporal layer identifier numbers for nine consecutive frames according to some embodiments. [Figure 12]
[0026] 1 is a flow diagram illustrating an exemplary method for video decoding, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0015]
[0027] According to common practice, the various features illustrated in the drawings are not necessarily drawn to scale and like reference numerals may be used to denote like features throughout the specification and figures.
[0016]
[0028] This disclosure describes signaling downsampling filter types in a luma-to-chroma (CfL) intra-prediction mode of video coding. A video stream having a sequence of frames is received. The sequence of frames includes one or more key frames, each having a respective downsampling filter type. According to a determination that a current frame corresponds to a first key frame of the one or more key frames, a syntax element associated with a first downsampling filter type associated with the first key frame is retrieved from the video bitstream. To obtain a downsampled frame, luma blocks (e.g., blocks of pixels) of the current frame and a predefined set of frames immediately following the current frame are downsampled using the first downsampling filter type. Chroma blocks (e.g., blocks of pixels) of the current frame are predicted based on the downsampled frame.
[0017] Exemplary Systems and Devices
[0029] 1 is a block diagram illustrating a communication system 100 according to some embodiments. Communication system 100 includes a source device 102 and multiple electronic devices 120 (e.g., electronic devices 120-1 through 120-m) communicatively coupled to each other via one or more networks. In some embodiments, communication system 100 is a streaming system for use with video-enabled applications, such as, for example, video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0018]
[0030] Source device 102 includes a video source 104 (e.g., a camera component or media storage) and an encoder component 106. In some embodiments, video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). Encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from video source 104 may be of a high data volume compared to encoded video bitstream 108 generated by encoder component 106. Because encoded video bitstream 108 has a lower data volume (less data) compared to the video stream from the video source, encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared to the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., configured to transmit uncompressed video data to network(s) 110).
[0019]
[0031] The one or more networks 110 represent any number of networks that convey information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wireline (wired) and / or wireless communication networks. The one or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0020]
[0032] The one or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from the source device 102). The server system 112 includes a coder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the coder component 114 includes an encoder component and / or a decoder component. In various embodiments, the coder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the coder component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different encoding standards and / or methodologies to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video bitstream 108.
[0021]
[0033] In some embodiments, server system 112 functions as a media-aware network element (MANE). For example, server system 112 may be configured to prune encoded video bitstream 108 to accommodate potentially different bitstreams to one or more of electronic devices 120. In some embodiments, a MANE is provided separately from server system 112.
[0022]
[0034] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, decoder component 122 is configured to decode encoded video data 116 to generate an outgoing video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include media storage). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.
[0023]
[0035] The source device and / or the plurality of electronic devices 120 may be referred to as “terminal devices” or “user devices.” In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smartphone, tablet, or laptop), a wearable device, a videoconferencing device, and / or other types of electronic devices.
[0024]
[0036] In an exemplary operation of communication system 100, source device 102 transmits encoded video bitstream 108 to server system 112. For example, source device 102 may code a stream of pictures captured by the source device. Server system 112 may receive encoded video bitstream 108 and decode and / or encode encoded video bitstream 108 using coder component 114. For example, server system 112 may apply coding to the video data, which is more optimal for network transmission and / or storage. Server system 112 may transmit encoded video data 116 (e.g., one or more coded video bitstreams) to one or more of electronic devices 120. Each electronic device 120 may decode the encoded video data 116 to recover the video pictures and, optionally, display them.
[0025]
[0037] In some embodiments, the transmission described above is a unidirectional data transmission. The unidirectional data transmission may be utilized in media serving applications, etc. In some embodiments, the transmission described above is a bidirectional data transmission. The bidirectional data transmission may be utilized in video conferencing applications, etc. In some embodiments, the coded video bitstream 108 and / or the coded video data 116 are coded and / or decoded according to any of the video coding / compression standards described herein, such as HEVC, VVC, and / or AV1.
[0026]
[0038] 2A is a block diagram illustrating exemplary elements of the encoder component 106, according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a different device than the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCB, or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. Video data may be provided as multiple individual pictures that, when viewed in sequence, impart motion. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art can readily understand the relationship between pixels and samples. The following discussion focuses on samples.
[0027]
[0039] The encoder component 106 is configured to code and / or compress pictures of a source video sequence into a coded video sequence 216 in real time or under other time constraints required by an application. Enforcing an appropriate coding rate is one function of the controller 204. In some embodiments, the controller 204 controls and is operatively coupled to other functional units as described below. Parameters set by the controller 204 may include rate control-related parameters (e.g., picture skip, quantizer, and / or λ value for rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller 204, as they may pertain to the encoder component 106 being optimized for a certain system design.
[0028]
[0040] In some embodiments, the encoder component 106 is configured to operate in a coding loop. In a simplified example, the coding loop includes a source coder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and one or more reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a manner similar to a (remote) decoder (when the compression between the symbols and the coded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory 208. Because decoding of the symbol stream yields bit-exact results independent of the decoder location (local or remote), the contents in the reference picture memory 208 are also bit-exact between the local and remote encoders. In this way, the predictive portion of the encoder interprets the same sample values as reference picture samples that the decoder would interpret when using prediction during decoding. This principle of reference picture synchrony (and the resulting drift if synchrony cannot be maintained, for example due to channel errors) is known to those skilled in the art.
[0029]
[0041] The operation of decoder 210 may be the same as that of a remote decoder, such as decoder component 122, which is described in detail below in connection with Figure 2B. However, with brief reference to Figure 2B, because symbols are available and the encoding / decoding of symbols into a coded video sequence by entropy coder 214 and parser 254 may be lossless, the entropy decoding portion of decoder component 122, including buffer memory 252 and parser 254, may not be fully implemented in local decoder 210.
[0030]
[0042] An observation that can be made at this point is that decoder techniques other than parsing / entropy decoding that are present in a decoder necessarily must also be present in a corresponding encoder in a substantially equivalent functional form. For this reason, the disclosed subject matter focuses on decoder operation. A description of the encoder techniques may be omitted, since they are the inverse of the decoder techniques that are comprehensively described. Only in some areas is further detail required, which is provided below.
[0031]
[0043] As part of its operation, source coder 202 may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from a video sequence designated as reference frames. In this manner, coding engine 212 codes differences between pixel blocks of the input frame and pixel blocks of reference frame(s) that may be selected as prediction reference(s) for the input frame. Controller 204 may manage the coding operations of source coder 202, including, for example, setting parameters and subgroup parameters used to encode the video data.
[0032]
[0044] The decoder 210 decodes coded video data of frames that may be designated as reference frames based on symbols created by the source coder 202. The operation of the coding engine 212 may advantageously be a lossy process. When the coded video data is decoded in a video decoder (not shown in FIG. 2A ), the reconstructed video sequence may be a replica of the source video sequence with some errors. The decoder 210 replicates the decoding process that may be performed on the reference frames by a remote video decoder, which may cause the reconstructed reference frames to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores copies of the reconstructed reference frames that have common content as the reconstructed reference frames that will be obtained by the remote video decoder (without transmission errors).
[0033]
[0045] The predictor 206 may perform the predictive search for the coding engine 212. That is, for a new frame to be coded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that can serve as suitable predictive references for the new picture. The predictor 206 may operate on a sample block-by-pixel block basis to find suitable predictive references. In some cases, as determined by the search results obtained by the predictor 206, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory 208.
[0034]
[0046] The output of all the above-mentioned functional units may undergo entropy coding in entropy coder 214. Entropy coder 214 turns the symbols produced by the various functional units into a coded video sequence by losslessly compressing them according to techniques known to those skilled in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).
[0035]
[0047] In some embodiments, the output of the entropy coder 214 is coupled to a transmitter. The transmitter may be configured to buffer the coded video sequence(s) created by the entropy coder 214 to prepare them for transmission over a communication channel 218, which may be a hardware / software link to a storage device that will store the coded video data. The transmitter may be configured to merge the coded video data from the source coder 202 with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown). In some embodiments, the transmitter may transmit additional data along with the coded video. The source coder 202 may include such data as part of the coded video sequence. The additional data may include other forms of redundant data, such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, supplemental enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.
[0036]
[0048] The controller 204 may manage the operation of the encoder component 106. During coding, the controller 204 may assign a coded picture type to each coded picture, which may affect the coding technique applied to the respective picture. For example, a picture may be assigned as an intra picture (I picture), a predicted picture (P picture), or a bidirectionally predicted picture (B picture). An intra picture may be coded and decoded without using other frames in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those skilled in the art are aware of these variations of I pictures and their respective applications and characteristics, and therefore, they will not be repeated here. A predicted picture may be coded and decoded using intra prediction or inter prediction, using at most one motion vector and reference index to predict sample values for each block. A bidirectionally predicted picture may be coded and decoded using intra or inter prediction, using at most two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predicted picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0037]
[0049] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be nonpredictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0038]
[0050] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture to be coded / decoded, called the current picture, is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture may be coded by a vector called a motion vector. A motion vector points to a reference block in a reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are in use.
[0039]
[0051] Encoder component 106 may perform coding operations in accordance with a given video coding technique or standard, such as any of those described herein. In doing so, encoder component 106 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.
[0040]
[0052] 2B is a block diagram illustrating exemplary elements of the decoder component 122, according to some embodiments. The decoder component 122 in FIG. 2B is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter 256 and configured to transmit data to the display 124 (e.g., via a wired or wireless connection).
[0041]
[0053] In some embodiments, the decoder component 122 includes a receiver coupled to the channel 218 and configured to receive data from the channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more coded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each coded video sequence is independent of the other coded video sequences. Each coded video sequence may be received from the channel 218, which may be a hardware / software link to a storage device that stores the coded video data. The receiver may receive the coded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver may separate the coded video sequence from the other data. In some embodiments, the receiver receives additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the decoder component 122 to decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0042]
[0054] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra-picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. In some embodiments, the decoder component 122 is implemented at least partially in software.
[0043]
[0055] The buffer memory 252 is coupled intermediate the channel 218 and the parser 254 (e.g., to eliminate network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 internal to the decoder component 122 (e.g., configured to handle playout timing), a separate buffer memory is provided external to the decoder component 122 (e.g., to eliminate network jitter). When receiving data from a storage / forwarding device of sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory 252 may not be required or may be small. For use with best-effort packet networks such as the Internet, the buffer memory 252 may be required and may be relatively large, advantageously of adaptive size, and may be implemented at least in part in an operating system or similar element (not shown) external to the decoder component 122.
[0044]
[0056] Parser 254 is configured to reconstruct symbols 270 from the coded video sequence. These symbols may include, for example, information used to manage the operation of decoder component 122 and / or information for controlling a rendering device such as display 124. The control information for the rendering device(s) may be in the form of, for example, a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). Parser 254 parses (entropy decodes) the coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 254 may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to that group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. Parser 254 may also extract information from the coded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0045]
[0057] The reconstruction of symbols 270 may involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.), and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by parser 254. The flow of such subgroup control information between parser 254 and the following units is not shown for clarity.
[0046]
[0058] Beyond the functional blocks already mentioned, the decoder component 122 may be conceptually subdivided into several functional units, which are described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be, at least partially, integrated with each other. However, for purposes of describing the disclosed subject matter, the conceptual subdivision into the following functional units will be maintained.
[0047]
[0059] The scaler / inverse transform unit 258 receives the quantized transform coefficients as well as control information (such as which transform to use, block size, quantization factor, and / or quantization scaling metric) from the parser 254 as symbol(s) 270. The scaler / inverse transform unit 258 may output blocks containing sample values that may be input to the aggregator 268.
[0048]
[0060] In some cases, the output samples of the scaler / inverse transform unit 258 relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 may generate blocks of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the current (partially reconstructed) picture from the current picture memory 264. The aggregator 268 may add, on a sample-by-sample basis, the prediction information generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258.
[0049]
[0061] In other cases, the output samples of the scalar / inverse transform unit 258 relate to inter-coded and potentially motion-compensated blocks. In such cases, the motion-compensated prediction unit 260 can access the reference picture memory 266 to fetch samples used for prediction. After motion-compensating the fetched samples according to symbols 270 related to the block, these samples can be added by the aggregator 268 to the output of the scalar / inverse transform unit 258 (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory 266 from which the motion-compensated prediction unit 260 fetches prediction samples can be controlled by motion vectors. The motion vectors can be available to the motion-compensated prediction unit 260, for example, in the form of symbols 270 that can have X, Y, and reference picture components. Motion compensation can also include interpolation of sample values fetched from the reference picture memory 266 when sub-sample exact motion vectors are in use, motion vector prediction mechanisms, and the like.
[0050]
[0062] The output samples of aggregator 268 may undergo various loop filtering techniques in loop filter unit 256. Video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video bitstream and made available to loop filter unit 256 as symbols 270 from parser 254, but may also be responsive to meta-information obtained during decoding of a coded picture or previous portion (in decoding order) of a coded video sequence, as well as to previously reconstructed and loop-filtered sample values.
[0051]
[0063] The output of the loop filter unit 256 may be a sample stream that may be output to a render device such as the display 124, as well as stored in the reference picture memory 266 for use in future inter-picture prediction.
[0052]
[0064] Some coded pictures, once fully reconstructed, may be used as reference pictures for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 254), the current reference picture may become part of reference picture memory 266, and a fresh current picture memory may be reallocated before beginning reconstruction of a subsequent coded picture.
[0053]
[0065] Decoder component 122 may perform decoding operations according to a predetermined video compression technology, which may be documented in a standard, such as any of the standards described herein. A coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in that it follows the syntax of the video compression technology or standard as specified in the video compression technology document or standard, and in particular, in a profile document therein. Also, for compliance with some video compression technologies or standards, the complexity of a coded video sequence may be within limits defined by the level of the video compression technology or standard. In some cases, the level restricts a maximum picture size, a maximum frame rate, a maximum reconstruction sample rate (e.g., measured in megasamples per second), a maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted through a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.
[0054]
[0066] 3 is a block diagram illustrating a server system 112, according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., a CPU, a GPU, and / or a DPU). In some embodiments, the control circuit includes one or more field programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuits).
[0055]
[0067] The network interface(s) 304 may be configured to interface with one or more communication networks (e.g., wireless networks, wireline networks, and / or optical networks). The communication networks may be local, wide-area, metropolitan, vehicular and industrial, real-time, delay-tolerant, etc. Examples of communication networks include local area networks such as Ethernet, wireless LANs, cellular networks, to include GSM, 3G, 4G, 5G, LTE, etc.; TV wireline or wireless wide-area digital networks, to include cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial, to include CANbus; etc. Such communications may be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., CANbus to some CANbus devices), or bidirectional (e.g., with other computer systems using local or wide-area digital networks). Such communications may include communications to one or more cloud computing networks.
[0056]
[0068] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device(s) 310 may include one or more of a keyboard, a mouse, a trackpad, a touchscreen, a data glove, a joystick, a microphone, a scanner, a camera, etc. The output device(s) 308 may include one or more of an audio output device (e.g., a speaker), a visual output device (e.g., a display or monitor), etc.
[0057]
[0069] Memory 314 may include high-speed random-access memory (such as DRAM, SRAM, DDR RAM, and / or other random-access solid-state memory devices) and / or non-volatile memory (such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Memory 314 optionally includes one or more storage devices located remotely from control circuitry 302. Memory 314, or alternatively, the non-volatile solid-state memory device(s) within memory 314, comprises a non-transitory computer-readable storage medium. In some embodiments, memory 314, or the non-transitory computer-readable storage medium of memory 314, stores the following programs, modules, instructions, and data structures, or a subset or superset thereof: an operating system 316, which includes procedures for handling various basic system services and for performing hardware-dependent tasks; a network communications module 318 used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); A coding module 320 for performing various functions related to encoding and / or decoding data, such as video data. In some embodiments, the coding module 320 is an instance of the coder component 114. The coding module 320 may include, but is not limited to: a decoding module 322 for performing various functions related to decoding the encoded data, such as those previously described with respect to the decoder component 122; an encoding module 340 for performing various functions related to encoding data, such as those previously described with respect to the encoder component 106; and A picture memory 352 for storing pictures and picture data, e.g., for use with the coding module 320. In some embodiments, the picture memory 352 includes one or more of the reference picture memory 208, the buffer memory 252, the current picture memory 264, and the reference picture memory 266.
[0058]
[0070] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to the parser 254), a transform module 326 (e.g., configured to perform various functions previously described with respect to the scalar / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to the loop filter 256).
[0059]
[0071] In some embodiments, encoding module 340 includes a code module 342 (e.g., configured to perform various functions previously described with respect to source coder 202 and / or coding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described with respect to predictor 206). In some embodiments, decoding module 322 and / or encoding module 340 include a subset of the modules shown in Figure 3. For example, a shared prediction module is used by both decoding module 322 and encoding module 340.
[0060]
[0072] Each of the above-identified modules stored in memory 314 corresponds to a set of instructions for performing functions described herein. The above-identified modules (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules; thus, various subsets of these modules may be combined or possibly rearranged in various embodiments. For example, coding module 320 optionally does not include separate decoding and encoding modules, but rather uses the same set of modules to perform both sets of functions. In some embodiments, memory 314 stores a subset of the above-identified modules and data structures. In some embodiments, memory 314 stores additional modules and data structures not described above, such as an audio processing module.
[0061]
[0073] In some embodiments, the server system 112 includes a web or Hypertext Transfer Protocol (HTTP) server, a File Transfer Protocol (FTP) server, and web pages and applications implemented using Common Gateway Interface (CGI) scripts, the PHP Hypertext Preprocessor (PHP), Active Server Pages (ASP), Hypertext Markup Language (HTML), Extensible Markup Language (XML), Java, JavaScript, Asynchronous JavaScript and XML (AJAX), XHP, Javelin, Wireless Universal Resource Files (WURFL), and the like.
[0062]
[0074] 3 illustrates a server system 112 according to some embodiments, although FIG. 3 is intended as a functional illustration of various features that may be present in one or more server systems, rather than a structural schematic of the embodiments described herein. In practice, and as will be recognized by those skilled in the art, items shown separately may be combined and some items may be separated. For example, some items shown separately in FIG. 3 may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement server system 112 and how features are allocated among them will vary from implementation to implementation and, optionally, depend in part on the amount of data traffic the server system handles during peak usage periods as well as during average usage periods.
[0063] Exemplary Intra Prediction Modes
[0075] 4 illustrates nominal angles for directional intra prediction, according to some embodiments. In some implementations of intra prediction, directional intra modes may be further extended to a finer-granularity angle set to further exploit the greater spatial redundancy in directional textures. For example, the VP9 coding format supports eight directional modes corresponding to angles between 45 and 207 degrees. In some embodiments, these eight directional modes are configured to provide eight nominal angles, referred to as V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, as shown in FIG. 4.
[0064]
[0076] For each nominal angle, a predefined number (e.g., 7) of finer angles may be added. With such expansion, a larger total number (e.g., 56 in this example) of directional angles may be available for intra prediction, corresponding to the same number of predefined directional intra modes. In some embodiments, the prediction angle may be represented by the nominal intra angle plus an angle delta. In the particular example above with seven finer angle directions for each nominal angle, the angle delta may be −3 to 3 multiplied by a step size of 3 degrees. To implement directional prediction modes in a general way, the directional intra prediction modes are implemented with a joint directional predictor that projects each pixel to a reference subpixel location and interpolates the reference pixel with a two-tap bilinear filter.
[0065]
[0077] In some embodiments, a predefined number of non-directional intra-prediction modes may also be predefined and made available. For example, five non-directional intra-prediction modes may be specified: DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. In some embodiments, the non-directional intra-prediction modes are also known as non-directional smooth intra-prediction modes.
[0066]
[0078] Prediction of samples of a particular block under these exemplary non-directional modes is shown in Figure 5. Figure 5 shows an exemplary 4-pixel by 4-pixel block 502, which is predicted by samples from the block's upper and / or left neighboring lines. A current pixel 510 in block 502 may correspond to a sample 504 located above the current pixel 510, an upper-left sample 506 located at the intersection of the upper and left neighboring lines, and a sample 508 located to the left of the current pixel 510. In DC intra-prediction mode, the average of the left and upper neighboring samples is used as the predictor for the block to be predicted. In PAETH intra-prediction mode, the upper reference sample, the left reference sample, and the upper-left reference sample are fetched, and then the value closest to (above + left - above left) is set as the predictor for the pixel to be predicted. The SMOOTH, SMOOTH_V, and SMOOTH_H intra-prediction modes predict block 502 using quadratic interpolation in the vertical or horizontal direction, or an average in both directions. The above-described non-directional intra-prediction modes are merely exemplary. It will be apparent to those skilled in the art that other adjacent line or non-directional modes may be selected and / or combined with the predicting sample to predict a particular sample in a prediction block, and are also contemplated.
[0067]
[0079] The encoder's selection of a particular intra-prediction mode from the above directional or non-directional modes at various coding levels (e.g., picture, slice, block, unit, etc.) may be signaled in the bitstream. In some embodiments, the exemplary eight nominal directional modes along with the five non-angle smooth modes (a total of 13 options) may be signaled first. In some embodiments, if the signaled mode is one of the eight nominal angle intra-modes, an index is further signaled to indicate the selected angle delta relative to the corresponding signaled nominal angle. In some embodiments, all intra-prediction modes may be indexed together for signaling (e.g., 56 directional modes + 5 non-directional modes, resulting in 61 intra-prediction modes). In some embodiments, the exemplary 56 or other number of directional intra-prediction modes may be implemented with a joint directional predictor that projects each sample of a block to a reference sub-sample location and interpolates the reference sample with a 2-tap bilinear filter.
[0068]
[0080] 6 is a block diagram illustrating a luma-to-chroma (CfL) prediction process 600 (performed, for example, by the prediction module 328) according to some embodiments. CfL is an intra-prediction mode in which chroma pixels are modeled as linear functions of simultaneous reconstructed luma pixels. CfL prediction is expressed as follows: CfL(α)=α×L AC +DC (1)
[0069]
[0081] In equation (1), L AC denotes the alternating current (AC) contribution of the luma component, α denotes the parameters of the linear model, and DC denotes the direct current (DC) contribution of the chroma component.
[0070]
[0082] Figure 6 shows that in a CfL prediction process 600, reconstructed luma pixels 602 are subsampled (604) (e.g., downsampled) to the chroma resolution and then averaged (606) to obtain an average value 608. The subsampled luma pixels have their average value 608 subtracted (610) from them to form the AC contributions 612 of the luma component. The AC contributions 612 of the luma component are then multiplied (614) by a scaling parameter α 616 (defined in equation (1) above) to generate scaled AC contributions 617 of the luma component. The scaled AC contributions 617 of the luma component are then added (618) to DC contribution predictions of the chroma components (620) to obtain chroma prediction samples (CfL predictions 622) for the predicted chroma block.
[0071]
[0083] In some embodiments, instead of requiring the decoder to calculate scaling parameters to approximate the chroma AC components from the AC contributions, the parameter α is determined (e.g., by the prediction module 328) based on the original chroma pixels and signaled in the bitstream. This reduces decoder complexity and results in more accurate prediction. As for the DC contribution of the chroma components, it is calculated using intra DC mode, which is sufficient for most chroma content and has a mature and fast implementation.
[0072]
[0084] In some embodiments, when some luma samples of a co-located luma block are outside the picture boundary, these luma samples may be padded, and the padded luma samples may be used to calculate the luma mean (block 606). Figure 7 is a block diagram of luma samples inside and outside the picture boundary, according to some embodiments. The outer picture luma samples 702 may be padded by addressing the value of the nearest available sample in the current block.
[0073]
[0085] In CfL mode, the luma subsampling process is combined with the average subtraction process, as shown in Figure 6. In this way, not only is the formula simplified, but also the subsampling division and the corresponding rounding error are eliminated. Equation (2) corresponds to the combination of both processes, which simplifies to Equation (3). Both Equation (2) and Equation (3) use integer division. M x N is a matrix of pixels in the luma plane.
[0074]
number
[0075]
[0086] Based on the supported chroma subsampling, x ×S y It can be shown that η∈{1,2,4} and that since both M and N are powers of two, M×N is also a power of two.
[0076]
[0087] For example, in the context of 4:2:0 chroma subsampling, instead of applying a box filter, the proposed technique only requires summing the four reconstructed luma pixels that match the chroma pixels, i.e., 4-tap
[0077]
number
[0078]
[0088] There may be different YUV formats according to different chroma downsampling phases. Figure 8 shows different chroma downsampling formats according to some embodiments. Different chroma formats define different downsampling grids (phases) for different color components. For the 4:2:0 format, there are two different common downsampling formats, namely 4:2:0 MPEG1 or 4:2:0 MPEG2, as shown in Figure 8.
[0079]
[0089] In some embodiments, the current luma downsampling filter in AV1 applies equation (4) to derive the reconstructed samples of luma.
[0080]
number
[0081]
[0090] The downsampling filter in AV1 may assume a chroma downsampling format corresponding to the 4:2:0 MPEG1 downsampling format, as shown in Figure 9. In some embodiments, multiple downsampling filters may be supported. In at least some of these embodiments, the filter type may be signaled, such as in a high-level syntax. In some embodiments, the multiple filters may include one or more 4-tap filters, one or more 6-tap filters, and another 4-tap filter in AVI.
[0082]
[0091] In entropy coding, syntax is binarized into 1s and 0s. Figure 10 shows various binarization schemes and their corresponding codes (e.g., binary symbols or bins) according to some embodiments. Generally, a binarization scheme defines a unique mapping of syntax element values to a sequence of binary symbols (e.g., bins), which may also be interpreted in terms of a binary code tree.
[0083]
[0092] Several different binarization processes are used in HEVC, including k-th truncated Rice (TRk), k-th exponential-Golomb (EGk), and fixed-length (FL) binarization. Some of these forms of binarization, including the truncated unary (TrU) method as 0-th TRk binarization, were also used in H.264 / AVC. These various methods of binarization can be described in terms of how they signal the unsigned value N. An example is also provided in Figure 10.
[0084]
[0093] Unary coding involves signaling a bin string of length N+1, where the first N bins are 1 and the last bin is 0. The decoder searches for 0 to determine when the syntax element is complete. In the TrU scheme, truncation is initiated for the largest possible value cMax1 of the syntax element being decoded.
[0085]
[0094] A k-th truncated rice is a parameterized Rice code consisting of a prefix and a suffix. The prefix is a truncated unary string of value N>>k, where the largest possible value is cMax. The suffix is a fixed-length binary representation of the least significant bin of N, where k denotes the number of least significant bins. When k=0, truncated rice is equivalent to truncated unary binarization.
[0086]
[0095] The k-th order Exponential-Golomb code is a robust, nearly optimal, prefix-free code for geometrically distributed sources with unknown or varying dispersion parameters. Each codeword has length L N +1 unary prefix and length L N +k suffix, where L N =log2((N>>k)+1).
[0087]
[0096] The fixed length code uses a fixed length bin string with length log2(cMax+1) and with the most significant bin signaled before the least significant bin.
[0088]
[0097] In some embodiments, the binarization scheme (e.g., process) is selected based on the type of syntax element. In some embodiments, the binarization scheme is selected depending on the value of a previously processed syntax element (e.g., the binarization of coeff_abs_level_remaining depends on the previously decoded coefficient level) or depending on a slice parameter indicating whether some modes are enabled. For example, the binarization of the partition mode, so-called part_mode, depends on whether asymmetric motion partitions are enabled. Most of the syntax elements use the binarization processes listed above, or some combination of them (e.g., cu_qp_delta_abs uses TrU(prefix)+EG0(suffix)). However, some syntax elements (e.g., part_mode and intra_chroma_pred_mode) use custom binarization processes.
[0089]
[0098] 12 is a flow diagram illustrating a method 1200 for video decoding according to some embodiments. Method 1200 may be implemented in a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 1200 is implemented by executing instructions stored in memory (e.g., memory 314) of the computing system.
[0090]
[0099] The system receives (1202) a video stream having a sequence of frames. The sequence of frames includes one or more key frames. Each key frame has a respective downsampling filter type. In accordance with determining that the current frame corresponds to a first key frame of the one or more key frames, the system retrieves (1204) from the video bitstream a syntax element (e.g., a signal element) associated with a first downsampling filter type associated with the first key frame (e.g., the syntax element is binarized to 1s and 0s). The system downsamples (1206) a luma block (e.g., a luma component, a block of luma pixels) of the current frame and a predefined set of frames immediately following the current frame using the first downsampling filter type to obtain a downsampled frame. In some embodiments, the predefined set of frames are frames immediately following the current frame and immediately preceding a second key frame of the one or more key frames. The system predicts (1208) chroma blocks (e.g., chroma components, blocks of chroma pixels) of the current frame (e.g., and a predefined set of frames immediately following the current frame) based on the downsampled frame.
[0091]
[0100] In some embodiments, the downsampling filter type is signaled in each key frame. In some embodiments, the downsampling filter type is signaled in an Instantaneous Decoder Refresh (IDR) frame. In some embodiments, the downsampling filter type is signaled in each intra frame. In some embodiments, the downsampling filter type is signaled in some selected key frames. In some embodiments, the downsampling filter type is signaled in some selected intra frames.
[0092]
[0101] In some embodiments, the downsampling filter type is signaled in a selected interframe. For example, the downsampling filter type is signaled in a selected interframe associated with a temporal layer identifier (ID) lower than a given threshold (e.g., a threshold corresponding to layer 1, layer 2, layer 3, or layer 4). Figure 11 shows an example including nine consecutive frames of a video bitstream with four temporal layers, according to some embodiments. The arrows indicate how the frames reference other frames. For example, frame 4 is a frame using bi-prediction that references frames 0 and 8, and frame 3 is a frame using bi-prediction that references frames 2 and 4. In some embodiments, a video bitstream can include thousands or tens of thousands of frames with tens or hundreds of temporal layers, each with a respective temporal layer ID.
[0093]
[0102] According to some embodiments, the downsampling filter (or filter type) is signaled using one or more binarization methods (eg, binarization schemes, FIG. 10).
[0094]
[0103] In some embodiments, the downsampling filter type is signaled in a fixed-length coding. For example, if there are N filter types, a fixed length of M bits is used to signal the N filters, where 2M is the smallest integer greater than or equal to N.
[0095]
[0104] In some embodiments, the downsampling filter type is signaled in the variable length coding. The downsampling filter type is binarized and signaled.
[0096]
[0105] In some embodiments, the first bin for signaling the downsample filter is for indicating whether the filter type is a 6-tap filter.
[0097]
[0106] In some embodiments, the semantics of the first bin of the downsample filter depend on whether the current frame is detected as screen content. For example, if the current frame is detected as screen content, the first bin for signaling the downsample filter is used to indicate whether the co-located luma samples are used directly without filtering. If the current frame is detected as non-screen content, the first bin for signaling the downsample filter is used to indicate whether the filter type is a 6-tap filter.
[0098]
[0107] In some embodiments, the downsample filter type may be binarized using a unary codeword (e.g., code), as shown in FIG.
[0099]
[0108] In some embodiments, the downsample filter type may be binarized using a truncated unary codeword (e.g., code), as shown in FIG.
[0100]
[0109] In some embodiments, the downsample filter type may be binarized using a truncated rice codeword (e.g., code), as shown in FIG.
[0101]
[0110] In some embodiments, the downsample filter type may be binarized using an Exponential-Golomb codeword (e.g., code), as shown in FIG.
[0102]
[0111] Table 1 shows an example binary table when there are three filter types.
[0103] [Table 1]
[0104]
[0112] 12 shows some logical stages in a particular order, but order-independent stages may be rearranged and other stages may be combined or separated. Any rearrangement or other grouping not specifically described will be apparent to one of ordinary skill in the art, and therefore the ordering and grouping presented herein is not exhaustive. Moreover, it should be recognized that the stages may be implemented in hardware, firmware, software, or any combination thereof.
[0105]
[0113] Reference will now be made to some exemplary embodiments.
[0106]
[0114] (A1) In one aspect, some embodiments include a method for video decoding (e.g., method 1200). In some embodiments, the method is implemented in a computing system (e.g., server system 112) having memory and control circuitry. In some embodiments, the method is implemented in a coding module (e.g., coding module 320). The method includes: (i) receiving a video stream having a sequence of frames, the sequence of frames including one or more key frames, each key frame having a respective downsampling filter type; (ii) according to a determination that a current frame corresponds to a first key frame of the one or more key frames, extracting from the video bitstream a syntax element (e.g., a signal element) associated with a first downsampling filter type associated with the first key frame (e.g., the syntax element is binarized to 1 and 0); (iii) downsampling luma blocks (e.g., luma components, blocks of luma pixels) of the current frame and a predefined set of frames immediately following the current frame using the first downsampling filter type to obtain downsampled frames; and (iv) predicting (e.g., using prediction module 344) chroma blocks (e.g., chroma components, blocks of chroma pixels) of the current frame (e.g., and the predefined set of frames immediately following the current frame) based on the downsampled frames. In some embodiments, the predefined set of frames is the frame that immediately follows the current frame and immediately precedes the second keyframe of the one or more keyframes.
[0107]
[0115] (A2) In some embodiments of A1, the current frame includes a plurality of color components, and the method further includes downsampling a first color component among the plurality of color components of the current frame to obtain a first downsampled color component, and predicting other color components among the plurality of color components based on the first downsampled color component.
[0108]
[0116] (A3) In some embodiments of A1 or A2, the one or more key frames include an intraframe.
[0109]
[0117] (A4) In some embodiments of any of A1-A3, the one or more key frames correspond to every intra-frame of the video stream.
[0110]
[0118] (A5) In some embodiments of any of A1-A4, the one or more key frames include an Instantaneous Decoder Refresh (IDR) frame. An IDR frame is a special type of I-frame. An IDR frame specifies that frames following the IDR frame cannot reference frames preceding it.
[0111]
[0119] (A6) In some embodiments of any of A1-A5, the one or more key frames include one or more inter frames, each of the one or more inter frames being associated with a temporal layer ID that is lower than a predetermined threshold.
[0112]
[0120] (A7) In some embodiments of any of A1-A6, the first downsampling filter type is a filter type corresponding to a luma-to-chroma (CfL) prediction mode for the video stream. In some embodiments, the first downsampling filter type is a filter type corresponding to a prediction mode that uses one color component to predict another color component, and downsampling is required with respect to one or more color components. In some embodiments, the first downsampling filter type is a filter type corresponding to a prediction mode where luma is replaced with one particular color component (e.g., R) and chroma is replaced with another particular color component (e.g., G or B).
[0113]
[0121] (A8) In some embodiments of any of A1-A7, the syntax elements have fixed length coding.
[0114]
[0122] (A9) In some embodiments of A8, the one or more key frames are associated with N downsampling filter types, the fixed length coding corresponds to M bits, and N and M are in the relationship N≦2 M are integers that satisfy the following:
[0115]
[0123] (A10) In some embodiments of any of A1-A9, the syntax elements have variable length coding.
[0116]
[0124] In some embodiments of any of A1-A10, the syntax element includes a first attribute having an attribute value indicating whether the first downsampling filter type is a 6-tap filter.
[0117]
[0125] (A12) In some embodiments of any of A1-A11, the syntax element includes a first attribute having an attribute value determined based on whether the current frame corresponds to detected screen content (e.g., detected by encoding module 340).
[0118]
[0126] (A13) In some embodiments of A12, when the current frame corresponds to detected screen content, the attribute value indicates whether co-located luma samples are used without applying a downsampling filter. When the current frame does not correspond to detected screen content, the attribute value indicates whether the downsampling filter type is a 6-tap filter.
[0119]
[0127] (A14) In some embodiments of any of A1-A13, the syntax elements are binarized using unary codes (see, for example, FIG. 10).
[0120]
[0128] (A15) In some embodiments of any of A1-A14, the syntax element is binarized using a truncated unary code (see, for example, FIG. 10).
[0121]
[0129] (A16) In some embodiments of any of A1-A15, the syntax element is binarized using a truncated rice code (see, for example, FIG. 10).
[0122]
[0130] (A17) In some embodiments of any of A1-A16, the syntax elements are binarized using Exponential-Golomb coding (see, for example, FIG. 10).
[0123]
[0131] In another aspect, some embodiments include a computing system (e.g., server system 112) including control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, wherein the memory stores one or more sets of instructions configured to be executed by the control circuitry, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1-A16 above).
[0124]
[0132] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by control circuitry of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1-A16 above).
[0125]
[0133] Although terms such as "first," "second," etc. may be used herein to describe various elements, it will be understood that these elements are not to be limited by these terms and are merely used to distinguish one element from another.
[0126]
[0134] The terminology used herein is for the purpose of describing particular embodiments only and does not limit the scope of the claims. As used in the description of these embodiments and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. It will also be understood that the term "and / or," as used herein, refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will further be understood that the terms "comprises" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0127]
[0135] The term "if," as used herein, may be interpreted to mean "when," or "upon," or "in response to determining," or "in accordance with a determination," or "in response to detecting" that a stated condition precedent is true, depending on the context. Similarly, the phrase "if determined [the stated condition precedent is true]," or "if [the stated condition precedent is true]," or "when [the stated condition precedent is true]," may be interpreted to mean "upon determining," or "in response to determining," or "in accordance with a determination," or "upon detecting," or "in response to detecting" that a stated condition precedent is true, depending on the context.
[0128]
[0136] The foregoing description has been set forth with reference to specific embodiments for purposes of explanation. However, the above exemplary description is not intended to be exhaustive or to limit the scope of the claims to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of operation and practical application, thereby enabling others skilled in the art.
Claims
1. 1. A method for video decoding implemented in a computing system having one or more processors and a memory, the method comprising: receiving a video stream having a sequence of frames, the sequence of frames including one or more key frames, each key frame having a respective downsampling filter type; in response to determining that the current frame corresponds to a first key frame of the one or more key frames; Extracting from the video stream a syntax element associated with a first downsampling filter type associated with the first keyframe; downsampling a luma block of the current frame and a predefined set of frames immediately following the current frame using the first downsampling filter type to obtain a downsampled frame; predicting chroma blocks of the current frame based on the downsampled frame; A method comprising:
2. the current frame includes a plurality of color components, and the method comprises: downsampling a first color component of the plurality of color components of the current frame to obtain a first downsampled color component; predicting other color components of the plurality of color components based on the first downsampled color component; The method of claim 1 further comprising:
3. The method of claim 1 , wherein the one or more key frames comprise an intraframe.
4. The method of claim 1 , wherein the one or more key frames correspond to every intra-frame of the video stream.
5. The method of claim 1 , wherein the one or more key frames comprise an instantaneous decoder refresh (IDR) frame.
6. The method of claim 1 , wherein the one or more key frames include one or more inter frames, each of the one or more inter frames being associated with a temporal layer ID that is lower than a predetermined threshold.
7. The method of claim 1 , wherein the first downsampling filter type is a filter type corresponding to a luma-to-chroma (CfL) prediction mode for the video stream.
8. The method of claim 1 , wherein the syntax elements have fixed length coding.
9. the one or more keyframes are associated with N downsampling filter types; the fixed length coding corresponds to M bits; N and M are in the relationship N≦2 M are integers that satisfy The method of claim 8.
10. The method of claim 1 , wherein the syntax elements have variable length coding.
11. The method of claim 1 , wherein the syntax element includes a first attribute having an attribute value indicating whether the first downsampling filter type is a 6-tap filter.
12. the syntax element includes a first attribute having an attribute value determined based on whether the current frame corresponds to detected screen content; The method of claim 1.
13. When the current frame corresponds to detected screen content, the attribute value indicates whether co-located luma samples are used without applying a downsampling filter; When the current frame does not correspond to detected screen content, the attribute value indicates whether the downsampling filter type is a 6-tap filter. The method of claim 12.
14. The method of claim 1 , wherein the syntax elements are binarized using unary codes.
15. The method of claim 1 , wherein the syntax elements are binarized using truncated unary codes.
16. The method of claim 1 , wherein the syntax elements are binarized using a truncated rice code.
17. The method of claim 1 , wherein the syntax elements are binarized using Exponential-Golomb coding.
18. a control circuit; Memory and one or more sets of instructions stored in the memory and configured for execution by the control circuitry; wherein the one or more sets of instructions receiving a video stream having a sequence of frames, the sequence of frames including one or more key frames, each key frame having a respective downsampling filter type; in response to determining that the current frame corresponds to a first key frame of the one or more key frames; Extracting from the video stream a syntax element associated with a first downsampling filter type associated with the first keyframe; downsampling a luma block of the current frame and a predefined set of frames immediately following the current frame using the first downsampling filter type to obtain a downsampled frame; predicting chroma blocks of the current frame based on the downsampled frame; 1. A computing system comprising instructions for performing
19. The computing system of claim 18 , wherein the one or more key frames comprise an intraframe.
20. 1. A non-transitory computer-readable storage medium storing one or more sets of instructions configured for execution by a computing device having control circuitry and a memory, the one or more sets of instructions comprising: receiving a video stream having a sequence of frames, the sequence of frames including one or more key frames, each key frame having a respective downsampling filter type; in response to determining that the current frame corresponds to a first key frame of the one or more key frames; Extracting from the video stream a syntax element associated with a first downsampling filter type associated with the first keyframe; downsampling a luma block of the current frame and a predefined set of frames immediately following the current frame using the first downsampling filter type to obtain a downsampled frame; predicting chroma blocks of the current frame based on the downsampled frame; 1. A non-transitory computer-readable storage medium comprising instructions for performing
Citation Information
Patent Citations
Multi-resolution video coding and decoding
JP2004266794A
Chroma block prediction method and device
JP2021535652A
Simplifying Cross-Component Linear Models
JP2022500967A
Simplified cross component prediction
US20210092395A1
Downsampling filter type for chroma blending mask generation
WO2021063419A1