Systems and methods for improved recursive intra region partitioning
By adopting different division methods for chrominance blocks and luminance blocks, the problem of low encoding efficiency in the prior art is solved, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN202480005143.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2024-04-18
- Publication Date
- 2025-07-11
AI Technical Summary
When processing chromaticity blocks and luminance blocks, existing video encoding technologies fail to effectively utilize the differences in texture characteristics between the two, resulting in low encoding efficiency.
Different division methods are used to encode the chromaticity block and the luminance block to limit the further division of the chromaticity blocks to reduce the encoding overhead.
The encoding efficiency is improved through different division methods, and the further division of chroma blocks is reduced, thereby reducing the bandwidth requirement of encoding operations.
Smart Images

Figure CN120303937A_ABST
Abstract
Description
Related Applications
[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 597,597, entitled "Improvement for Recursive Intra Region Partitioning", filed on November 9, 2023, and is a continuation application of and claims the benefit of priority of U.S. Patent Application No. 18 / 633,360, entitled "Systems and Methods for Improved Recursive Intra Region Partitioning", filed on April 11, 2024. Technical Field
[0002] The disclosed embodiments generally relate to video coding and decoding, including but not limited to systems and methods for coding and decoding chrominance blocks and luminance blocks using different partitions. Background Art
[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices transmit and receive digital video data over a communication network or otherwise convey digital video data, and / or store digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited memory resources of the storage device, video coding can be used to compress video data according to one or more video coding standards before transmitting or storing the video data. Video coding and decoding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services.
[0004] Video coding typically uses prediction methods that exploit the redundancy inherent in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard designed to be a successor to HEVC. The ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the confirmed version 1.0.0 with Specification Errata 1 was released. SUMMARY OF THE INVENTION
[0005] The present disclosure particularly describes a set of methods for video (image) compression, and more particularly relates to block partitioning, intra-frame prediction, and partitioning of chrominance blocks and luminance blocks. Some embodiments include using different partitioning for chrominance blocks and luminance blocks within an encoded region encoded using an intra-frame prediction mode, and / or when the size of the encoded region meets a size threshold, using different partitioning for chrominance blocks and luminance blocks within the encoded region. The advantage of using different partitioning for chrominance blocks is to reduce overhead by restricting further partitioning of chrominance blocks, since chrominance blocks typically have less texture than luminance blocks.
[0006] According to some embodiments, a method of video decoding includes: (i) receiving a video bitstream including a plurality of encoded blocks; (ii) identifying an encoded region including two or more of the plurality of encoded blocks based on a first indicator in the video bitstream, wherein each encoded block in the encoded region is encoded using an intra-frame prediction mode; (iii) applying a first partitioning method to luminance blocks in the encoded region; (iv) applying a second partitioning method to chrominance blocks in the encoded region, wherein the second partitioning method is different from the first partitioning method; and (v) reconstructing two or more encoded blocks of the encoded region using the first partitioning method and the second partitioning method.
[0007] According to some embodiments, a method for video encoding includes: (i) receiving video data including a plurality of coding blocks; (ii) identifying a coding region including two or more of the plurality of coding blocks, wherein each coding block in the coding region is to be encoded in an intra prediction mode; and (iii) encoding two or more coding blocks of the coding region into a video bitstream by applying a first partitioning method to the luminance blocks in the coding region and a second partitioning method to the chrominance blocks in the coding region, wherein the second partitioning method is different from the first partitioning method.
[0008] According to some embodiments, a method for processing visual media data includes: (i) obtaining a source video sequence including a plurality of frames; and (ii) performing a conversion between the source video sequence and a video bitstream of the visual media data, wherein the bitstream includes: (a) a plurality of encoded blocks corresponding to the plurality of frames; and (b) an indicator indicating a coding region of a frame among the plurality of frames, wherein the coding region is composed of blocks encoded in an intra prediction mode, wherein the luminance blocks encoded in the video bitstream are partitioned according to a first partitioning method type, and wherein the chrominance blocks encoded in the video bitstream are partitioned according to a second partitioning method type different from the first partitioning method type.
[0009] According to some embodiments, a computing system such as a streaming system, a server system, a personal computer system, or other electronic device is provided. The computing system includes a control circuit and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component. According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions for execution by the computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.
[0010] Thus, devices and systems having methods for encoding and decoding video are disclosed. Such methods, devices, and systems may supplement or replace conventional methods, devices, and systems for video encoding and / or decoding. The features and advantages described in the specification are not necessarily all-inclusive, and in particular, given the drawings, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Additionally, it should be noted that the language used in the specification is primarily selected for readability and guidance purposes and is not necessarily selected to depict or limit the subject matter described herein. Description of the Drawings
[0011] For a more detailed understanding of the present disclosure, a more specific description may be made with reference to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the accompanying drawings only illustrate the relevant features of the present disclosure and are therefore not necessarily considered restrictive, as those skilled in the art will understand that other valid features are permitted in the specification after reading the present disclosure.
[0012] Figure 1 is a block diagram illustrating an example communication system according to some embodiments.
[0013] Figure 2A is a block diagram illustrating example elements of an encoder component according to some embodiments.
[0014] Figure 2B is a block diagram illustrating example elements of a decoder component according to some embodiments.
[0015] Figure 3 is a block diagram illustrating an example server system according to some embodiments.
[0016] Figure 4A 、 Figure 4B 、 Figure 4C and Figure 4D illustrate examples of partitioning an encoded block into regions according to some embodiments.
[0017] Figure 5 illustrates different types of encoded block partitioning according to some embodiments.
[0018] Figure 6A illustrates an example video decoding process according to some embodiments.
[0019] Figure 6B illustrates an example video encoding process according to some embodiments.
[0020] According to conventional practice, the various features illustrated in the accompanying drawings need not be drawn to scale, and throughout the specification and the drawings, like reference numerals may be used to represent like features. Detailed Description
[0021] The present disclosure describes video / image compression techniques, including using different partitioning methods for chrominance blocks and luminance blocks in a coding region when encoding a coding block in the coding region with a first predefined prediction mode and / or when the coding region meets a size threshold. The first predefined prediction mode may be an intra coding mode, an inter coding mode, and / or a hybrid of the intra coding mode and the inter coding mode. When partitioning a block (e.g., recursively or using a predefined partitioning pattern) into one or more sub-blocks of equal or smaller size, a flag may be signaled to indicate whether the chrominance block can be further partitioned. The advantage of using different partitioning methods for chrominance blocks is to reduce overhead (e.g., fewer coding operations and bandwidth are required) by restricting further partitioning of chrominance blocks because chrominance blocks generally have less texture than luminance blocks. Example Systems and Devices
[0022] Figure 1 FIG. 6 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, used with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0023] The source device 102 includes a video source 104 (e.g., a camera component or a media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. Compared with the video stream from the video source 104, the video stream from the video source 104 may have a high data volume. Since the encoded video bitstream 108 has a lower data volume (less data) compared with the video stream from the video source, the encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared with the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video data to the network 110).
[0024] One or more networks 110 represent any number of networks that convey information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0025] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content such as an encoded video stream from the source device 102). The server system 112 includes an encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the encoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the encoder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the encoder component 114 is configured to decode the encoded video stream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video stream 108. In some embodiments, the server system 112 functions as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to trim the encoded video stream (108) to crop potentially different streams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.
[0026] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0027] The source device and / or the multiple electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.
[0028] In an example operation of communication system 100, source device 102 transmits encoded video bitstream 108 to server system 112. For example, source device 102 may encode a stream of pictures captured by the source device. Server system 112 receives encoded video bitstream 108 and may use encoder component 114 to decode and / or encode encoded video bitstream 108. For example, server system 112 may apply an encoding that is more optimized for network transmission and / or storage to the video data. Server system 112 may transmit encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more of electronic devices 120. Each electronic device 120 may decode encoded video data 116 to recover and optionally display video pictures.
[0029] Figure 2A FIG. is a block diagram illustrating example elements of encoder component 106 according to some embodiments. Encoder component 106 receives video data (e.g., a source video sequence) from video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is part of a device different from encoder component 106). Video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that produce motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. Those of ordinary skill in the art can readily understand the relationship between pixels and samples.
[0030] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, the encoder component 106 is configured to perform a conversion between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. Parameters set by the controller 204 may include parameters related to rate control (e.g., picture skipping, quantizer, and / or λ values of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as they may pertain to the encoder component 106 optimized for a particular system design.
[0031] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data (when the compression between the symbols and the encoded video bitstream is lossless). This reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream results in a bit-exact result independent of the decoder location (local or remote), the content in this reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. Thus, when prediction is used during decoding, the prediction part of the encoder interprets the same sample values as the decoder interprets as reference picture samples.
[0032] The operation of the decoder 210 may be the same as the operation of a remote decoder (such as the decoder component 122), which is described in detail below in connection with Figure 2B which. However, briefly referring to Figure 2B , since the symbols are available and the entropy encoder 214 and the parser 254 encoding / decoding the symbols into the encoded video sequence can be lossless, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.
[0033] Except for the parsing / entropy decoding present in the decoder, the decoder techniques described herein may also exist in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operation. The description of encoder techniques may be abbreviated as they may be the inverse process of decoder techniques.
[0034] As part of its operation, source encoder 202 may perform motion-compensated predictive coding that predictively encodes an input frame by referring to one or more previously encoded frames designated as reference frames from a video sequence. In this way, encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of one or more reference frames that may be selected as one or more predictive references for the input frame. Controller 204 may manage the encoding operation of source encoder 202, including, for example, setting parameters and sub-group parameters for encoding video data.
[0035] Decoder 210 may decode the encoded video data of a frame that may be designated as a reference frame based on the symbols created by source encoder 202. The operation of encoding engine 212 may advantageously be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 2A (not shown herein)), the reconstructed video sequence may be a replica of the source video sequence with some errors. Decoder 210 replicates the decoding process that may be performed by a remote video decoder on the reference frame and may cause the reconstructed reference frame to be stored in reference picture memory 208. In this way, encoder component 106 locally stores a copy of the reconstructed reference frame that has the same content (in the absence of transmission errors) as the reconstructed reference frame that would be obtained by the remote video decoder.
[0036] Predictor 206 may perform a prediction search on encoding engine 212. That is, for a new frame to be encoded, predictor 206 may search reference picture memory 208 for sample data (as candidate reference pixel blocks) that may be used as an appropriate predictive reference for the new picture or certain metadata such as reference picture motion vectors, block shapes, etc. Predictor 206 may operate on a per-pixel-block sample basis to find an appropriate predictive reference. As determined by the search results obtained by predictor 206, the input picture may have a predictive reference extracted from a plurality of reference pictures stored in reference picture memory 208.
[0037] The outputs of all of the above functional units may be entropy encoded in entropy encoder 214. Entropy encoder 214 converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those of ordinary skill in the art (such as Huffman coding, variable length coding, and / or arithmetic coding).
[0038] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer one or more encoded video sequences created by the entropy encoder 214 to prepare them for transmission via a communication channel 218, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to merge the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown). In an embodiment, the transmitter may transmit additional data along with the encoded video. The video encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data, such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set segments, and the like.
[0039] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a particular encoded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). Intra pictures may be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those of ordinary skill in the art are aware of these variations of I pictures and their respective applications and characteristics, and thus are not repeated herein. Predictive pictures may be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures may be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multi-predictive pictures may use more than two reference pictures and associated metadata to reconstruct a single block.
[0040] Source pictures may typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and encoded on a block-by-block basis. Blocks may be predictively encoded with reference to other (already encoded) blocks determined by the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively encoded, or they may be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be non-predictively encoded with reference to one previously encoded reference picture, either via spatial prediction or via temporal prediction. Blocks of a B picture may be non-predictively encoded with reference to one or two previously encoded reference pictures, either via spatial prediction or via temporal prediction.
[0041] Video can be captured as multiple source pictures (video pictures) in a time series. Intra-picture prediction (usually abbreviated as intra-frame prediction) exploits the spatial correlation within a given picture, and inter-picture prediction exploits the (temporal or other) correlation between pictures. In an example, a particular picture under encoding / decoding, referred to as the current picture, is partitioned into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. In the case of using multiple reference pictures, the motion vector points to the reference block in the reference picture and can have a third dimension identifying the reference picture.
[0042] The encoder component 106 can perform encoding operations according to a predetermined video coding technique or standard (such as any described herein). In its operation, the encoder component 106 can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0043] Figure 2B is a block diagram illustrating example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter unit 256 and configured to transmit data to the display 124 (e.g., via a wired or wireless connection).
[0044] In some embodiments, the decoder component 122 includes a receiver coupled to the channel 218 and configured to receive data from the channel 218 (e.g., via a wired or wireless connection). The receiver can be configured to receive one or more encoded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence can be received from the channel 218, which can be a hardware / software link to a storage device storing the encoded video data. The receiver can receive the encoded video data with other data, such as encoded audio data and / or auxiliary data streams, which can be forwarded to their respective entities using an entity (not depicted). The receiver can separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data can be included as part of one or more encoded video sequences. The additional data can be used by the decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0045] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaling / inverse transform unit 258, intra picture prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, decoder component 122 may be implemented at least partially in software.
[0046] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to counter network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 internal to decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided external to decoder component 122 (e.g., to counter network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, buffer memory 252 may not be required, or this buffer memory 252 may be small. For use on a best-effort packet network such as the Internet, buffer memory 252 may be necessary, which may be relatively large and / or have an adaptive size, and may be implemented at least partially in an operating system or similar element (not shown) external to decoder component 122.
[0047] The parser 254 is configured to reconstruct symbols 270 from an encoded video sequence. The symbols can include, for example, information for managing the operation of decoder components 122, and / or information for controlling rendering devices such as display 124. The control information for one or more rendering devices can be, for example, in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set segments (not depicted). The parser 254 can parse (entropy decode) the encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technology or standard, and can follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the set. Subgroups can include Group of Pictures (GOP), pictures, slices, strips, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser 254 can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0048] The reconstruction of symbols 270 can involve multiple different units, depending on the type of the encoded video picture or a portion thereof (such as: inter and intra pictures, inter and intra blocks) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information, which is parsed by the parser 254 from the encoded video sequence. For clarity, this subgroup control information flow between the parser 254 and the multiple units is not depicted below.
[0049] The decoder component 122 can be conceptually subdivided into multiple functional units. In some implementations, many of these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, the conceptual subdivision of the functional units is maintained here.
[0050] The Scaling / Inverse Transform Unit 258 receives quantized transform coefficients and control information (such as the transform to be used, block size, quantization factor, and / or quantization scaling matrix) as one or more symbols 270 from the parser 254. The Scaling / Inverse Transform Unit 258 can output a block including sample values that can be input to the aggregator 268.
[0051] In some cases, the output samples of the scaling / inverse transform unit 258 belong to intra-coded blocks; that is: blocks that do not use prediction information from previously reconstructed pictures, but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 can use the already reconstructed surrounding information obtained from the current (partially reconstructed) picture in the current picture memory 264 to generate a block having the same size and shape as the block under reconstruction. The aggregator 268 can add, on a per-sample basis, the prediction information already generated by the intra-picture prediction unit 262 to the output sample information provided by the scaling / inverse transform unit 258.
[0052] In other cases, the output samples of the scaling / inverse transform unit 258 belong to inter-coded and possibly motion-compensated blocks. In such cases, the motion compensation prediction unit 260 can access the reference picture memory 266 to obtain samples for prediction. After motion-compensating the obtained samples according to the symbol 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaling / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The address from which the motion compensation prediction unit 260 obtains the prediction samples from within the reference picture memory 266 can be controlled by a motion vector. The motion vector can be used in the form of the symbol 270 for the motion compensation prediction unit 260, and the symbol 270 can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values obtained from the reference picture memory 266 when using sub-sampled accurate motion vectors, motion vector prediction mechanisms, and the like.
[0053] The output samples of the aggregator 268 can be processed by various loop filtering techniques in the loop filter unit 256. The video compression technique can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and are available to the loop filter unit 256 as the symbol 270 from the parser 254, but can also be in response to meta-information obtained during the decoding of the previous (in decoding order) parts of the encoded picture or the encoded video sequence, and in response to the previously reconstructed and loop-filtered sample values. The output of the loop filter unit 256 can be a sample stream that can be output to a rendering device such as the display 124 and stored in the reference picture memory 266 for future inter-picture prediction use.
[0054] Once some encoded pictures are reconstructed, they can be used as reference pictures for future prediction. Once an encoded picture is reconstructed and the encoded picture has been identified as a reference picture (e.g., identified by parser 254), the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0055] Decoder component 122 can perform decoding operations according to a predetermined video compression technique that can be recorded in a standard (such as any of the standards described herein). The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that it complies with the syntax of the video compression technique or standard as specified in the video compression technique document or standard and particularly in the profile document therein. Also, in order to conform to some video compression technique or standard, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the assumptions of the hypothetical reference decoder (HRD) specification and metadata signaled in the encoded video sequence through HRD buffer management can further limit the limits set by the level.
[0056] Figure 3 is a block diagram illustrating a server system 112 according to some embodiments. Server system 112 includes control circuitry 302, one or more network interfaces 304, memory 314, user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, control circuitry 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuitry includes a field programmable gate array, a hardware accelerator, and / or an integrated circuit (e.g., an application specific integrated circuit).
[0057] The network interface 304 can be configured to interface with one or more communication networks (e.g., wireless, wired, and / or optical networks). The communication network can be local, wide area, metropolitan area, vehicular, and industrial, real-time, delay-tolerant, etc. Examples of communication networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANBus, etc. Such communication can be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., to other computer systems using local or wide area digital networks). Such communication can include communication to one or more cloud computing networks.
[0058] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 can include one or more of the following: keyboard, mouse, touchpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. The output device 308 can include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or monitors), etc.
[0059] The memory 314 can include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices remote from the control circuit 302. Alternatively, the memory 314 or the non-volatile solid-state storage device within the memory 314 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: · An operating system 316, which includes programs for handling various basic system services and for performing hardware-related tasks; · A network communication module 318, which is used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); · A codec module 320 for performing various functions regarding encoding and / or decoding data (e.g., video data). In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following: ο A decoding module 322, which is configured to perform various functions regarding decoding encoded data, such as those previously described with respect to decoder component 122; and ο An encoding module 340, which is configured to perform various functions regarding encoding data, such as those previously described with respect to encoder component 106; and · A picture memory 352, which is configured to store pictures and picture data, for example, for use by the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.
[0060] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to parser 254), a transform module 326 (e.g., configured to perform various functions previously described with respect to the scaling / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to the loop filter unit 256).
[0061] In some embodiments, the encoding module 340 includes an encoding module 342 (e.g., configured to perform various functions previously described with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 a subset of the modules shown. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.
[0062] Each of the above-identified modules stored in the memory 314 corresponds to a set of instructions for performing the functions described herein. The above-identified modules (e.g., instruction sets) need not be implemented as separate software programs, routines, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but rather uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the above-identified modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above.
[0063] Although Figure 3 a server system 112 according to some embodiments is illustrated, however Figure 3Rather, it is intended to be a functional description of various features that may exist in one or more server systems, rather than a structural schematic of the embodiments described herein. In practice, items shown separately may be combined and some items may be separated. For example, Figure 3 Some of the items shown separately in Figure 3 may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement server system 112, and how the features are distributed among them, will vary from one implementation to another and may optionally depend in part on the amount of data traffic processed by the server system during peak usage periods as well as during average usage periods. Example Coding and Decoding Techniques
[0064] The encoding processes and techniques described below may be performed at the devices and systems described above (e.g., source device 102, server system 112, and / or electronic device 120). Hereinafter, a block (or sub-block) refers to an encoding block having an encoding block (such as a super block, or a maximum coding unit, or a coding tree block), a prediction block, a transform block, or a filtering unit. For example, a sub-block of block A refers to a block whose area is completely contained within block A.
[0065] Hereinafter, a block region refers to a specific block-shaped region containing one or more blocks. A block size group refers to the group to which the current block belongs. Blocks having multiple sizes may belong to the same group. A block size group is a set of multiple block sizes. For example, multiple block sizes that are similar to each other (e.g., the number of samples is similar, or the difference between the width and / or height is close) may be grouped into the same block size group.
[0066] Turning to block partitioning for encoding and decoding, the general partitioning may start from a base block (e.g., a super block or a root node) and may follow a predefined set of rules, partitioning structure, and / or scheme. The partitioning may be hierarchical and recursive. After partitioning or dividing the base block using any one or a combination of the example partitioning procedures or other procedures described below, a set of final partitions or encoding blocks may be obtained. Each of these partitions may be at a certain level in the partitioning hierarchy and may have various shapes. Each of the partitions may be referred to as an encoding block (CB). Such partitions are called encoding blocks because they can form units, basic encoding / decoding decisions can be made for these units, and encoding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the encoding block partitioning structure of the tree.
[0067] A coding block can be a luminance coding block or a chrominance coding block. The hierarchical structures of all color channels can be collectively referred to as a coding tree unit (CTU). The partitioning pattern or structure for various color channels in a CTU can be the same or can be different. In some embodiments, the partitioning tree scheme or structure for the luminance channel and the chrominance channel can be different (the luminance channel and the chrominance channel can have separate coding tree structures). When applying separate coding partitioning tree structures or patterns, the luminance channel can be divided into luminance CBs by one coding partitioning tree structure, and the chrominance channel can be divided into chrominance CBs by another coding partitioning tree structure. In some embodiments of the partitioning structure, the CTU size can be set to 128×128 luminance samples, corresponding to two 64×64 chrominance sample blocks (when considering and using exemplary chrominance subsampling).
[0068] A region or coding region can be used to refer to any level in any of the above-described partitioning schemes or any other partitioning scheme not specifically described above. Thus, a region can be a frame, a slice, a superblock, a macroblock, a subblock, a prediction block, etc. For example, a region can be any partitioning level of a recursive partitioning scheme. A region can be at a leaf level or a non-leaf level of a particular partitioning scheme. A leaf-level region is a region that is not further partitioned. On the other hand, a non-leaf-level region is further partitioned into at least two sub-regions, each sub-region can be at a leaf level, or can be at a non-leaf level and thus can be further partitioned. A leaf-level region is predicted as a whole using a particular prediction mode. For example, a leaf-level region can be encoded inter-frame or intra-frame. Optionally, if an intra-inter prediction mode is allowed, the leaf-level region can additionally be encoded using an intra-inter coding mode. The intra-inter coding / prediction mode refers to a coding mode that uses both intra-prediction methods and inter-prediction methods to generate a prediction block. For example, a prediction mode that derives a prediction block as a weighted sum of an intra-prediction block and an inter-prediction block.
[0069] In some embodiments, when a coding region includes coding blocks that are all intra-coded (e.g., coded using an intra-prediction mode), the luminance component and the chrominance component within the block region can be coded using different block partitioning modes, and optionally, the partitioning modes of the luminance component and the chrominance component within the region can be signaled separately. In some embodiments, if a coding region is not all intra-coded, the luminance component and the chrominance component of the region can follow the same partitioning mode, and optionally, the partitioning mode used for the luminance component and the chrominance component of the region can be signaled jointly and shared.
[0070] Figure 4AIllustrated is an example of applying a first partitioning method to a luminance block and a second partitioning method to a chrominance block in the coding region based on a prediction mode (e.g., an intra prediction mode, an inter prediction mode, or a mixture of an intra prediction mode and an inter prediction mode) for coding the coding region. Figure 4A Shown is an example where a top region 402 (e.g., a superblock), as an example partitioning scheme, is partitioned into regions or blocks at four levels or depths labeled as level 1 to level 4. The leaf level regions (sometimes also referred to as leaf partition tree nodes) include regions 412, 416, 418, 422, 424, 432, and 434. In Figure 4A it, the top region 402 includes a signaling syntax element (sometimes also referred to as a flag hereinafter), which indicates that a mixed intra prediction mode and inter prediction mode are used to code the coding blocks in the top region 402. The presence of the prediction mode flag of the top region 402 indicates that one or more first coding blocks (e.g., a first region) in the top region 402 are coded in the intra mode, and one or more second coding blocks (e.g., a second region) in the top region 402 are coded in the inter mode. The top region 402 is partitioned into two regions, region 410 and region 420, both of which are partitions at level 2 and at a depth of 1 from the top region 402.
[0071] In some embodiments, the signaling syntax element is a region type flag (such as intra_region_flag, inter_region_flag, mixed_region_flag, intra - inter_region_flag, or other region type flags), which indicates the type of prediction mode for coding all coding blocks within the coding region. For example, intra_region_flag indicates that all coding blocks included in the corresponding region in a frame of an inter prediction type (signaled by a higher - level syntax such as frame - level syntax) are intra - coded (e.g., coded using an intra prediction mode). Similarly, inter_region_flag indicates that all coding blocks included in the corresponding region are inter - coded (e.g., coded using an intra prediction mode). In contrast, mixed_region_flag indicates that some coding blocks in the corresponding region are intra - coded and some are inter - coded, while intra - inter_region_flag indicates that all coding blocks included in the corresponding region are coded as a weighted sum of one or more inter prediction blocks and one or more intra prediction blocks.
[0072] Figure 4A Regions including "mixed", "intra", or "inter" labels in it represent regions with corresponding region type flags. In contrast, Figure 4AIn the blocks illustrated in the figure, the label indicating the region type flag may optionally be absent, or the decoder does not need to detect the presence of the region type flag for the corresponding sub-region or block (e.g., skips parsing). In Figure 4A In Figure 4A , the top region 402 includes a mixed_region_flag. The presence of the mixed_region_flag in the top region 402 indicates that one or more first encoded blocks (e.g., the first region) in the top region 402 are intra-coded, and one or more second encoded blocks (e.g., the second region) in the top region 402 are inter-coded. In this example, the top region 402 can be divided into two regions, region 410 and region 420, both of which are partitions at level 2 and have a depth of 1 from the top region 402. For example, the top region 402 is divided according to a predetermined partitioning scheme.
[0073] Region 420 has an intra_region_flag, indicating that all encoded blocks within this region 420 are intra-coded. Region 420 is further divided into region 412 and region 414. Region 414 (a level 3 partition with a depth of 2 from the top region 402) is further divided into region 416 and region 418, which are partitions at level 4 and have a depth of 3 from the top region 402. 412, 414, 416, and 418 may not have intra_region_flags because they are all partitions of region 420 and have already been marked for intra-coding at region 420. Therefore, the decoding component will optionally not perform any additional determination of intra_region_flags when parsing any partition below region 420 (including regions 412, 414, 416, and 418). Optionally, regions 412, 416, and 418, which are also leaf partitions, may not include any other prediction mode indicators because they are intra-coded, as indicated by the presence of the intra_region_flags at region 420.
[0074] Region 410 has a mixed_region_flag, indicating that one or more first encoded blocks (e.g., the first region) in region 410 are intra-coded, and one or more second encoded blocks (e.g., the second region) in region 410 are inter-coded. Region 410 is further divided into region 422, region 424, and region 426, which are level 3 partitions. Region 422 has an inter_region_flag, indicating that all encoded blocks within region 422 are inter-coded, and region 424 has an intra_region_flag, indicating that all encoded blocks within region 424 are intra-coded.
[0075] There is an inter_region_flag in region 426, indicating that all coding blocks within region 422 are inter-frame coded. Additionally, region 426 is divided into two level-4 partitions: region 432 and region 434. Region 432 and region 434 may not have inter_region_flags as they are both marked as inter-frame coded at region 426. In some embodiments, flags are used to indicate that a region or block is coded in an intra-inter coding mode. The intra-inter coding mode refers to a coding mode that generates a predicted block using both intra prediction and inter prediction. For example, a prediction mode that derives a predicted block as a (e.g., weighted sum) of an intra-predicted block and an inter-predicted block.
[0076] Regions 422, 424, 432, 434, 412, 416, and 418 are leaf partitions. Optionally, the decoder does not determine whether there are any region type flags in the leaf partitions, and / or reads the leaf level prediction mode indicators of these leaf partitions to determine their respective prediction modes. For example, leaf partitions under a region with a mixed_region_flag may optionally not include region type flags, but instead include prediction mode indicators for the decoder to determine the prediction mode of the leaf partition. Conversely, since regions 432, 434 are leaf partitions under an inter-frame coded region, region type flags and prediction mode indicators may optionally not be signaled for these regions (e.g., all coding blocks in regions 432, 434 are inferred to be inter-frame coded blocks). Similarly, since regions 416 and 418 are leaf partitions under an intra-frame coded region, region type flags and prediction mode indicators may optionally not be signaled for these regions (e.g., all coding blocks in regions 432 and 434 are inferred to be intra-frame coded blocks).
[0077] Figure 4B The figure illustrates an example partitioning pattern corresponding to the partitioning scheme described above with respect to Figure 4A For example, the top region 402 is vertically divided into two second-level regions 410 and 420 of equal size. The second-level region 420 is further horizontally divided into two third-level regions 412 and 414 of equal size. Region 414 is further horizontally divided into two fourth-level regions 416 and 418 of equal size. The second-level region 410 is further divided into three third-level regions 422, 424, and 426 (e.g., by Figure 5(as illustrated by partition 506 in). Region 426 is further horizontally divided into two equal-sized regions 432 and 434, which are level 4 partitions. In this example, the diagonal shaded regions 420, 412, 414, 416, 418, and 424 are all intra-coded, while the cross-hatched regions 422, 426, 432, and 434 are all inter-coded. In this example, the mixed_region_flag of region 402 can optionally be signaled with a predefined value to indicate that region 402 includes both coded blocks that are intra-coded and coded blocks that are inter-coded. Similarly, an intra_region_flag can be signaled for each of the regions 420, 412, 414, 416, 418, and 424 that are all intra-coded. For region 420, the intra_region_flag is signaled to indicate that all subsequent partitions are intra-coded. Thus, regardless of whether regions 412, 414, 416, 418 are leaf-level partitions, they do not include lower-level intra_region_flags. This can help reduce signaling overhead. Similarly, an inter_region_flag can be used to indicate that regions 422, 426, 432, and 434 are all inter-coded. For the top region 402 and region 410, the mixed_region_flag can be used to indicate that some coded blocks are inter-coded while other coded blocks are intra-coded.
[0078] Figure 4A Each coded block in the coded regions depicted can be a luminance coded block (also referred to as a "luma block") or a chrominance coded block (also referred to as a "chroma block"). In some embodiments, a first partitioning type is applied to luma blocks, while a second partitioning type different from the first partitioning type is applied to corresponding chrominance blocks (e.g., co-located chrominance blocks).
[0079] Figure 5 Illustrates various partitioning types and partition structures according to some embodiments. Figure 5 The partitioning types and / or structures illustrated in can be used in conjunction with the regions and flags previously described with respect to Figure 4A and Figure 4B described. A predefined example 10-way partitioning structure allows for recursive partitioning to form a partitioning tree. The root block can start from a predefined level (e.g., starting from a base block at the 128×128 or 64×64 level). Figure 5The partitions 501, 502, 503, 504, 505, 506, 507, 508, 509, and 510 shown in Figure 5 include various 2:1 / 1:2 rectangular partitions and 4:1 / 1:4 rectangular partitions. The partition types can also include partitions from a three-way partitioning scheme, which can be implemented vertically, as shown in partition 511, or horizontally, as shown in partition 512. Although Figure 5 the example partition ratios for partitions 511 and 512 are shown as 1:2:1 in
[0080] Figure 5 Figure 5 it, other ratios can be used. Figure 4B The partitioning illustrated for region 402 in Figure 4C can refer to the partitioning of the luma blocks in region 402, and the corresponding partitioning for the chroma blocks is shown in partitioning 450c of
[0081] Since region 420 is an intra region, further partitioning of the chroma blocks in region 420 (420c) is not allowed. In some embodiments, high-level syntax is signaled into the bitstream (and received at the decoder) to indicate whether chroma blocks within the intra region (optionally inter-intra) will be further partitioned. Thus, the high-level syntax provides an option to allow partitioning of chroma blocks, rather than always restricting the partitioning of chroma blocks. Figure 4B the partitioning illustrated in region 402 of Figure 4DThe corresponding partitioning for chrominance blocks is shown in partition 452c of the region. For example, chrominance blocks may be allowed to share the luminance block partitioning up to the depth of level 3 partitioning, but further partitioning beyond this depth is not allowed for chrominance blocks. In partition 452c, the chrominance blocks in region 420 are partitioned into regions 412c and 414c, but partitioning of chrominance blocks corresponding to the luminance blocks in regions 416, 418, 432, and 434 is not allowed. For example, compared to luminance blocks, chrominance blocks may have less texture, so chrominance blocks do not need to be partitioned to the same depth as luminance blocks.
[0082] In some embodiments, both luminance blocks and chrominance blocks within the intra-frame region are further partitioned, but for further partitioned chrominance blocks, only a subset of restricted partition types is allowed. In some embodiments, for chrominance blocks within the intra-frame region (optionally resulting from the partitioning of an inter-frame), only a limited set of partition patterns is allowed (e.g., horizontal (e.g., partition 508) or vertical binary partitioning (e.g., partition 507) or quadtree partitioning (e.g., partition 510)). For example, non-uniform 4-way partitioning (e.g., resulting in the partitions 513, 514, 515, and 516 shown in Figure 5 or ternary partitioning (e.g., resulting in the partitions 511 and 512 shown in Figure 5 is not allowed for chrominance blocks, but non-uniform 4-way partitioning or ternary partitioning is allowed for luminance blocks.
[0083] In some embodiments, if the block width in the intra-frame region is greater than the block height (e.g., regions 416 and 418), then for chrominance blocks in this intra-frame region (optionally resulting from the partitioning of an inter-frame), only vertical partitioning (e.g., Figure 5 the partition 507 shown in Figure 5in the partition 508). In some embodiments, if the block width in the intra region is equal to the block height (e.g., regions 732 and 734), for the chrominance blocks within that intra region, only horizontal (e.g., partition 508) or vertical binary partitioning (e.g., partition 507) or quadtree partitioning (e.g., partition 510) is allowed. In some embodiments, for chrominance blocks, the subset of allowed partitioning types depends on the luma partitioning mode. For example, if quadtree partitioning (e.g., partition 510) is used to partition a luma block, only horizontal partitioning (e.g., partition 508) is allowed for the corresponding chrominance block; however, if ternary partitioning (e.g., partitions 511 and 512) is used to partition a luma block, only vertical partitioning (e.g., partition 507) is allowed for the corresponding chrominance block. For example, when the luma block and the chrominance block have different partitioning types, a flag for signaling the corresponding partitioning mode must be signaled. To reduce overhead, for chrominance blocks, only a subset of partitioning modes is allowed, providing some flexibility in partitioning chrominance blocks.
[0084] In some embodiments, when an encoding region is partitioned into multiple sub-regions, if the block width or block height of a sub-region is equal to or less than a threshold (e.g., a length threshold or an area threshold), the syntax indicating whether to encode all the encoding blocks within the encoding region using a predefined prediction mode is not parsed, but rather is derived as a default value. For example, the size threshold can be a sample size length of 4, and optionally, the default value of the predefined prediction mode is a mixture of intra prediction mode and inter prediction mode. For example, if a 32×32 block is partitioned into four regions with sizes 4×32, 8×32, 16×32, and 4×32 due to the first region and the last region having a width (or height) of 4, the syntax indicating whether all the encoding blocks within the encoding region are in the predefined prediction mode is not parsed, but rather is derived as a default value (e.g., a mixture of intra prediction mode and inter prediction mode), attempting to derive whether the predefined prediction mode is intra, inter, or a mixture. For example, when the encoding region is smaller than a certain size threshold, all the encoding blocks within that region are either all intra-encoded or all inter-encoded. For example, because the encoding region is very small, the overhead saved by not signaling the prediction mode may also be very small.
[0085] In some embodiments, when dividing an encoding region into multiple sub-regions, if the block width, block height, the product of the block width and the block height, the maximum value of the block width and the block height, or the minimum value of the block width and the block height of a sub-region is equal to or greater than a threshold (e.g., a length threshold or an area threshold), then the syntax indicating whether all sub-blocks within the first block are in a predefined prediction mode is not parsed, but rather exported as a default value. For example, the default value of the predefined prediction mode is a mixture of intra prediction mode and inter prediction mode. For example, the threshold is greater than or equal to 32. In some embodiments, when the encoding region is too large (e.g., greater than the threshold), the prediction mode is not signaled, but rather the prediction mode is inferred as hybrid encoding, and the encoding region is allowed to be divided into smaller regions that can subsequently be signaled. When the encoding region is too small, the prediction mode is also not signaled, but rather the prediction mode is inferred as hybrid encoding, but the encoding region is not allowed to be further divided.
[0086] In some embodiments, whether the current block is within an intra region provides context for a probability model that is used to encode a flag characterizing the block partition type of the current block. The partition types include Figure 5 any of the example partitions shown. For example, the probability of a particular block partition type depends on whether the encoding region is an inter region or an intra region, and considering the region flag can improve the context for entropy encoding and also improve the efficiency of entropy encoding.
[0087] Figure 6A is a flowchart illustrating a method 600 for decoding video according to some embodiments. The method 600 can be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions for execution by the control circuit. In some embodiments, the method 600 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.
[0088] The system receives (602) a video bitstream including a plurality of encoded blocks. The system identifies (604) an encoding region including two or more of the plurality of encoded blocks based on a first indicator in the video bitstream, wherein each encoded block in the encoding region is encoded in an intra prediction mode. The system applies (606) a first partitioning method to the luminance blocks in the encoding region. The system applies (608) a second partitioning method to the chrominance blocks in the encoding region, wherein the second partitioning method is different from the first partitioning method. The system uses the first partitioning method and the second partitioning method to reconstruct (610) two or more of the encoded blocks in the encoding region. In some embodiments, method 600 further includes the various partitioning embodiments described above.
[0089] Figure 6BFIG. is a flow chart of a method 650 for encoding video according to some embodiments. The method 650 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions for execution by the control circuit. In some embodiments, the method 650 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.
[0090] The system receives (652) video data including a plurality of video blocks. The system identifies (654) an encoding region including two or more video blocks of the plurality of encoded blocks, wherein each video block in the encoding region will be encoded in an intra prediction mode. The system encodes two or more video blocks of the encoding region into a video bitstream by: adopting a first partitioning manner for the luminance blocks of the encoding region and a second partitioning manner for the chrominance blocks of the encoding region, wherein the second partitioning manner is different from the first partitioning manner. As previously mentioned, the encoding process can be a mirror of the decoding process (e.g., the partitioning embodiments described above). For the sake of brevity, these details are not repeated here.
[0091] Although Figure 6A and Figure 6B a number of logical stages are illustrated in a particular order, stages that do not depend on order can be reordered, and other stages can be combined or split. Certain reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, so the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that the stages can be implemented in hardware, firmware, software, or any combination thereof.
[0092] Now turning to some example embodiments.
[0093] (A1)In one aspect, some embodiments include a method for video decoding (e.g., method 600). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) receiving a video bitstream including a plurality of coded blocks; (ii) identifying a coded region including two or more of the plurality of coded blocks based on a first indicator in the video bitstream, wherein each coded block in the coded region is coded in an intra prediction mode; (iii) applying a first partitioning scheme to the luma blocks in the coded region; (iv) applying a second partitioning scheme to the chroma blocks in the coded region, wherein the second partitioning scheme is different from the first partitioning scheme; and (v) reconstructing two or more of the coded blocks in the coded region using the first partitioning scheme and the second partitioning scheme. For example, when a first block is recursively partitioned into one or more equally sized or smaller sub-blocks, at least one flag is received at the decoder side to indicate whether all sub-blocks within the first block are coded in a first predefined prediction mode. The first predefined prediction mode may be an intra coding mode, and / or an inter coding mode, and / or a mixture of an intra coding mode and an inter coding mode. As an example, when a parsed syntax indicates that all blocks within a block area are intra coded blocks, the luma blocks within the intra area may be further partitioned, but further partitioning of the chroma blocks within the block area is not allowed. In some embodiments, the first partitioning scheme and the second partitioning scheme are applied in accordance with the identification of the coded region (e.g., when the coded region is identified or in response to the identification of the coded region). For example, if the coded region is not identified, the first partitioning scheme and the second partitioning scheme may not be applied.
[0094] (A2)In some embodiments of A1, applying the second partitioning scheme to the chroma blocks in the coded region includes: applying the first partitioning scheme for luma blocks to the chroma blocks to a first depth, and prohibiting further partitioning of the chroma blocks after exceeding the first depth. As an example, when a parsed syntax indicates that all blocks within a block area are intra coded blocks, the luma blocks within the intra area may be further partitioned, and the chroma blocks within the block area share the luma block partitioning up to a certain depth and then further partitioning of the chroma blocks is not allowed.
[0095] (A3)In some embodiments of A1 or A2, the second partitioning scheme for the chroma blocks in the coded region is selected from a subset of the partitioning types available for partitioning luma blocks. For example, when a parsed syntax indicates that all blocks within a block area are intra coded blocks, both the luma blocks and the chroma blocks within the intra area may be further partitioned, but for the chroma blocks to be further partitioned, only a subset of the partitioning types is allowed.
[0096] (A4)In some embodiments of A3, a subset of partition types is identified based on one or more of the block size, prediction mode, and partition type of a luminance block. For example, for a chrominance block within an intra region in an inter frame, only horizontal or vertical binary partitioning or quadtree partitioning is allowed. For example, if the block width is greater than the block height, only vertical partitioning is allowed for a chrominance block within an intra region in an inter frame. For example, if the block height is greater than the block width, only horizontal partitioning is allowed for a chrominance block within an intra region in an inter frame. For example, if the block width is equal to the block height, only horizontal or vertical binary partitioning or quadtree partitioning is allowed for a chrominance block within an intra region in an inter frame. For example, the subset of allowed partition types depends on the luminance partition mode.
[0097] (A5)In some embodiments of any one of A1 to A4, the method further includes: partitioning a frame of a video bitstream to obtain a second coding region; and when the second coding region has a size that meets one or more criteria, deriving a first element indicating whether all coding blocks within the second coding region are coded in a predefined prediction mode. For example, when a first block is partitioned into multiple sub-blocks, if the block size of the first block is equal to or greater than a threshold, instead of parsing the syntax indicating whether all sub-blocks within the first block are in a predefined prediction mode, it is derived as a default value. The block size can refer to but is not limited to the block width, block height, the product of the block width and the block height, and the maximum of the block width and the block height. As an example, the default value indicates a mixture of intra coding mode and inter coding mode. In some embodiments, based on the determined fact that the second coding region has a size that meets one or more criteria, the first element indicating whether all coding blocks within the second coding region are coded in a predefined prediction mode is derived based on the encoded information (e.g., at the decoder). In some embodiments, based on the determined fact that the second coding region has a size that does not meet the one or more criteria, the first element is signaled and parsed from the video bitstream.
[0098] (A6)In some embodiments of A5, when the length of the second coding region is less than a threshold length, the second coding region has a size that satisfies the one or more criteria. For example, when dividing a first block into multiple sub-blocks, if the block width or height of a sub-block is equal to or less than a threshold, the syntax indicating whether all sub-blocks within the first block are in a predefined prediction mode is not parsed, but instead is derived as a default value. As an example, the threshold is set to 4. As an example, the default value indicates a mixture of intra-coding mode and inter-coding mode. As an example, if a 32×32 block is partitioned into four sub-blocks of sizes 4×32, 8×32, 16×32, and 4×32, in this case, since there are two sub-blocks with a sub-block width / height of 4, the syntax indicating whether all sub-blocks within the first block are in a predefined prediction mode is not parsed, but instead is derived as a default value. In some embodiments, based on determining that the coding blocks within the second coding region have a size that does not satisfy the one or more criteria, a first element indicating whether all coding blocks within the second coding region are encoded in a predefined prediction mode is derived.
[0099] (A7)In some embodiments of A5, when the area of the second coding region is greater than a threshold area, the second coding region has a size that satisfies the one or more criteria.
[0100] (A8)In some embodiments of any one of A1 to A7, applying a second partitioning method to the chrominance blocks in the coding region includes forgoing partitioning of the chrominance blocks. For example, when a parsed syntax indicates that all blocks within a block region are intra-coded blocks, the luma blocks within the intra-region may be further partitioned, but further partitioning of the chrominance blocks within the block region is not allowed.
[0101] (A9)In some embodiments of any one of A1 to A8, the method further includes: partitioning an inter-frame of a video bitstream to obtain a coding region; and determining whether to restrict further partitioning of the chrominance blocks within the coding region based on a second indicator in the video bitstream. For example, a high-level syntax is received at the decoder side to indicate whether the chrominance blocks within the intra-region of an inter-frame can be further partitioned.
[0102] (A10)In some embodiments of any one of A1 to A9, the context for signaling the block partitioning type of the corresponding coding blocks in the coding region is based on whether the coding region is encoded using an intra-prediction mode. For example, the context for signaling the block partitioning type may depend on whether the current block is within an intra-region.
[0103] (B1)In another aspect, some embodiments include a method of video encoding (e.g., method 650). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) receiving video data including a plurality of coded blocks; (ii) identifying a coded region including two or more of the plurality of coded blocks, wherein each coded block in the coded region is to be encoded in an intra prediction mode; and (iii) encoding two or more coded blocks of the coded region into a video bitstream by applying a first partitioning method to the luminance blocks in the coded region and a second partitioning method to the chrominance blocks in the coded region, wherein the second partitioning method is different from the first partitioning method.
[0104] (B2)In some embodiments of B1, applying the second partitioning method to the chrominance blocks in the coded region includes: applying the first partitioning method for the luminance blocks to the chrominance blocks to a first depth, and foregoing partitioning of the chrominance blocks after the first depth.
[0105] (B3)In some embodiments of B1 or B2, the second partitioning method for the chrominance blocks in the coded region is selected from a subset of the partitioning types available for partitioning luminance blocks.
[0106] (B4)In some embodiments of B3, the subset of the partitioning types is selected based on one or more of the block size, prediction mode, and partitioning type of the luminance blocks.
[0107] (B5)In some embodiments of any one of B1 to B4, the one or more sets of instructions further include instructions for: partitioning a frame of the video data to identify a second coded region; and when the second coded region has a size that meets one or more criteria, foregoing signaling a first element indicating whether all coded blocks within the second coded region are encoded in a predefined prediction mode.
[0108] (B6)In some embodiments of B5, the second coded region has a size that meets the one or more criteria when the length of the second coded region is less than a threshold length.
[0109] (B7)In some embodiments of any one of B1 to B6, the one or more sets of instructions further include instructions for: partitioning an inter-frame of the video data to obtain a coded region; and signaling a second indicator in the video bitstream to indicate whether further partitioning of the chrominance blocks within the coded region is restricted.
[0110] In some embodiments of any one of B1 to B7, the one or more sets of instructions further include instructions for encoding information related to the block partition type of the corresponding encoded block in the encoded region based on encoding the encoded region in the intra prediction mode.
[0111] In some embodiments, the above methods and embodiments are applied to an encoded region having a first prediction mode (e.g., intra prediction mode, inter prediction mode, or intra - inter prediction mode).
[0112] (C1) In another aspect, some embodiments include a method for processing visual media data. In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) obtaining a source video sequence including a plurality of frames; and (ii) performing a conversion between the source video sequence and a video bitstream of visual media data, where the video bitstream includes: (a) a plurality of encoded blocks corresponding to the plurality of frames; and (b) an indicator indicating an encoded region of a frame in the plurality of frames, where the encoded region is composed of blocks encoded in the intra prediction mode, where the luminance blocks encoded in the video bitstream are partitioned according to a first partitioning type, and where the chrominance blocks encoded in the video bitstream are partitioned according to a second partitioning type different from the first partitioning type.
[0113] (C2) In some embodiments of C1, the video bitstream further includes encoded information related to the block partition type of the corresponding encoded block in the encoded region based on encoding the encoded region in the intra prediction mode.
[0114] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, where the memory stores one or more sets of instructions configured to be executed by the control circuit, and the one or more sets of instructions include instructions for performing any method described herein (e.g., A1 to A10 and B1 to B8, C1 and C2 above).
[0115] In another aspect, some embodiments include a non - volatile computer - readable storage medium storing one or more sets of instructions, which are executed by a control circuit of a computing system, and the instructions include instructions for performing any method described herein (e.g., A1 - A10, B1 - B8, C1 and C2 above).
[0116] Unless otherwise specified, any syntactic element described herein can be High-Level Syntax (HLS). As used herein, HLS is signaled at a level higher than a block. For example, HLS can correspond to a sequence level, a frame level, a slice level, or a tile level. As another example, an HLS element can be signaled in a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), an Adaptive Parameter Set (APS), a slice header, a picture header, a tile header, and / or a CTU header.
[0117] It will be understood that although the terms “first,” “second,” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should also be understood that when used in this specification, the terms “comprises” and / or “comprising” indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0118] As used herein, the term “when” can be interpreted to mean “if” or “at the time of” or “in response to determining” or “in accordance with determining” or “in response to detecting,” depending on the context, provided that the stated precondition is true. Similarly, depending on the context, the phrase “if it is determined that [the stated precondition is true]” or “if [the stated precondition is true]” or “when [the stated precondition is true]” can be interpreted to mean “at the time of determining” or “in response to determining” or “in accordance with determining” or “at the time of detecting” or “in response to detecting” that the stated precondition is true. As used herein, N refers to a variable number. Unless explicitly stated, different instances of N may refer to the same number (e.g., the same integer value, such as the number 2) or different numbers.
[0119] For purposes of explanation, the above description has been presented with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Given the above teachings, many modifications and variations are possible. The embodiments were chosen and described in order to best explain the operating principles and the practical application, thereby enabling others skilled in the art to understand.
Claims
1. A video decoding method performed at a computing system having a memory and one or more processors, characterized in that, The method includes: Receiving a video bitstream including a plurality of coded blocks; Based on a first indicator in the video bitstream, identifying a coded region including two or more of the plurality of coded blocks, wherein each coded block in the coded region is coded in an intra prediction mode; According to identifying the coded region: Applying a first partitioning method to the luma blocks in the coded region; and Applying a second partitioning method to the chroma blocks in the coded region, wherein the second partitioning method is different from the first partitioning method; and Using the first partitioning method and the second partitioning method to reconstruct the two or more coded blocks of the coded region.
2. The method according to claim 1, wherein Applying the second partitioning method to the chroma blocks in the coded region includes: applying the first partitioning method for the luma blocks to the chroma blocks to a first depth, and prohibiting further partitioning of the chroma blocks beyond the first depth.
3. The method according to claim 1, characterized in that, The second partitioning method for the chroma blocks in the coded region is selected from: a subset of partitioning types that can be used to partition the luma blocks.
4. The method according to claim 3, wherein Based on one or more of the block size, prediction mode, and partitioning type of the luma blocks, identifying the subset of the partitioning types.
5. The method according to claim 1, characterized in that, Further includes: Partitioning a frame of the video bitstream to obtain a second coded region; And When the second coded region has a size that meets one or more criteria, deriving a first element that indicates whether all coded blocks within the second coded region are coded in a predefined prediction mode.
6. The method according to claim 5, wherein When the length of the second coded region is less than a threshold length, the second coded region has a size that meets the one or more criteria.
7. The method according to claim 5, characterized in that, When the area of the second coded region is greater than a threshold area, the second coded region has a size that meets the one or more criteria.
8. The method according to claim 1, characterized in that, Applying the second partitioning method to the chroma blocks in the coded region includes abandoning partitioning of the chroma blocks.
9. The method according to claim 1, wherein Further includes: Partitioning an inter prediction frame of the video bitstream to obtain the coded region; And Based on a second indicator in the video bitstream, determining whether to limit further partitioning of chroma blocks within the coded region.
10. The method according to claim 1, characterized in that, The context for signaling the block partitioning type of the corresponding coded blocks in the coded region depends on whether the coded region is coded in the intra prediction mode.
11. A computing system, characterized in that, Includes: A control circuit; A memory; And One or more sets of instructions stored in the memory and configured to be executed by the control circuit, the one or more sets of instructions including instructions for: Receiving video data including a plurality of coded blocks; Identifying a coded region including two or more of the plurality of coded blocks, wherein each coded block in the coded region will be coded in an intra prediction mode; And Encoding two or more video blocks of the coded region into a video bitstream by: using a first partitioning method for the luma blocks of the coded region and a second partitioning method for the chroma blocks of the coded region, wherein the second partitioning method is different from the first partitioning method.
12. The computing system according to claim 11, wherein Applying a second partitioning method to the chrominance blocks of the coding region includes: applying the first partitioning method for the luminance blocks to the chrominance blocks to a first depth, and after exceeding the first depth, abandoning further partitioning of the chrominance blocks.
13. The computing system according to claim 11, wherein The second partitioning method for the chrominance blocks in the coding region is selected from: a subset of partitioning types that can be used to partition the luminance blocks.
14. The computing system according to claim 13, wherein Based on one or more of the block size, prediction mode, and partitioning type of the luminance block, a subset of the partitioning types is selected.
15. The computing system according to claim 11, wherein The one or more sets of instructions further include instructions for: Partitioning a frame of the video data to identify a second coding region; and When the second coding region has a size that meets one or more criteria, abandoning signaling a first element that indicates whether all coding blocks within the second coding region are coded in a predefined prediction mode.
16. The computing system according to claim 15, wherein When the length of the second coding region is less than a threshold length, the second coding region has a size that meets the one or more criteria.
17. The computing system according to claim 11, wherein The one or more sets of instructions further include instructions for: Partitioning an inter-predicted frame of the video data to obtain the coding region; and Signaling a second indicator in the video bitstream to indicate whether further partitioning of chrominance blocks within the coding region is restricted.
18. The computing system according to claim 11, wherein The one or more sets of instructions further include instructions for encoding information related to the block partitioning type of the corresponding coding blocks in the coding region, based on the coding region being coded in the intra prediction mode.
19. A non-volatile computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores one or more sets of instructions configured to be executed by a computing device having a control circuit and a memory, the one or more sets of instructions including instructions for: Obtaining a source video sequence including a plurality of frames; And Performing a conversion between the source video sequence and a video bitstream of visual media data, where the video bitstream includes: A plurality of coded blocks corresponding to the plurality of frames; And An indicator indicating a coding region of a frame among the plurality of frames, where the coding region consists of blocks coded in the intra prediction mode; Wherein the luminance blocks coded in the video bitstream are partitioned according to a first partitioning method type, and Wherein the chrominance blocks coded in the video bitstream are partitioned according to a second partitioning method type different from the first partitioning method type.
20. The non-volatile computer-readable storage medium according to claim 19, wherein The video bitstream further includes coded information related to the block partitioning type of the corresponding coding blocks in the coding region, based on the coding region being coded in the intra prediction mode.
Citation Information
Cited By
Video slice coding control method, entropy encoder, equipment and readable storage medium
CN121000872A