Video decoding method, apparatus, and computer program, and video encoding method
By applying geometric transformations to groups of samples within a picture, the coding performance is enhanced by adapting to varying local texture patterns, improving compression efficiency.
Patent Information
- Application Number
- JP2025519507
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-04
- Filing Date
- 2023-10-05
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2043-10-05
AI Technical Summary
Existing video coding technologies are limited by the fixed orientation of input pictures, which can restrict the overall coding performance due to varying local texture patterns within a picture.
Applying geometric transformations, such as horizontal and vertical flips, and rotations, to groups of samples within a picture rather than the entire picture, allowing for different orientations based on local texture patterns.
Enhances coding performance by optimizing the orientation of different groups of samples within a picture, leading to improved compression efficiency.
Smart Images

Figure 2025535039000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure describes embodiments generally related to video coding. [Background technology]
[0002] The background description provided herein is intended to generally present the context for the present disclosure. The work of the presently named inventors, to the extent that that work is described in this background section, and any aspect of the description that may not otherwise qualify as prior art at the time of filing, is not admitted expressly or implicitly as prior art to the present disclosure.
[0003] Image / video compression can help transmit image / video data between different devices, storage, and networks with minimal quality loss. In some examples, video codec technology can compress video based on spatial and temporal redundancy. As an example, video codecs can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. As another example, video codecs can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation can be represented by a motion vector (MV). Summary of the Invention
[0004] Aspects of the present disclosure include methods and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit.
[0005] According to an aspect of the present disclosure, a method of video decoding is provided. In the method, a video bitstream including a current picture of the video is received. A first group of samples and a second group of samples in the current picture are determined. A first geometric transformation is determined for the first group of samples in the current picture, and a second geometric transformation is determined for the second group of samples in the current picture. The first geometric transformation is configured to adjust an orientation of the first group of samples in the current picture. The second geometric transformation is different from the first geometric transformation and is configured to adjust an orientation of the second group of samples in the current picture. The picture is reconstructed, where the first group of samples is reconstructed based on the determined first geometric transformation, and the second group of samples is reconstructed based on the determined second geometric transformation.
[0006] As an example, one or more rows of a coding tree unit (CTU) in the current picture are determined as the first group of samples.
[0007] As an example, one or more columns of coding tree units (CTUs) in the current picture are determined as the first group of samples.
[0008] For example, one or more rows of coding tree units (CTUs) in a current picture are determined, each CTU row among the one or more rows of the CTUs is divided into a plurality of portions, and a first portion among the plurality of portions of each CTU row among the one or more rows of the CTUs is determined as a first group of samples.
[0009] For example, one or more columns of coding tree units (CTUs) in a current picture are determined, and each column of CTUs among the one or more columns of CTUs is divided into a plurality of portions. A first portion of the plurality of portions of each column of CTUs among the one or more columns of CTUs is determined as a first group of samples.
[0010] As an example, each CTU row of one or more rows of CTUs may be divided based on coding information in the received video bitstream.
[0011] For example, a set of geometry transformation units (GTUs) in the current picture is determined, each of which includes a sample in the current picture, and the set of GTUs is determined as a first group of samples.
[0012] By way of example, the first geometric transformation is determined as one or a combination of a vertical flip, a horizontal flip, and a rotation to be performed on the first group of samples.
[0013] In one aspect, the vertical flipping is configured to flip the first group of samples in one of an upward and downward direction, the horizontal flipping is configured to flip the first group of samples in one of a left and right direction, and the rotation is configured to rotate the first group of samples by one of 90 degrees, 180 degrees, and 270 degrees.
[0014] In one aspect, predicted samples of a first group of samples are determined, the orientation of the predicted samples of the first group of samples is reversed based on an inverse geometric transformation that is opposite to the first geometric transformation, and loop filtering is performed on the reversed-oriented predicted samples of the first group of samples.
[0015] According to another aspect of the present disclosure, an apparatus is provided. The apparatus includes a processing circuit. The processing circuit may be configured to perform any of the described video decoding / encoding methods. For example, the processing circuit is configured to receive a video bitstream including a current picture of the video. The processing circuit is configured to determine a first geometric transformation for a first group of samples in the current picture and a second geometric transformation for a second group of samples in the current picture. The first geometric transformation is configured to adjust an orientation of the first group of samples in the current picture. The second geometric transformation is different from the first geometric transformation and is configured to adjust an orientation of the second group of samples in the current picture. The processing circuit is configured to reconstruct the picture, wherein the first group of samples is reconstructed based on the determined first geometric transformation and the second group of samples is reconstructed based on the determined second geometric transformation.
[0016] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the described video decoding / encoding methods.
[0017] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a video processing system (100). [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4] 1 shows a flowchart illustrating a decoding process according to some embodiments of the present disclosure. [Figure 5] 1 shows a flowchart illustrating an encoding process according to some embodiments of the present disclosure. [Figure 6] FIG. 1 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0019] 1 shows a block diagram of a video processing system 100 in some examples. The video processing system 100 is an example of an application of the disclosed subject matter, namely, a video encoder and video decoder in a streaming environment. The disclosed subject matter can be similarly applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0020] The video processing system (100) includes a capture subsystem (113) that may include, for example, a video source (101), such as a digital camera, that generates an uncompressed stream of video pictures (102). By way of example, the stream of video pictures (102) includes samples captured by the digital camera. The stream of video pictures (102) is represented by a bold line to emphasize its high data volume compared to the encoded video data (104) (or coded video bitstream) and may be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (104) (or coded video bitstream) is represented by a thin line to emphasize its lower data volume compared to the stream of video pictures (102) and may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as the client subsystems 106 and 108 of FIG. 1, can access the streaming server 105 to retrieve copies 107 and 109 of the encoded video data 104. The client subsystem 106 may include a video decoder 110, for example, in an electronic device 130. The video decoder 110 decodes the incoming copy 107 of the encoded video data and generates an outgoing stream 111 of video pictures that can be rendered on a display 112 (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data 104, 107, and 109 (e.g., a video bitstream) may be encoded according to a particular video coding / compression standard. An example of such a standard is ITU-T Recommendation H.265.By way of example, a video coding standard under development is commonly known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in connection with VVC.
[0021] It should be noted that electronic devices 120 and 130 may include other components (not shown). For example, electronic device 120 may include a video decoder (not shown), and electronic device 130 may similarly include a video encoder (not shown).
[0022] 2 shows an example block diagram of a video decoder (210). The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used in place of the video decoder (110) in the example of FIG. 1.
[0023] The receiver (231) may receive one or more coded video sequences, e.g., contained in a bitstream, to be decoded by the video decoder (210). In an embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the coded video data. The receiver (231) may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver (231) may separate the coded video sequences from other data. To combat network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter "parser (220)"). In certain applications, the buffer memory 215 is part of the video decoder 210. In others, it can be external to the video decoder 210 (not shown). In still other applications, there can be a buffer memory (not shown) external to the video decoder 210, e.g., to combat network jitter, plus another buffer memory 215 within the video decoder 210, e.g., to manipulate playback timing. When the receiver 231 is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory 215 may not be required or may be small. For use with best-effort packet networks such as the Internet, the buffer memory 215 may be required, but it can be relatively large and advantageously adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder 210.
[0024] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and, potentially, information for controlling a rendering device, such as a render device (212) (e.g., a display screen) that is not an essential part of the electronic device (230) but may be coupled to the electronic device (230) as shown in FIG. 2. Control information for the rendering device may take the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, context-dependent or non-context-dependent arithmetic coding, etc. The parser (220) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (220) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.
[0025] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to generate symbols (221).
[0026] The reconstruction of the symbols (221) can have many different units depending on the type of coded video picture or portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. Which units are included and how may be controlled by subgroup control information parsed by the parser (220) from the coded video sequence. The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.
[0027] Beyond the functional blocks already described, the video decoder (210) may be conceptually subdivided into a number of functional units, which are described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0028] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients as symbols (221) from the parser (220) along with control information including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. The scalar / inverse transform unit (251) can output blocks containing sample values that can be input to an aggregator (255).
[0029] In some cases, the output samples of the scaler / inverse transformer (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258), for example, buffers a partially reconstructed and / or fully reconstructed current picture. The aggregator (255), in some cases, adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transformer unit (251).
[0030] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to an inter-coded, and potentially motion-compensated, block. In such cases, the motion-compensated prediction unit (253) may access the reference picture memory (257) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) related to the block, the samples may be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) fetches prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (253), for example, in the form of symbols (221), which may have X, Y, and reference picture components. Motion compensation can also include interpolation of sample values fetched from the reference picture memory (257) when sub-sample accurate motion vectors are used, as well as motion vector prediction mechanisms.
[0031] The output samples of the aggregator (255) may undergo various loop filtering techniques in a loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression may also respond to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, and may also respond to previously constructed loop-filtered sample values.
[0032] The output of the loop filter unit (256) can be a sample stream that can be output to a render device (212) and further stored in a reference picture memory (257) for use in future inter-picture prediction.
[0033] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and any unused current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.
[0034] The video decoder (210) may perform decoding operations in accordance with a given video compression technology or standard, such as ITU-T Recommendation H.265. A coded video sequence may conform to the syntax prescribed by the video compression technology or standard in use, in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and a profile documented in the video compression technology or standard. Specifically, a profile may select specific tools from all tools available in the video compression technology or standard as the only tools available for use under that profile. Compliance also requires that the complexity of the coded video sequence be within the boundaries defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained through a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.
[0035] In embodiments, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may also be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0036] 3 shows an example block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmitting circuit). The video encoder (303) can be used in place of the video encoder (303) in the example of FIG. 1.
[0037] The video encoder (303) receives video samples from a video source (301) (which, in the example of FIG. 3, is not part of the electronic device (620)) that can capture video images to be coded by the video encoder (303). In other examples, the video source (301) is part of the electronic device (320).
[0038] The video source (301) may provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCB, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (301) may be a storage device storing prepared video. In a video conferencing system, the video source (301) may be a camera capturing local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed in sequence, impart motion. The pictures themselves may be organized as a spatial array of pixels, each of which may have one or more samples depending on the sampling structure, color space, etc., in use. This specification focuses hereinafter on samples.
[0039] According to an embodiment, the video encoder (303) may code and compress pictures of a source video sequence into a coded video sequence (343) in real time or under any other time constraints as needed. Imposing an appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) controls and is operatively coupled to other functional units, as described below. The coupling is not shown for clarity. Parameters set by the controller (350) may include parameters related to rate control (picture skip, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured with other appropriate functions related to the video encoder (303) optimized for a particular system design.
[0040] In some embodiments, the video encoder (303) is configured to operate in a coding loop. As an overly simplified description, the coding loop may include a source coder (330) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data in a manner similar to what a (remote) decoder would also generate. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-exact results independent of the location (local or remote) of the decoder, the contents of the reference picture memory (334) are also bit-perfect between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values as the decoder would "see" when using prediction during decoding. This basic principle of reference picture synchronicity (and the resulting drift when synchronicity cannot be maintained, for example due to channel errors) is also used in several related techniques.
[0041] The operation of the "local" decoder (333) can be the same as a "remote" decoder, such as the video decoder (210), already described in detail above in conjunction with Figure 2. Referring also momentarily to Figure 2, however, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (233), given the availability of symbols and the encoding / decoding of the symbols into a coded video sequence by the entropy coder (345) and parser (220) can be lossless.
[0042] In embodiments, decoder techniques, with the exception of parsing / entropy decoding, present in a decoder are present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the operation of the decoder. Descriptions of encoder techniques may be omitted, as they are the inverse of the decoder techniques described generically. To the extent specified, more detailed descriptions are provided below.
[0043] In operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as "reference pictures." In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of the reference pictures that may be selected as predictive references for the input picture.
[0044] The local video decoder (333) may decode coded video data of pictures that may be designated as reference pictures based on symbols generated by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence is typically a copy of the source video sequence, with some errors. The local video decoder (333) may replicate the decoding process that may be performed by the video decoder on the reference pictures, causing the reconstructed reference pictures to be stored in a reference picture cache (334). In this way, the video encoder (303) may locally store copies of reconstructed reference pictures that have content in common with reconstructed reference pictures that would be obtained by a far-end video decoder (without transmission errors).
[0045] The predictor (335) may perform a prediction search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) may search the reference picture memory (334) for specific metadata, such as reference picture motion vectors, block shapes, or sample data (as candidate reference pixel blocks) that can serve as suitable prediction references for the new picture. The predictor (335) may operate on a sample block-by-pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).
[0046] The controller (350) may manage the coding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0047] The output of all of the above functional units may undergo entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.
[0048] The transmitter (340) may buffer the coded video sequence produced by the entropy coder (345) to prepare it for transmission over a communication channel (360), which can be a hardware / software link to a storage device that stores the coded video data. The transmitter (340) may also merge the coded video data from the video coder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0049] The controller (350) may manage the operation of the video encoder (303). During coding, the controller (350) may assign a particular coding picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following picture types:
[0050] Intra pictures (I-pictures) can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow various types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0051] Predictive Pictures (P-pictures) can be coded and decoded by intra-prediction or inter-prediction, which uses motion vectors and reference indices to predict the sample values of each block.
[0052] Bi-directionally Predictive Pictures (B-pictures) can be coded and decoded using intra- or inter-prediction, which uses two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple-predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0053] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to each picture of the blocks. For example, blocks of an I-picture may be non-predictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P-picture may be predictively coded by spatial prediction or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded by spatial prediction or temporal prediction with reference to one or two previously coded reference pictures.
[0054] The video encoder (303) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. During its operation, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the coded video data may conform to a syntax defined by the video coding technique or standard being used.
[0055] In embodiments, the transmitter (340) may transmit additional data along with the encoded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0056] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. As an example, a particular picture being encoded / decoded, called the current picture, is partitioned into blocks. If a block in the current picture is similar to a reference block in a previously coded reference picture in the video that is still buffered, the block in the current picture may be coded by a vector called a motion vector. A motion vector points to a reference block within a reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.
[0057] In some embodiments, bi-prediction techniques may be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures are used, e.g., a first reference picture and a second reference picture, both of which precede the current picture in decoding order (but may be past and future, respectively, in display order) in the video. A block in the current picture may be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. The block is predictable by a combination of the first and second reference blocks.
[0058] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.
[0059] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. By way of example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0060] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In some embodiments, the video encoders (103) and (203) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (103) and (203) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.
[0061] This disclosure includes aspects related to applying geometric transformations to input source and reconstructed samples during image / video encoding and decoding.
[0062] In related designs of video / image coding schemes, the orientation of the input pictures may be fixed. Adaptive picture flipping coding methods are capable of coding pictures with different orientations.
[0063] In an exemplary method of picture-level geometry transformation, three types of geometry transforms (also called "geometric transform(s)") may be applied at the picture level: horizontal flip, vertical flip, and 180° rotation. Coding information may be signaled in the slice header to indicate whether and which geometry transform is applied.
[0064] The encoder can encode pictures based on each geometric transformation and no transformation, and select one of the geometric transformations (e.g., horizontal flip, vertical flip, and 180° rotation) and no transformation that corresponds to the lowest RD cost. To accelerate the encoder, a set of coding tools and partitioning selections can be skipped in interim encoding, and pictures with the selected transformation can be coded with all features.
[0065] The decoder can restore the reconstructed picture (or the reconstructed picture) by inverse transforming the reconstructed picture according to the type of geometric transformation parsed.
[0066] In the related picture-level geometry transform design, the orientation of the entire input picture can be changed, i.e., it is a picture-level geometry transform, and the best (or selected) orientation can depend on the input picture content. However, the input picture may have different local texture patterns, and each local texture may prefer a different orientation. The picture-level restriction of the input picture orientation may limit the overall coding performance of the geometry transform design.
[0067] In the present disclosure, a geometric transformation (also referred to as a "geometric transform") may be applied to a group of input samples rather than to an entire picture. Different groups of input samples may have different geometric transformations (or geometric transforms) applied to them. As an example, a first group of samples and a second group of samples in a current picture are determined. A first geometric transformation (or a first geometric transform) may be determined for the first group of samples in the current picture, and a second geometric transformation (or a second geometric transform) may be determined for the second group of samples in the current picture. The first geometric transformation may be configured to adjust an orientation (or position) of the first group of samples in the current picture. The second geometric transformation is different from the first geometric transformation and is configured to adjust an orientation (or position) of the second group of samples in the current picture. As an example, the first geometric transformation may be one of a horizontal flip (or horizontal flipping), a vertical flip (or vertical flipping), and a rotation. The second geometric transformation may be another one of a horizontal flip, a vertical flip, and a rotation. By way of example, the first geometric transformation may be configured to rotate a first group of samples by a first angle, and the second geometric transformation may be configured to rotate a second group of samples by a second angle, where the first angle may be different from the second angle.
[0068] In one aspect, the geometric transformation may be applied to a group of input samples, which may include one or more rows (or columns) of a coding tree unit (CTU). For example, one or more rows or columns of a coding tree unit (CTU) in a current picture may be determined as a first group of samples.
[0069] In one aspect, the geometric transformation may be applied to a group of input samples that includes a portion of one or more rows / columns of a coding tree unit (CTU). For example, a portion of each of one or more rows or columns of a coding tree unit (CTU) in a current picture may be determined as a first group of samples.
[0070] In one embodiment, each CTU row and / or column may be divided into two portions, and each portion may include multiple CTUs in the same row and / or column. For example, each CTU row among one or more rows of CTUs may be divided into multiple portions. A first portion among the multiple portions of each CTU row among one or more rows of CTUs may be determined as a first group of samples. For example, each CTU column among one or more columns of CTUs may be divided into multiple portions, each of the multiple portions including at least one CTU or at least a portion of a CTU.
[0071] In one aspect, the location at which the CTU rows and / or columns are divided may further be signaled for each row and / or column. For example, each CTU row among one or more rows of CTUs may be divided based on coding information in the received video bitstream. For example, each CTU column among one or more columns of CTUs may be divided based on coding information in the received video bitstream. The location information may be signaled in a high-level syntax such as a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a picture header, a tile header, or a CTU header.
[0072] In one aspect, a geometry transform may be applied to a group of input samples, which may include a group of geometry (or geometric) transform units (GTUs). The size of the GTUs may be either predefined or explicitly signaled, and the size may be greater than or equal to the maximum size of a transform unit (TU).
[0073] For example, a set of geometry transform units (GTUs) in the current picture may be determined, each of which may include each sample of the current picture. The set of GTUs may be determined as a first group of samples.
[0074] The GTU width (or height) can be a different multiple of the CTU width (height). For example, the GTU size can be 256 x 256, while the CTU size can be 128 x 128. As another example, the GTU size can be 128 x 128, while the CTU size can be 256 x 256.
[0075] By way of example, a GTU can be larger or smaller than a GTU.
[0076] When the GTU size (e.g., width, height) is explicitly signaled, the GTU size may be signaled in a high-level syntax (HLS). For example, the GTU size may be signaled in any of the VPS, SPS, PPS, APS, slice header, picture header, tile header, or CTU header.
[0077] In one embodiment, the geometric transformations may include, but are not limited to, up / down flips (vertical flipping), left / right flips (horizontal flipping), and 90 / 180 / 270 degree rotations.
[0078] For example, vertical flipping may be configured to flip the first group of samples in one of an upward and downward direction, horizontal flipping may be configured to flip the first group of samples in one of a left and right direction, and rotation may be configured to rotate the first group of samples by one of 90 degrees, 180 degrees, and 270 degrees.
[0079] In one aspect, the geometric transformation can be any combination of transformation operations, such as a plurality of the above transformation operations. As an example, the geometric transformation may include a combination of a vertical flip and a 90-degree rotation.
[0080] As an example, one or a combination of vertical flipping, horizontal flipping, and rotation may be determined as the first group of samples.
[0081] In one aspect, the type of geometry transformation may be signaled in the bitstream, for example, at the coding block level, at the slice / picture level, or at the sequence level.
[0082] In one aspect, an inverse geometric transform (or inverse geometric transform) may be applied before loop filtering (eg, deblocking, SAO, and ALF) is applied to each group of input samples.
[0083] For example, prediction samples of a first group of samples may be determined. The orientations (or positions) of the prediction samples of the first group of samples may be reversed based on an inverse geometric transformation that is opposite to the first geometric transformation. Loop filtering may be performed on the orientation-reversed prediction samples of the first group of samples.
[0084] 4 shows a flowchart illustrating a process (400) according to an embodiment of the present disclosure. The process (400) may be used in a video decoder. In various embodiments, the process (400) is performed by a processing circuit, such as a processing circuit performing the functions of the video decoder (110), a processing circuit performing the functions of the video decoder (210), etc. In some embodiments, the process (400) is implemented by software instructions, such that the processing circuit performs the process (400) when the processing circuit executes the software instructions. The process (400) begins at (S401) and proceeds to (S410).
[0085] At (S410), a video bitstream containing a current picture of the video is received.
[0086] At (S420), a first group of samples and a second group of samples in the current picture are determined.
[0087] At (S430), a first geometric transform is determined for a first group of samples in the current picture, and a second geometric transform is determined for a second group of samples in the current picture, the first geometric transform being configured to adjust an orientation of the first group of samples in the current picture, and the second geometric transform being different from the first geometric transform and configured to adjust an orientation of the second group of samples in the current picture.
[0088] At (S440), the picture is reconstructed, where a first group of samples is reconstructed based on the determined first geometric transformation and a second group of samples is reconstructed based on the determined second geometric transformation.
[0089] As an example, one or more rows of a coding tree unit (CTU) in the current picture are determined as the first group of samples.
[0090] As an example, one or more columns of coding tree units (CTUs) in the current picture are determined as the first group of samples.
[0091] For example, one or more rows of coding tree units (CTUs) in a current picture are determined, each CTU row among the one or more rows of the CTUs is divided into a plurality of portions, and a first portion among the plurality of portions of each CTU row among the one or more rows of the CTUs is determined as a first group of samples.
[0092] For example, one or more columns of coding tree units (CTUs) in a current picture are determined, and each column of CTUs among the one or more columns of CTUs is divided into a plurality of portions. A first portion of the plurality of portions of each column of CTUs among the one or more columns of CTUs is determined as a first group of samples.
[0093] As an example, each CTU row of one or more rows of CTUs may be divided based on coding information in the received video bitstream.
[0094] For example, a set of geometry transformation units (GTUs) in the current picture is determined, each of which includes a sample in the current picture, and the set of GTUs is determined as a first group of samples.
[0095] By way of example, the first geometric transformation is determined as one or a combination of a vertical flip, a horizontal flip, and a rotation to be performed on the first group of samples.
[0096] In one aspect, the vertical flipping is configured to flip the first group of samples in one of an upward and downward direction, the horizontal flipping is configured to flip the first group of samples in one of a left and right direction, and the rotation is configured to rotate the first group of samples by one of 90 degrees, 180 degrees, and 270 degrees.
[0097] In one aspect, predicted samples of a first group of samples are determined, the orientation of the predicted samples of the first group of samples is reversed based on an inverse geometric transformation that is opposite to the first geometric transformation, and loop filtering is performed on the reversed-oriented predicted samples of the first group of samples.
[0098] The process then proceeds to (S499) and ends.
[0099] Process 400 may be adapted as appropriate. Steps of process 400 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0100] 5 shows a flowchart illustrating a process (500) according to an embodiment of the present disclosure. The process (500) may be used in a video encoder. In various embodiments, the process (500) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), etc. In some embodiments, the process (500) is implemented by software instructions, such that the processing circuit performs the process (500) when the processing circuit executes the software instructions. The process (500) begins at (S501) and proceeds to (S510).
[0101] At (S510) a first group of samples and a second group of samples in the current picture are determined.
[0102] At (S520), a first geometric transform is performed on a first group of samples in the current picture, and a second geometric transform is performed on a second group of samples in the current picture, the first geometric transform being configured to adjust an orientation of the first group of samples in the current picture, and the second geometric transform being different from the first geometric transform and configured to adjust an orientation of the second group of samples in the current picture.
[0103] At (S530), the picture is coded, where a first group of samples is coded based on a first geometric transformation performed and a second group of samples is coded based on a second geometric transformation performed.
[0104] The process then proceeds to (S599) and ends.
[0105] Process 500 may be adapted as appropriate. Steps of process 500 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0106] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 6 illustrates a computer system (600) suitable for implementing certain embodiments of the disclosed subject matter.
[0107] Computer software can be coded in any suitable machine code or computer language that can be subjected to mechanisms such as assembly, compilation, linking, etc. to generate code containing instructions that can be executed by one or more central processing units (CPUs), graphics processing units (GPUs), etc. directly or through interpretation, microcode execution, etc.
[0108] The instructions may be executable by various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.
[0109] 6 for computer system (600) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components described in the exemplary embodiment of computer system (600).
[0110] The computer system 600 may include certain human interface input devices. Such human interface input devices may respond to input by one or more users through, for example, tactile input (e.g., keyboard, swipe, dataglove motion), audio input (e.g., voice, claps), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).
[0111] The input human interface devices may include one or more of a keyboard (601), a mouse (602), a trackpad (603), a touchscreen (610), a data glove (not shown), a joystick (605), a microphone (606), a scanner (607), and a camera (608) (only one of each is shown).
[0112] The computer system 600 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen 610, data gloves (not shown), or joystick 605; however, haptic feedback devices that do not function as input devices may also exist), audio output devices (e.g., speakers 609, headphones (not shown)), visual output devices (e.g., CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capabilities, each with or without haptic feedback capabilities, some of which may provide two-dimensional visual output or output in more than three dimensions via means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0113] The computer system (600) may also include human-accessible storage devices and their associated media, such as CD / DVD or similar media (621), including CD / DVD ROM / RW (620), thumb drives (622), removable hard disks or solid state drives (623), legacy magnetic media such as tape and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.
[0114] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.
[0115] The computer system 600 may also include an interface 654 to one or more communications networks 655. Networks may be, for example, wireless, wireline, or optical. Networks may also be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet; wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wireline or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicle and factory networks including CAN bus. Certain networks generally require an external network interface adapter attached to a particular general-purpose digital port or peripheral bus 649 (e.g., a USB port on the computer system 600). Others are generally integrated into the core of the computer system 600 by attachment to a system bus as described below (e.g., an Ethernet network interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, computer system 600 can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV) or one-way transmit-only (e.g., a CAN bus to a specific CAN bus device), or it can be two-way to other computer systems using, for example, a local or wide-area digital network. Specific protocols or protocol stacks can be used with each of the networks and network interfaces described above.
[0116] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core 640 of the computer system 600.
[0117] The core (640) may include one or more central processing units (CPUs) (641), graphics processing units (GPUs) (642), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (643), task-specific hardware accelerators (644), graphics adapters (650), etc. These devices may be connected through a system bus (648), along with read-only memory (ROM) (645), random access memory (RAM) (646), internal mass storage devices such as internal non-user-accessible hard drives, SSDs, etc. (647). In some computer systems, the system bus (648) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus (648) directly or through a peripheral bus (649). In an example, a display (610) may be connected to the graphics adapter (650). Architectures for peripheral buses include PCI and USB.
[0118] The CPU (641), GPU (642), FPGA (643), and accelerator (644) can execute specific instructions that, in combination, can constitute the above-mentioned computer code. The computer code can be stored in ROM (645) or RAM (646). Temporary data can also be stored in RAM (646), while persistent data can be stored, for example, in an internal mass storage device (647). Rapid storage and retrieval from any of the memory devices can be enabled through the use of cache memory. Cache memory can be closely associated with one or more of the CPU (641), GPU (642), mass storage device (647), ROM (645), RAM (646), etc.
[0119] The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.
[0120] By way of example, and not limitation, a computer system having the architecture (600), and specifically the core (640), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage devices previously introduced, in addition to specific storage devices of the core (640) that are non-transitory in nature, such as the core's internal mass storage (647) or ROM (645). Software implementing various embodiments of the present disclosure can be stored on such devices and executable by the core (640). The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause the core (640), and specifically the processor (including a CPU, GPU, FPGA, etc.) therein, to perform particular processes or portions of particular processes described herein, including defining data structures stored in RAM (646) and modifying such data structures according to software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (644)) that can operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software can encompass logic, where appropriate, and vice versa. References to computer-readable media can encompass circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0121] The use of "at least one of" or "one of" in this disclosure is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). The use of "one of" does not exclude any combination of the listed elements, where applicable, such as when the elements are not mutually exclusive.
[0122] While this disclosure has described several exemplary embodiments, alternatives, permutations, and various substitute equivalents exist and are included within the scope of this disclosure. Thus, it will be apparent to those skilled in the art that numerous systems and methods, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.
[0123] [Incorporated by reference] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 413,927, filed October 6, 2022, and entitled "CTU-Row Based Geometric Transform," and U.S. Patent Application No. 18 / 376,821, filed October 4, 2023, and entitled "CTU-ROW BASED GEOMETRIC TRANSFORM," the disclosures of which are incorporated herein by reference in their entireties.
Claims
1. 1. A method of video decoding performed by a decoder, comprising: receiving a video bitstream including a current picture of the video; determining a first group of samples and a second group of samples in the current picture; determining a first geometric transformation for a first group of the samples in the current picture and a second geometric transformation for a second group of the samples in the current picture, the first geometric transformation being configured to adjust an orientation of the first group of samples in the current picture, and the second geometric transformation being different from the first geometric transformation and configured to adjust an orientation of the second group of samples in the current picture; reconstructing the current picture, wherein a first group of samples is reconstructed based on the determined first geometric transformation and a second group of samples is reconstructed based on the determined second geometric transformation; A method having the following.
2. Determining the first group of samples comprises: determining one or more rows of a coding tree unit (CTU) in the current picture as the first group of samples; The method of claim 1.
3. Determining the first group of samples comprises: determining one or more columns of coding tree units (CTUs) in the current picture as the first group of samples; The method of claim 1.
4. Determining the first group of samples comprises: determining one or more rows of coding tree units (CTUs) in the current picture; Dividing each CTU row of the one or more rows of the CTU into a plurality of portions; determining a first portion of the plurality of portions of each CTU row of one or more rows of the CTU as the first group of samples; Including, The method of claim 1.
5. Determining the first group of samples comprises: determining one or more sequences of coding tree units (CTUs) within the current picture; Dividing each CTU string of the one or more strings of CTUs into a plurality of portions; determining a first portion of the plurality of portions of each CTU sequence of one or more sequences of the CTU as the first group of samples; Including, The method of claim 1.
6. The dividing step is Separating each CTU row among the one or more rows of the CTU based on coding information in the received video bitstream. The method of claim 4.
7. Determining the first group of samples comprises: determining a set of geometry transformation units (GTUs) in the current picture, each GTU in the set including a respective sample of the current picture; determining the set of GTUs as the first group of samples; Including, The method of claim 1.
8. Determining the first geometric transformation comprises: determining the first geometric transformation as one or a combination of a vertical flip, a horizontal flip, and a rotation performed on the first group of samples. The method of claim 1.
9. the vertical flipping is configured to flip the first group of samples in one of an upward direction and a downward direction; the horizontal flipping is configured to flip the first group of samples in one of a leftward direction and a rightward direction; the rotating is configured to rotate the first group of samples by one of 90 degrees, 180 degrees, and 270 degrees. The method of claim 8.
10. Reconstructing the first group of samples comprises: determining predicted samples for the first group of samples; reversing the orientation of the prediction samples of the first group of samples based on an inverse geometric transformation that is reversible to the first geometric transformation; performing loop filtering on the reversed-direction prediction samples of the first group of samples; Including, The method of claim 1.
11. 11. A method according to claim 1, comprising: a processing circuit and a memory storing instructions, the processing circuit being configured to perform the method according to any one of claims 1 to 10 by executing the instructions stored in the memory. Device.
12. A computer program which, when executed on a computer, causes the computer to carry out the method according to any one of claims 1 to 10.
13. 1. A method of video encoding performed by an encoder, comprising: determining a first group of samples and a second group of samples in the current picture; performing a first geometric transformation on a first group of the samples in the current picture and a second geometric transformation on a second group of the samples in the current picture, the first geometric transformation being configured to adjust an orientation of the first group of samples in the current picture, and the second geometric transformation being different from the first geometric transformation and configured to adjust an orientation of the second group of samples in the current picture; encoding the current picture, wherein the first group of samples is encoded based on the first geometric transformation performed and the second group of samples is encoded based on the second geometric transformation performed; A method having the following.
Citation Information
Patent Citations
Position Resetting of Prediction Residual Blocks in Video Coding
JP2016521070A
Video signal encoding / decoding method and device therefor
JP2022537767A
Image decoding method using residual information in image coding system, and device for same
US20220264104A1
Image processing device and image processing method
WO2020145381A1