Derivation of Transform Kernel for Improving Inter-Cycling Blocks Using Intra-Prediction Mode Information

By using intra-frame prediction mode information to derive the transform kernel in inter-frame coding blocks, and combining intra-frame prediction and inter-frame prediction techniques, the problem of insufficient efficiency and quality of inter-frame prediction in existing technologies is solved, and more efficient video coding is achieved.

CN122095623APending Publication Date: 2026-05-26TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2025-03-07
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies struggle to effectively combine intra-frame prediction mode information with inter-frame prediction when deriving the transform kernel, resulting in deficiencies in coding efficiency and quality.

Method used

By deriving the transform kernel using intra-prediction mode information in inter-frame coding blocks, and combining intra-frame prediction and inter-frame prediction techniques, such as Inter-Frame Intra-Frame Joint Prediction (CIIP) and Geometric Partitioning (GPM), coding efficiency can be improved.

Benefits of technology

It improves the efficiency and quality of video coding by combining in-frame and out-of-frame prediction mode information, optimizing the derivation process of the transform kernel, and enhancing coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122095623A_ABST
    Figure CN122095623A_ABST
Patent Text Reader

Abstract

This disclosure provides a video decoding method. For example, receiving an encoded video stream. The encoded video stream includes encoded information from multiple images. Based on the encoded information, determining that a current block in the current image is encoded using an inter-frame prediction mode, and generating prediction samples of the current block in the inter-frame prediction mode based at least in part on reference blocks in a reference image different from the current image. Furthermore, obtaining an intra-frame prediction mode associated with the current block in the current image. Determining at least one transform kernel based on the intra-frame prediction mode associated with the current block. Reconstructing the current block based on the at least one transform kernel.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] References merged

[0002] This application claims priority to U.S. Patent Application No. 19 / 072,848, filed March 6, 2025, which in turn claims priority to U.S. Provisional Application No. 63 / 563,168, filed March 8, 2024. The entire disclosure of the earlier applications is incorporated herein by reference. Technical Field

[0003] This disclosure describes, in general, various aspects relating to video encoding and decoding. Background Technology

[0004] The background description provided herein is intended to present the general context of this disclosure. The work of the currently attributed inventors, to the extent described in the background section, and in aspects that may not conform to the prior art at the time of application, is neither expressly nor impliedly acknowledged as belonging to the prior art of this disclosure.

[0005] Image / video compression can help transmit image / video data across different devices, storage, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial and temporal redundancy. For instance, a video codec can use a technique called intra-frame prediction, which compresses images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current image in the current reconstruction to predict samples. In another example, a video codec can use a technique called inter-frame prediction, which compresses images based on temporal redundancy. For example, inter-frame prediction can use previously reconstructed images to predict samples in the current image through motion compensation. Motion compensation can be indicated by motion vectors (MV). Summary of the Invention

[0006] This disclosure includes aspects of bitstreams, methods, and apparatus for video encoding / decoding. In some examples, an apparatus for video encoding / decoding includes processing circuitry.

[0007] One aspect of this disclosure provides a video decoding method. For example, receiving an encoded video stream. The encoded video stream includes encoded information from multiple images. Based on the encoded information, determining that a current block in the current image is encoded using an inter-frame prediction mode, and generating prediction samples of the current block in the inter-frame prediction mode based at least in part on reference blocks in a reference image different from the current image. Furthermore, obtaining an intra-frame prediction mode associated with the current block in the current image. Determining at least one transform kernel based on the intra-frame prediction mode associated with the current block. Reconstructing the current block based on the at least one transform kernel.

[0008] Another aspect of this disclosure provides a video coding method. For example, determining to encode a current block in a current image using an inter-frame prediction mode. Predicted samples of the current block are generated, at least partially based on reference blocks in a reference image different from the current image. An intra-frame prediction mode associated with the current block in the current image is obtained. At least one transform kernel is determined based on the intra-frame prediction mode associated with the current block. The current block is encoded into bits in a bitstream based on the at least one transform kernel.

[0009] Another aspect of this disclosure provides a method for processing visual media data. In this method, a conversion between a bitstream of visual media data and a visual media file is performed according to a format rule. In an example, the bitstream carries encoded information from multiple images. The format rule specifies that an inter-frame prediction mode is used to encode a current block in the current image, generating a prediction sample of the current block in the inter-frame prediction mode based at least in part on a reference block in a reference image different from the current image. The format rule also specifies obtaining an intra-frame prediction mode associated with the current block in the current image, determining at least one transform kernel based on the intra-frame prediction mode associated with the current block, and reconstructing the current block based on the at least one transform kernel.

[0010] This disclosure also provides a video encoding apparatus. The video encoding apparatus includes processing circuitry configured to perform any of the described video encoding methods.

[0011] This disclosure also provides a video decoding method. This method includes any method performed by a video decoding apparatus.

[0012] This disclosure also provides a non-volatile computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the described video decoding / encoding methods. Attached Figure Description

[0013] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, wherein:

[0014] Figure 1 This is a schematic diagram of an example block diagram of a communication system (100).

[0015] Figure 2 This is a schematic diagram of an example block diagram of the decoder.

[0016] Figure 3 This is a schematic diagram of an example block diagram of an encoder.

[0017] Figure 4 A diagram showing some examples of the 24 angles used in the Geometric Partitioning (GPM) pattern.

[0018] Figure 5 Example diagram showing possible dividing edges for angle index 3.

[0019] Figure 6 An example table showing how intra-frame prediction modes are mapped to a set of low-frequency non-separable transforms (LFNSTs) is presented.

[0020] Figure 7 A diagram showing the blocks in the GPM mode in the example.

[0021] Figure 8 A diagram showing the blocks in the GPM mode in the example.

[0022] Figure 9 A flowchart outlining some aspects of the decoding methods disclosed herein is shown.

[0023] Figure 10 A flowchart outlining the coding methods for some aspects of this disclosure is shown.

[0024] Figure 11 This is a schematic diagram of a computer system. Detailed Implementation

[0025] Figure 1 Block diagrams of some examples of video processing systems (100) are shown. The video processing system (100) is an application example of the disclosed subject matter, video encoders and video decoders in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0026] The video processing system (100) includes an acquisition subsystem (113) that may include a video source (101) such as a digital camera, which creates an uncompressed video image stream (102). In an embodiment, the video image stream (102) includes samples captured by a digital camera. The video image stream (102) is depicted as a thick line to emphasize the high data volume of the video image stream compared to encoded video data (104) (or encoded video bitstream). The video image stream (102) may be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or carry out aspects of the disclosed subject matter as described in more detail below. Compared to the video image stream (102), the encoded video data (104) (or the encoded video bitstream (104)) is depicted as a thin line to emphasize the lower data volume of the encoded video data (104) (or the encoded video bitstream (104)), which can be stored on a streaming server (105) for future use. At least one streaming client subsystem, such as Figure 1 Client subsystems (106) and (108) in the application can access a streaming server (105) to retrieve copies (107) and (109) of encoded video data (104). Client subsystem (106) may include, for example, a video decoder (110) in an electronic device (330). The video decoder (110) decodes the incoming copy (107) of the encoded video data and produces an output video image stream (111) that can be displayed on a display (112) (e.g., a screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), video data (107), and video data (109) (e.g., a video stream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T H.265. In embodiments, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and this application can be used in the context of the VVC standard.

[0027] It should be noted that the electronic devices (120) and (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).

[0028] Figure 2This is an example block diagram of a video decoder (210). The video decoder (210) may be disposed in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., receiving circuitry). The video decoder (210) may be used in place of... Figure 1 The video decoder (110) in the embodiment.

[0029] The receiver (231) may receive at least one encoded video sequence, such as a bitstream, to be decoded by the video decoder (210). According to one aspect, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective user entities (not indicated). The receiver (231) may separate the encoded video sequences from other data. To prevent network jitter, a buffer (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as the "parser (220)"). In some applications, the buffer (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be located externally to the video decoder (210) (not shown). In other cases, a buffer memory (not shown) may be located externally to the video decoder (210) to prevent network jitter, for example, and another buffer memory (215) may be configured internally to handle broadcast timing, for example. When the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer memory (215) may not be necessary, or it may be made smaller. Of course, for use on packet networks such as the Internet, a buffer memory (215) may also be required; this buffer memory may be relatively large and adaptive in size, and may be at least partially implemented in an operating system or a similar component (not shown) external to the video decoder (210).

[0030] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (210) and potential information for controlling display devices (212), such as displays that are not part of the electronic device (230) but may be coupled to it. Figure 2As shown in the figure. The control information for the display device may be a parameter set fragment (not shown) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (220) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set of at least one subgroup of pixels in the subgroup of pixels in the encoded video sequence based on at least one parameter corresponding to a group. The subgroup may include a Group of Pictures (GOP), image, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (220) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0031] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0032] Depending on the type of encoded video image or a portion thereof (e.g., inter-frame and intra-frame images, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by the subgroup control information parsed from the encoded video sequence by the parser (220). For the sake of brevity, the flow of such subgroup control information between the parser (220) and the various units described below is not described.

[0033] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0034] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantization transform coefficients as symbols (221) and control information from the parser (220), including the transform method used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block containing sample values, which can be input into the aggregator (255).

[0035] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed image, but can use predictive information from a previously reconstructed portion of the current image. Such predictive information may be provided by the intra-image prediction unit (252). In some cases, the intra-image prediction unit (252) uses reconstructed information extracted from the current image buffer (258) to generate a surrounding block of the same size and shape as the block being reconstructed. For example, the current image buffer (258) buffers a partially reconstructed current image and / or a fully reconstructed current image. In some cases, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.

[0036] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit (253) can access the reference image memory (257) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (221), these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (referred to in this case as residual samples or residual signals) to generate output sample information. The motion compensation prediction unit (253) can obtain the predicted samples from the address in the reference image memory (257) under motion vector control, and the motion vector is available to the motion compensation prediction unit (253) in the form of the symbols (221), which, for example, include X, Y and reference image components. Motion compensation may also include interpolation of sample values ​​extracted from the reference image memory (257) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.

[0037] The output samples of the aggregator (255) can be employed by various loop filtering techniques in the loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video stream), and these parameters can be used as symbols (221) from the parser (220) in the loop filter unit (256). Video compression techniques may also respond to metadata obtained during decoding of a previous (in decoding order) portion of the encoded image or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0038] The output of the loop filter unit (256) can be a sample stream, which can be output to a display device (212) and stored in a reference image memory (257) for subsequent inter-frame image prediction.

[0039] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded image corresponding to the current image has been fully reconstructed and the encoded image (by, for example, the parser (220)) is identified as the reference image, the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.

[0040] The video decoder (210) performs decoding operations according to a predetermined video compression technology or standard, such as ITU-T H.265. The encoded video sequence may conform to the syntax specified by the video compression technology or standard in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the configuration file recorded in the video compression technology or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technology or standard as the only tools available under said configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference image size, etc. In some cases, the limitations set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.

[0041] In one aspect, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be a portion of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0042] Figure 3 This is an example block diagram of a video encoder (303). The video encoder (303) is disposed in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) can be used in place of... Figure 1 The video encoder (103) in the embodiment.

[0043] The video encoder (303) can be obtained from the video source (301) (not Figure 3 In one embodiment, a portion of the electronic device (320) receives video samples, the video source being capable of capturing video images to be encoded by a video encoder (303). In another embodiment, the video source (301) is a portion of the electronic device (320).

[0044] A video source (301) can provide a sequence of source videos in the form of a digital video sample stream, which will be encoded by a video encoder (303). This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit…), any color space (e.g., BT.601 YCrCb, RGB…), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual images, which are given motion when viewed sequentially. The images themselves can be constructed as spatial pixel arrays, where each pixel can include at least one sample, depending on the sampling structure, color space, etc., used. The following focuses on describing samples.

[0045] According to one aspect, the video encoder (303) can encode and compress images of a source video sequence into an encoded video sequence (343) in real time or under any other required time constraints. Implementing an appropriate encoding rate is a function of the controller (350). In some aspects, the controller (350) controls and is functionally coupled to other functional units described below. For simplicity, coupling is not indicated in the figures. Parameters set by the controller (350) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be used with other suitable functions related to the video encoder (303) optimized for a particular system design.

[0046] In one aspect, the video encoder (303) operates within an encoding loop. As a simplified description, in an embodiment, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (333) embedded within the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference image memory (334). Since decoding of the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference image memory (334) also correspond bit-accurately between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same sample values ​​that the decoder will "see" when using the prediction during decoding. This fundamental principle of reference image synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related techniques.

[0047] The operation of the “local” decoder (333) can be combined with, for example, the above-mentioned... Figure 2 The video decoder (210) is described in detail as the same as the "remote" decoder. However, a further brief reference is provided. Figure 2 When symbols are available and the entropy encoder (345) and parser (220) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (210), including the buffer (215) and parser (220), may not be fully implemented in the local decoder (333).

[0048] In one aspect, any decoder technique other than parsing / entropy decoding present in the decoder also exists in the corresponding encoder in the same or substantially the same functional form. Accordingly, this application focuses on decoder operation. The description of encoder techniques can be simplified because encoder techniques are inverses of the fully described decoder techniques. More detailed descriptions are provided in certain areas below.

[0049] During operation, in some embodiments, the source encoder (330) may perform motion-compensated predictive coding. The motion-compensated predictive coding predictively encodes the input image with reference to at least one previously encoded image from the video sequence designated as a "reference image." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.

[0050] The local video decoder (333) can decode encoded video data of an image that can be designated as a reference image, based on symbols created by the source encoder (330). The operation of the encoding engine (332) can be a lossy process. When the encoded video data can be decoded by the video decoder (333), Figure 3 When the source video sequence (not shown) is decoded, the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in a reference image memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.

[0051] The predictor (335) can perform a prediction search against the encoding engine (332). That is, for a new image to be encoded, the predictor (335) can search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. The predictor (335) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, based on the search results obtained by the predictor (335), it can be determined that the input image can have prediction references obtained from multiple reference images stored in the reference image memory (334).

[0052] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0053] The outputs of all the above-mentioned functional units can be entropy encoded in the entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.

[0054] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission via a communication channel (360), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0055] The controller (350) manages the operation of the video encoder (303). During encoding, the controller (350) can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types:

[0056] Intraframe images (I-images) can be encoded and decoded without using any other images in the sequence as prediction sources. Some video codecs allow different types of intraframe images, including, for example, Independent Decoder Refresh (IDR) images.

[0057] Predictive images (P-images) can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses motion vectors and reference indices to predict sample values ​​for each block.

[0058] Bidirectional predictive images (B-images) can be encoded and decoded using intra-frame or inter-frame prediction, which uses two motion vectors and a reference index to predict sample values ​​for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata to reconstruct a single block.

[0059] The source image is typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and each block is encoded sequentially. These blocks can be predictively coded with reference to other (already coded) blocks, determined based on the coding assignment of the corresponding images applied to the blocks. For example, blocks of an I-image can be non-predictively coded, or the blocks can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be predictively coded with reference to a previously coded reference image via spatial prediction or temporal prediction. Blocks of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial prediction or temporal prediction.

[0060] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) can perform various compression operations, including predictive coding operations utilizing temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0061] In one aspect, the transmitter (340) may transmit additional data while transmitting encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, redundant images and slices, other forms of redundant data, SEI messages, VUI parameter set fragments, etc.

[0062] The acquired video can be presented as multiple source images (video images) in a time series. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In an embodiment, a specific image being encoded / decoded is segmented into blocks, referred to as the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. This motion vector points to the reference block in the reference image, and when multiple reference images are used, the motion vector may have a third dimension that identifies the reference image.

[0063] In some aspects, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. Specifically, the block can be predicted using a combination of the first and second reference blocks.

[0064] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.

[0065] According to some aspects disclosed in this application, predictions such as inter-frame image prediction and intra-frame image prediction are performed on a block-by-block basis. For example, according to the HEVC standard, images in a video image sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the images have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU comprises three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Furthermore, each CTU can be further subdivided into at least one coding unit (CU) using a quadtree. For example, a 64×64 pixel CTU can be subdivided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In embodiments, each CU is analyzed to determine the prediction type used for the CU, such as inter-frame prediction or intra-frame prediction. Furthermore, depending on temporal and / or spatial predictability, the CU is divided into at least one prediction unit (PU). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one aspect, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. Taking a luma prediction block as an example, a prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0066] It should be noted that any suitable technology can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In one aspect, at least one integrated circuit can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In another aspect, at least one processor executing software instructions can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210).

[0067] Some aspects of this disclosure provide techniques for transform kernel derivation in inter-frame coded blocks using intra-frame prediction mode information.

[0068] In some respects, intra-frame prediction information is available for inter-coded blocks.

[0069] According to one aspect of this disclosure, intra-frame prediction and inter-frame prediction can be appropriately combined using coding techniques. One coding technique for combining intra-frame prediction and inter-frame prediction is called combined inter-intra-prediction (CIIP), also known as multi-hypothesis intra-inter prediction. For example, CIIP can combine one type of intra-frame prediction and one type of combined prediction. In an example, when the CU is in combined mode, a specific flag is represented by a signal indicating the intra-frame mode. When this specific flag is true, an intra-frame mode can be selected from an intra-frame candidate list. For the luma component, the intra-frame candidate list is derived from four intra-frame prediction modes, such as DC mode, planar mode, horizontal mode, and vertical mode, and the size of the intra-frame mode candidate list depends on the block shape and can be 3 or 4. In an example, when the CU width is greater than twice the CU height, the horizontal mode is removed from the intra-frame mode candidate list, and when the CU height is greater than twice the CU width, the vertical mode is removed from the intra-frame mode candidate list. In some embodiments, intra-frame prediction is performed based on the intra-frame prediction mode selected by the intra-frame mode index, and inter-frame prediction is performed based on the merge index. A weighted average is used to merge the intra-frame prediction results and the inter-frame prediction results. For the chroma component, in some examples, copies of the intra-frame prediction mode information and / or inter-frame prediction mode information from the luma component can be used without additional signaling.

[0070] In some embodiments, the weights used to merge intra-frame and inter-frame prediction results can be determined in an appropriate manner. In the example, when DC or planar mode is selected, or when the coded block (CB) width or height is less than 4, equal weights are used for inter-frame and intra-frame prediction results. In another example, for CBs with a width and height greater than or equal to 4, when horizontal / vertical mode is selected, the CB is first vertically / horizontally divided into four equal-area regions. Each region has a weight set, denoted as (…). ), where i is from 1 to 4. In the example, the first weight set ( ) = (6, 2), second weight set ( ) = (5, 3), Third weight set ( ) = (3, 5) and the fourth weight set ( ) = (2, 6) is applied to the corresponding region. For example, the first weight set ( The fourth weight set is used for the region closest to the reference sample, while the fifth weight set is used for the region closest to the reference sample. The region furthest from the reference sample is used. Then, the combined prediction result can be calculated by summing the two weighted predictions and shifting them right by 3 bits.

[0071] Furthermore, when adjacent CBs are intra-coded, the intra-prediction mode of the intra-prediction assumptions used for the predictor can be saved for intra-mode coding of subsequent adjacent CBs.

[0072] According to another aspect of this disclosure, an inter-frame coded block can be divided into at least two partitions. These at least two partitions can be predicted using different prediction information. In one example, one partition can be predicted using inter-frame prediction, and the other partition can be predicted using intra-frame prediction. In another example, these at least two partitions can be predicted using inter-frame prediction with different motion information.

[0073] In some examples (e.g., VVC), a technique called geometric partition mode (GPM) is used. Specifically, in VVC, GPM is used for inter-frame prediction blocks. In the examples, GPM is only applied to CUs of 8×8 or larger. As a merging mode, GPM can be represented by signals using CU-level flags. Other merging modes alongside it include regular merging mode, merge with motion vector difference (MMVD) mode, inter-intra-interjoint prediction (CIIP) mode, and sub-block merging mode.

[0074] When using GPM mode on a CU, one of several partitioning methods is employed, dividing the CU into two geometrically shaped partitions using dividing edges. In some examples, 64 different partitioning methods are used. A partitioning method can be distinguished by 24 angles (non-uniformly quantized between 0° and 360°), with up to four dividing edges relative to the center of the CU at each angle. Dividing edges are lines that intersect the boundary of the CU and divide it into two partitions.

[0075] Figure 4 The diagram shows some examples of the 24 angles used in GPM. These angles can be identified using angle indices, such as angle indices 0 to 23 used in some examples.

[0076] Figure 5 A graph showing the possible dividing edges for angle index 3 in the example is provided. Figure 5 In this context, angle index 3 can have four possible splitting edges. It should be noted that for some angle indices, there can be three possible splitting edges corresponding to each angle index.

[0077] In some examples, each geometric partition in the CU uses its own motion for inter-frame prediction. In these examples, only unidirectional prediction is allowed for each partition; that is, each partition has one motion vector and one reference image index. Unidirectional prediction motion constraints are used to ensure similarity to bidirectional prediction, employing two motion-compensated predictions for each CU.

[0078] In some examples, when GPM is used for the current CU, signals are further used to indicate the geometric partition index (e.g., indicating angles and edges) and two merge indexes (one merge index per partition). In the example, the number of maximum GPM candidate sizes is explicitly indicated by signals at the stripe level, and the syntax binarization result used for the GPM merge index is specified.

[0079] It should also be noted that in some examples, the two partitions of an inter-coded block can be encoded separately using inter-frame prediction and intra-frame prediction. Intra-frame prediction information can be considered as intra-frame prediction information associated with the inter-coded block.

[0080] It should also be noted that in some codecs, buffers are used to store intra-prediction information for the CUs. For each CU, whether inter-coded or intra-coded, intra-prediction information is derived and stored in the buffer. In some examples, the intra-prediction information in the buffer can be used for transform kernel derivation. Transform kernel derivation can be used for the primary transform and / or secondary transform.

[0081] In some codec examples (e.g., VVC), multiple transform types can be used in the main transform, such as Type-2 DCT (DCT-2), Type-7 DST (DST-7), Type-8 DCT (DCT-8), etc. In some examples, a technique called multiple transform selection (MTS) can be used. In some examples, explicit MTS can use signals to explicitly indicate the selection of the transform kernel. In another example, implicit MTS can implicitly derive the selection of the transform kernel. In some aspects, explicit MTS can be applied to both intra-coded blocks and inter-coded blocks, while implicit MTS can be used only for intra-coded blocks. In one aspect, in explicit MTS, the selection of DST-7 / DCT-8 is indicated by explicit signaling of the transform type. In the other aspect, in implicit MTS, the transform type is selected based on encoded information known to both the encoder and decoder, without the need for transform type signaling.

[0082] In some examples, in explicit MTS, an index (e.g., represented by `mts_idx`) is used at the end of the CU-level syntax to indicate the transform types used for horizontal and vertical transforms. In these examples, the value of `mts_idx` ranges from 0 to 4. For instance, a value of 0 indicates DCT-2 for both horizontal and vertical transforms; a value of 1 indicates DST-7 for both horizontal and vertical transforms; a value of 2 indicates DCT-8 for both horizontal and vertical transforms; a value of 3 indicates DST-7 for both horizontal and vertical transforms; and a value of 4 indicates DCT-8 for both horizontal and vertical transforms.

[0083] In some examples, a secondary transform can be performed after the primary transform. For instance, the low-frequency non-separable transform (LFNST) is a non-separable transform that can be applied to the upper-left low-frequency region of the primary transform coefficients. In some examples (e.g., VVC), LFNST can be applied to an intra-coded block using DCT-2 as the primary transform. The transform kernel defined in the LFNST can include multiple transform sets, such as the four transform sets in VVC. In some examples, the selection of which transform set from the four LFNST sets (e.g., denoted by lfnstSetIdx) depends on the intra-prediction mode (e.g., denoted by intraPredMode).

[0084] Figure 6 An example of mapping LFNST sets to tables of intra-frame prediction modes is shown.

[0085] This disclosure provides techniques for video compression, including the derivation of transform kernels for inter-frame coded blocks using intra-frame prediction mode information. For example, for a current block in a current image encoded at least partially using an inter-frame prediction mode (which generates prediction samples based on a reference image different from the current image), an intra-frame prediction mode associated with the current block of the current image is obtained. At least one transform kernel can be determined based on the intra-frame prediction mode associated with the current block. The current block is then encoded / decoded based on this at least one transform kernel.

[0086] In some aspects, at least one primary transform kernel or at least one non-primary transform kernel (e.g., at least one secondary transform kernel) of a block in inter-frame and intra-frame joint prediction (CIIP) mode can be derived based on the block's intra-frame prediction mode information. When a block is in CIIP mode, the block can be predicted based on a combination of the block's inter-frame prediction information and the block's intra-frame prediction information. At least one transform kernel can be derived based on the block's intra-frame prediction information.

[0087] In some embodiments, at least one transform kernel is derived using the intra-prediction mode of a block in CIIP mode (used to generate the intra-prediction portion of that block). It should be noted that the intra-prediction mode of a block in CIIP mode can be obtained using any suitable technique for deriving at least one transform kernel. In one example, the intra-prediction mode is a predefined intra-prediction mode. In another example, the intra-prediction mode is an intra-prediction mode represented by a signal. In yet another example, the intra-prediction mode is an intra-prediction mode derived on the decoder side.

[0088] It should also be noted that any suitable method that can derive at least one master core or non-master core based on intra-frame prediction mode information can be used to derive the master transform core or non-master transform core of a block in CIIP mode.

[0089] In some examples, the method of deriving the primary transform kernel or non-primary transform kernel using intra-frame prediction mode information of intra-coded blocks can be used to derive at least one primary transform kernel or non-primary transform kernel of a block in a CIIP mode using intra-frame prediction mode information of a block in a CIIP mode. For example, Figure 6 The table in the table can be used to derive at least one transform kernel of a block in CIIP mode based on the intra-prediction information of the block in CIIP mode.

[0090] In some embodiments, intra-prediction modes are derived by comparing the template costs (e.g., SAD, SATD, etc.) of candidate intra-prediction modes of blocks in a CIIP mode. The derived intra-prediction modes are used to determine at least one primary transform kernel or non-primary transform kernel using a predefined mapping table. In the example, the predefined mapping table maps various intra-prediction modes to at least one primary transform kernel and / or at least one secondary transform kernel.

[0091] In the example, for a candidate intra-prediction mode, the candidate intra-prediction mode is applied to a template of the current block to derive candidate reconstructed samples for that template. This template is, for example, at least one row of neighboring samples above the current block, at least one column of neighboring samples to the left of the current block, a combination of one row of neighboring samples above the current block and one column of neighboring samples to the left of the current block, or neighboring samples in the L-shaped region at the top left corner of the current block. In the example, the candidate reconstructed samples of this template can be compared with the already reconstructed samples of this template (the template that was reconstructed when reconstructing the current block) to calculate the sum of absolute differences (SAD) as the SAD template cost of the candidate intra-prediction mode.

[0092] In another example, the sum of absolute transformed difference (SATD) is calculated based on candidate reconstructed samples of the template and the reconstructed samples of the template. For example, the Hadamard transform is applied to the difference between the candidate reconstructed samples of the template and the reconstructed samples of the template to obtain transform coefficients in the frequency domain, and the SATD is calculated based on the transform coefficients in the frequency domain as the SATD template cost of the candidate intra-frame prediction mode.

[0093] In the example, based on the template cost of the candidate intra-prediction modes (e.g., SAD, SATD, etc.), the candidate intra-prediction mode with the minimum template cost is used as the intra-prediction mode to select the transform set (e.g., at least one primary transform kernel and / or non-primary transform kernel) using a predefined mapping table.

[0094] On one hand, a list of candidate intra-prediction modes is derived based on the template costs of each candidate intra-prediction mode, wherein the template costs (e.g., SAD, SATD, etc.) of each candidate intra-prediction mode have been sorted. Then, a candidate intra-prediction mode is selected from this list, and a transform set (e.g., at least one primary transform kernel, at least one secondary transform kernel, etc.) is determined using the selected intra-prediction mode and a predefined mapping table. In the example, a candidate intra-prediction mode selected from the candidate intra-prediction modes can be signaled in the bitstream using syntax (such as indicating the index of the candidate intra-prediction mode in the list).

[0095] In another example, the candidate intra-prediction mode selected in the list is implicitly determined by comparing other cost metrics among the candidate intra-prediction modes in the list. In this example, the mean removal SAD cost value is calculated for the top two candidate intra-prediction modes in the list, and a candidate intra-prediction mode with a lower mean removal SAD cost value can be used to select a transform set (e.g., at least one primary transform kernel and / or a non-primary transform kernel) based on a predefined mapping table.

[0096] In some embodiments, in CIIP mode, intra-prediction modes are derived by comparing the occurrence counts of neighboring blocks using each intra-prediction mode in a predefined region adjacent to the current block. The derived intra-prediction modes are used (e.g., according to a predefined mapping table) to determine at least one transform kernel (e.g., at least one primary transform kernel, at least one secondary transform kernel, etc.).

[0097] In some examples, the intra-prediction pattern that appears most frequently in a predefined region adjacent to the current block is used to select the transform kernel (e.g., at least one primary transform kernel, at least one secondary transform kernel, etc.).

[0098] In some examples, a list of intra-prediction modes ordered by frequency of occurrence in a predefined region adjacent to the current block is obtained. One intra-prediction mode in the list is used (e.g., according to a predefined mapping table) to determine at least one transform kernel (e.g., at least one primary transform kernel, at least one secondary transform kernel, etc.). In the examples, the intra-prediction mode used to determine the transform kernel can be signaled in the bitstream using syntax (e.g., an index indicating the intra-prediction mode used in the list).

[0099] In another example, the intra prediction mode used is implicitly determined by comparing other cost metrics (e.g., template cost, etc.) between intra prediction modes in the list.

[0100] In some embodiments, when a specified intra-prediction mode is used in CIIP mode, the transform kernel is derived using the decoder-side intra-prediction mode derivation method.

[0101] In the example, the specified intra-prediction mode is planar mode. For example, when reconstructing the current block in CIIP mode using planar mode, the intra-prediction mode of the current block is derived using the decoder-side intra-prediction mode derivation method, and the derived intra-prediction mode is used to derive the transform kernel (e.g., at least one primary kernel, at least one secondary transform kernel, etc.) according to a predefined mapping table.

[0102] In another example, the specified intra-prediction mode is either a horizontal planar mode or a vertical planar mode. For example, when using a horizontal planar mode or a vertical planar mode in CIIP mode to reconstruct the current block in CIIP mode, the intra-prediction mode for the current block is derived using a decoder-side intra-prediction mode derivation method, and the derived intra-prediction mode is used to derive the transform kernel (e.g., at least one primary kernel, at least one secondary transform kernel, etc.) according to a predefined mapping table.

[0103] In another example, the specified intra-prediction mode is DC mode. For example, when reconstructing the current block in CIIP mode using DC mode, the intra-prediction mode of the current block is derived using the decoder-side intra-prediction mode derivation method, and the derived intra-prediction mode is used to derive the transform kernel (e.g., at least one primary kernel, at least one secondary transform kernel, etc.) according to a predefined mapping table.

[0104] In some embodiments, a predefined intra-frame mode (e.g., a planar mode) is used to determine the transform kernel. For example, when the current block is encoded in CIIP mode, regardless of the intra-frame prediction mode used for the reconstruction of the current block, the predefined intra-frame prediction mode is used to derive the transform kernel (e.g., at least one primary kernel, at least one secondary transform kernel, etc.) based on a predefined mapping table.

[0105] In some aspects, at least one primary transform kernel or at least one non-primary transform kernel (e.g., at least one secondary transform kernel) of a block in GPM mode can be derived based on intra-prediction mode information associated with the block. When a block is in GPM mode, the block is further partitioned into two parts using non-rectangular boundaries, such as a first part and a second part. In some examples, the first part uses inter-frame prediction, and the second part uses intra-frame prediction. In examples, at least one primary transform kernel or at least one non-primary transform kernel (e.g., at least one secondary transform kernel) of the block in GPM mode can be derived based on intra-prediction mode information associated with the second part.

[0106] In some embodiments, the intra-prediction mode used to generate the intra-prediction signal is used to determine the transform kernel. For example, a block in a GPM mode includes a first part using inter-frame prediction and a second part performing intra-prediction based on an intra-prediction mode. The intra-prediction mode of the second part of the block in the GPM mode is used to derive at least one transform kernel. It should be noted that the intra-prediction mode for deriving the second part of at least one transform kernel can be obtained by any suitable technique. In one example, the intra-prediction mode is a predefined intra-prediction mode. In another example, the intra-prediction mode is an intra-prediction mode represented by a signal. In yet another example, the intra-prediction mode is an intra-prediction mode derived on the decoder side.

[0107] It should also be noted that any suitable method that can derive at least one master core or non-master core based on intra-frame prediction mode information can be used to derive the master transform core or non-master transform core of a block in GPM mode.

[0108] In some examples, the method of deriving the primary transform kernel or non-primary transform kernel using intra-frame prediction mode information of intra-coded blocks can be used to derive at least one primary transform kernel or non-primary transform kernel of a block in GPM mode based on the intra-frame prediction mode information of the block in GPM mode. For example, Figure 6 The tables in the table can be used to derive at least one subtransform kernel of a block in GPM mode based on intra-frame prediction information of the block (e.g., the intra-coded portion of the block).

[0109] In some embodiments, intra-prediction modes are derived by comparing the template costs (e.g., SAD, SATD, etc.) of each candidate intra-prediction mode in a block of GPM mode. The derived intra-prediction modes are used to determine at least one primary transform kernel or non-primary transform kernel according to a predefined mapping table. In the example, the predefined mapping table maps various intra-prediction modes to at least one primary transform kernel and / or at least one secondary transform kernel.

[0110] In the example, based on the template cost of each candidate intra-prediction mode (e.g., SAD, SATD, etc.), the candidate intra-prediction mode with the lowest template cost is used as the intra-prediction mode associated with the current block to select a transform set (e.g., at least one primary transform kernel and / or a non-primary transform kernel) according to a predefined mapping table.

[0111] On one hand, a list of candidate intra-prediction modes, sorted by template cost (e.g., SAD, SATD, etc.), is obtained based on the template cost of the candidate intra-prediction modes. Then, one candidate intra-prediction mode from the list is selected to determine the transform set (e.g., at least one primary transform kernel, at least one secondary transform kernel, etc.) according to a predefined mapping table. In the example, the selected candidate intra-prediction mode can be represented in the bitstream using a syntax (e.g., indicating the index of the candidate intra-prediction mode in the list) as a signal.

[0112] In another example, the selected candidate intra-prediction mode in the list is implicitly determined by comparing other cost metrics among the candidate intra-prediction modes in the list. In this example, mean-demeaned SAD cost values ​​are calculated for the top two candidate intra-prediction modes in the list, and a transform set (e.g., at least one primary transform kernel and / or a non-primary transform kernel) can be selected based on a predefined mapping table using a candidate intra-prediction mode with a lower mean-demeaned SAD cost value.

[0113] In some examples, a block in GPM mode can have multiple candidate LFNST / NSPT transform kernel sets. In some examples, multiple candidate intra-prediction modes are associated with the block in GPM mode, each mapped to a different LFNST / NSPT transform kernel set. A signal can be used to represent a syntax indicating which candidate intra-prediction mode is used to determine the LFNST / NSPT transform kernel set used in encoding / decoding. For example, to determine that two candidate intra-prediction modes are associated with the block in GPM mode, these two candidate intra-prediction modes can be appropriately ordered in a list, and then a block-level flag (1 bit) can be signaled in the bitstream to indicate which candidate intra-prediction mode in the list is used to determine the mapped LFNST / NSPT transform kernel set for encoding / decoding the block.

[0114] In some embodiments, an intra-prediction mode is derived by comparing the occurrence counts of neighboring blocks using various intra-prediction modes in a predefined region adjacent to the current block in GPM mode. The derived intra-prediction mode is used (e.g., according to a predefined mapping table) to determine at least one transform kernel (e.g., at least one primary transform kernel, at least one secondary transform kernel, etc.).

[0115] In some examples, the intra-prediction pattern that appears most frequently in a predefined region adjacent to the current block is used to select the transform kernel (e.g., at least one primary transform kernel, at least one secondary transform kernel, etc.).

[0116] In some examples, a list of intra-prediction modes ordered by frequency of occurrence in a predefined region adjacent to the current block is derived. One intra-prediction mode in the list is used (e.g., according to a predefined mapping table) to determine at least one transform kernel (e.g., at least one primary transform kernel, at least one secondary transform kernel, etc.). In the examples, the intra-prediction mode used to determine the transform kernel can be signaled in the bitstream using syntax (e.g., an index indicating the intra-prediction mode used in the list).

[0117] In another example, the intra-prediction mode used is implicitly determined by comparing other cost metrics (e.g., template cost, etc.) of each intra-prediction mode in the list.

[0118] In some embodiments, an intra-prediction mode dedicated to a GPM mode with two partitions is used to determine the transform kernel.

[0119] In the example, a parallel intra-frame (prediction) mode associated with the partition boundary is used, such as... Figure 7 As shown in the image.

[0120] Figure 7A diagram of the current block (710) is shown in some examples. The current block (710) is encoded using GPM mode. The current block (710) is partitioned into a first part (711) and a second part (712) by partition boundaries (720). The first part (711) uses inter-frame coding, while the second part (712) uses intra-frame coding. Figure 7 In the example, the intra-prediction mode can be determined based on the partition boundary (720), for example, whose angle is parallel to the partition boundary (720), as indicated by the direction line (721). The intra-prediction mode corresponding to the direction line (721) can be referred to as the parallel intra-prediction mode of the block (710) in the GPM mode.

[0121] It should be noted that in some examples, portions of a block in GPM mode can be entirely intra-coded or entirely inter-coded. For example, a block in GPM mode may be partitioned into a first portion and a second portion by a partition boundary. The first portion is inter-coded using first motion information, while the second portion is inter-coded using second motion information different from the first motion information. A parallel intra-frame (predictive) mode can be derived from this partition boundary, and this parallel intra-frame (predictive) mode can be used to determine at least one transform kernel (e.g., at least one primary transform kernel, at least one secondary transform kernel, etc.) (e.g., according to a predefined mapping table).

[0122] In another example, a vertical intra-frame mode associated with the partition boundary is used, such as... Figure 8 As shown.

[0123] Figure 8 A diagram of the current block (810) is shown in some examples. The current block (810) is encoded in GPM mode. The current block (810) is partitioned into a first part (811) and a second part (812) by a partition boundary (820). The first part (811) is inter-coded, while the second part (812) is intra-coded. Figure 8 In the example, the intra-prediction mode can be determined based on the partition boundary (820), for example, its angle being perpendicular to the partition boundary (820), as indicated by the direction line (821). The intra-prediction mode corresponding to the direction line (821) can be referred to as the vertical intra-prediction mode of the block (810) in the GPM mode.

[0124] It should be noted that in some examples, portions of a block in GPM mode can be entirely intra-coded or entirely inter-coded. For example, a block in GPM mode can be partitioned into a first portion and a second portion by partition boundaries. The first portion is inter-coded using first motion information, and the second portion is inter-coded using second motion information different from the first motion information. The vertical intra-frame (predictive) mode can be derived from the partition boundaries, and the vertical intra-frame (predictive) mode can be used to determine at least one transform kernel (e.g., at least one primary transform kernel, at least one secondary transform kernel, etc.) (e.g., according to a predefined mapping table).

[0125] In some embodiments, a predefined intra-frame mode (e.g., planar mode) is used to determine the transform set (e.g., at least one primary transform kernel and / or at least one secondary transform kernel). For example, when the current block is encoded in GPM mode, regardless of the partition boundary orientation or the intra-frame prediction mode used for reconstruction of a portion of the current block, a predefined intra-frame prediction mode (e.g., planar mode) is used to derive the transform kernel (e.g., at least one primary kernel, at least one secondary transform kernel, etc.) based on a predefined mapping table.

[0126] Figure 9 A flowchart outlining one aspect of the method (900) of this disclosure is shown. The method (900) can be used in a video decoder. In various aspects, the method (900) is executed by processing circuitry, such as processing circuitry that performs the functions of the video decoder (110), processing circuitry that performs the functions of the video decoder (210), etc. In some aspects, the method (900) is implemented by software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the method (900). The process begins at (S901) and proceeds to (S910).

[0127] At (S910), the encoded video stream is received. The encoded video stream includes encoded information for multiple images.

[0128] At (S920), based on the encoded information, it is determined that the current block in the current image is encoded using an inter-frame prediction mode, wherein the prediction samples of the current block in the inter-frame prediction mode are generated at least partially based on a reference block in a reference image, which is different from the current image.

[0129] At (S930), the intra-frame prediction mode associated with the current block in the current image is obtained.

[0130] At (S940), at least one transform kernel is determined based on the intra-prediction mode associated with the current block.

[0131] At (S950), the current block is reconstructed based on the at least one transformation kernel.

[0132] In some examples, the transform coefficients of the current block are decoded from the encoded information. At least one inverse transform is performed on the transform coefficients based on this at least one transform kernel to compute residual values ​​for the samples in the current block. The samples of the current block are reconstructed based on these residual values ​​and the prediction results for the current block, which are at least partially based on a reference block in a reference image.

[0133] In some examples, the current block is encoded using at least one of the inter-intra-frame joint prediction (CIIP) mode and the geometric partitioning mode (GPM). When the current block is encoded in inter-intra-frame joint prediction (CIIP) mode, the CIIP mode includes an intra-prediction mode for generating intra-prediction portions of the prediction samples. When the current block is encoded in geometric partitioning mode (GPM), one partition of the current block is encoded in an intra-prediction mode.

[0134] In some examples, at least one transform kernel is determined based on a predefined mapping table that maps intra-prediction modes to at least one transform kernel.

[0135] In one example, the intra-prediction mode is a predefined intra-prediction mode. In another example, the intra-prediction mode is an intra-prediction mode represented by a signal, such as one obtained based on a signal in the encoded video bitstream that indicates the intra-prediction mode. In yet another example, the intra-prediction mode is obtained from an intra-prediction mode derived on the decoder side.

[0136] In some examples, template cost values ​​are calculated separately for multiple candidate intra-prediction modes. The intra-prediction mode with the lowest template cost value is selected from the multiple candidate intra-prediction modes.

[0137] In some examples, template cost values ​​are calculated for multiple candidate intra-prediction modes. A list is formed. Based on the template cost values, this list includes at least two candidate intra-prediction modes from the multiple candidate intra-prediction modes. The at least two candidate intra-prediction modes in this list are sorted according to their template cost values. Intra-prediction modes are selected from this list. In the example, a syntax is decoded from the encoded video stream, which indicates the index in the list, and an intra-prediction mode is selected from the list according to this syntax.

[0138] In another example, a second cost value (different from the template cost value) is calculated for at least two candidate intra-prediction modes in the list. An intra-prediction mode is then selected from the list based on the second cost value.

[0139] In the example, the first and second candidate intra-prediction modes are sorted. A flag is decoded from the encoded video stream. Based on this flag, one of the first and second candidate intra-prediction modes is selected as the intra-prediction mode.

[0140] The at least one transform core may include any one of the following: at least one primary transform core, at least one secondary transform core, at least one low-frequency non-separable transform (LFNST) core and / or at least one non-separable primary transform core.

[0141] In some examples, the frequency of occurrence of intra-prediction modes used by adjacent blocks in a predefined region of the current block is counted. The intra-prediction mode with the highest frequency of occurrence is selected from the used intra-prediction modes.

[0142] In some examples, the occurrence counts of intra-prediction modes used by adjacent blocks within a predefined region of the current block are counted. A list is formed based on these occurrence counts, including at least two of these used intra-prediction modes, ordered by their occurrence counts. An intra-prediction mode is then selected from this list.

[0143] In one example, a syntax is decoded from the encoded video stream, indicating an index in a list, and an intra-prediction mode is selected from the list based on this syntax. In another example, cost values ​​are calculated for at least two used intra-prediction modes in the list. An intra-prediction mode is then selected from the list based on a second cost value.

[0144] In some examples, when the current block is encoded using at least one of the following modules: planar mode, planar horizontal mode, planar vertical mode, DC mode, and / or predefined intra-frame mode, the intra-frame prediction mode is derived from the decoder-side intra-frame prediction mode derivation.

[0145] In some examples, the current block is encoded in Geometric Partitioning (GPM) mode. The intra-prediction mode is determined by an angle parallel or perpendicular to the current block partition boundary.

[0146] Then, the process proceeds to (S999) and terminates.

[0147] Method (900) can be adjusted as appropriate. At least one step in method (900) can be modified and / or omitted. At least one additional step can be added. Any suitable execution order can be used.

[0148] Figure 10A flowchart outlining one aspect of the method (1000) of this disclosure is shown. The method (1000) can be used in a video encoder. In various aspects, the method (1000) is executed by processing circuitry, such as processing circuitry that performs the functions of the video encoder (103), processing circuitry that performs the functions of the video encoder (303), etc. In some aspects, the method (1000) is implemented by software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the method (1000). The process begins at (S1001) and proceeds to (S1010).

[0149] At (S1010), it is determined that the current block in the current image will be encoded using the inter-frame prediction mode.

[0150] At (S1020), a predicted sample of the current block is generated based at least in part on a reference block in a reference image that is different from the current image.

[0151] At (S1030), the intra-frame prediction mode associated with the current block in the current image is obtained.

[0152] At (S1040), at least one transform kernel is determined based on the intra-prediction mode associated with the current block.

[0153] At (S1050), the current block is encoded into bits in the bitstream based on the at least one transform kernel.

[0154] In some examples, at least one transformation is performed on the residual values ​​of samples in the current block according to at least one transformation kernel to calculate transform coefficients, the residual values ​​being calculated based on the predicted samples and original values ​​of samples in the current block. The transform coefficients are encoded as bits in the bitstream.

[0155] In some examples, the inter-frame prediction mode is at least one of inter-intra-frame joint prediction (CIIP) mode and geometric partitioning mode (GPM). When the current block is encoded in inter-intra-frame joint prediction (CIIP) mode, the CIIP mode includes an intra-frame prediction mode for generating prediction samples. When the current block is encoded in geometric partitioning mode (GPM), a partition of the current block is encoded in intra-frame prediction mode.

[0156] In some examples, at least one transform kernel is determined based on a predefined mapping table that maps intra-prediction modes to at least one transform kernel.

[0157] In the example, the intra-prediction mode is a predefined intra-prediction mode. In another example, the intra-prediction mode is obtained based on the intra-prediction mode derived from the decoder side.

[0158] In some examples, the syntax indicating the intra-frame prediction mode is encoded as at least one bit in the bitstream.

[0159] In some examples, template cost values ​​are calculated separately for multiple candidate intra-prediction modes. From these candidate intra-prediction modes, the intra-prediction mode with the lowest template cost value is determined.

[0160] In some examples, the template cost value for each of the multiple candidate intra-prediction modes is calculated separately. A list is formed based on the template cost values, comprising at least two candidate intra-prediction modes from the multiple candidate intra-prediction modes, ordered by their template cost values. An intra-prediction mode is then selected from this list.

[0161] In some examples, the syntax is encoded as at least one bit in the bitstream that indicates the index of the intra-prediction mode in the list.

[0162] In some examples, a second cost value is calculated for each of the at least two candidate intra-prediction modes in the list, and an intra-prediction mode is selected from the list based on the second cost value.

[0163] In the example, one of the first candidate intra-prediction mode and the second candidate intra-prediction mode is selected as the intra-prediction mode. A flag is encoded into the bitstream that indicates one of the first candidate intra-prediction mode and the second candidate intra-prediction mode.

[0164] In some examples, at least one transform core includes any one of at least one primary transform core, at least one secondary transform core, at least one low-frequency non-separable transform (LFNST) core, and / or at least one non-separable primary transform core.

[0165] In some examples, the frequency of occurrence of intra-prediction modes used by adjacent blocks in a predefined region of the current block is counted. The intra-prediction mode with the highest frequency of occurrence is selected from the used intra-prediction modes.

[0166] In some examples, the occurrence counts of intra-prediction modes used by adjacent blocks in a predefined region of the current block are counted. A list is formed based on these occurrence counts, including at least two intra-prediction modes from the used intra-prediction modes, ordered by their occurrence counts. An intra-prediction mode is then selected from this list.

[0167] In the example, a syntax is encoded as at least one bit in the bitstream, which indicates the index of the intra-prediction mode selected from the list.

[0168] In another example, cost values ​​are calculated for at least two intra-prediction modes used in the list, and an intra-prediction mode is selected from the list based on the cost values.

[0169] In some examples, when encoding the current block using at least one of a planar mode, a planar horizontal mode, a planar vertical mode, a DC mode, and / or a predefined intra-frame mode, the intra-frame prediction mode is determined based on decoder-side intra-frame prediction mode derivation.

[0170] In some examples, the current block is encoded using Geometric Partitioning (GPM). The intra-prediction mode is determined by the angle parallel or perpendicular to the partition boundary of the current block.

[0171] Then, the process proceeds to (S1099) and terminates.

[0172] Method (1000) can be adapted as appropriate. At least one step in method (1000) can be modified and / or omitted. At least one additional step can be added. Any suitable execution order can be used.

[0173] According to one aspect of this disclosure, a method for processing visual media data is provided. In this method, a conversion between a bitstream of visual media data and a visual media file is performed according to a format rule. For example, the bitstream may be a bitstream decoded / encoded in any of the decoding and / or encoding methods described herein. The format rule may specify at least one constraint on the bitstream and / or at least one method performed by the decoder and / or encoder.

[0174] In the example, the bitstream carries encoded information from multiple images. The format rule specifies that the current block in the current image is encoded using an inter-frame prediction mode, generating prediction samples of the current block in the inter-frame prediction mode based at least in part on reference blocks in a reference image different from the current image. The format rule also specifies obtaining an intra-frame prediction mode associated with the current block in the current image, determining at least one transform kernel based on the intra-frame prediction mode associated with the current block, and reconstructing the current block based on the at least one transform kernel.

[0175] The above-described techniques can be implemented as computer software using computer-readable instructions and physically stored on at least one computer-readable medium. For example, Figure 11 A computer system (1100) suitable for implementing certain aspects of the disclosed subject matter is shown.

[0176] Computer software can be coded using any appropriate machine code or computer language. These machine codes or computer languages ​​can be assembled, compiled, linked, and other mechanisms to create code that includes instructions. These instructions can be executed directly by at least one computer central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, or other methods.

[0177] The instructions can be executed on various types of computers or their components, including personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0178] Figure 11 Examples of components of the computer system (1100) are shown herein and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing various aspects of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement on any one or combination of the components shown in the examples of the computer system (1100).

[0179] The computer system (1100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from at least one human user via, for example, tactile input (e.g., key presses, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0180] The input human-machine interface device may include at least one of the following (only one of each is depicted): keyboard (1101), mouse (1102), touchpad (1103), touch screen (1110), data glove (not shown), joystick (1105), microphone (1106), scanner (1107), camera (1108).

[0181] The computer system (1100) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of at least one human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (1110), data gloves (not shown), or joystick (1105), but may also include tactile feedback devices not used as input devices), audio output devices (e.g., speakers (1109), headphones (not shown)), visual output devices (e.g., screens (1110), virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown). Screens (1110) include CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capability and each with or without tactile feedback capability. Some of these screens are capable of outputting two-dimensional visual output or more than three-dimensional output in a manner such as stereoscopic output.

[0182] The computer system (1100) may also include human-accessible storage devices and their associated media, such as optical media (1121) including CD / DVD ROM / RW (1120) with media such as CD / DVD, thumb drives (1122), removable hard disks or solid-state drives (1123), conventional magnetic media such as magnetic tapes and floppy disks (not shown), devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown), etc.

[0183] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other volatile signals.

[0184] The computer system (1100) may also include an interface (1154) to at least one communication network (1155). The network may be, for example, wireless, wired, or optical. The network may further be a local area network (LAN), wide area network (WAN), metropolitan area network (MAN), vehicular and industrial network, real-time network, delay-tolerant network, etc. Examples of networks include LANs such as Ethernet, wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), wired or wireless wide area digital TV networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicular and industrial networks (including CAN bus), etc. Some networks typically require external network interface adapters attached to certain general-purpose data ports or peripheral buses (1149) (e.g., USB ports of the computer system (1100)); other networks are typically integrated into the core of the computer system (1100) via connections to system buses (e.g., Ethernet interfaces connected to PC computer systems or cellular network interfaces connected to smartphone computer systems). Using any of these networks, the computer system (1100) can communicate with other entities. Such communication can be unidirectional (receive-only, e.g., broadcasting TV), unidirectional (transmit-only, e.g., CAN bus to certain CAN bus devices), or bidirectional, e.g., to other computer systems using a local area network or a wide area digital network. As mentioned above, certain protocols and protocol stacks can be used on each of these networks and network interfaces.

[0185] The aforementioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the kernel (1140) of the computer system (1100).

[0186] The kernel (1140) may include at least one central processing unit (CPU) (1141), a graphics processing unit (GPU) (1142), a dedicated programmable processing unit (1143) in the form of a field-programmable gate area (FPGA), a task-specific hardware accelerator (1144), a graphics adapter (1150), etc. These devices, along with read-only memory (ROM) (1145), random access memory (1146), and internal mass storage (1147) such as an internal non-user-accessible hard disk drive, SSD, etc., may be connected via a system bus (1148). In some computer systems, the system bus (1148) may be accessed as at least one physical connector to allow for expansion by adding CPUs, GPUs, etc. Peripheral devices may be directly attached or attached to the kernel's system bus (1148) via a peripheral bus (1149). In this example, a screen (1110) may be connected to a graphics adapter (1150). Peripheral bus architectures include PCI, USB, etc.

[0187] The CPU (1141), GPU (1142), FPGA (1143), and accelerator (1144) can execute certain instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM (1145) or RAM (1146). Transient data can also be stored in RAM (1146), while permanent data can be stored, for example, in internal mass storage (1147). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with at least one CPU (1141), GPU (1142), mass storage (1147), ROM (1145), RAM (1146), etc.

[0188] Computer-readable media may have computer code thereon for performing operations of various computer implementations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or they may be of types known and available to those skilled in the art of computer software.

[0189] By way of example and not limitation, a computer system having an architecture (1100), particularly a kernel (1140), may be functionally provided by at least one processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in at least one tangible computer-readable medium. Such a computer-readable medium may be a medium associated with user-accessible mass storage as described above, and some storage of the kernel (1140) having non-volatile properties, such as internal kernel mass storage (1147) or ROM (1145). Software implementing various aspects of this disclosure may be stored in such a device and executed by the kernel (1140). Depending on specific needs, the computer-readable medium may include at least one memory device or chip. The software may cause the kernel (1140) and, in particular, the processor therein (including a CPU, GPU, FPGA, etc.) to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (1146) and modifying such data structures according to software-defined processes. In addition, or as an alternative, a computer system may provide functionality as a result of logic hard-wired in circuitry (e.g., an accelerator (1144)) or otherwise embodied, which may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may cover logic, and vice versa. Where appropriate, references to computer-readable media may cover circuitry (such as integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. This disclosure covers any suitable combination of hardware and software.

[0190] The use of “at least one of” or “one of” in this disclosure is intended to include any one or a combination of the listed elements. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include either A or B or (A and B). The use of “one of” does not exclude any combination of the listed elements where applicable, such as when the elements are not mutually exclusive.

[0191] While this disclosure has described several examples of aspects, there are variations, substitutions, and various alternative equivalents that fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are therefore within its spirit and scope.

[0192] The above disclosure also covers the features mentioned below. These features can be combined in various ways, and are not limited to the combinations mentioned below.

[0193] (1) A video decoding method, comprising: receiving an encoded video stream, the encoded video stream including encoded information of a plurality of images; determining, based on the encoded information, that a current block in a current image is encoded using an inter-frame prediction mode, and generating a prediction sample of the current block in the inter-frame prediction mode based at least in part on a reference block in a reference image different from the current image; obtaining an intra-frame prediction mode associated with the current block in the current image; determining at least one transform kernel based on the intra-frame prediction mode associated with the current block; and reconstructing the current block based on the at least one transform kernel.

[0194] (2) The method according to feature (1), wherein the reconstruction includes: decoding the transform coefficients of the current block from the encoded information; performing at least one inverse transform on the transform coefficients based on the at least one transform kernel to calculate the residual value of the sample in the current block; and reconstructing the sample of the current block based on the residual value and the prediction result of the current block, the prediction result of the current block being obtained at least in part based on the reference block in the reference image.

[0195] (3) The method according to any one of features (1) to (2), wherein the current block is encoded in at least one of inter-frame joint prediction (CIIP) mode and geometric partitioning mode (GPM). When the current block is encoded in inter-frame joint prediction (CIIP) mode, the CIIP mode includes the intra prediction mode for generating the intra prediction portion of the prediction sample. When the current block is encoded in geometric partitioning mode (GPM), the partitions of the current block are encoded in the intra prediction mode.

[0196] (4) The method according to any one of features (1) to (3), wherein determining the at least one transform kernel comprises: determining the at least one transform kernel according to a predefined mapping table that maps the intra-frame prediction mode to the at least one transform kernel.

[0197] (5) The method according to any one of features (1) to (4), wherein obtaining the intra prediction mode includes at least one of the following: obtaining the intra prediction mode as a predefined intra prediction mode; obtaining the intra prediction mode based on a signal in the encoded video bitstream that indicates the intra prediction mode; and / or obtaining the intra prediction mode according to an intra prediction mode derived from the decoder side.

[0198] (6) The method according to any one of features (1) to (5), wherein obtaining the intra-prediction mode comprises: calculating the template cost value of each of a plurality of candidate intra-prediction modes respectively; and determining the intra-prediction mode having the minimum template cost value from the plurality of candidate intra-prediction modes.

[0199] (7) The method according to any one of features (1) to (6), wherein obtaining the intra-prediction mode comprises: calculating the template cost value of each of the plurality of candidate intra-prediction modes respectively; forming a list including at least two candidate intra-prediction modes from the plurality of candidate intra-prediction modes according to the template cost value, wherein the at least two candidate intra-prediction modes are sorted according to the template cost value in the list; and selecting the intra-prediction mode from the list.

[0200] (8) The method according to any one of features (1) to (7), wherein the selection comprises: decoding a syntax from the encoded video stream, the syntax indicating an index in the list; and selecting the intra-frame prediction mode from the list according to the syntax.

[0201] (9) The method according to any one of features (1) to (8), wherein the selection comprises: calculating a second cost value for each of the at least two candidate intra-prediction modes in the list; and selecting the intra-prediction mode from the list based on the second cost value.

[0202] (10) The method according to any one of features (1) to (9), wherein obtaining the intra-prediction mode comprises: sorting a first candidate intra-prediction mode and a second candidate intra-prediction mode in a list; decoding a flag from the encoded video bitstream; and selecting a mode from the first candidate intra-prediction mode and the second candidate intra-prediction mode as the intra-prediction mode based on the flag.

[0203] (11) The method according to any one of features (1) to (10), wherein the at least one transform core comprises at least one of the following: a primary transform core; a secondary transform core; a low-frequency non-separable transform (LFNST) core; and / or a non-separable primary transform core.

[0204] (12) The method according to any one of features (1) to (11), wherein obtaining the intra-prediction mode comprises: counting the number of occurrences of the intra-prediction mode used by adjacent blocks in a predefined region of the current block; and determining the intra-prediction mode that occurs most frequently from the used intra-prediction modes.

[0205] (13) The method according to any one of features (1) to (12), wherein obtaining the intra-prediction mode comprises: counting the number of occurrences of intra-prediction modes used by adjacent blocks in a predefined region of the current block; forming a list including at least two intra-prediction modes among the used intra-prediction modes based on the number of occurrences, wherein the at least two intra-prediction modes are sorted in the list according to the number of occurrences; and selecting the intra-prediction mode from the list.

[0206] (14) The method according to any one of features (1) to (13), wherein the selection comprises: decoding a syntax from the encoded video stream, the syntax indicating an index in the list; and selecting the intra-frame prediction mode from the list according to the syntax.

[0207] (15) The method according to any one of features (1) to (14), wherein the selection comprises: calculating cost values ​​for the at least two used intra-prediction modes in the list respectively; and selecting the intra-prediction mode from the list based on the cost values.

[0208] (16) The method according to any one of features (1) to (15), wherein obtaining the intra-prediction mode comprises: obtaining the intra-prediction mode based on decoder-side intra-prediction mode derivation when the current block is encoded using at least one of the following modes: plane mode, plane horizontal mode, plane vertical mode, DC mode and / or predefined intra-prediction mode.

[0209] (17) The method according to any one of features (1) to (16), wherein the current block is encoded in a geometric partitioning mode (GPM), and obtaining the intra-prediction mode comprises: determining the intra-prediction mode having an angle parallel to or perpendicular to the partitioning boundary of the current block.

[0210] (18) A video coding method, comprising: determining to encode a current block in a current image using an inter-frame prediction mode; generating a prediction sample of a sample of the current block based at least in part on a reference block in a reference image different from the current image; obtaining an intra-frame prediction mode associated with the current block in the current image; determining at least one transform kernel based on the intra-frame prediction mode associated with the current block; and encoding the current block into bits in a bitstream based on the at least one transform kernel.

[0211] (19) The method according to feature (18), wherein the encoding includes: performing at least one transformation on the residual value of the sample in the current block according to the at least one transformation kernel to calculate transformation coefficients, the residual value being calculated based on the predicted sample and the original value of the sample in the current block; and encoding the transformation coefficients into bits in the bitstream.

[0212] (20) The method according to any one of features (18) to (19), wherein the inter-frame prediction mode is at least one of inter-intra-frame joint prediction (CIIP) mode and geometric partitioning mode (GPM). When the current block is encoded in inter-intra-frame joint prediction (CIIP) mode, the CIIP mode includes the intra-frame prediction mode for generating the intra-frame prediction portion of the prediction sample. When the current block is encoded in geometric partitioning mode (GPM), a partition of the current block is encoded in the intra-frame prediction mode.

[0213] (21) The method according to any one of features (18) to (20), wherein determining the at least one transform kernel comprises: determining the at least one transform kernel according to a predefined mapping table, the mapping table mapping the intra-frame prediction mode to the at least one transform kernel.

[0214] (22) The method according to any one of features (18) to (21), wherein obtaining the intra prediction mode includes at least one of the following: obtaining the intra prediction mode as a predefined intra prediction mode; and / or obtaining the intra prediction mode according to an intra prediction mode derived from the decoder side.

[0215] (23) The method according to any one of features (18) to (22), wherein the method further comprises: encoding the syntax indicating the intra-prediction mode into at least one bit in the bitstream.

[0216] (24) The method according to any one of features (18) to (23), wherein obtaining the intra-prediction mode comprises: calculating the template cost value of each of a plurality of candidate intra-prediction modes; and determining the intra-prediction mode having the minimum template cost value from the plurality of candidate intra-prediction modes.

[0217] (25) The method according to any one of features (18) to (24), wherein obtaining the intra-prediction mode comprises: calculating the template cost value of each of a plurality of candidate intra-prediction modes respectively; forming a list including at least two candidate intra-prediction modes from the plurality of candidate intra-prediction modes based on the template cost value, wherein the at least two candidate intra-prediction modes are sorted according to the template cost value in the list; and selecting the intra-prediction mode from the list.

[0218] (26) The method according to any one of features (18) to (25) further includes: encoding a syntax into at least one bit in the bitstream, the syntax indicating the index of the intra-frame prediction mode in the list.

[0219] (27) The method according to any one of features (18) to (26), wherein the selection comprises: calculating a second cost value for each of the at least two candidate intra-prediction modes in the list; and selecting the intra-prediction mode from the list based on the second cost value.

[0220] (28) The method according to any one of features (18) to (27), wherein obtaining the intra-prediction mode comprises: selecting one of a first candidate intra-prediction mode and a second candidate intra-prediction mode as the intra-prediction mode; and encoding a flag in the bitstream, the flag indicating one of the first candidate intra-prediction mode and the second candidate intra-prediction mode.

[0221] (29) The method according to any one of features (18) to (28), wherein the at least one transform core comprises at least one of the following: a primary transform core; a secondary transform core; a low-frequency non-separable transform (LFNST) core; and / or a non-separable primary transform core.

[0222] (30) The method according to any one of features (18) to (29), wherein obtaining the intra-prediction mode comprises: counting the number of occurrences of the intra-prediction mode used by adjacent blocks in a predefined region of the current block; and determining the intra-prediction mode that occurs most frequently from the used intra-prediction modes.

[0223] (31) The method according to any one of features (18) to (30), wherein obtaining the intra-prediction mode comprises: counting the number of occurrences of intra-prediction modes used by adjacent blocks in a predefined region of the current block; forming a list including at least two intra-prediction modes among the used intra-prediction modes based on the number of occurrences; sorting the at least two intra-prediction modes in the list based on the number of occurrences; and selecting the intra-prediction mode from the list.

[0224] (32) The method according to any one of features (18) to (31) further includes: encoding a syntax into at least one bit in the bitstream, the syntax indicating the index of the intra-frame prediction mode in the list.

[0225] (33) The method according to any one of features (18) to (32), wherein the selection comprises: calculating the cost value of each of the at least two intra-prediction modes in the list; and selecting the intra-prediction mode from the list based on the cost value.

[0226] (34) The method according to any one of features (18) to (33), wherein obtaining the intra-prediction mode comprises: obtaining the intra-prediction mode by derivation of the decoder-side intra-prediction mode when the current block is encoded using at least one of a plane mode, a plane horizontal mode, a plane vertical mode, a DC mode and / or a predefined intra-prediction mode.

[0227] (35) The method according to any one of features (18) to (34), wherein encoding the current block in a geometric partitioning mode (GPM) and obtaining the intra-prediction mode comprises: determining the intra-prediction mode having an angle parallel to or perpendicular to the partitioning boundary of the current block.

[0228] (36) A method for processing visual media data, comprising: processing a bitstream including the visual media data according to a format rule, wherein: the bitstream carries a plurality of images; and the format rule specifies: encoding a current block in a current image using an inter-frame prediction mode, generating a prediction sample of the current block in the inter-frame prediction mode based at least in part on a reference block in a reference image different from the current image; obtaining an intra-frame prediction mode associated with the current block in the current image; determining at least one transform kernel based on the intra-frame prediction mode associated with the current block; and reconstructing the current block based on the at least one transform kernel.

[0229] (37) A video decoding apparatus, including a processing circuit configured to perform the method according to any one of features (1) to (17).

[0230] (38) An apparatus for video encoding, comprising processing circuitry configured to perform the method according to any one of features (18) to (35).

[0231] (39) A non-volatile computer-readable storage medium for storing instructions, wherein the instructions, when executed by at least one processor, cause the at least one processor to perform the method according to any one of features (1) to (36).

Claims

1. A video decoding method, characterized in that, include: Receive an encoded video stream, wherein the encoded video stream includes encoded information of multiple images; Based on the encoded information, it is determined that the current block in the current image is encoded using the inter-frame prediction mode, and the prediction sample of the current block in the inter-frame prediction mode is generated at least in part based on the reference block in the reference image that is different from the current image. Obtain the intra-frame prediction mode associated with the current block in the current image; At least one transform kernel is determined based on the intra-prediction mode associated with the current block; and The current block is reconstructed based on the at least one transformation kernel.

2. The method according to claim 1, characterized in that, The reconstruction includes: Decode the transform coefficients of the current block from the encoded information; Based on the at least one transform kernel, at least one inverse transform is performed on the transform coefficients to calculate the residual values ​​of the samples in the current block; and The sample of the current block is reconstructed based on the residual value and the prediction result of the current block, wherein the prediction result of the current block is obtained at least in part based on the reference block in the reference image.

3. The method according to any one of claims 1 to 2, characterized in that: When the current block is encoded in an inter-frame and intra-frame joint prediction (CIIP) mode, the CIIP mode includes the intra-frame prediction mode for generating the intra-frame prediction portion of the prediction sample. and When the current block is encoded in Geometric Partitioning (GPM) mode, a partition of the current block is encoded in the intra-prediction mode.

4. The method according to any one of claims 1 to 3, characterized in that, Determining the at least one transform kernel includes: The at least one transform kernel is determined according to a predefined mapping table, which maps the intra-frame prediction mode to the at least one transform kernel.

5. The method according to any one of claims 1 to 4, characterized in that, Obtaining the intra-frame prediction mode includes at least one of the following: Obtain the intra-frame prediction mode, wherein the intra-frame prediction mode is a predefined intra-frame prediction mode; The intra-frame prediction mode is obtained based on the signal in the encoded video stream that indicates the intra-frame prediction mode. and / or The intra-prediction mode is obtained based on the intra-prediction mode derived from the decoder side.

6. The method according to any one of claims 1 to 4, characterized in that, Obtaining the intra-frame prediction mode includes: Calculate the template cost value for each of the multiple candidate intra-frame prediction modes; and The intra-prediction mode with the minimum template cost value is determined from the plurality of candidate intra-prediction modes.

7. The method according to any one of claims 1 to 4, characterized in that, Obtaining the intra-frame prediction mode includes: Calculate the template cost value for each of the multiple candidate intra-frame prediction modes; A list is formed based on the template cost value, comprising at least two candidate intra-prediction modes from the plurality of candidate intra-prediction modes, wherein the at least two candidate intra-prediction modes are sorted according to the template cost value; and Select the intra-frame prediction mode from the list.

8. The method according to claim 7, characterized in that, The selection includes: Decode a syntax from the encoded video stream, the syntax indicating an index in the list; and Select the intra-frame prediction mode from the list according to the syntax.

9. The method according to claim 7, characterized in that, The selection includes: Calculate the second cost value for each of the at least two candidate intra-prediction modes in the list; and The intra-frame prediction mode is selected from the list based on the second cost value.

10. The method according to any one of claims 1 to 4, characterized in that, Obtaining the intra-frame prediction mode includes: Sort the first and second candidate intra-prediction modes in the list. Decode a flag from the encoded video stream; and Based on the flag, one of the first candidate intra-prediction mode and the second candidate intra-prediction mode is selected as the intra-prediction mode.

11. The method according to any one of claims 1 to 10, characterized in that, The at least one transformation kernel includes at least one of the following: Main transform kernel; Secondary transformation kernel; Low-frequency non-separable transform (LFNST) core; and / or Non-separable master transform kernel.

12. The method according to any one of claims 1 to 4, characterized in that, Obtaining the intra-frame prediction mode includes: Count the number of occurrences of intra-prediction modes used by adjacent blocks in the predefined region of the current block; and The intra prediction mode that appears most frequently is determined from the intra prediction modes used.

13. The method according to any one of claims 1 to 4, characterized in that, Obtaining the intra-frame prediction mode includes: The number of times the intra-prediction mode used by adjacent blocks in the predefined region of the current block is counted; A list is formed based on the occurrence count, comprising at least two intra-prediction modes from the used intra-prediction modes, wherein the at least two intra-prediction modes are ordered in the list according to the occurrence count; and Select the intra-frame prediction mode from the list.

14. A video encoding method, characterized in that, include: Determine to use inter-frame prediction mode to encode the current block in the current image; A predicted sample of the current block is generated, at least in part, based on a reference block in a reference image that is different from the current image. Obtain the intra-frame prediction mode associated with the current block in the current image; At least one transform kernel is determined based on the intra-prediction mode associated with the current block; and The current block is encoded into bits in the bitstream based on the at least one transform kernel.

15. A visual media data processing method, characterized in that, include: The bitstream including the visual media data is processed according to a format rule, wherein: The stream carries multiple images; and The formatting rules specify: The current block in the current image is encoded using an inter-frame prediction mode, and a prediction sample of the current block in the inter-frame prediction mode is generated based at least in part on a reference block in a reference image that is different from the current image. Obtain the intra-frame prediction mode associated with the current block in the current image; At least one transform kernel is determined based on the intra-prediction mode associated with the current block; and The current block is reconstructed based on the at least one transformation kernel.