Interaction between transform partitioning and primary / secondary transform type selection
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2025-09-16
- Publication Date
- 2026-04-17
AI Technical Summary
Existing video coding technologies face challenges in efficiently compressing and decoding video data while maintaining image quality, particularly in applications with varying distortion tolerances and bandwidth requirements, due to limitations in intra-prediction and motion compensation techniques.
The proposed method involves a transform partitioning scheme and selection of primary and secondary transform types, which includes dividing data blocks into specific partition types and applying appropriate transform types based on predefined sets, enhancing entropy coding efficiency and reducing the bit rate.
This approach improves video coding efficiency by optimizing transform partitioning and selection, leading to reduced data requirements and improved image quality, suitable for various video applications with different distortion tolerances.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 175,897, filed April 16, 2021, and U.S. Non-Provisional Patent Application No. 17 / 568,275, filed January 4, 2022, both of which are incorporated by reference in their entireties.
[0002] This disclosure describes a set of advanced video coding techniques. More specifically, the disclosure The proposed technique is a transform partitioning scheme and primary Including the interaction between the / secondary transform type selection. [Background technology]
[0003] The background discussion provided herein is intended to generally present the context for the present disclosure. The inventors' research has been carried out insofar as it is described in this Background Art section. and other descriptions that may not be admitted as prior art at the time of filing of this application. This disclosure, together with the above aspects, is not admitted expressly or impliedly as prior art to the present disclosure.
[0004] Video coding and decoding involves inter-picture prediction with motion compensation. Uncompressed digital video contains a series of pictures. Each picture can contain, for example, 1920x1080 luminance samples and associated full sampling The sequence of pictures has spatial dimensions of chromatic or subsampled chrominance samples. Fixed or variable picture rates (or frames) of, for example, 60 pictures per second or 60 frames per second. Uncompressed video can have a bit rate (also called a frame rate). For example, a 1920x1080 pixel resolution, 60 frames per second, frame rate of 1000 frames per second and 4:2:0 chroma at 8 bits per pixel per color channel Video with subsampling requires a bandwidth approaching 1.5 Gbit / s. Such a video would require over 600 GByte of storage space.
[0005] One goal of video coding and video decoding is to compress and decode an uncompressed input. Compression can be the reduction of redundancy in a video signal. This can help reduce image quality by more than two orders of magnitude in some cases. Lossless compression means that an exact copy of the original signal is reproduced. This refers to a technique that allows a signal to be reconstructed from the original signal that has been compressed by a lossy compression process. This means that the original video information is not fully preserved during coding and cannot be fully recovered during decoding. When lossy compression is used, the reconstructed The reconstructed signal may not be identical to the original signal, but the distortion between the original and the reconstructed signal may be This means that even though there may be some loss of information, the reconstructed signal is still sufficient to be useful for the intended purpose. In the case of video, lossy compression is widely used in many applications. The amount of distortion that can be tolerated depends on the application, e.g., certain consumer video streaming applications. Users of commercial applications may tolerate higher distortion than users of film and television broadcast applications. The compression ratios achievable by the filtering algorithms are chosen to reflect different distortion tolerances. That is, generally, the higher the strain tolerance, the higher the loss and Coding algorithms that result in compression ratios are possible.
[0006] Video encoders and decoders perform various functions, such as motion compensation, Fourier transform, quantization, Techniques from several broad categories and steps, including quantization and entropy coding You can use the technique.
[0007] Video codec technology can include a technique known as intracoding. In intra-coding, sample values are derived from a previously reconstructed reference picture. It is represented without reference to samples or other data. In some video codecs, it is represented as a picture. is spatially subdivided into blocks of samples. Every block of samples is an intra-sampled image. If a picture is coded in the JPEG2000 mode, the picture can be called an intra-picture. Intra pictures and their derivatives, such as independent decoder refresh pictures. The decoder state can be reset and therefore as the first picture in the encoded video bitstream and video session, can be used as a still image. Then, the sample of the block after intra prediction can be transformed into the frequency domain, and the transform coefficients so generated can be used as the entropy Before coding, the image can be quantized. Intra prediction is performed by sampling in the pre-transform domain. In some cases, the smaller the DC value after conversion, and the greater the AC The smaller the number, the better the accuracy of a given quantization step to represent the block after entropy coding. This reduces the number of bits required in the clip size.
[0008] For example, conventional intracoders such as those known from MPEG-2 generation coding techniques. However, some newer video compression techniques do not use intra prediction. The technique may be used to detect, for example, intra-coding, which is obtained when encoding and / or decoding spatial neighbors. the surrounding samples that precede in decoding order the block of data being decoded or intra-decoded. coding / decoding of blocks based on the block data and / or metadata Such techniques are hereafter referred to as "intra-prediction" techniques. In some cases, intra prediction involves using only references from the current picture being reconstructed. Note that the reference data from the first reference picture is used, and not from other reference pictures. .
[0009] Intra prediction can take many different forms. Two or more such techniques If available for a given video coding technique, the technique used is One or more intra prediction modes can be used for a particular codec. In certain cases, a mode may have submodes and / or or may be associated with various parameters, such as mode / submode information and video The intra-coding parameters of the blocks can be coded individually or collectively. The codeword for a given mode, submode, and / or The codeword used for the parameter combination is determined by the codeword used via intra prediction. This can affect the coding efficiency, so the codewords are The entropy coding technique used to convert the image into a do.
[0010] A specific mode of intra prediction was introduced in H.264 and improved in H.265, the joint search model. (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS) Newer coding techniques have further improved it. In general, intra prediction The predictor block can be formed using neighboring sample values that have been The available values of a particular set of neighboring samples along a particular direction and / or line are used to determine the predictor block. A reference to the direction used is coded in the bitstream. It can be predicted or it can be predicted.
[0011] Referring to FIG. 1A, the bottom right shows the 35 intra-motion modes specified in H.265. The nine predictor directions specified in H.265 correspond to the 33 possible predictor directions in H.265 (corresponding to the 33 angle modes in the The point where the arrows converge (101) is the sample being predicted. The arrows represent the neighboring samples that are used to predict the 101 samples from them. For example, the arrow (102) indicates that the sample (101) is adjacent to one or more neighboring samples. The arrow (10 3) indicates that the sample (101) is shifted from one or more adjacent samples to the lower left of the sample (101); This indicates a forecast angle of 22.5 degrees from the horizontal.
[0012] Still referring to FIG. 1A, at the top left is a 4×4 sample of positive (indicated by the thick dashed line). A square block (104) is depicted. The square block (104) contains 16 samples. , respectively, "S", its position in the Y dimension (e.g., row index), and its position in the X dimension ( For example, sample S21 is labeled by the Y-dimension ( It is the second sample (from the top) and the first sample (from the left) in the X dimension. Sample S44 is the fourth sample in both the Y and X dimensions in block (104). The block size is 4x4 samples, so S44 is in the bottom right. An example of a reference sample is also shown. The reference sample is R, for block (104). Each pixel is labeled with its Y position (e.g., row number) and X position (column number). In both H.264 and H.265, prediction samples that neighbor the block being reconstructed are used.
[0013] The intra-picture prediction in block 104 is performed by predicting adjacent sub-pictures according to the signaled prediction direction. It may start by copying the reference sample values from the coded sample. The resulting video bitstream indicates the prediction direction of the arrow (102) for this block 104. Signaling is included, i.e., the sample is derived from one or more predicted samples to the top right, Assume that the projection is at a 45 degree angle from the horizontal direction. In such a case, samples S41, S32, and S2 3, S14 is predicted from the same reference sample R05. Then sample S44 is predicted from the reference sample Predicted from R08.
[0014] In certain cases, the orientation is divided evenly by 45 degrees to calculate the reference sample. When this is not possible, the values of multiple reference samples may be combined, for example by interpolation. .
[0015] The number of possible directions has increased as video coding technology continues to develop. In .264 (2003), for example, nine different directions are available for intra prediction. increased to 33 in H.265 (2013), and JEM / VVC / BMS is up to 65 at the time of this disclosure. It can support the most appropriate intra prediction direction. Experimental studies have been carried out to determine the directionality of the image using specific techniques of entropy coding. Accepting all specific bit penalties, their most appropriate direction is coded with a small number of bits. Furthermore, the direction itself can be used in intra-prediction of the decoded neighboring blocks. In some cases, it can be predicted from the adjacent directions.
[0016] Figure 1B shows the increasing number of prediction directions in various coding techniques that have evolved over time. To illustrate, we present a schematic diagram (180) showing the 65 intra prediction directions according to JEM.
[0017] Bits representing intra-prediction direction in a coded video bitstream The mapping to the prediction direction may vary depending on the video coding technology, e.g. For example, a simple direct mapping of prediction direction to intra prediction mode yields the codeword, the most likely These may range from complex adaptive schemes including highly variable modes and similar techniques. In all cases, it is statistically more likely to occur in video content than in any other particular orientation. There may be a specific direction of low intro prediction. Therefore, in a well-designed video coding technique, these less likely Directions are represented by more bits than more likely directions.
[0018] Inter-picture prediction, or inter-prediction, can be based on motion compensation. Compensation involves the use of sample data from a previously reconstructed picture or part of it (reference picture). After the data is spatially shifted in the direction indicated by the motion vector (hereafter MV), It can be used to predict newly reconstructed pictures or picture portions (e.g., blocks). In some cases, the reference picture may be the same as the picture currently being reconstructed. , it may have two dimensions X and Y, or three dimensions, the third dimension being (similar to the time dimension) An indication of the reference picture to be used (similar to a reference picture).
[0019] Some video compression techniques use the current MV that can be applied to specific areas of the sample data. from other MVs, e.g., spatially adjacent to the area being reconstructed and preceding the current MV in decoding order. It can be predicted from other MVs related to other areas of the sample data. By using a MV coding algorithm, we can code MVs by relying on the removal of redundancy in correlated MVs. This can significantly reduce the overall amount of data required to MV prediction can work effectively for example in the case of natural video. When coding an input video signal derived from a camera (such as Areas larger than the applicable area move in a similar direction in the video sequence. There is a statistical likelihood that the MVs in the adjacent areas may be derived. As a result, the desired motion vector can be predicted using the same motion vector. The actual MV of a given area will be similar or identical to the MV predicted from the surrounding MVs. Furthermore, after entropy coding, the MV is predicted from one or more neighboring MVs. is less than the number of bits that would be used if the In some cases, MV prediction can be performed using the original signal (i.e., the number of samples) This can be an example of lossless compression of a signal (i.e., MV) derived from a video stream. In other cases, e.g. due to rounding errors when computing predictors from several surrounding MVs Additionally, the MV prediction itself may be lossy.
[0020] H.265 / HEVC (ITU-T Rec.H.265, "High Efficiency Video Coding", December 2016 ) describes various MV prediction mechanisms. Among the many MV prediction mechanisms specified by H.265, the following are Described below is a technique hereafter referred to as "spatial merging."
[0021] Specifically, referring to FIG. 2, the current block (201) is The coder assumes that the block is predictable from a previous block of the same size but spatially shifted. Instead of coding the MV directly, we divide the MV into A0, A1, and and one of the five surrounding samples represented by B0, B1, and B2 (202 to 206, respectively). Using linked MVs, one or more reference pictures and associated metadata can be used to For example, it can be derived from the last reference picture (in decoding order). MV prediction uses predictors from the same reference picture as the neighboring blocks. It is possible. Summary of the Invention [Means for solving the problem]
[0022] The present disclosure relates to methods, apparatus, and computers for video encoding and / or video decoding. Various embodiments of a readable storage medium are described.
[0023] According to one aspect, embodiments of the present disclosure provide a method for encoding / decoding video data in a decoder. The present invention provides a method for encoding coded video bits of a data block. receiving a coded video bitstream from the coded video bitstream; extracting a transformation partition type associated with the data block; The partition type is the division pattern for dividing data blocks into conversion blocks. belongs to a subset of a predefined set of transformation partition types, each of which specifies signaled in the coded video bitstream in response to The transformation type of the transformation associated with the transformation block split from the data block, such as wherein the transformation type is selected from a first predefined set of transformation types. Steps that perform inverse transformations on transformation blocks according to the step and transformation type. Includes steps.
[0024] According to another aspect, an embodiment of the present disclosure provides a method for encoding / decoding video data. The method includes receiving a coded video bitstream of data blocks. and receiving the coded video bitstream and decoding the video data. extracting a transformation partition type associated with the data block; Transformation partitions that belong to a subset of a predefined set of partition types. 2. Extracting data blocks from the coded video bitstream in response to the type of Extracting a transformation type associated with a transformation block; In response to a conversion partition type that does not belong to a predefined set of types, and identifying the conversion type by default for the tab block.
[0025] According to another aspect, an embodiment of the present disclosure provides a method for encoding / decoding video data. The method includes receiving a coded video bitstream of data blocks. receiving a data block from the coded video bitstream; Extracting a transformation type of a transformation associated with a transformation block; The coded video bitstream is then processed in response to a transform type belonging to a predefined set. A stream that extracts the transformation partition type associated with a data block. Includes steps.
[0026] According to another aspect, an embodiment of the present disclosure provides a method for video encoding and / or video decoding. The apparatus includes a memory for storing instructions and a processor in communication with the memory. When the processor executes the instructions, the processor may cause the device to perform video decoding and / or video The apparatus is configured to perform the above method for video encoding.
[0027] According to yet another aspect, an embodiment of the present disclosure provides a method for video decoding and / or video encoding. for video decoding and / or video encoding when executed by a computer A non-transitory computer-readable medium storing instructions for causing a computer to perform the above method. to provide.
[0028] These and other aspects and their implementations are set forth in the drawings, specification, and claims. This will be explained in more detail below.
[0029] Further features, characteristics and advantages of the subject matter disclosed in the following detailed description and accompanying drawings are set forth. and various effects become more apparent. [Brief explanation of the drawings]
[0030] [Figure 1A] FIG. 10 is a schematic diagram of an example subset of intra-prediction directional modes. [Figure 1B] FIG. 2 illustrates exemplary intra-prediction directions. [Figure 2] FIG. 1 is a schematic diagram illustrating a current block and its surrounding spatial merge candidates for motion vector prediction in one example. [Figure 3] FIG. 3 is a schematic diagram illustrating a simplified block diagram of a communication system (300) according to an exemplary embodiment. [Figure 4] FIG. 4 is a schematic diagram illustrating a simplified block diagram of a communication system (400) according to an exemplary embodiment. [Figure 5] FIG. 2 is a schematic diagram illustrating a simplified block diagram of a video decoder according to an example embodiment. [Figure 6] FIG. 1 is a schematic diagram illustrating a simplified block diagram of a video encoder according to an example embodiment. [Figure 7] FIG. 2 is a block diagram illustrating a video encoder according to another example embodiment. [Figure 8] FIG. 2 is a block diagram illustrating a video decoder according to another example embodiment. [Figure 9] FIG. 10 illustrates directional intra-prediction modes according to an exemplary embodiment of the present disclosure. [Figure 10] FIG. 10 illustrates a non-directional intra-prediction mode according to an exemplary embodiment of this disclosure. [Figure 11] 1 illustrates a recursive intra-prediction mode according to an exemplary embodiment of the present disclosure. [Figure 12] 1 illustrates transform block partitioning and scanning of intra-predicted blocks according to an exemplary embodiment of the present disclosure. [Figure 13] 1 illustrates transform block partitioning and scanning of inter-predicted blocks according to an exemplary embodiment of this disclosure. [Figure 14] 1 illustrates a low frequency non-separable transformation process according to an exemplary embodiment of the present disclosure. [Figure 15] FIG. 1 illustrates various baseline-based intra-prediction schemes, according to an exemplary embodiment of the present disclosure. [Figure 16] FIG. 1 illustrates a non-recursive block division scheme according to an exemplary embodiment of the present disclosure. [Figure 17] 1 shows a flowchart according to an embodiment of the present disclosure. [Figure 18] FIG. 1 is a schematic diagram illustrating a computer system according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0031] FIG. 3 illustrates a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication systems (300) communicate with each other via, for example, a network (350). For example, the communication system (300) includes a network (350) 3, the system includes a first pair of terminal devices (310) and (320) interconnected via a The first pair of terminal devices 310 and 320 may perform unidirectional transmission of data. For example, The terminal device (310) transmits the data to the other terminal device (320) via the network (350). (e.g., of a stream of video pictures captured by the terminal device (310)) The encoded video data may be coded into one or more codes. The terminal device (320) can transmit the video in the form of a coded video bitstream. receives coded video data from the network (350) and The video data is decoded to restore the video picture, and the video data is then processed according to the restored video data. One-way data transmission is implemented in media serving applications, etc. It can be applied.
[0032] In another example, the communication system (300) may be implemented, for example, during video conferencing applications. A second pair of terminal devices (330) and (340) for performing bidirectional transmission of the coded video data. For bidirectional data transmission, in one example, each of the terminal devices 330 and 340 includes The terminal device communicates with the other terminal device of the terminal devices (330) and (340) via the network (350). for transmitting to the terminal device (e.g., streaming of video pictures captured by the terminal device) Each of the terminal devices 330 and 340 may code video data (streams). The terminal also receives the coding transmitted by the other of the terminals (330) and (340). The video data is received, and the coded video data is decoded to generate video pictures. and displaying the video picture on an accessible display device according to the restored video data. The image may be displayed.
[0033] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) are servers, personal computers, and The principles underlying this disclosure may be implemented as a personal computer, a smart phone, or the like. The applicability of the present disclosure is not so limited. Laptop computers, tablet computers, media players, wearables It may be implemented in a computer, dedicated video conferencing equipment, etc. The network includes, for example, a terminal device (310) including a wired (wired connection) and / or a wireless communication network. ), (320), (330) and (340) to transmit coded video data between The communication network (350) 9 represents the number or type of network. , packet-switched channels, and / or other types of channels. Typical networks include telecommunications networks, local area networks, and wide area networks. For the purposes of this discussion, a network (35 0) architecture and topology are incorporated herein by reference unless expressly described herein. may not be important to the operation of
[0034] FIG. 4 illustrates an example of an application of the disclosed subject matter in a video streaming environment. The disclosed subject matter is useful in, for example, video conferencing, digital digital media including live TV, games, virtual reality, CDs, DVDs, memory sticks, etc. The present invention may be equally applied to other video-enabled applications, including storage of compressed video on media, etc.
[0035] A video streaming system transmits a stream of uncompressed video pictures or images. The video source (401) may include, for example, a digital camera, for creating a video (402). In one example, a video capture subsystem (413) may be included that can capture video pictures. The stream (402) contains samples recorded by the digital camera of the video source 401. The video picture stream (402) is encoded video data (404) (or to emphasize the high data volume compared to the 4, which includes a video encoder (403) coupled to a video source (401). The video encoder (403) can be processed by an electronic device (420) including: To enable or implement aspects of the disclosed subject matter as described in more detail, hardware The encoded video data may include hardware, software, or a combination thereof. The data (404) (or encoded video bitstream (404)) is a set of uncompressed video pictures. The data volume is shown in thin lines to emphasize the low volume compared to the original stream (402). and then forwards it to a streaming server (405) for future use or to a downstream video device (Figure The client subsystems (406) and (407) of FIG. 8) One or more streaming client subsystems, such as The server (405) is accessed to retrieve copies (407) and (408) of the encoded video data (404). The client subsystem (406) can acquire, for example, an electronic device (4 The video decoder (410) may include a video decoder (410) within the encoding unit (30). Decode the input copy (407) of the decoded video data and display it uncompressed. 12) Rendering onto a display screen or other rendering device (not shown) The video decoder 410 creates an output stream of video pictures (411) that can be , may be configured to perform some or all of the various functions described in this disclosure. In a streaming system, encoded video data (404), (407), and (409) are ) (e.g., video bitstream) according to a specific video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265. The video coding standard under development is called Versatile Video Coding (VVC). The subject matter of the disclosure may be used in the context of VVC and other video coding standards. It can be used.
[0036] Electronic devices 420 and 430 may include other components (not shown). Note that, for example, the electronic device (420) may include a video decoder (not shown). Alternatively, the electronic device (430) may also include a video encoder (not shown).
[0037] FIG. 5 shows a block diagram of a video decoder (510) according to any embodiment of the present disclosure. The video decoder (510) can be included in an electronic device (530). The video decoder (510) may include a receiver (531) (e.g., a receiving circuit). can be used in place of the video decoder (410) in the example of FIG.
[0038] The receiver (531) receives one or more codes to be decoded by the video decoder (510). In the same or another embodiment, the video sequences may be received one at a time. The coded video sequences can be decoded, and each coded video sequence The decoding of each video sequence is independent of other coded video sequences. A sequence may be associated with multiple video frames or images. The encoded video sequence is received from a channel (501), which is Hardware / software link to storage device that stores the encoded video data, or The receiver (531) may be a streaming source that transmits encoded video data. The encoded video data can be transferred to a respective processing circuit (not shown). together with other data such as audio data and / or auxiliary data streams The receiver (531) can receive the coded video sequence from other data. To counteract network jitter, a buffer memory (515) is provided for the receiver. (531) and the entropy decoder / parser (520) (hereafter referred to as "parser (520)"). In certain applications, the buffer memory (515) may be located between the video decoder (5 10). In other applications, the buffer memory (515) may be implemented as part of a video decoder. In still other applications, e.g., To counteract network jitter, a buffer memory ( (not shown) may also be present, for example, a video decoder (51) to process playback timing. There may be another buffer memory (515) inside the receiver (531). From a controllable storage / forwarding device or over an isosynchronous network When receiving data from a network, the buffer memory (515) may not be necessary or can be made smaller. Best-effort packet networks such as the Internet A buffer memory (515) of sufficient size may be required for use with the , its size can be relatively large. Such buffer memories are implemented with adaptive sizes. may be implemented by an operating system or similar component external to the video decoder (510). The system may be at least partially implemented in a computer (not shown).
[0039] The video decoder (510) extracts symbols (521) from the coded video sequence. The parser (520) may be included to recover the symbol categories. Information used to manage the operation of the audio decoder (510) and potentially as shown in Figure 5. , which may or may not be an integral part of the electronic device (530), a rendering device such as a display (512) (e.g., a display screen) that can be coupled to the and information for controlling the rendering device(s). The control information is provided as Supplementary Enhancement Information (SEI) messages or Video Usability Information (VUI) parameters. The parser 520 may be in the form of a meter set fragment (not shown). ) to parse / entropy decode the coded video sequence received by The coding of the entropy coded video sequence can be performed by It can be according to a coding technique or standard, such as variable length coding, Huffman coding, according to various principles, including coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may be a parser for a coded video sequence. from a video decoder based on at least one parameter corresponding to the subgroup. Extract a set of subgroup parameters for at least one of the subgroups of pixels in Subgroups include Groups of Pictures (GOP), pictures, tiles, and slides. block, macroblock, coding unit (CU), block, transform unit (TU), The parser (520) can also include a coded From the video sequence, transform coefficients (e.g., Fourier transform coefficients), quantization parameter values, , motion vectors, and other information may also be extracted.
[0040] The parser (520) reads the data received from the buffer memory (515) to create symbols (521). It is possible to perform entropy decoding / parsing operations on captured video sequences. Cut.
[0041] The reconstruction of the symbols (521) is carried out by the timing of the coded video picture or part thereof. (interpicture and intrapicture, interblock and intrablock, etc.) may contain multiple different processing or functional units, depending on the number of processors, processor types, and other factors. The included units and how the units are included are determined by the parser (520). The subgroup control information parsed from the coded video sequence is used to Between the parser (520) and the following processing or functional units: Such subgroup control information flow is not shown for the sake of simplicity.
[0042] In addition to the functional blocks already mentioned, the video decoder (510) includes the following: It can be conceptually subdivided into several functional units, as shown below. In a working implementation, many of these functional units interact closely with each other, However, various aspects of the disclosed subject matter may be integrated with one another, at least in part. To clearly explain the functionality, the following disclosure adopts a conceptual subdivision into functional units. do.
[0043] The first unit is the Scaler / Inverse Transform Unit (551). (551) is a block diagram of the quantized transform coefficients, as well as information indicating which type of inverse transform to use. control information including block size, quantization coefficients / parameters, quantization scaling matrix, etc. It can receive one or more symbols (521) from the parser (520). The inverse transformation unit (551) provides sample values that can be input to the aggregator (555). It is possible to output blocks that can be
[0044] In some cases, the output samples of the scaler / inverse transform (551) are intra-coded. a block that is reconstructed, i.e., does not use prediction information from a previously reconstructed picture. , a block that can use prediction information from a previously reconstructed part of the current picture. Such prediction information may be related to the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) , the surrounding blocks that have already been reconstructed and stored in the current picture buffer (558). This information may be used to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) may contain, for example, a partially reconstructed current picture. and / or buffer the fully reconstructed current picture. In some implementations, for each sample, the intra prediction unit (552) generates The prediction information is converted into output sample information provided by the scaler / inverse transform unit (551). It may be added.
[0045] In other cases, the output samples of the scaler / inverse transform unit (551) are inter-coded. In such cases, the The motion compensation prediction unit (553) accesses the reference picture memory (557) to The samples used for picture prediction can be fetched from the block associated with After motion compensation of the samples fetched according to symbol (521), these samples The aggregator (555) scales / descales the sampled data to generate the output sample information. The output of the conversion unit (551) can be added to the residual sample or residual signal), from which a motion compensation prediction unit (553) derives prediction samples. The address in the reference picture memory (557) to be fetched is, for example, the X component, the Y component (shift and a motion compensated prediction unit in the form of a symbol (521) which may have a reference picture component (temporal). The motion compensation can be controlled by the motion vectors available in the unit (553). Also, when subsample accurate motion vectors are used, the reference picture memory ( 557), and the motion vector prediction mechanism It may also be associated with, for example:
[0046] The output samples of the aggregator (555) are variously filtered in the loop filter unit (556). Video compression technology can be applied to various loop filtering techniques (coding contained in a coded video sequence (also called a coded video bitstream) It is controlled by the parameters included in the loop as a symbol (521) from the parser (520). The filter unit (556) may include in-loop filter technology. However, the coded pictures or coded video sequences (Firstly) respond to meta-information obtained during the decoding of the previous part, but also to previously restored and routed It can also respond to loop-filtered sample values, as explained in more detail below. As such, several types of loop filters use loop filter units in various orders. It may be included as part of part 556.
[0047] The output of the loop filter unit (556) may be output to a rendering device (512). and reference picture memory (557) for use in future inter-picture prediction. ) can also be stored in
[0048] A particular coded picture, once fully reconstructed, will be used for future interpictures. For example, the current picture can be used as a reference picture for pixel prediction. The coded picture is fully reconstructed and the coded picture is Once identified as a picture (e.g., by the parser (520)), the current picture buffer The reference picture memory (558) can be part of the reference picture memory (557) and can store unused current pictures. The picture buffer is reallocated before starting reconstruction of the next coded picture. It is possible.
[0049] The video decoder (510) may be a predetermined standard adopted in, for example, ITU-T Rec. H.265. The decoding operation may be performed in accordance with the video compression techniques of of a video compression technology or a profile documented in the syntax of a standard and video compression technology In the sense that it adheres to both, the coded video sequence is used It may conform to a syntax specified by a video compression technology or standard. A profile defines a video compression tool that can only be used under that profile. Ability to select specific tools from all available tools in a technology or standard To comply with the standard, the complexity of the coded video sequence must be adjusted to accommodate the video compression It can be within a range defined by the level of the technology or standard. In some cases, the level , maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., per second) (measured in megasamples), maximum reference picture size, etc. The limits set by the IEEE 802.11a standard are, in some cases, based on the Hypothetical Reference Decoder (HRD) specification and the Metadata for HRD buffer management signaled in the captured video sequence Therefore, it can be further restricted.
[0050] In some exemplary embodiments, the receiver (531) receives additional (Redundant) data may be received. The additional data may be coded The additional data may be included as part of the video sequence. and / or to more accurately recover the original video data, a video decoder (51 0). Additional data may be used, for example, to measure time, space, or signal noise. in the form of signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc. could be.
[0051] FIG. 6 shows a block diagram of a video encoder (603) according to an exemplary embodiment of the present disclosure. The video encoder (603) may be included in an electronic device (620). The video encoder (603) may further include a receiver (640) (e.g., a transmission circuit). This can be used instead of the video encoder (403) in the example.
[0052] The video encoder (603) is configured to encode the video data to be encoded by the video encoder (603). A video source (601) (in the example of FIG. 6, an electronic In another example, the video samples may be received from a device (not part of the device 620). The source (601) may be implemented as part of an electronic device (620).
[0053] A video source (601) is a stream to be coded by a video encoder (603). The source video sequence can be encoded in any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, Any color space (e.g., BT.601 Y CrCb, RGB, XYZ...), and any The appropriate sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4) The media serving system may provide the data in the form of a digital video sample stream that can be used to In the system, the video source (601) can store previously prepared video. In a video conferencing system, the video source (601) may store local video It can be a camera that captures information as a video sequence. It may be presented as a number of separate pictures or images that give the impression of movement when viewed. The field may be organized as a spatial array of pixels, each pixel representing the sampling It may contain one or more samples depending on the structure, color space, etc. The following explanation is based on the relationship between pixels and samples. It's focused.
[0054] According to some exemplary embodiments, the video encoder (603) performs, in real time: or any other time constraint required by the application. The image can be coded and compressed into a coded video sequence (643). Enforcing the proper coding rate constitutes one function of the controller (650). In some embodiments, the controller (650) is operatively coupled to other functional units. and may control other functional units as described below. Not shown. Parameters set by the controller 650 include rate Control-related parameters (picture skip, quantizer, lambda value of rate-distortion optimization method) ), picture size, Group of Pictures (GOP) layout, maximum motion vector search The controller 650 may be optimized for a particular system design. The video encoder (603) may be configured to have other suitable functionality associated with it. can.
[0055] In some exemplary embodiments, the video encoder (603) performs the following in the coding loop: As an overly simplified explanation, in one example, the coding The loop is executed by a source coder (630) (e.g., an input picture to be coded and a and (one or more) reference pictures to generate symbols, such as a symbol stream. and a (local) decoder (633) built into the video encoder (603). The decoder (633) may include an embedded decoder 633 for entropy. Process video streams coded by source coder 630 without coding Even if it is processed, the symbols will be reconstructed and created by the (remote) decoder. (In the video compression technology considered in the subject matter of the disclosure, Any compression between the symbols and the coded video bitstream is lossless. The reconstructed sample stream (sample data) is stored in the reference picture memory ( The symbol stream is decoded at the decoder location (local or remote). This leads to bit-accurate results regardless of the code in the reference picture memory (634). The content is also bit-accurate between the local and remote encoders. In other words, the prediction part of the encoder is what the decoder "sees" when using the prediction during decoding. The reference picture samples "see" exactly the same sample values that would be used in the reference picture. Synchronization of the architecture (and if synchronization cannot be maintained due to, for example, channel errors) This basic principle of the resulting drift is used to improve coding quality. It is used for this purpose.
[0056] The operation of the "local" decoder (633) has already been described in detail above in conjunction with FIG. , can be the same as the operation of a "remote" decoder such as the video decoder (510). Simply by reference, however, the symbols are available and the entropy coder (645 ) and encoding of symbols into the coded video sequence by a parser (520). Since the decoding can be lossless, the video decoder includes a buffer memory (515) and a parser (520). The entropy decoding part of the coder (510) is performed in a local decoder (633) within the encoder. Some may not be fully implemented.
[0057] At this point, it is possible to say that, excluding parsing / entropy decoding, which can only exist in the decoder, Any decoder technology will necessarily also have substantially the same functionality in the corresponding encoder. Therefore, the disclosed subject matter is decoder-operated. We may focus on the operation, which is similar to the decoding part of the encoder. The description of the encoder technique is omitted since it is the reverse of the decoder technique which is comprehensively described. Only in certain areas or aspects will a more detailed description of the encoder be provided below. vinegar.
[0058] During operation, in some example implementations, the source coder (630) generates "reference pictures" One or more previously coded pictures from a video sequence specified as motion-compensated predictive coding, which predictively codes an input picture by referring to the motion-compensated predictive coding In this way, the coding engine (632) may and the pixel blocks that can be selected as prediction reference(s) to the input picture (1 The difference (or residual) in the color channels between the pixel blocks of the reference picture(s) is computed. The term "residual" and its adjective form "residual difference" are used interchangeably. obtain.
[0059] The local video decoder (633) receives the symbols created by the source coder (630). The coded video data of a picture that can be designated as a reference picture based on The operation of the coding engine (632) is advantageously It may be a lossless process. When a video sequence can be decoded by a video decoder (not included), the reconstructed video sequence usually has some It can be a replica of the source video sequence with some errors. The decoder (633) is a decoding process that can be performed by a video decoder on a reference picture. and store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) can transmit the A reconstructed reference picture having a common content with a reconstructed reference picture obtained by A copy of the data can be stored locally (without transmission errors).
[0060] The predictor (635) can perform predictive searches for the coding engine (632). That is, for a new picture to be coded, the predictor (635) (Candidate reference pixel blocks) that can serve as suitable prediction references for a picture. Sample data (as blocks) or reference pictures. Specific data such as motion vectors, block shapes, etc. The reference picture memory (634) can be searched for metadata. ) is calculated for each pixel block against the sample block to find the appropriate prediction reference. In some cases, the search results obtained by the predictor (635) As determined by the result, the input picture is stored in a reference picture memory (634). A picture can have prediction references derived from multiple reference pictures.
[0061] The controller (650) may, for example, control the parameters used to encode the video data. The coding behavior of the source coder (630), including the setting of the data and subgroup parameters. You can manage your work.
[0062] The output of all the aforementioned functional units is entropy coded in the entropy coder (645). The entropy coder (645) is a Huffman coder. Reversible coding of symbols according to techniques such as coding, variable length coding, arithmetic coding, etc. Compression converts the symbols produced by the various functional units into coded video. Convert it to an audio sequence.
[0063] The transmitter (640) receives the coded signal generated by the entropy coder (645). Buffering the video sequence and preparing it for transmission over a communication channel (660) The communication channel (660) can include a storage device for storing the encoded video data. The transmitter (640) may be a hardware / software link to the video codec. The coded video data from the data processor (603) is transmitted to other data to be transmitted, e.g. Coded audio data and / or auxiliary data streams (sources not shown) can be merged with the original (not included).
[0064] The controller (650) can manage the operation of the video encoder (603). During loading, the controller (650) assigns a specific coding to each coded picture. You can assign a coded picture type to each picture. This can affect the coding techniques that can be applied. For example, pictures are often For example, it may be assigned as one of the following picture types:
[0065] Intra-pictures (I-pictures) are pictures that use any other picture in the sequence as a prediction source. Some video codecs may be capable of encoding and decoding without requiring any special hardware. For example, different types of inputs, including independent decoder refresh ("IDR") pictures, Those skilled in the art will understand these variations of I-pictures and their equivalents. Recognize the uses and characteristics of each.
[0066] A predicted picture (P-picture) uses at most one moving image to predict the sample values of each block. using intra prediction or inter prediction using vectors and reference indices, It may be capable of being encoded and decoded.
[0067] Bidirectionally predicted pictures (B pictures) use up to two bins to predict the sample values of each block. Uses intra or inter prediction using one motion vector and reference index Similarly, multiple predicted images may be encoded and decoded as a single block. Use of three or more reference images and associated metadata for rock reconstruction. This can be done.
[0068] A source picture generally consists of multiple sample coding blocks (e.g., 4x4 , 8x8, 4x8, or 16x16 sample blocks), and The blocks can be coded based on the coding applied to the respective picture. Referencing other (already coded) blocks as determined by block assignments For example, the blocks of an I picture can be coded non-predictively. A new block can be coded or refer to an already coded block in the same picture. The pixels of a P picture can be coded predictively (spatial prediction or intra prediction). Each block is predicted via spatial prediction with reference to one previously coded reference picture. Blocks in B pictures may be coded predictively via temporal prediction or as a , by spatial prediction with reference to one or two previously coded reference pictures; or predictively coded via temporal prediction. A picture may be subdivided into other types of blocks for other purposes. The division of blocks and other types of blocks is performed in the same manner as described in more detail below. They may or may not follow a method.
[0069] The video encoder (603) encodes a predetermined video coding technique such as ITU-T Rec. H.265. The coding operation may be performed in accordance with a technique or standard. The video encoder (603) exploits temporal and spatial redundancies in the input video sequence. Various compression operations can be performed, including predictive coding operations. The coded video data is encoded according to the video coding technology or standard used. Therefore, it can conform to the specified syntax.
[0070] In some exemplary embodiments, the transmitter (640) may transmit additional The source coder (630) may transmit such data as coded data. Additional data may be included as part of the video sequence. , other forms of redundant data such as redundant pictures and slices, SEI messages, and VUI parameters It may include set fragments, etc.
[0071] Video may be captured as multiple source pictures (video pictures) in chronological order. Intra-picture prediction (often abbreviated as intra-prediction) is the process of predicting the Inter-picture prediction exploits the spatial correlation between pictures, while inter-picture prediction exploits the temporal or other correlation between pictures. For example, a specific picture being coded / decoded, called the current picture, is used as a block. The blocks in the current picture may be divided into blocks. If the reference block is similar to a reference block in a reference picture that is still buffered, the motion vector A motion vector can be coded by a vector called a motion vector. If multiple reference pictures are used, the reference picture is There can be a third dimension that distinguishes.
[0072] In some exemplary embodiments, bi-prediction techniques may be used for inter-picture prediction. According to such bi-prediction techniques, both the current picture in the video are predicted in decoding order. The first reference point (which may be in the past or future in display order) Two reference pictures are used, such as the first reference picture and the second reference picture. The block is a first motion vector pointing to a first reference block in a first reference picture. and a second motion vector pointing to a second reference block in the second reference picture. The block can be coded by the first reference block and the second reference block. The combination allows for cooperative prediction.
[0073] Additionally, merge mode technology improves coding efficiency in inter-picture prediction. It may also be used to
[0074] According to some exemplary embodiments of the present disclosure, inter-picture prediction and intra-picture prediction are Prediction, such as block prediction, is performed on a block-by-block basis. For example, For compression, pictures in the image are divided into coding tree units (CTUs), The CTUs within a channel may have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. In general, a CTU consists of three parallel coding tree blocks (CTBs), i.e., one Each CTU may contain one or more coding units (C For example, a 64x64 pixel CTU can be recursively divided into 64x64 pixel It can be divided into one CU or four CUs of 32x32 pixels. Each of the one or more CUs may be further divided into four CUs of 16x16 pixels. In an embodiment, each CU supports various prediction types, such as inter prediction type and intra prediction type. The CU can be analyzed during encoding to determine which of the CUs to encode. The image is divided into one or more prediction units (PUs) depending on the image's predictability and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed on a prediction block basis. The division of CUs into PUs (or PBs for different color channels) can be performed in various spatial patterns. The luma PB or chroma PB may include a matrix of sample values (e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0075] FIG. 7 shows a diagram of a video encoder (703) according to another exemplary embodiment of the present disclosure. The video encoder (703) is configured to: receives a processing block (e.g., a prediction block) of sample values and converts the processing block into a to a coded picture that is part of a loaded video sequence. The exemplary video encoder (703) is configured to: (403) can be used instead.
[0076] For example, the video encoder (703) may process blocks such as prediction blocks of 8x8 samples. The video encoder (703) then receives a matrix of sample values for the block. Uses Rewrite Optimization (RDO) to ensure that processing blocks are best coded using it. The processing determines whether the mode is intra mode, inter mode, or bi-predictive mode. If it is decided that a block is to be coded in intra mode, the video encoder ( 703) uses intra prediction techniques to encode processing blocks into coded pictures. and determines that the processing block is coded in inter mode or bi-predictive mode. If so, the video encoder (703) uses inter-prediction or bi-prediction techniques, respectively. The processing blocks may be encoded into a coded picture using the In the embodiment, as a submode of inter-picture prediction, a motion vector outside the predictor One or more motion vector predictions without the benefit of the coded motion vector components. A merge mode derived from the probes may be used. There may be motion vector components applicable to the current block. The da (703) includes a mode determination module, etc., for determining the prediction mode of the processing block. 7. The above components may include components not explicitly shown in FIG.
[0077] In the example of FIG. 7, the video encoders (703) are connected to each other as shown in the exemplary configuration of FIG. an inter-encoder (730), an intra-encoder (722), and a residual calculator (72 3), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entry Includes a rpy encoder (725).
[0078] The inter-encoder (730) processes the samples of the current block (e.g., the processing block). and compares the block with one or more reference blocks (e.g., blocks in the previous and subsequent pictures in the order of presentation) and inter-prediction information ( For example, redundant information from inter-coding techniques, motion vectors, and merge mode information) and generating an inter prediction result ( In some examples, the reference pixel may be configured to calculate a predicted block. The architecture is shown as residual decoder 728 in FIG. 7, as described in more detail below. 6. The decoder 633 is encoded using the decoder unit 633 incorporated in the exemplary encoder 620 of FIG. The decoded reference picture is a decoded reference picture that is decoded based on the video information.
[0079] The intra-encoder (722) decodes the samples of the current block (e.g., the processing block). and compares the block with already coded blocks in the same picture; Generate transformed quantized coefficients, and possibly intra prediction information (e.g., one or more The intra-prediction direction information (intra-prediction direction information by the intra-encoding technique) is also generated. The intra encoder (722) performs intra prediction based on the intra prediction information and the reference block in the same picture. Based on this, an intra prediction result (eg, a predicted block) may be calculated.
[0080] The general-purpose controller (721) determines general-purpose control data and performs a bidding based on the general-purpose control data. The video encoder (703) may be configured to control other components of the video encoder (703). The controller (721) determines the prediction mode of the block and switches based on the prediction mode. For example, if the prediction mode is an intra mode, the general The controller (721) controls the switch (726) to select the value used by the residual calculator (723). It selects the intra mode result for the image and controls the entropy encoder (725) to Intra prediction information is selected and included in the bitstream, and If the description mode of the block is the intermode, the general-purpose controller (721) (726) to select the inter prediction result for use by the residual calculator (723). , and controls the entropy encoder (725) to select the inter prediction information and This allows for target prediction information to be included in the bitstream.
[0081] The residual calculator (723) calculates the residuals of the received block and the intra-encoder (722) or - The difference between the prediction result and the block selected from the encoder (730) (residual data) The residual encoder (724) may be configured to encode the residual data to calculate For example, the residual encoder (724) may be configured to generate transform coefficients from residual data. The data may be configured to transform the data from the spatial domain to the frequency domain to generate transform coefficients. , the transform coefficients are subjected to a quantization process to obtain quantized transform coefficients. In some embodiments, the video encoder (703) also includes a residual decoder (728). 728) is configured to perform the inverse transform and generate decoded residual data. The residual data is then applied by an intra-encoder (722) and an inter-encoder (730). For example, the inter-encoder (730) can use the decoded residual data A decoded block can be generated based on the inter prediction information and the inter prediction information. The intra encoder (722) decodes the residual data based on the decoded residual data and the intra prediction information. The decoded blocks can be used to generate the decoded picture. The decoded pictures are buffered in a memory circuit (not shown) to generate a picture. and can be used as a reference picture.
[0082] The entropy encoder (725) generates a bitstream containing the encoded blocks. and configured to perform entropy coding. The tropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) may include general control data, selected prediction information ( For example, intra-prediction information and inter-prediction information, residual information, and other suitable information are The video stream can be configured to include either inter or bi-predictive modes. When coding a block in the merge submode of There is.
[0083] FIG. 8 shows a diagram of an exemplary video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) receives the coded video data that is part of the coded video sequence. and receives the coded picture and decodes the coded picture to obtain a reconstructed picture. In one example, the video decoder (810) is configured to generate the video It can be used in place of the decoder (410).
[0084] In the example of FIG. 8, the video decoders (810) communicate with each other as shown in the exemplary configuration of FIG. The entropy decoder (871), the inter-decoder (880), and the residual decoder ( 873), a reconstruction module (874), and an intra-decoder (872).
[0085] The entropy decoder (871) extracts the coded image from the coded picture. and restoring specific symbols representing syntax elements of which the extracted pictures are composed. Such symbols can be used, for example, to indicate the mode in which the block is coded ( For example, an intra mode, an inter mode, a bi-predictive mode, a merged submode, or another submode. (sub-mode), used for prediction by the intra-decoder (872) or inter-decoder (880) Predictive information (e.g., information) that can identify the particular sample or metadata to be Prediction information (inter-prediction information, inter-prediction information), e.g. residual information in the form of quantized transform coefficients. In one example, when the prediction mode is an inter mode or a bi-prediction mode, - Prediction information is provided to the inter decoder (880), and the prediction type is an intra prediction type. In some cases, intra prediction information is provided to the intra decoder (872). It may be subjected to quantization and provided to a residual decoder (873).
[0086] The inter decoder (880) receives the inter prediction information and performs a The inter prediction unit 100 may be configured to generate inter prediction results using the 100-bit approximation.
[0087] The intra decoder (872) receives the intra prediction information and performs a The predictive results may be generated using the
[0088] The residual decoder (873) performs inverse quantization to extract the inverse quantized transform coefficients and The residual decoder may be configured to process the coefficients to transform the residual from the frequency domain to the spatial domain. The reader (873) also uses specific control information (to include the quantization parameter (QP)). In some cases, this information can be provided by the entropy decoder (871) (which (Data paths are not shown as there may only be a small amount of control information.)
[0089] The reconstruction module (874) generates, in the spatial domain, the residual image as output by the residual decoder (873). the residuals (possibly due to an inter-prediction module or an intra-prediction module) The prediction result is combined with the reconstructed output as part of the reconstructed video. The image may be arranged to form a reconstructed block that forms part of the selected picture. Other suitable operations, such as deblocking operations, may be performed to improve the quality of the received signal. Please note that:
[0090] Video encoders (403), (603), and (703), and video decoders (410), ( It should be noted that 510) and 810) may be implemented using any suitable technique. In some exemplary embodiments, the video encoders (403), (603), and (703) and video decoders (410), (510), and (810) using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603), and video decoders (410), (510), and (810) Software instructions It may be implemented using one or more processors that execute the
[0091] Returning to the intra prediction process, a block (e.g., a luma or chroma prediction block, or is a coding block if it is not further divided into prediction blocks) , to generate a prediction block, the adjacent line, the next adjacent line, or one other or multiple lines, or a combination thereof. The residual between the actual block and the predicted block is transformed and then quantized. Various intra prediction modes may be made available, and The parameters related to the code selection and other parameters are signaled in the bitstream. Various intra prediction modes may be used, for example, to predict one or more samples. Prediction samples are selected from multiple line positions, predicting one or more lines. It may relate to direction and other special intra prediction modes.
[0092] For example, a set of intra prediction modes (interchangeably referred to as "intra modes") may be: It may include a predetermined number of directional intra-prediction modes, as described above in connection with the example implementation of FIG. As such, these intra prediction modes determine the prediction of the samples predicted within a particular block. The out-block samples may correspond to a predetermined number of directions from which they are selected. In an exemplary implementation, eight main directional modes can be supported and predefined, corresponding to angles from 45 degrees to 207 degrees relative to the horizontal axis.
[0093] Some other implementations of intra prediction utilize more diverse space in directional textures. To further exploit inter-redundancy, directional intra-mode has finer granularity. For example, the eight-angle implementation above can be expanded to include a set of angles, as shown in FIG. , V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, each of which may be configured to provide eight nominal angles A predetermined number (for example, 7) of finer angles may be added to the above. Therefore, a larger total number (e.g., this In the example, 56) direction angles can be used for intra prediction. The predicted angles are the nominal intra angles. Each nominal angle can be expressed as a degree + angle delta. In the particular example above, the angle delta may be -3 to 3 times the step size of 3 degrees .
[0094] In some implementations, instead of or in addition to the directional intra modes described above, A number of non-directional intra-prediction modes may also be predefined and made available. Five non-directional intra modes, called low-speed intra prediction modes, are specified. These non-directional intra-mode prediction modes are specifically DC, PAETH, and SMOOTH. These exemplary intra modes are sometimes called SMOOTH_V, SMOOTH_H, and SMOOTH_V intra modes. The prediction of samples of a particular block under directional mode is shown in Figure 10. As an example, 10 shows a 4×4 block predicted by samples from the upper and / or left adjacent lines. 10 shows a block 1002. A particular sample 1010 in the block 1002 may correspond to a sample 1004 directly above the sample 1010 in the above-neighboring line of the block 1002, a sample 1006 above and to the left of the sample 1010 as the intersection of the above-neighboring line and the left-neighboring line, and a sample 1008 directly to the left of the sample 1010 in the left-neighboring line of the block 1002. In an exemplary DC intra-prediction mode, an average of the left and above-neighboring samples 1008 and 1004 may be used as a predictor for the sample 1010. In an exemplary PAETH intra-prediction mode, the above, left, and above-left reference samples 1004, 1008, and 1006 may be fetched, and then any value among these three reference samples that is closest to (above + left - above-left) may be set as the predictor for the sample 1010. In the example of the SMOOTH_V intra prediction mode, the sample 1010 may be predicted by quadratic interpolation in the vertical direction of the upper-left neighboring sample 1006 and the left neighboring sample 1008. For the example of the SMOOTH_H intra prediction mode, the sample 1010 may be predicted by quadratic interpolation in the horizontal direction of the upper-left neighboring sample 1006 and the upper neighboring sample 1004. For the example of the SMOOTH intra prediction mode, the sample 1010 may be predicted by an average of quadratic interpolations in the vertical and horizontal directions. The above implementations of the non-directional intra modes are provided merely as non-limiting examples. Other implementations may be used. adjacent lines, and other non-directional selection of samples, and specific samples within the predicted block A method of combining prediction samples to predict the
[0095] At various coding levels (picture, slice, block, unit, etc.) A particular intra prediction mode by the encoder from the above directional or non-directional modes. The selection of the mode may be signaled in the bitstream. , the exemplary eight nominal direction modes may be signaled first, along with five non-angle smooth modes (a total of 13 options). The signaled modes are then adjusted to the eight nominal angles. If the mode is one of the degree intra modes, the selected angle delta is An index is further signaled to indicate the nominal angle being signaled. In this example implementation, all intra-prediction modes are all unified for signaling purposes. Together (for example, 56 directional modes plus 5 non-directional modes for a total of 61 intra predictions) Generates a mode) may be indexed.
[0096] In some example implementations, the example 56 or other number of directional intra-prediction modes are: Each sample in the block is projected to a reference subsample position and filtered by a 2-tap bilinear filter. This can be implemented using a unified directional predictor that interpolates the reference samples using the
[0097] In some implementations, the decaying spatial correlation with the fiducial on the edge is captured. To achieve this, an additional filter mode can be designed, called the FILTER INTRA mode. In these modes, the intra-prediction reference samples for some patches in the block are As a result, in-block prediction samples can be used in addition to out-block samples. These modes are, for example, predefined and include at least the luma block (or only the luma block). ) can be made available for intra prediction. A predefined number of filters (e.g., 5) can be used. It is possible to pre-design intramodes, each of which corresponds to a sample in a 4x2 patch, for example. An n-tap filter (e.g., 7 taps) reflects the correlation between its n neighbors. In other words, the weighting coefficients of an n-tap filter are It may depend on the position. Take an example of 8x8 block, 4x2 patch, 7-tap filtering, as shown in Figure 1. As shown in FIG. 1, an 8×8 block 1102 may be divided into eight 4×2 patches. These patches The patches are shown as B0, B1, B2, B3, B4, B5, B6, and B7 in Figure 11. Each patch is shown as R0 to R7 in Figure 11. The seven neighboring patches can then be used to predict the sample in the current patch. For patch B0, all neighbors may already be reconstructed. However, for other patches, some of the neighbors are in the current block and therefore cannot be reconstructed. may not be built, in which case the predicted values of the nearest neighbors are used as a reference. For example, all neighbors of patch B7 shown in Figure 11 have not been reconstructed. The predicted samples of the neighboring elements of patch B7 are used instead.
[0098] In some implementations of intra prediction, one color component may predict one or more other color components. The color components can be any of the components in the YCrCb, RGB, XYZ color space, etc. For example, a luma component (e.g., a luma base) can be used, called chroma from luma, or CfL. It is possible to perform prediction of chroma components (e.g., chroma blocks) from the quasi-samples. In some example implementations, much of the cross-color prediction is done from luma to chroma. For example, chroma samples within a chroma block must match the reconstructed It can be modeled as a linear function of the luma samples. CfL prediction is performed as follows: This can be done. CfL(α)=α×L AC +DC (1)
[0099] In the formula, L AC represents the AC contribution of the luma component, α represents the parameters of the linear model, and DC is , which represent the DC contribution of the chroma components. For example, the AC components are obtained for each sample of the block, The DC component is obtained for the entire block. Specifically, the reconstructed luma samples are The image may be subsampled to the desired luma resolution, and then the average luma value (the luma DC) is subtracted from each luma value. The AC contributions of the luma can then be used in the linear mode of equation (1) to predict the AC values of the chroma components. In order to approximate or predict the chroma AC components from the luma AC contributions, the decoder needs to calculate scaling parameters. Instead, an exemplary CfL implementation determines the parameter α based on the original chroma samples. This allows decoding to be performed quickly and efficiently. This reduces the complexity of the decoder and results in a more accurate prediction. In some exemplary implementations, the chrominance is calculated using intra DC mode in the chrominance content. obtain.
[0100] In some example implementations of the baseline, multi-line intra prediction may be used. In these implementations, two or more reference lines are available for selection in intra prediction. The encoder may determine the reference lines used to generate the intra predictor. The reference line index is signaled before the intra prediction mode. The most likely case is when a non-zero reference line index is signaled. Only the higher modes are allowed. Referring to Figure 15, there are four reference lines (from Reference 0 to Reference 3) are shown, each reference line has six segments with a top left reference sample. , that is, segments A to F (shown as 1502 to 1512). and F may be padded with the nearest samples from segments B and E, respectively.
[0101] Next, a transform of the residual is performed on either the intra-predicted block or the inter-predicted block. For the purpose of performing the transform, the transform coefficients may be quantized. Both inter-coded and inter-coded blocks are encoded as Multiple transform blocks (the term "unit transform" is usually used to describe a collection of three color channels) It is sometimes used interchangeably as "conversion unit", e.g., "code The "coding unit" consists of the luma coding block and the chroma coding block. In some implementations, the coding blocks (or prediction blocks) A maximum decomposition depth of 1000 bits (blocks) may be specified (the term "coding block" is used interchangeably with "coding block"). (This may be used interchangeably with "block"). For example, such divisions may not exceed two levels. The division of predicted blocks into transform blocks is divided into intra-predicted blocks and inter-predicted blocks. However, in some implementations, Such a division may be similar between intra-predicted and inter-predicted blocks. .
[0102] In some example implementations, for intra-coding blocks, the transform party The transformation can be done so that all transform blocks have the same size, and the transform blocks are The intra-coding blocks are coded in raster scan order. An example of replacement block partitioning is shown in Figure 12. Specifically, Figure 12 shows the coding block The lock 1202 is split into the same blocks via a mid-level quadtree split 1204, as shown by 1206. The block size is divided into 16 transform blocks. An exemplary raster scan order is shown by the ordered arrows in FIG.
[0103] In some example implementations, and for inter-coding blocks, transform units The splitting is done recursively up to a given number of levels (e.g., two levels). As shown in FIG. 13, the partitioning can be performed at any level for any subpartition. In particular, Figure 13 shows that block 1302 is a quadtree of four 1304, one of which further performs four second-level transformations. While the division of other sub-blocks stops after the first level, Here is an example that results in a total of seven transform blocks of different sizes: The exemplary raster scan order is further illustrated by the ordered arrows in FIG. FIG. 13 shows an exemplary implementation of quadtree decomposition of square transform blocks with up to two levels. However, in some production implementations, the transformation partitioning is 1:1 (square), 1:2 / 2:1 and 1:4 / 4:1 conversion block shapes and sizes must be in the range of 4x4 to 64x64. In some example implementations, the coding block may be For 64x64 and below, the transform block partitioning is The block can only be applied to the luma component (which is the same as the coding block under that condition). Otherwise, if the width or height of the coding block is greater than 64, the luma Both the coding block and the chroma coding block are min(W,64) Implicit division into multiples of min(H,64) × min(H,64) and min(W,32) × min(H,32) transform blocks. It can be done.
[0104] In some example implementations, as shown in FIG. 16, a coding block or Figure 1 further illustrates another alternative exemplary scheme for dividing a prediction block into transform blocks. As shown in Figure 16, instead of using recursive transformation partitioning, we A predefined set of partition types is coded according to the block's transformation type. In the particular example shown in FIG. 16, one of six exemplary division types is applied to the block. One may be applied to divide a coding block into a variable number of transform blocks. Such a scheme can be applied to either coding blocks or prediction blocks. In this disclosure, the term "partition type" generally refers to a block (e.g., a predicted block or coding block) is partitioned, "Transform partition type", "Prediction block partition type", or "Code It can also refer to a "conversion partition type." For the explanation under "Application Types", the same concept is used in "Coding Block Partitions" This also applies to "type" and vice versa.
[0105] More specifically, the division scheme of FIG. 16 provides the following for any given transformation type: This method provides up to six partition types for all coding blocks or A prediction block may be assigned a transform type based on, for example, a rate-distortion cost. In the example, the partition type assigned to a coding block or a prediction block is The transformation type of the prediction block or the moving block can be determined based on the transformation type of the prediction block. As shown by the four partition types, a particular partition type determines the partitioning of the transform block. It can accommodate various sizes and patterns (or division types). An exemplary correspondence between the rate-distortion cost and the type of A large number indicating the type of transformation that can be assigned to a coding block or a prediction block based on the They are shown below with their letter labels.
[0106] · PARTITION_NONE: Allocate a transformation size equal to the block size.
[0107] PARTITION_SPLIT: Transformation with width 1 / 2 of the block size and height 1 / 2 of the block size. Allocate size.
[0108] PARTITION_HORZ: A conversion size with the same width as the block size and half the height of the block size. Assign a size.
[0109] PARTITION_VERT: A conversion area with a width half the block size and the same height as the block size. Assign a size.
[0110] PARTITION_HORZ4: A conversion part with the same width as the block size and a height of 1 / 4 of the block size. Assign a size.
[0111] PARTITION_VERT4: A conversion area with a width of 1 / 4 of the block size and the same height as the block size. Assign a size.
[0112] In the above example, all the partition types shown in Figure 16 are This is not a limitation, but is merely an example. In this case, the mixed transform block size is determined by the partitioning of the block in a particular partition type (or pattern). may be used for the selected transform block.
[0113] Each of the above obtained transform blocks may then be subjected to a primary transform. The Mali transform essentially moves the residual in the transform block from the spatial domain to the frequency domain. In some implementations of actual primary transforms, the above example extended coding block To support block partitioning, multiple transform sizes (one for each dimension of the two dimensions) are used. ranges from 4 points to 64 points) and transform shape (square; width / height ratio 2:1 / 1:2 , and 4:1 / 1:4 rectangles) are acceptable.
[0114] Turning to the actual primary transformation, in some example implementations, a 2D transformation process The process is performed by using a hybrid transform kernel (which is, for example, the next The example may include the use of a 1-D transform (which may consist of a different 1-D transform for each element). Conversion kernels may include, but are not limited to: a) 4-point, 8-point, 1 a) 6-point, 32-point, and 64-point DCT-2; b) 4-point, 8-point, and 16-point DCT-2 a) Asymmetric DSTs (DST-4, DST-7) and their flipped versions; b) 4-point, 8-point The choice of transformation kernel used for each dimension is , can be based on a rate-distortion (RD) criterion. For example, DCT-2 and asymmetric The basis functions of DST's are listed in Table 1. [Table 1]
[0115] In some example implementations, a hybrid transform card for a particular primary transform implementation is used. The availability of the kernel can be based on the transform block size and the prediction mode. Examples are listed in Table 2. For chroma components, the conversion type selection is performed in an implicit way. In the example, for the intra prediction residual, the transform type, as specified in FIG. can be selected according to the intra prediction mode. The transform type for a luma block follows the transform type selection of the co-located luma block. Therefore, for the chroma components, the conversion timing is There is no signaling of the type. [Table 2] [Table 3]
[0116] In some implementations, a secondary transform may be performed on the primary transform coefficients. For example, as shown in FIG. 14, a reduction The secondary transform known as LFNST (Low Frequency Non-Separable Transform) is Between primary transform and quantization (at the encoder), and between inverse quantization and inverse primary transform Essentially, LFNST can be applied between the secondary transform and the To proceed, we first select a portion of the primary transform coefficients, e.g., the low frequency portion (hence the transform block). The LFNST can be a "reduced" version of the complete set of primary transform coefficients of the lock. In the example, depending on the transform block size, we have a 4x4 non-separable transform or an 8x8 non-separable transform. A transform that can be applied, e.g., for small transform blocks (e.g., min(width, height)<8), 4x4 LFNST is applied to the transform blocks, and 8x4 LFNST is applied to the transform blocks (e.g., min(width, height)>8). For example, if an 8x8 transform block is subjected to a 4x4 LFNST, , only the low frequency 4x4 portion of the 8x8 primary transform coefficients undergoes a further secondary transform.
[0117] As specifically shown in FIG. 14, the transform block may be 8×8 (or 16×16). Therefore, the forward primary transform 1402 of the transform block is an 8x8 (or 16x16) primary transform. The forward LFNS generates a transform coefficient matrix 1404, where each square unit represents a 2x2 (or 4x4) portion. The input to T does not have to be, for example, the entire 8x8 (or 16x16) primary transform coefficients. For example, a 4x4 (or 8x8) LFNST can be used for the secondary transform. As shown in portion (top left) 1406, the 4×4 (or 8×8) low-frequency Only the wavenumber primary transformation coeffects can be used as input to LFNST. The remaining part of the secondary transform coefficient matrix does not need to undergo a secondary transform. After the secondary transformation, the portion of the primary transformation co-effect affected by LFNST is The remaining part (e.g., the unshaded part of matrix 1404) that is not affected by LFNST becomes the retransform coefficient. The remaining part (the part that is not included) retains the corresponding primary transform coefficients. The remaining portion not subject to the secondary transform may be set to all zero coefficients.
[0118] An example application of the non-separable transform used in LFNST is described below. To use this, a 4×4 input block X (e.g., the shaded area of the primary transformation matrix 1404 in FIG. 14) is (representing the 4x4 low frequency part of the primary transform coefficient block such as min 1406) as follows: It can be expressed as:
number
[0119] This 2D input matrix is first linearized or converted into a vector in the exemplary order:
number
number
[0120] The non-separable transformation of a 4x4 LFNST is then
number
number
number
[0121] The above exemplary LFNST is based on a direct matrix multiplication approach for applying a non-separable transform and as a result is performed in a single pass without multiple iterations. In some further exemplary implementations, the dimensions of the non-separable transform matrix (T) of the 4×4 LFNST example can be further reduced to minimize the computational complexity and memory space requirements for storing the transform coefficients. Such an implementation can be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (where N is 4×4 = 16 in the above example, but could be equal to 64 for an 8×8 block) dimensional vector to an R dimensional vector in a different space where N / R (R < N) represents a dimensionality reduction factor. Thus, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows.
Equation
[0122] In Equation 7, the R rows of the transform matrix are the reduced R basis of the N dimensional space. Thus, the transform converts an input vector or an output vector of the reduced R dimension of N dimensions. Thus, as shown in FIG. 14 the secondary transform coefficients 1408 transformed from the primary coefficients 1406 are reduced in dimension by a factor of coefficient or N / R. The three squares around 1408 in FIG. 14 may be padded with zeros.
[0123] The inverse transform matrix of an RTS can be the transpose of its forward transform. For a more diverse description of the LFNST, the example reduction factor is 4. can be applied, and therefore the 64x64 direct non-separable transformation matrix is given by In addition, in some implementations, the input primary coefficients are reduced to a 16x64 direct matrix. A portion of may be linearized into the input vector of the LFNST, rather than the entirety of Only a portion of the input primary transform coefficients may be linearized into the above X vector. So, of the four 4x4 quadrants of the 8x8 primary transform coefficient matrix, the bottom right (high frequency coefficients) is can be excluded, and only the other three quadrants use a predetermined scan order rather than a 48x1 vector. In such an implementation, the non-separable transformation matrix can be further reduced from 16x64 to 16x48.
[0124] Therefore, the exemplary reduced 48x16 inverse RST matrix can be used at the decoder side to produce an 8x8 Can generate the top-left, top-right, and bottom-left 4x4 quadrants of the core (primary) transform coefficients Specifically, instead of a 16x64RST with the same transformation set configuration, a further reduced 16 When a ×48RST matrix is applied, the non-separable secondary transformation is the bottom right 4×4 block 48 vectorized from three 4x4 quadrant blocks of 8x8 primary coefficient blocks excluding matrix elements as input. In such an implementation, the omitted bottom right 4x4 The primary transform coefficients are ignored in the secondary transform. This further reduced transform is This vector is converted to a 16x1 output vector, which is a 4x4 matrix to fill 1408 in Figure 14. The three squares of secondary transform coefficients surrounding 1408 are zero-padded. That's fine.
[0125] With the help of such a reduction in the dimensions of the RST, it is possible to store all the LFNST matrices in memory. In the above example, for example, the memory usage is less than that of the implementation without dimensionality reduction. This can be reduced from 10 KB to 8 KB with reasonably little performance degradation compared to the previous state.
[0126] In some implementations, to reduce complexity, LFNST is used to LFNST may be further restricted to be applicable only if all coefficients outside the primary transform coefficient portion (e.g., outside 1406 portion of 1404 in FIG. 14) are insignificant. When used, all primary-only transform coefficients (e.g., the primary coefficient matrix in Figure 14) 1404) can be close to 0. Such a limit is imposed by the lowest This allows for the adjustment of LFNST index signaling at the Additional coefficient scans may be required to check for significant coefficients at specific locations when there are no Some implementations avoid the worst case scenario of LFNST (multiplying by 100 per pixel). The constraints (in terms of computation) can limit the non-separable transforms of 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In these cases, for other sizes less than 16, the last significant scan position must be less than 8 if LFNST is applied. For blocks with shapes 4xN and Nx4 and N>8, the above restriction no longer means that LFNST is applied only once to the top-left 4x4 region. Since all primary-only coefficients are 0 when LFNST is applied, the number of operations required for the primary transform is reduced in such cases. From the encoder's perspective, the quantization of coefficients can be simplified when LFNST transforms are tested. Rate-distortion optimized quantization (RDO) can be applied to the first 16 coefficients (in scan order) at most. and the remaining coefficients can be set to zero.
[0127] In some example implementations, the available RST kernels are: may be specified as a set of transformations containing non-separable transformation matrices of A total of four transformation sets and two non-separable transformation rows per transformation set used in LFNST. There can be sequences (kernels). These kernels may be pre-trained offline. , and thus data-driven. Offline trained transform kernels are used for encoding / decoding. stored in memory or on an encoding or decoding device for use during processing The selection of the transform set during the encoding or decoding process can be hard-coded into The mapping from intra prediction modes to transform sets can be determined by the intra prediction modes. The mapping can be predefined. An example of such a predefined mapping is shown in Table 4. For example, as shown in Table 4, three cross-component linear model (CCLM) modes (INTRA_LT_CCLM, INT RA_T_CCLM or INTRA_L_CCLM) is the current block (i.e., 81<=pred When used with ModeIntra<=83, transform set 0 is used for the current chroma block. For each transformation set, the selected non-separable secondary transformation candidates may be selected as follows: It can be further specified by an explicitly signaled LFNST index. For example, the index is synchronized in the bitstream after the transform coefficients, once per intra CU. It can be gunned down. [Table 4]
[0128] LFNST is implemented as follows: In the above implementation, all coefficients outside the first coefficient subgroup or portion are The LFNST index is restricted to be applicable only if it is not significant. The coding may depend on the position of the least significant coefficient. It can be text coding, but it does not depend on the intra prediction mode, and only the first bin is coded. Furthermore, LFNST supports both intra and inter coding. Can be applied to both intra CUs in a slice, as well as both luma and chroma When dual tree is enabled, luma and chroma LFNST indexes are calculated separately. Inter-slice (dual tree disabled) signaling is possible. If there is a single LFNST index, it is signaled and used for both luma and chroma. It is used.
[0129] In some example implementations, an intra-subpartitioning (ISP) mode is selected. If selected, performance will be poor even if RST is applied to all feasible partition blocks. Performance gains are likely to be limited, so LFNST is disabled and the RST index is signaled. Furthermore, disabling RST for ISP prediction residuals reduces the coding complexity. In some further implementations, multiple linear regression intra prediction ( When MIP mode is selected, LFNST is also disabled and the RST index is not signaled. It is not necessary.
[0130] Due to the existing maximum transform size limit (e.g., 64x64), large CUs ( or any other predetermined size representing the maximum transform block size) is implicitly split (e.g., Considering that the LFNST index search is performed in a certain number of decoding packets, Data buffering can be increased by a factor of four for the pipeline stage. Therefore, in some implementations, the maximum size allowed for an LFNST is limited to, for example, 64x64. In some implementations, LFNST is only valid with DCT2 as the primary transform. It can be made into.
[0131] In some other implementations, for example, with three kernels in each set, e.g. , we define 12 sets of secondary transforms for the luma component, called the Intra Secondary Transform (IST) ) is provided. Intra-mode dependent indexes are used for transformation set selection. The kernel selection within the set may be based on signaled syntax elements. When either DCT2 or ADST is used as both the horizontal and vertical primary transform, In some implementations, a 4x4 separation can be enabled according to the block size. You can choose between a non-separable transform or an 8x8 non-separable transform. _height)<8, 4x4IST can be chosen. For larger blocks, 8 x8IST can be used. Here, tx_width&tx_height are the conversion blocks. The input to the IST is a set of low frequency primaries in a zigzag scan order. It may also be a conversion coefficient.
[0132] Various transformations in the video coding or decoding process, e.g., within a residual block The primary transform or primary transform coefficients process the secondary transform of the block of samples. If only separable transformations are used, one of the transformations will be Orientational texture patterns such as edges that are substantially away from the vertical As mentioned above, some In an exemplary implementation, one or more separations are performed for the secondary transform of the primary transform coefficients. Impossible transformation designs can be used.
[0133] Transform block partitioning and transform type applied to the divided transform blocks can be interrelated. For example, The conversion type depends on the specific partition type. For example, the transform partitioning illustrated in Figure 16 and described above may be more suitable. The scheme uses non-recursive partition types as compared to recursive partitions such as those previously described in FIG. All available conversion types are presented along with all available partition patterns. For transformation blocks partitioned under the transformation partition type (e.g., transformation partition type in Figure 16), If allowed, the encoder can select which transform party to use to obtain the transform block. What partitioning type to use and how to allocate it for each partitioned transformation block When deciding which transformation type to use, it is important to perform optimization over a large parameter space. In practice, a certain set of conversion types will generally be needed for a particular type. Some conversions may be more appropriate for some partition types than others. Various implementations of the Transform Partition Type and Transform Type may consider interactions between them. Consider and use the following to limit the types of conversions allowed for a particular partition type: Similarly, we have arrived at a method to restrict the partition types allowed for certain conversion types. Such an implementation is particularly suited to non-recursive transformation partitioning. In the situation where the transform partition pattern and transform type of each divided transform block are used, When determining the loop selection, it may allow a reduction in the optimization space for the encoder.
[0134] These example implementations may be used separately, in any order, or in any manner. In the discussion above and below, "coded blocks" are used. The terms "picture block," "coding block," and the like refer to a picture unit on which prediction or transformation is performed. The coding block may be a luma coding block. In some situations, the coding The coded / coded block may refer to a predicted block. The term refers to the width or height of a coding block, or the maximum width and height, or the maximum width and height. Minimum height, or area size (width * height), or aspect ratio (width:height, or height :width).
[0135] Multiple candidate primary conversion types
[0136] In one embodiment, there can be multiple candidate primary transformation types for a block. The block may contain transformation blocks resulting from the partition. The type selection and / or signaling is based on a predefined set of transformation partition types. A predefined set of transformation partition types can be limited to the available transformations. A larger set of interchangeable partition types (e.g., the partition type flag in Figure 16) In other words, the transformation partition type of the block The primary partition is converted only if the partition belongs to a predefined set of partition types. The choice of transformation is made and signaled, instead of, for example, the transformation partition. If the type is of any other type, it is defaulted rather than selected and signaled. The default conversion type can be used.
[0137] In one implementation, the primary transform type is a discrete cosine transform (DCT) type 1 to a DCT Asymmetric Discrete Sine Transform (ADST) Type 8; Discrete Sine Transform (DST) Type 1 to DST Type 8 Line graph transformation (LGT); or at least one of the Karhunen-Loeve transformation (KLT) It may include one.
[0138] In one implementation, the predefined set of transformation partition types is, for example, Among the various conversion partition types, only PARTITION_NONE is included, i.e., conversion The block size is equal to the prediction block (or coding block) size. The primary transformation type is only available if it belongs to this predefined set. can be selected and / or signaled.
[0139] In one implementation, the number of partitions is also determined by a predefined number of conversion partition types. For example, the number of conversion partitions of a particular type can be considered to determine the selected set. Primary transformation type only if the number of partitions is below a predefined threshold A predefined set of transformation partition types from which the In one implementation, the predefined threshold is an integer between 1 and 16. That's fine.
[0140] Multiple possible secondary conversion types
[0141] In one embodiment, there may be multiple candidate secondary transformation types for a block. The block may contain the transform blocks resulting from the partition. The type selection and / or signaling is based on a predefined set of transformation partition types. A predefined set of conversion partition types is available. A larger set of possible transformation partition types (e.g., the transformation partition types in Figure 16) In other words, the transformation participants of a block A partition type is converted only if it belongs to a predefined set of partition types. Instead, the choice of the transformation is made and signaled, e.g. If the partition type is of another type, it is not selected and signaled. A default secondary transformation type may be used, or a secondary transformation may be performed. It's not necessary.
[0142] In one implementation, the secondary transform type may include a KLT, which is a transform with a different kernel. It may be configured.
[0143] In one implementation, the predefined set of transformation partition types is, for example, Of the various conversion partition types, only PARTITION_NONE may be included, i.e. The transform block size is equal to the prediction block (or coding block) size. Therefore, a secondary conversion type can be added only if it belongs to this predefined set. The type can be selected and / or signaled.
[0144] In one implementation, the number of partitions is also determined by a predefined number of conversion partition types. For example, the number of conversion partitions of a particular type can be considered to determine the selected set. Secondary transformation type only if the number of partitions is below a predefined threshold A predefined set of transformation partition types from which the In one implementation, the predefined threshold is an integer between 1 and 16. That's fine.
[0145] In one implementation, the selection and / or signaling of the secondary transform type is performed by the transform party. The combination may be based on a combination of the transformation type and the primary transformation type. a conversion partition type in a predefined set of conversion partition types; and a primary transform type in a predefined set of transform types, for example The secondary transformation type is a transformation partition type of PARTITION_NONE, and the block Select only if the primary transform type used in the block is DCT or ADST and / or Instead, the conversion partition type may need to be signaled, e.g. If it is of any other type, it will be used as the default secondary rather than being selected and signaled. A secondary transformation type may be used, or no secondary transformation may be performed.
[0146] Transformation-related signaling
[0147] In this disclosure, the order of transformation-related syntax elements / parameters is taken into consideration when signaling Various signaling mechanisms are disclosed that aim to increase efficiency.
[0148] In one embodiment, the conversion partition type information is a primary / secondary conversion type. The primary / secondary transform type selection can be signaled before the transform selection information. The conversion partition is a predefined set of conversion partition types, e.g. PARTIT It should only be signaled when it belongs to ION_NONE. If the partition does not belong to the predefined set of conversion partition types, The primary / secondary transform type selection may not need to be signaled. Instead, The primary and secondary conversion types are introduced as predefined default conversion types. It may be issued.
[0149] In one embodiment, the primary / secondary conversion type selection information is a conversion partition. It can be signaled before the type information. In this case, the conversion partition type selection and The signaling and / or the signaling may depend on the primary / secondary conversion type selection information.
[0150] In one implementation, the primary conversion type belongs to a predefined set of conversion types. Conversion partition type information may only need to be signaled if Otherwise, the conversion partition type information may not need to be signaled. As an example, the predefined set of transform types is DCT type 1 through DCT type 8, ADST, These may include, but are not limited to, DST Type 1 to DST Type 8, LGT, and KLT.
[0151] In one implementation, the secondary conversion type belongs to a predefined set of conversion types. Transform partition type information may only need to be signaled if: The predefined set of transformation types includes, but is not limited to, the predefined KLT interfaces. It may contain a specific KLT with a kernel associated with it. If the secondary transformation type does not belong to the predefined set of transformation types, Partition type information may not need to be signaled. Instead, the transformation parameter The partition type information is a predefined default conversion partition, such as PARTITION_NONE. It may be derived as a partition type.
[0152] 17 illustrates an example method 1700 for decoding video data. The method 1700 includes the following steps: receiving a coded video bitstream for a data block at step 1710; and extracting, from the coded video bitstream, a transform partition type associated with the data block at step 1720. Step 1730, the conversion partition type converts the data block to the conversion block. Pre-defined conversion partition types that specify the partitioning patterns for each block. In response to belonging to a subset of the defined set, the coded video bits are Transformation blocks separated from data blocks as signaled in the data stream extracting a transformation type of a transformation associated with the transformation, The transformation blocks are organized according to the steps and transformation types, which belong to a first predefined set of types. and performing an inverse transformation on the lock.
[0153] In embodiments of the present disclosure, any step and / or action may be performed in any quantity or amount as desired. Two or more steps and / or actions may be combined or arranged in a sequence. can be executed in parallel.
[0154] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the methods (or embodiments), the encoder, and the decoder may include processing circuitry (e.g., , one or more processors, or one or more integrated circuits) In one example, one or more processors may be stored in a non-transitory computer-readable medium. The embodiments of the present disclosure are applicable to luma blocks or chroma blocks. may be used.
[0155] The techniques described above involve the use of computer programs physically stored on one or more computer-readable media. It may be implemented as computer software using computer readable instructions, for example: FIG. 18 illustrates a computer system suitable for implementing certain embodiments of the disclosed subject matter. This indicates the year (1800).
[0156] Computer software is software that runs on one or more computer central processing units (CPUs). processing unit (GPU) and graphics processing unit (GPU) Contains instructions that can be executed directly by the microprocessor, or through interpretation and execution of microcode, etc. Assembly, compilation, linking, or similar mechanisms to create the code It may be coded using any suitable machine code or computer language that is amenable to do.
[0157] The instructions may be transmitted to, for example, a personal computer, a tablet computer, a server, a smartphone, or the like. Various types of computers including smartphones, game consoles, Internet of Things devices, etc. It may be executed on a computer or components thereof.
[0158] The components shown in FIG. 18 for computer system 1800 are illustrative in nature. and the scope of use or functionality of computer software implementing embodiments of the present disclosure. The configuration of components is not intended to imply any limitation on the Any one of the components shown in the exemplary embodiment of the computer system (1800) No dependency or requirement relating to combinations should be construed.
[0159] The computer system (1800) includes a specific human interface input device. Such a human interface input device may be, for example, a tactile input (keypad strokes, swipes, data glove movements, etc.), voice input (voice, clapping, etc.), visual One or more human users may be able to sense the presence of the object through sensory input (such as gestures), olfactory input (not shown), or other sensory input. It can respond to inputs from the user using a human interface device. music, ambient sounds, etc.), images (scanned images, photographic images obtained from still image cameras, etc. ), video (two-dimensional video, three-dimensional video including stereoscopic video, etc.), etc., It is also possible to incorporate specific media that are not necessarily directly related to the intended input.
[0160] Input human interface devices include a keyboard (1801), a mouse (1802), Rack pad (1803), touch screen (1810), data glove (not shown), One of the following: stick (1805), microphone (1806), scanner (1807), camera (1808) The device may include one or more (one of each is shown).
[0161] The computer system (1800) also includes a specific human interface output device. Such human interface output devices may include, for example, haptic output, It may stimulate one or more of the human user's senses through sound, light, and smell / taste. Human interface output devices such as tactile output devices (e.g., touch screen (1810), data glove (not shown), or joystick (1805) haptic feedback device that provides haptic feedback but does not function as an input device (possibly a chair), audio output device (speaker (1809), headphones (not shown) visual output devices (CRT screens, LCD screens, plasma screens, OLED screens (each with or without touchscreen input capability, each with or without touchscreen input capability) Some of these may provide two-dimensional visual output, or screen-based visual feedback. It may be possible to output more than three dimensions by means of telegraphic output or other means. a screen (1810) including a virtual reality glasses (not shown), a holographic display The equipment may include a power supply, a power source ...
[0162] The computer system (1800) also includes human-accessible storage devices and their associated Optical media including CD / DVD ROM / RW (1820) with media such as CD / DVD 1821), thumb drive (1822), removable hard drive or solid state drive legacy magnetic media such as EVE (1823), tape and floppy disks (not shown), Includes dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown) It is possible.
[0163] Those skilled in the art will also understand the concepts of "computer-readable media" used in connection with the presently disclosed subject matter. It is understood that the term "transmission medium" does not encompass transmission media, carrier waves, or other transitory signals. It should be.
[0164] The computer system (1800) also provides connectivity to one or more communication networks (1855). The network may include, for example, a wireless, wired, or optical network. The network can be further classified as local, wide area, metropolitan, vehicular and industrial, It can be real-time, delay tolerant, etc. Examples of networks include Ethernet Security including local area networks such as Wi-Fi, GSM, 3G, 4G, 5G, LTE, etc. Television, including regional networks, cable, satellite, and terrestrial broadcast television Wired or wireless wide area digital networks, including CANbus, for vehicles and industrial applications. A particular network is usually connected to a network (e.g., a USB port on a computer system (1800)). Any external network attached to a specific general-purpose data port or peripheral bus (1849) Other networks typically require an interface adapter, as described below. It is integrated into the core of the computer system (1800) by attaching it to the system bus ( For example, an Ethernet interface to a PC computer system or a smartphone computer. cellular network interfaces to computer systems). The computer system (1800) can communicate with other entities using either Such communications may be one-way, receive only (e.g., television broadcast), or one-way, transmit only (e.g., CANbus to a specific CANbus device), or bidirectional, e.g., local or wide area Communication to other computer systems using a real digital network. The protocols and protocol stacks, as explained above, are It can be used for both the network interface and the network interface.
[0165] the human interface device, the human-accessible storage device, and The network interface is connected to the core (1840) of the computer system (1800). It is possible.
[0166] The core (1840) may include one or more central processing units (CPUs) (1841), a graphics processing unit (GPU), or a (GPU) (1842), Field Programmable Gate Array (FPGA) (1843) dedicated programmable processing units, hardware accelerators for specific tasks (1844 ), a graphics adapter (1850), etc. These devices can Read-Only Memory (ROM) (1845), Random Access Memory (1846), User Accessible Internal mass storage devices such as internal hard drives, SSDs, etc. (1847) cannot be used with the system. Some computer systems may have additional CPUs, GPUs, A system bus in the form of one or more physical plugs to allow expansion by PUs etc. The peripheral devices can access the core's system bus (1848). They may be directly connected or may be connected via a peripheral bus (1849). The peripheral bus architecture allows the computer (1810) to be connected to a graphics adapter (1850). The architecture includes PCI, USB, etc.
[0167] CPU (1841), GPU (1842), FPGA (1843), and accelerator (1844) are combined The computer is capable of executing the specific instructions that together constitute the aforementioned computer code. The code can be stored in ROM (1845) or RAM (1846). The transition data is stored in RAM (1846). Alternatively, permanent data may be stored, for example, on an internal mass storage device (1847). High speed storage and retrieval of any memory device is performed by one or more CPUs (1841), GPUs (1842) , mass storage devices (1847), ROM (1845), RAM (1846), etc. This can be made possible through the use of cache memory.
[0168] The computer-readable medium may include computer code for performing various computer-implemented operations. The media and computer code may be specially may be designed and constructed by a person skilled in the art of computer software. The material may be of any known and available type.
[0169] As a non-limiting example, a computer system (1800) having an architecture, particularly The processor (1840) is a processor (including CPU, GPU, FPGA, accelerator, etc.) the processor is software embodied in one or more tangible computer readable media Such a computer-readable medium can provide functionality as a result of executing the program. The user-accessible mass storage devices introduced above, as well as the core internal mass storage devices, Related to specific storage devices of the core (1840) of a non-transitory nature, such as memory (1847) and ROM (1845) The software implementing the various embodiments of the present disclosure may be stored on such media. The computer-readable medium may be stored in such a device and executed by the core (1840). , which may include one or more memory devices or chips depending on specific needs. The software is responsible for the cores (1840), specifically the processors (CPU, GPU, and FP) within them. (including GA, etc.) to define the data structure stored in RAM (1846) and modifying such data structures according to a process defined by the software, Performing a particular process or a particular part of a particular process described herein Additionally or alternatively, the computer system may be The software is used to perform a specific process or a specific part of a specific process. hardwired into circuitry or otherwise operable in place of or in conjunction with software. Providing functionality as a result of embodied logic (e.g., accelerators (1844)) References to software include logic, where appropriate. Where appropriate, reference to a computer-readable medium may also refer to a medium for execution. A circuit (such as an integrated circuit (IC)) that stores software for The present disclosure may include a circuit that embodies logic for the encompasses any suitable combination of hardware and software.
[0170] While this disclosure has described several exemplary embodiments, modifications that fall within the scope of this disclosure are possible. , substitutions, and various alternative equivalents. Although not shown or described, it embodies the principles of the present disclosure and therefore It will be appreciated that numerous systems and methods can be devised that fall within the spirit and scope. Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplemental Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: Conversion unit PU: Prediction Unit CTU: Coding Tree Unit CTB: coding tree block PB: Predicted Block HRD: Hypothetical Reference Decoder SNR: Signal-to-Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: cathode ray tube LCD: Liquid crystal display OLED: Organic Light Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area SSD: Solid State Drive IC: Integrated Circuit HDR: High Dynamic Range SDR: Standard Dynamic Range JVET: Joint Video Exploration Team MPM: Most Probable Mode WAIP: Wide-angle Intra Prediction CU: Coding Unit PU: Prediction Unit TU: Conversion unit CTU: Coding Tree Unit PDPC: Position-dependent prediction combination ISP: Intra-subpartition SPS: Sequence parameter settings PPS: Picture Parameter Set APS: Adaptive Parameter Set VPS: Video Parameter Set DPS: Decoding Parameter Set ALF: Adaptive Loop Filter SAO: Sample Adaptive Offset CC-ALF: Cross-component adaptive loop filter CDEF: Constrained Directional Enhancement Filter CCSO: Cross-component sample offset LSO: Local Sample Offset LR: Loop Recovery Filter AV1:AOMedia Video1 AV2:AOMedia Video2 [Explanation of symbols]
[0171] 101 Samples 102 Arrow 103 Arrow 104 Square Blocks Block 201 300 Communication Systems 310 Terminal Equipment 320 Terminal Equipment 330 Terminal Equipment 350 Network 400 Communication Systems 401 Video Source 402 Stream 403 Video Encoder 404 Video Data, Video Bitstream 405 Streaming Server 406 Client Subsystem 407 Input copy, video data 410 Video Decoder 411 Output Stream 412 Display 413 Video Capture Subsystem 420 Electronic equipment 430 Electronic equipment 501 Channel 510 Video Decoder 512 Rendering devices, displays 515 buffer memory 520 Parser 521 Symbol 530 Electronic equipment 531 Receiver 551 Reverse conversion unit 552 Intra-picture prediction unit 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Picture Memory 558 Picture Buffer 601 Video Sources 603 Video Coder, Video Encoder 620 Electronic equipment 630 Source Coder 632 Coding Engine 633 Local Video Decoder, Decoding Unit 634 Reference Picture Cache, Reference Picture Memory 635 Predictor 640 Transmitter 643 Video Sequences 645 Entropy Coder 650 Controller 660 Communication Channels 703 Video Encoder 721 General-purpose controller 722 Intra Encoder 723 Residual Calculator 724 Residual Encoder 725 Entropy Encoder 726 Switch 728 Residual Decoder 730 Interencoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Reconstruction Module 880 Interdecoder 1002 blocks 1004 Upper sample 1006 Upper left sample 1008 Left Sample 1010 samples Block 1102 1202 coded blocks 1204 Quadtree splitting Block 1302 1304 Quadtree Subblocks 1402 Forward Primary Transform 1404 Primary Transform Coefficient Matrix 1406 Primary coefficients, shaded area 1408 Secondary Conversion Factors 1800 Computer Systems 1801 keyboard 1802 Mouse 1803 Trackpad 1805 Joystick 1806 Mike 1807 Scanner 1808 Camera 1809 Audio Output Device Speaker 1810 touchscreen 1821 Optical media 1822 thumb drive 1823 Solid State Drive 1840 Core 1843 Field Programmable Gate Area (FPGA) 1844 Hardware Accelerator 1845 Read-Only Memory (ROM) 1846 Random Access Memory 1847 Core Internal Mass Storage 1848 System Bus 1849 Peripheral Bus 1850 graphics adapter 1854 Interface 1855 Communication Network
Claims
1. A method for video processing, The steps include determining the conversion partition type associated with the data block, The steps include encoding a first syntax element indicating the conversion partition type into a video bitstream, The conversion partition type belongs to a subset of a predefined set of conversion partition types, each specifying a partition pattern for dividing the data block into subblocks, and in response that the number of conversion partitions associated with each conversion partition type in the subset of the predefined set of conversion partition types is less than or equal to a predefined threshold, A step of determining the transformation type of a transformation associated with a subblock separated from the data block, wherein the transformation type belongs to a first predefined set of transformation types. The steps include encoding a second syntax element indicating the conversion type into the video bitstream, In response that the conversion partition type does not belong to the subset of the predefined set of conversion partition types, A step of determining the transformation type of the transformation associated with the subblock separated from the data block as the default transformation type, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream and Methods that include...
2. The transformation is a primary transformation, and the first predefined set of transformation types is: Discrete cosine transform (DCT) type 1 to DCT type 8, Asymmetric Discrete Sine Transform (ADST), Discrete sine transform (DST) type 1 to DST type 8, Line graph transformation (LGT), and The method according to claim 1, comprising the Karunen-Löwe transformation (KLT).
3. The method according to claim 1 or 2, wherein the subset of the predefined set of conversion partition types consists of PARTITION_NONE for no conversion block partitioning.
4. The method according to any one of claims 1 to 3, wherein the predefined threshold includes an integer between 1 and 16.
5. The method according to any one of claims 1 to 4, wherein the transformation is a secondary transformation, and the first predefined set of transformation types includes KLT.
6. The method according to any one of claims 1 to 5, wherein the subset of the predefined set of conversion partition types consists of PARTITION_NONE for no conversion block partitioning.
7. The method according to any one of claims 1 to 6, wherein the number of conversion partitions associated with each conversion partition type in the subset of the predefined set of conversion partition types is less than or equal to a predefined threshold.
8. The method according to any one of claims 1 to 7, wherein the conversion type further indicates that the type of primary conversion associated with the conversion block belongs to a second predefined set of conversion types.
9. The method according to claim 8, wherein the subset of the predefined set of conversion partition types consists of PARTITION_NONE for no conversion block partitioning, and the second predefined set of conversion types consists of DCT and ADST.
10. A method for video processing, The steps include determining the transformation type associated with a subblock of a data block, The steps include encoding a first syntax element indicating the conversion type into a video bitstream, If the aforementioned conversion type belongs to a subset of a predefined set of conversion types, The steps include determining the conversion partition type associated with the data block, The steps include encoding a second syntax element indicating the conversion partition type into the video bitstream, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream, Methods that include...
11. If the conversion type does not belong to the predefined set of conversion partition types, The method of claim 10, further comprising the step of determining that the conversion partition type associated with the data block is a predefined default conversion partition type, wherein the predefined default conversion partition type includes PARTITION_NONE.
12. The aforementioned transformation is a primary transformation, and the predefined set of transformation types is: Discrete cosine transform (DCT) type 2, Asymmetric Discrete Sine Transform (ADST), DCT Type 1 to DCT Type 8, Discrete sine transform (DST) type 1 to DST type 8, Line graph transformation (LGT), or The method according to claim 10 or 11, comprising the Karunen-Löwe transformation (KLT).
13. The method according to any one of claims 10 to 12, wherein the transformation is a secondary transformation, and the predefined set of transformation types includes a KLT having a kernel associated with a predefined KLT index.
14. A method for transmitting a video bitstream generated by a method of encoding a conversion block into a video bitstream, The method for encoding the conversion block into a video bitstream is: The steps include determining the conversion partition type associated with the data block, The steps include encoding a first syntax element indicating the conversion partition type into a video bitstream, The conversion partition type belongs to a subset of a predefined set of conversion partition types, each specifying a partition pattern for dividing the data block into subblocks, and in response that the number of conversion partitions associated with each conversion partition type in the subset of the predefined set of conversion partition types is less than or equal to a predefined threshold, A step of determining the transformation type of a transformation associated with a subblock separated from the data block, wherein the transformation type belongs to a first predefined set of transformation types. The steps include encoding a second syntax element indicating the conversion type into the video bitstream, In response that the conversion partition type does not belong to the subset of the predefined set of conversion partition types, A step of determining the transformation type of the transformation associated with the subblock separated from the data block as the default transformation type, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream and Methods that include...
15. A method for storing a video bitstream generated by a method of encoding a conversion block into a video bitstream, The method for encoding the conversion block into a video bitstream is: The steps include determining the conversion partition type associated with the data block, The steps include encoding a first syntax element indicating the conversion partition type into a video bitstream, The conversion partition type belongs to a subset of a predefined set of conversion partition types, each specifying a partition pattern for dividing the data block into subblocks, and in response that the number of conversion partitions associated with each conversion partition type in the subset of the predefined set of conversion partition types is less than or equal to a predefined threshold, A step of determining the transformation type of a transformation associated with a subblock separated from the data block, wherein the transformation type belongs to a first predefined set of transformation types. The steps include encoding a second syntax element indicating the conversion type into the video bitstream, In response that the conversion partition type does not belong to the subset of the predefined set of conversion partition types, A step of determining the transformation type of the transformation associated with the subblock separated from the data block as the default transformation type, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream and Methods that include...
16. A method for transmitting a video bitstream generated by a method of encoding a conversion block into a video bitstream, The method for encoding the conversion block into a video bitstream is: The steps include determining the transformation type associated with a subblock of a data block, The steps include encoding a first syntax element indicating the conversion type into a video bitstream, If the aforementioned conversion type belongs to a subset of a predefined set of conversion types, The steps include determining the conversion partition type associated with the data block, The steps include encoding a second syntax element indicating the conversion partition type into the video bitstream, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream, Methods that include...
17. A method for storing a video bitstream generated by a method for encoding a conversion block into a video bitstream, The method for encoding the conversion block into a video bitstream is: The steps include determining the transformation type associated with a subblock of a data block, The steps include encoding a first syntax element indicating the conversion type into a video bitstream, If the aforementioned conversion type belongs to a subset of a predefined set of conversion types, The steps include determining the conversion partition type associated with the data block, The steps include encoding a second syntax element indicating the conversion partition type into the video bitstream, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream, Methods that include...