Interaction between transformation partitioning and primary / secondary transformation type selection
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2025-09-16
- Publication Date
- 2026-08-04
AI Technical Summary
【0029】 以下の詳細な説明と添付の図面とから、開示されている保護対象のさらなる特徴、性質 及び様々な効果がより明らかになる。
Smart Images

Figure 0007900581000012 
Figure 0007900581000013 
Figure 0007900581000014
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 175,897, filed on April 16, 2021, and U.S. Non - Provisional Patent Application No. 17 / 568,275, filed on January 4, 2022, and the entire contents of both applications are incorporated herein by reference.
[0002] This disclosure describes a set of advanced video coding techniques. More specifically, the disclosed techniques include the interaction between transform partitioning schemes and primary / secondary transform type selection in video encoding and decoding.
Background Art
[0003] The description of the background art provided herein is intended to present the context of the present disclosure generally. The research of the inventors is not, insofar as that research is described in this background art section, and also not, in other respects, admitted as prior art at the time of filing of this application, either explicitly or implicitly, with respect to the present disclosure.
[0004] Video encoding and video decoding can be performed using inter - picture prediction with motion compensation. Uncompressed digital video consists of a series of pictures that can be, for example, 1920×1080 luminance samples and associated full - sampling or sub - sampled chrominance samples in spatial dimensions. A series of pictures can have a fixed or variable picture rate (or frame rate, also called) of, for example, 60 pictures per second or 60 frames per second. Uncompressed video is suitable for streaming or data processing such as, for example, 1920×1080 luminance samples and associated full - sampling or sub - sampled chrominance samples in spatial dimensions. A series of pictures can have a fixed or variable picture rate (or frame rate, also called) of, for example, 60 pictures per second or 60 frames per second. Uncompressed video is suitable for streaming or data processing such as, for example, 1920×1080 luminance samples and associated full - sampling or sub - sampled It has specific bitrate requirements for [something]. For example, a pixel resolution of 1920 x 1080, 60 frames per second. A frame rate of ms / second, and 4:2:0 chroma channel with 8 bits per pixel per color channel. Video with subsampling requires a bandwidth of nearly 1.5 Gbit / s. Videos like this require more than 600 GB of storage space.
[0005] One purpose of video coding and video decoding is to compress uncompressed input. This may reduce the redundancy of the video signal. Compression can meet the aforementioned bandwidth and / or storage space requirements. In some cases, this can help reduce the size by more than two orders of magnitude. Both lossless and lossy compression, and And combinations of these can be used. Lossless compression is the decoding of an exact copy of the original signal. This refers to a technique that allows for the reconstruction of the original signal from a compressed signal. (Lossy compression) This means that the original video information is not fully preserved during coding and cannot be fully recovered during decryption. This refers to the coding / decoding process. When using lossy compression, the recomposition process is performed. The reconstructed signal may not be identical to the original signal, but there may be distortion between the original signal and the reconstructed signal. Even with some information loss, the reconstructed signal is sufficient to be useful for its intended purpose. The file size decreases by the minute. In the case of video, lossy compression is widely used in many applications. The amount of distortion that can be tolerated depends on the application. For example, certain consumer video streaming applications These users may tolerate higher levels of distortion than users who use it for movies or television broadcasts. The compression ratio achievable by the compression algorithm is selected to reflect various distortion tolerances. Or it can be adjusted. That is, generally, the higher the distortion tolerance, the higher the loss and the higher This enables coding algorithms that provide a compression ratio.
[0006] Video encoders and video decoders include, for example, motion compensation, Fourier transform, quantization, And techniques from several broad categories and steps, including entropy coding. They can use magic.
[0007] Video codec technology can include techniques known as intracoding. In intracoding, sample values are taken from a previously reconstructed reference picture. Represented without referencing samples or other data. Some video codecs use picture The sample is spatially subdivided into blocks. All blocks of the sample are intra When coded in a specific mode, that picture can be called an intra picture. Intrapicture and their derivatives such as independent decoder refresh pictures. The command can be used to reset the decoder state, and therefore the decoder As the first picture in the video bitstream and video session, It can be used as a still image. Next, a sample of the block after intra prediction. This can be transformed into the frequency domain, and the resulting transformation coefficient can be expressed as entropy. It can be quantized before coding. Intra prediction is the sun in the pre-transformation region. This describes a technique to minimize the pull value. In some cases, the smaller the DC value after conversion, and the AC value... The smaller the number, the less a given quantization state is needed to represent the block after entropy coding. The number of bits required for the top size decreases.
[0008] For example, conventional intra coding, such as that known from MPEG-2 generation coding technology, does not use intra prediction. However, some newer video compression technologies attempt to code / decode blocks based on, for example, surrounding sample data and / or metadata that precede in decode order the block of intra-coded or intra-decoded data obtained during spatial adjacent coding and / or decoding. Such technologies are hereinafter referred to as "intra prediction" technologies. It should be noted that in at least some cases, intra prediction uses reference data only from the current picture being reconstructed and does not use reference data from other reference pictures. Intra prediction can take many different forms. If two or more of such technologies are available in a given video coding technology, the technologies used can be referred to as intra prediction modes. One or more intra prediction modes can be provided for a particular codec. In certain cases, a mode can have sub-modes and / or can be associated with various parameters. The intra coding parameters of the mode / sub-mode information and the video block can be coded individually or can be included together in the codeword of the mode. Which codeword to use for a given combination of mode, sub-mode, and / or parameters can affect the coding efficiency improvement via intra prediction and thus can also affect the entropy coding technology used to convert the codeword into the bitstream.
[0009] <http: / / www.w3.org / TR / PNG / >
[0010] Specific modes of intra prediction were introduced in H.264, improved in H.265, and further improved in newer coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Generally, in intra prediction, predictor blocks can be formed using available adjacent sample values. For example, available values of a specific set of adjacent samples along a specific direction and / or line can be copied into the predictor block. The reference to the direction used can be coded in the bitstream or can itself be predicted. along with newer coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). and further improved in newer coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Generally, in intra prediction, predictor blocks can be formed using available adjacent sample values. For example, available values of a specific set of adjacent samples along a specific direction and / or line can be copied into the predictor block. The reference to the direction used can be coded in the bitstream or can itself be predicted. For example, available values of a specific set of adjacent samples along a specific direction and / or line can be copied into the predictor block. The reference to the direction used can be coded in the bitstream or can itself be predicted.
[0011] Referring to FIG. 1A, shown in the lower right is a subset of nine predictor directions specified by 33 possible predictor directions of H.265 (corresponding to 33 of the 35 intra modes specified in H.265). The point (101) where the arrows converge represents the predicted sample. The arrows represent the directions in which adjacent samples are used to predict the sample at 101. For example, arrow (102) indicates that sample (101) is predicted from one or more adjacent samples at a 45-degree angle from the horizontal in the upper right direction. Similarly, arrow (103) indicates that sample (101) is predicted from one or more adjacent samples at a 22.5-degree angle from the horizontal in the lower left direction of sample (101). Referring to FIG. 1A, shown in the lower right is a subset of nine predictor directions specified by 33 possible predictor directions of H.265 (corresponding to 33 of the 35 intra modes specified in H.265). The point (101) where the arrows converge represents the predicted sample. The arrows represent the directions in which adjacent samples are used to predict the sample at 101. For example, arrow (102) indicates that sample (101) is predicted from one or more adjacent samples at a 45-degree angle from the horizontal in the upper right direction. Similarly, arrow (103) indicates that sample (101) is predicted from one or more adjacent samples at a 22.5-degree angle from the horizontal in the lower left direction of sample (101). Similarly, arrow (103) indicates that sample (101) is predicted from one or more adjacent samples at a 22.5-degree angle from the horizontal in the lower left direction of sample (101).
[0012] Further referring to FIG. 1A, in the upper left, a square block (104) of 4×4 samples (shown by the thick dashed line) is depicted. The square block (104) contains 16 samples. The square block (104) contains 16 samples. , respectively, "S", its position in the Y dimension (e.g., row index), and its position in the X dimension ( For example, it is labeled by the column index. For example, sample S21 is Y-dimensional ( This is the second sample (from the top) and the first sample (from the left) in dimension X. Similarly, Sample S44 is the fourth sample in both the Y and X dimensions within block (104). Since the block size is 4x4 samples, S44 is in the bottom right. Further examples of reference samples to follow are shown. The reference sample is R, block (104) It is labeled at its Y position (e.g., row number) and X position (column number). H.264 In both H.265 and H.265, predicted samples adjacent to the block being reconstructed are used.
[0013] Intrapicture prediction in block 104 follows the signaled prediction direction for adjacent images. You can start by copying a reference sample value from the sample. For example, coded The video bitstream indicates the predicted direction of arrow (102) for this block 104. This includes signaling, i.e., the sample is from one or more predicted samples to the upper right, water Assume the prediction is at a 45-degree angle from the horizontal direction. In such cases, samples S41, S32, S2 3. S14 is predicted from the same reference sample R05. Then, sample S44 is predicted from the reference sample Predicted from R08
[0014] In certain cases, to calculate the reference sample, the directions are divided evenly by 45 degrees. If the values cannot be separated, the values of multiple reference samples may be combined, for example, by interpolation. .
[0015] The number of possible directions has increased as video coding technology continues to develop. In .264 (2003), for example, nine different directions are available for intra-prediction. In H.265 (2013), this increased to 33, and JEM / VVC / BMS had a maximum of 65 at the time of this disclosure. It can support direction. It helps in identifying the most appropriate intra-predictive direction. Experimental research is being conducted, using specific entropy coding techniques to determine direction. Accepting a specific bit penalty, their most appropriate direction is coded with a small number of bits. It can be numbered. Furthermore, the direction itself can be used in the intra-prediction of the decoded adjacent blocks. In some cases, predictions can be made from adjacent directions.
[0016] Figure 1B shows the increasing number of prediction directions in various coding techniques that have developed over time. For illustrative purposes, a schematic diagram (180) showing 65 intra-prediction directions by JEM is provided.
[0017] Bits representing the intra-prediction direction in the coded video bitstream The mapping to the predicted direction may differ depending on the video coding technique, for example. Therefore, from a simple direct mapping of prediction direction versus intra-prediction mode, the codeword, the most possible This can extend to complex adaptive schemes, including highly efficient modes, and similar technologies. However, not all In this case, statistically, the likelihood of it occurring in video content is greater than in other specific directions. There may be specific directions for low intro prediction. The purpose of video compression is to reduce redundancy. Therefore, in well-designed video coding techniques, the less likely of these is A direction is represented by more bits than the more likely direction.
[0018] Interpicture prediction, or interpretation, may be based on motion compensation. For compensation, sample data from a previously reconstructed picture or part thereof (reference picture) After Ta is spatially shifted in the direction indicated by the motion vector (MV from now on), It can be used to predict a newly reconstructed picture or a portion of a picture (e.g., a block). In some cases, the reference picture may be the same as the picture currently being reconstructed. MV is It may have two dimensions X and Y, or three dimensions, the third dimension being (time dimension and class This is an instruction for a similar reference picture to be used.
[0019] Some video compression techniques have current MV that can be applied to specific areas of sample data. This is a sequence of MVs that is spatially adjacent to, for example, the area being reconstructed, and precedes the current MV in the decoding order. Therefore, it can be predicted from other MVs related to other areas of the sample data. By doing so, we code the MV by relying on the removal of redundancy in correlated MVs. This significantly reduces the overall amount of data required, thereby improving compression efficiency. It increases. MV prediction can work effectively for, for example, (knowing as natural video). When coding the input video signal derived from the camera (which is being used), a single MV Areas larger than the applicable area will move in the same direction in the video sequence. There is a statistical likelihood that, therefore, in some cases, it can be derived from the MV of the adjacent area. This is because it can be predicted using similar motion vectors. As a result, The actual MV in a given area is the same as or identical to the MV predicted from the surrounding MV. Furthermore, after entropy coding, the MV is predicted from (one or more) adjacent MVs. Fewer bits would be used if it were coded directly rather than being coded. It can be represented by an unknown number of bits. In some cases, MV prediction is made using the original signal (i.e., samples). This can be used as an example of lossless compression of a signal (i.e., MV) derived from a stream. In other cases, for example, due to rounding errors when calculating the predictor from some surrounding MVs. Furthermore, MV prediction itself can be irreversible.
[0020] H.265 / HEVC (ITU-T Rec.H.265, "High Efficiency Video Coding", December 2016 Various MV prediction mechanisms are described in ). Of the many MV prediction mechanisms specified by H.265, the following The technique described below will be referred to as "spatial merging" from now on.
[0021] Specifically, referring to Figure 2, the current block (201) is during the motion search process. The coda can predict from the previous block of the same size, which has been spatially shifted. Includes the detected sample. Instead of directly coding its MV, the MV is A0, A1, and And one of the five surrounding samples represented by B0, B1, and B2 (202 to 206 respectively) Using linked MVs, you can associate metadata with one or more reference pictures. For example, it can be derived from the last referenced picture (in the decoding order). In H.265, MV prediction uses predictors from the same reference picture used by the adjacent block. It is possible. [Overview of the Initiative] [Means for solving the problem]
[0022] This disclosure relates to methods, apparatus, and computers for video encoding and / or video decoding. Various embodiments of readable storage media will be described.
[0023] According to one embodiment of the present disclosure, the decoder encodes / decodes video data. This provides a method for doing so. This method involves coding video bits in data blocks. The steps involve receiving a video bitstream and coding the video bitstream. Steps to extract the conversion partition type associated with the data block, and conversion The partition type is a partitioning pattern for dividing data blocks into transformation blocks. Each belongs to a subset of a predefined set of conversion partition types that specify them. In response to this, it is signaled with a coded video bitstream. The transformation type associated with the transformation block separated from the data block. A step to extract the conversion type, wherein the conversion type is a first predefined set of conversion types. A step that performs an inverse transformation on a transformation block according to the step and transformation type. Includes steps.
[0024] In another aspect, the present disclosure provides a method for encoding / decoding video data. The method provides a coded video bitstream of data blocks. The steps involved in the coding process, and the video data from the coded video bitstream. Steps to extract the conversion partition type associated with the data block, and conversion A conversion partition belonging to a subset of a predefined set of partition types. In response to the type, the data blocks are extracted from the coded video bitstream. Steps to extract the conversion type associated with the conversion block, and the conversion partition In response to a conversion partition type that does not belong to a predefined set of types, data This includes the step of identifying the conversion type of the tab block by default.
[0025] In another aspect, the present disclosure provides a method for encoding / decoding video data. The method provides a coded video bitstream of data blocks. The steps to trust and the data blocks from the coded video bitstream Steps to extract the conversion type of the conversion associated with the conversion block, and the conversion type In response to the conversion type belonging to the predefined set, the coded video bits Extract the converted partition type associated with the data block from the stream. Includes steps.
[0026] In another aspect, one embodiment of the present disclosure is for video encoding and / or video decoding The device provides a memory for storing instructions and a processor for communicating with the memory. Includes. When the processor executes an instruction, the processor sends video decoding and / or video to the device. The system is configured to perform the above method for decoded encoding.
[0027] In yet another aspect, one embodiment of the present disclosure relates to video decoding and / or video encoding. When performed by a computer, for the purpose of video decoding and / or video encoding A non-temporary computer-readable medium that stores instructions to cause a computer to perform the above method. To provide.
[0028] Other embodiments and their implementations are described in the drawings, specification, and claims. I will explain in more detail.
[0029] Further features and properties of the protected subject matter disclosed can be found in the following detailed description and attached drawings. And various effects become clearer. [Brief explanation of the drawing]
[0030] [Figure 1A] This is a schematic diagram of an exemplary subset of intra predictive directionality modes. [Figure 1B] This is a diagram illustrating an example of an intra-prediction direction. [Figure 2] This is a schematic diagram showing the current block and potential spatial merge candidates around it for motion vector prediction in one example. [Figure 3] This is a schematic diagram showing a simplified block diagram of a communication system (300) according to one exemplary embodiment. [Figure 4] This is a schematic diagram showing a simplified block diagram of a communication system (400) according to one exemplary embodiment. [Figure 5] This is a schematic diagram showing a simplified block diagram of a video decoder according to one exemplary embodiment. [Figure 6] This is a schematic diagram showing a simplified block diagram of a video encoder according to one exemplary embodiment. [Figure 7] A block diagram showing a video encoder according to another exemplary embodiment. [Figure 8] A block diagram showing a video decoder according to another exemplary embodiment. [Figure 9] This figure shows a directional intra-prediction mode according to an exemplary embodiment of the present disclosure. [Figure 10] This figure shows a non-directional intra-predictive mode according to an exemplary embodiment of the present disclosure. [Figure 11] This disclosure illustrates a recursive intra-prediction mode according to an exemplary embodiment. [Figure 12] This disclosure illustrates the transformation block partitioning and scanning of intra-predictive blocks according to exemplary embodiments of this disclosure. [Figure 13] This disclosure illustrates transform block partitioning and scanning of interpredictive blocks according to exemplary embodiments of this disclosure. [Figure 14] This disclosure illustrates an exemplary embodiment of a low-frequency inseparable conversion process. [Figure 15] This figure shows an intra-prediction scheme based on various reference lines according to an exemplary embodiment of the present disclosure. [Figure 16] This figure shows a non-recursive block partitioning method according to an exemplary embodiment of the present disclosure. [Figure 17] A flowchart according to an embodiment of this disclosure is shown. [Figure 18] This is a schematic diagram showing a computer system according to an exemplary embodiment of the present disclosure. [Modes for carrying out the invention]
[0031] Figure 3 shows a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication system (300) communicates with each other, for example, via a network (350). It includes multiple terminal devices capable of doing so. For example, the communication system (300) is a network (350 Includes a first pair of terminal devices (310) and (320) interconnected via ). In the example in Figure 3, The first pair of terminal devices (310) and (320) can perform one-way data transmission. For example, Terminal device (310) transmits to other terminal device (320) via network (350) (For example, the stream of video pictures captured by the terminal device (310)) Video data can be coded. Encoded video data is one or more codes. It can be transmitted in the form of a processed video bitstream. Terminal device (320) It receives coded video data from the network (350) and coding The decrypted video data is restored to the video picture, and according to the restored video data... It can display video pictures. One-way data transmission is practical for media serving applications, etc. It can be done.
[0032] In another example, the communication system (300) may perform, for example, a video conferencing application. A second pair of terminal devices (330) and (34) perform bidirectional transmission of the transmitted video data. Including 0). For bidirectional data transmission, in one example, terminal devices (330) and (340) The terminal device connects to the other terminal device of terminal devices (330) and (340) via the network (350). For sending to a device (for example, the video picture captured by that terminal device) The video data of the Ream can be coded. Each terminal device of terminal devices (330) and (340) The device also transmits coded data from terminal devices (330) and (340) to the other terminal device. It receives the coded video data and decodes the coded video data to create a video picture. Restore the video and view the video picture on an accessible display device according to the restored video data. It can display "ya".
[0033] In the example in Figure 3, terminal devices (310), (320), (330), and (340) are servers, personal The underlying principles of this disclosure may be implemented as a computer and a smartphone. The applicability is not limited in this way. Embodiments of this disclosure include desktop computers, Laptop computers, tablet computers, media players, wearables This can be implemented in computers, dedicated video conferencing equipment, etc. Network (350) This includes, for example, terminal devices (310) including wired (wired connection) and / or wireless communication networks. ), (320), (330), and (340) any Represents a number or type of network. Communication network (350)9 is a circuit-switched channel Data can be exchanged over packet-switched channels and / or other types of channels. Typical networks include telecommunications networks, local area networks, and wide-area networks. Including networks and / or the Internet. For the purposes of this study, networks (35 The architecture and topology of 0) are described herein unless expressly described herein. This may not be important for its operation.
[0034] Figure 4 shows an example of the subject matter of disclosure, specifically a video streaming environment. The arrangement of the decoder and video decoder is shown. The subject of the disclosure is, for example, video conferencing, digital Digital media including television broadcasting, games, virtual reality, CDs, DVDs, memory sticks, etc. This can be equally applied to other video-related applications, including the storage of compressed video on a device.
[0035] Video streaming systems are uncompressed video picture or image streaming systems. A video source (401) for creating a video (402), such as a digital camera, may be used. It may include a video capture subsystem (413) capable of capturing video pictures. The stream (402) contains samples recorded by the digital camera of video source 401. Hmm. A video picture stream (402) is encoded video data (404) (or code To emphasize the high data volume compared to a processed video bitstream. It is shown in bold and includes a video encoder (403) coupled to a video source (401). It can be processed by an electronic device (420). The video encoder (403) is as follows: To enable or implement aspects of the subject matter of the disclosure as described in more detail, This may include wearables, software, or a combination thereof. Encoded video data A bitstream (404) (or encoded video bitstream (404)) is an uncompressed video bitstream. The thin line highlights the low data volume compared to the chat stream (402). For future use, the streaming server (405) or downstream video equipment (Figure) It can be directly stored in (not shown). The client subsystem (406) and (40) in Figure 4. 8) One or more streaming client subsystems, such as the streaming client subsystem, Access the server (405) and obtain copies (407) and (4) of the encoded video data (404). 09) can be obtained. The client subsystem (406) can obtain, for example, an electronic device (4 30) may include a video decoder (410). The video decoder (410) encodes The input copy (407) of the video data is decoded and displayed (4 12) Render on (for example, a display screen) or another rendering device (not shown). The video decoder 410 creates a video picture output stream (411) that can be used. It may be configured to perform some or all of the various functions described in this disclosure. In a streaming system, encoded video data (404), (407), and (409 )(for example, a video bitstream) according to a specific video coding / compression standard It can be encoded. Examples of these standards include ITU-T Recommendation H.265. So, the video coding standard under development is not a multi-purpose video coding standard (VVC). It is officially known. The subject of the disclosure is used in the context of VVC and other video coding standards. It can be used.
[0036] The electronic devices (420) and (430) may include other components (not shown). Please note: For example, the electronic device (420) may include a video decoder (not shown). The electronic device (430) may also include a video encoder (not shown).
[0037] Figure 5 shows a block diagram of a video decoder (510) according to any embodiment of the present disclosure described below. The video decoder (510) can be included in the electronic device (530). ) may include a receiver (531) (e.g., a receiving circuit). Video decoder (510) This can be used instead of the video decoder (410) in the example in Figure 4.
[0038] The receiver (531) receives one or more codes to be decoded by the video decoder (510). It can receive a recorded video sequence. In the same or another embodiment, one at a time The coded video sequence can be decoded, and each coded video sequence Kens' decoding is independent of other coded video sequences. Each video A sequence can be associated with multiple video frames or video images. Coding The video sequence is received from channel (501), and channel (501) is encoded Hardware / software link to storage device that stores the video data, or code It may be a streaming source that transmits converted video data. The receiver (531) The encoded video data can be transferred to the respective processing circuits (not shown). Along with the recorded audio data and / or other data such as auxiliary data streams The receiver (531) can receive the coded video sequence from other data. It can be separated. To counteract network jitter, the buffer memory (515) is used in the receiver. (531) and the entropy decoder / parser (520) (hereafter referred to as "parser (520)") It may be placed in between. For certain applications, the buffer memory (515) may be used as the video decoder (5 10) It can be implemented as part of the buffer memory (515) for video deco It can be separated from the (510) and located externally (not shown). In other applications, for example, To combat network jitter, a buffer memory is added outside the video decoder (510). There may be a video decoder (51) (not shown) to process the playback timing, for example. There may be another buffer memory (515) inside (0). The receiver (531) has sufficient bandwidth and From a controllable memory / transfer device, or from an isosynchronous network When receiving data from the source, is the buffer memory (515) unnecessary? It can be made smaller. Best-effort packet networks such as the internet In some cases, a sufficiently large buffer memory (515) may be required for use. Therefore, its size can be relatively large. Such buffer memory is implemented with adaptive sizing. This may be done, and the video decoder (510) may have an external operating system or similar requirement. It can be implemented at least partially in its basic form (not shown).
[0039] The video decoder (510) extracts symbols (521) from the coded video sequence. A parser (520) may be included to restore them. The categories of those symbols are Vide Information used to manage the operation of the Odecoder (510), and potentially as shown in Figure 5. In some cases, it may or may not be an essential part of the electronic device (530), but the electronic device (5 30) Rendering such as a display (512) (e.g., display screen) that can be coupled Includes information for controlling the rendering device. For (one or more) rendering devices. Control information is either supplemental extended information (SEI message) or video usability information (VUI) parameter. It may be in the form of a meter set fragment (not shown). Parser (520) The coded video sequence received by ) is parsed / entropy restored It is possible. The coding of an entropy-coded video sequence is video Coding techniques or standards may be followed, including variable-length coding and Hafma. It follows various principles, including arithmetic coding with or without context dependency. It can be considered that the parser (520) is the coded video sequence Based on at least one parameter corresponding to the subgroup, the video decoder Extract a set of subgroup parameters from at least one subgroup of pixels within the image. It can be released. Subgroups include Groups of Pictures (GOP), Pictures, Tiles, and Slides. S, macroblock, coding unit (CU), block, conversion unit (TU), pre It can include measurement units (PUs), etc. The parser (520) is also coded From the video sequence, the transformation coefficients (e.g., Fourier transform coefficients) and quantization parameter values are obtained. Furthermore, information such as motion vectors can also be extracted.
[0040] The parser (520) takes from buffer memory (515) to create a symbol (521). Entropy decoding / analysis operations can be performed on the captured video sequence. Cut.
[0041] The reconstruction of symbol (521) is the typography of the coded video picture or part thereof. (Interpicture and intrapicture, interblock and intrablock) (and other factors, including multiple different processing units or functional units) This is possible. The included units and how they are included are in the parser (520). Therefore, the subgroup control information parsed from the coded video sequence This can be controlled between the parser (520) and the following multiple processing units or functional units. The flow of such subgroup control information is not illustrated for the sake of brevity.
[0042] In addition to the functional blocks already mentioned, the video decoder (510) is described below. Thus, it can be conceptually subdivided into several functional units, under commercial constraints. In actual working implementations, many of these functional units interact closely with each other. They can be integrated with each other, at least partially. However, the various subjects of disclosure To clearly explain the functions, the following disclosures employ a conceptual subdivision into functional units. do.
[0043] The first unit is the scaler / inverse unit (551). (551) is the quantization transformation coefficient, as well as information indicating which type of inverse transformation to use, block Control information including quantization size, quantization coefficients / parameters, quantization scaling matrix, etc. The parser (520) can receive (one or more) symbols (521). Scaler / The inverse conversion unit (551) is equipped with sample values that can be input to the aggregator (555). It can output an L block.
[0044] In some cases, the output samples of the scaler / inverse transform (551) are intracoded. The block, i.e., does not use predictive information from the previously reconstructed picture. The current picture can use predictive information from previously reconstructed parts. This may be related to the . Such predictive information is provided by the intrapicture prediction unit (552). It can be provided in this way. In some cases, the intrapicture prediction unit (552) , the surrounding blocks that have already been reconstructed and stored in the current picture buffer (558) Using this information, you can generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) is, for example, a partially reconstructed current picture. and / or buffer the current picture after it has been completely reconstructed. The aggregator (555) In some implementations, an intra-prediction unit (552) is generated for each sample. The prediction information is converted to the output sample information provided by the scaler / inverse transform unit (551). It can be added.
[0045] In other cases, the output sample of the scaler / inverse unit (551) is intercoded. It may be related to a block that has been modified and potentially has motion compensation. The motion compensation prediction unit (553) accesses the reference picture memory (557) and interacts with the - Samples used for picture prediction can be fetched. Related to the block. After compensating for the motion of the samples fetched according to symbol (521), these samples To generate output sample information, the aggregator (555) scales / inverts the signal. It can be added to the output of the replacement unit (551) (the output of unit 551 is the residual sample (or may be called a residual signal). A motion-compensated prediction unit (553) then extracts a predicted sample from it. The address in the reference picture memory (557) to fetch is, for example, the X component, the Y component (shift (t), and motion compensation prediction in the form of a symbol (521) which may have a reference picture component (time). Knit(553) is available and can be controlled by motion vectors. Motion compensation is Also, when the precise motion vector of a subsample is used, the reference picture memory ( 557) may also include interpolation of sample values fetched from, and motion vector prediction mechanism It may also be associated with things like this.
[0046] The output samples from the aggregator (555) vary in the loop filter unit (556). Loop filtering techniques can be applied. Video compression technology is (coding A coded video sequence (also called a coded video bitstream) contains Controlled by parameters, the loop returns symbols (521) from the parser (520). The filter unit (556) may include in-loop filtering technology. However, the (decoding order) of the coded picture or coded video sequence. (Incidentally) In addition to responding to metadata obtained during the decryption of the previous part, it also responds to previously restored and It can also respond to filtered sample values. Further details are provided below. To that end, several types of loop filters can be used in various orders as loop filter units. It may be included as part of T556.
[0047] The output of the loop filter unit (556) is output to the rendering device (512). It can also be used for future interpicture prediction, and reference picture memory (557 This could be a sample stream that can also be stored in ).
[0048] Certain coded pictures, once fully reconfigured, will be used in future InterPictures It can be used as a reference picture for prediction. For example, the current picture corresponds The coded picture is fully restored and the coded picture is referenced When identified as a picture (for example, by parser (520)), the current picture bag Fa (558) can become part of the reference picture memory (557), and unused current P The crunch buffer is reallocated before starting the restoration of the next coded picture. It is possible.
[0049] The video decoder (510) is a predetermined standard adopted by standards such as ITU-T Rec.H.265. Decryption can be performed according to the video compression technology. Coded video sequence The syntax of video compression technology or standards and the documented profiles in video compression technology In the sense of complying with both, the coded video sequence is used It can conform to the syntax specified by the video compression technology or standard. Specifically, A profile is a tool that is only available for use under that profile, for video compression. You can select a specific tool from all the tools available in the technology or standard. The complexity of the coded video sequence in order to comply with the standard is a factor in video compression. It may be within the range defined by the level of technology or standard. In some cases, the level is , maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., per second) It limits the maximum reference picture size, etc. (measured by the megasample count). The limitations set may, in some cases, be based on the Virtual Reference Decoder (HRD) specification and encoding. In the video sequence, metadata for HRD buffer management is notified by the signal. Therefore, it may be further restricted.
[0050] In some exemplary embodiments, the receiver (531) receives additional data along with the encoded video. (Redundant) data may be received. Additional data may be coded (one or more). It may be included as part of the video sequence. Additional data will properly decode the data. In order to do so, and / or to more accurately restore the original video data, a video decoder (51 0) may be used. Additional data may include, for example, time, space, or signal noise. In the form of signal-to-noise ratio (SNR) extension layers, redundant slices, redundant pictures, forward error correction codes, etc. could be.
[0051] Figure 6 shows a block diagram of a video encoder (603) according to an exemplary embodiment of the present disclosure. The video encoder (603) may be included in the electronic device (620). The electronic device (620) transmits It may further include a transmitter (640) (e.g., a transmission circuit). A video encoder (603) is shown in Figure 4. This can be used as a substitute for the video encoder (403) in the example.
[0052] The video encoder (603) is coded by the video encoder (603) A video source (601) capable of capturing (one or more) video images (in the example in Figure 6, an electronic source). A video sample may be received from (not part of the device (620)). In another example, a video sample may be received from a video system. The (601) can be implemented as part of an electronic device (620).
[0053] The video source (601) is to be coded by the video encoder (603). - Video sequence to any appropriate bit depth (e.g., 8-bit, 10-bit, 12-bit) , ...), any color space (e.g., BT.601 Y CrCb, RGB, XYZ...), and any The appropriate sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4) should be used. It can be provided in the form of a digital video sample stream. In the stem, the video source (601) can store previously prepared videos. It can be a memory device. In a video conferencing system, the video source (601) is a local image. It could be a camera that captures information as a video sequence. The video data is sequential. It may be provided as multiple individual pictures or images that give movement when viewed. The body may be organized as a spatial array of pixels, and each pixel is the sampling used. Depending on the structure, color space, etc., one or more samples may be included. This makes it easy to understand the relationship between pixels and samples. The following explanation is based on the sample. It's the focus.
[0054] According to some exemplary embodiments, the video encoder (603) is used in real time. Or, under any other time constraints required by the application, the source video sequence The kucha can be coded and compressed into a coded video sequence (643). Enforcing an appropriate coding speed constitutes one of the functions of the controller (650). In some embodiments, the controller (650) is functionally coupled to other functional units. Furthermore, it can control other functional units as described below. For brevity, combine This is not illustrated. The parameters set by the controller (650) include rate Control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization method) (etc.), picture size, Group of Pictures (GOP) layout, maximum motion vector search This may include ranges, etc. The controller (650) is optimized for a specific system design. It can be configured to have other appropriate functions related to the video encoder (603). can.
[0055] In some exemplary embodiments, the video encoder (603) is in the coding loop. It can be configured to work. As an overly simplified explanation, one example is coding. The loop is source coder (630) (for example, input picture to be coded and (1 Based on one or more reference pictures, create symbols such as symbol streams. (It is responsible for the role of) and the (local) decoder (633) which is built into the video encoder (603) ) may include. The decoder (633) is an integrated decoder 633 which has entropy Process video steam coded by source coder 630 without coding. Even if we were to do that, the symbols would be reconstructed and created by the (remote) decoder. Sample data is created in a similar manner (the video compression techniques considered in the subject matter of the disclosure are as follows): Any compression between the symbols and the coded video bitstream is lossless. (This is possible). The reconstructed sample stream (sample data) is referenced in picture memory ( It is input to 634). Decoding of the symbol stream is performed at the decoder location (local or remote). Regardless of the number, this leads to bit-accurate results, so the code in reference picture memory (634) The content is also bit-accurate between the local encoder and the remote encoder. In other words, the predictive part of the encoder is what the decoder "sees" when using the prediction during decoding. The exact same sample value will be "viewed" as the reference picture sample. The synchronization of the Kucha (and, for example, in cases where synchronization cannot be maintained due to channel errors) In some cases, this basic principle (of the resulting drift) improves coding quality. It is used for [purpose].
[0056] The operation of the "local" decoder (633) has already been described in detail above, along with Figure 5. This may be the same as the operation of a “remote” decoder such as a video decoder (510). Figure 5 also shows a simplified version. Simply referencing, however, the symbol is available, and the entropy coder (645 ) and encoding of symbols into the video sequence coded by the parser (520) / Since decoding may be reversible, the video data includes buffer memory (515) and parser (520). The entropy decoding portion of the coder (510) is connected to the local decoder (633) within the encoder. In some cases, it may not be fully implemented.
[0057] At this point, we can say that parsing / entropy decoding, which can only exist within the decoder, is excluded. Any decoder technology will also necessarily be substantially identical in the corresponding encoder. This means that it may need to exist in a functional form. For this reason, the subject of disclosure is the decoder's dynamic The focus may be on the process, and this operation is similar to the decoding part of an encoder. Therefore, The explanation of encoder technology is the reverse of the comprehensive explanation of decoder technology, so it will be omitted. This is possible. A more detailed description of the encoder is shown below, but only in specific areas or embodiments. vinegar.
[0058] During operation, in some exemplary implementations, the source coder (630) uses a "reference picture". One or more previously coded pictures from the video sequence specified as Motion-compensated predictive coding that predictively codes the input picture by referring to the character. This may be done. In this way, the coding engine (632) takes the input picture The pixel block and which can be selected as (one or more) predictive references to the input picture (1 The difference (or residual) of the color channels between one or more pixel blocks of the reference picture. The term "residual" and its adjective form "residual difference" are used interchangeably. obtain.
[0059] The local video decoder (633) is a symbol created by the source coder (630). Based on this, the coded video data of the picture that can be designated as the reference picture The data can be decoded. The operation of the coding engine (632) is advantageous, non It may be a reversible process. The coded video data is shown in Figure 6. When it can be decoded with a video decoder (not available), the restored video sequence is usually several It can be a replica of the source video sequence with slight errors. Local video decoding Da(633) is a decoding process that can be performed by a video decoder on a reference picture. It is possible to duplicate the reference picture and store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) is controlled by the remote video decoder. The reconstructed reference picture obtained has the same content as the reconstructed reference picture. A copy of the file can be stored locally (without transmission errors).
[0060] The predictor (635) can perform predictive searches for the coding engine (632). In other words, in the case of a new picture to be coded, the predictor (635) will predict the new Can serve as a suitable predictive reference for pictures (candidate reference pixel block) (As a reference) Sample data or reference picture, motion vector, block shape, etc. The reference picture memory (634) can be searched to obtain metadata. Predictor (635 ) in order to find the appropriate predictive reference, sample blocks are used for each pixel block. It can operate in this way. In some cases, the search results obtained by the predictor (635) As determined by the result, the input picture is stored in the reference picture memory (634). It can have predictive references drawn from multiple reference pictures.
[0061] The controller (650) is used, for example, to encode video data. The coding behavior of the source coder (630), including the setting of data and subgroup parameters. It allows you to manage your work.
[0062] The output of all the aforementioned functional units is entropy within the entropy coder (645). You can receive coding. Entropy Coder (645) is Huffman Coddy Reversible symbols following techniques such as coding, variable-length coding, and arithmetic coding. Compression allows the symbols generated by various functional units to be coded into a video. Convert to an Osequence.
[0063] The transmitter (640) is coded by the entropy coder (645) The video sequence is buffered and prepared for transmission over the communication channel (660). The communication channel (660) is a storage device for storing encoded video data. It may also be a hardware / software link to the video code. The transmitter (640) is a video code The coded video data from DA(603) is transmitted along with other data, for example, Coded audio data and / or auxiliary data streams (sources are shown in the diagram) (It can be merged with items that are not included.)
[0064] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) assigns a specific code to each coded picture. You can assign a picture type to each picture, and it will have a different type for each picture. This can affect the coding techniques that can be applied. For example, pictures are often Alternatively, it may be assigned as one of the following picture types.
[0065] The intra-image (I-picture) can use any other image in the sequence as a prediction source. It may be possible to encode and decode without doing so. Some video codecs, for example, For example, different types of inputs, including independent decoder refresh ("IDR") pictures. To enable LaPicture. To a person skilled in the art, I will explain the variations of these pictures and their... They understand the uses and characteristics of each.
[0066] The predictive image (P-picture) uses up to one moving image to predict the sample value of each block. Using intra-prediction or inter-prediction that uses vectors and reference indices, It may also be something that can be encoded and decoded.
[0067] The bidirectional predictive image (B-picture) uses up to 2 to predict the sample values for each block. Use intra-prediction or inter-prediction that uses two motion vectors and a reference index. They may be able to be encoded and decoded. Similarly, multiple prediction images may be single-color To reconstruct the lock, use three or more reference images and associated metadata. It is possible.
[0068] Source pictures typically consist of multiple sample coding blocks (for example, each being 4x4). The space is spatially subdivided into blocks of 8x8, 4x8, or 16x16 samples, and each block contains It can be coded. Blocks are coded with the coding applied to each picture in the block. Refer to other (already coded) blocks as determined by the assignment. And it can be coded predictively. For example, a block of I picture can be coded unpredictably. It may be possible to reference an already coded block of the same picture. Therefore, it can be coded predictively (spatial prediction or intra prediction). P picture pixels Rubrock uses a previously coded reference picture to perform spatial predictions. It may be coded predictively by or via time prediction. The block of picture B is , referencing one or two previously coded reference pictures, by spatial prediction, Alternatively, it may be coded predictively via time prediction. Source picture or intermediate processing The picture may be subdivided into other types of blocks for other purposes. Coding blocks The division of blocks and other types of blocks is the same, as will be explained in more detail below. Sometimes the method is followed, and sometimes it is not.
[0069] The video encoder (603) uses a specified video coding technique such as ITU-T Rec.H.265. Coding operations can be performed according to techniques or standards. In that operation, video The O-encoder (603) utilizes temporal and spatial redundancy in the input video sequence. It can perform various compression operations, including predictive coding operations. Therefore, Coded video data is subject to the video coding technology or standard used. Therefore, it can conform to the specified syntax.
[0070] In some exemplary embodiments, the transmitter (640) transmits additional data along with the encoded video. Data can be transmitted. Source coder (630) has coded such data. It may be included as part of the video sequence. Additional data can be added to the time / space / SNR enhancement layer. other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameters It may include set fragments, etc.
[0071] The video can be captured as multiple source pictures (video pictures) in chronological order. Intra-picture prediction (often abbreviated as intra-prediction) is a method that predicts the value of a given picture. By utilizing spatial correlations, interpicture prediction predicts time or other correlations between pictures. Use it. For example, a specific picture being encoded / decoded, called the current picture, is blocked. It can be divided into blocks. The blocks in the current picture are previously coded in the video. If it resembles a reference block within a referenced picture that is still buffered, then the motion will It can be coded by a vector called a vector. The motion vector is a reference picture. This refers to a reference block within a picture, and if multiple reference pictures are used, the reference pictures are... It may have a third dimension for identification.
[0072] In some exemplary embodiments, a dual prediction technique is used for interpicture prediction. Yes, it is possible. According to such a biprediction technique, both current pixels in the video are in the decoding order. Continue chat (however, in terms of display order, each may be in the past or future) First reference point Two reference pictures are used, such as Kucha and a second reference picture. The block is a first motion vector that points to the first reference block in the first reference picture. And, by a second motion vector pointing to a second reference block within a second reference picture It can be coded. The block is a first reference block and a second reference block. By combining different factors, they can work together to make predictions.
[0073] Furthermore, merge mode technology improves coding efficiency in interpicture prediction. It may be used for that purpose.
[0074] According to some exemplary embodiments of this disclosure, interpicture prediction and intrapicture Predictions such as channel prediction are performed in blocks. For example, video picture sequence The pictures within the file are divided into coding tree units (CTUs) for compression, and the pictures CTUs within a chat can have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU consists of three parallel coding tree blocks (CTBs), i.e., one r It may include a maCTB and two chromaCTBs. Each CTU may contain one or more coding units (C U) can be recursively partitioned into a quadtree. For example, a 64x64 pixel CTU can be recursively partitioned into a quadtree. It can be divided into one CU or four CUs of 32x32 pixels. Each of the one or more can be further divided into four CUs of 16x16 pixels. In terms of implementation methods, each CU has various prediction types such as interpretation type and intraprediction type. The CU can be analyzed during encoding to determine its encoding. The CU is temporally predictable. Depending on the potential and / or spatial predictability, it is divided into one or more prediction units (PUs). This is possible. Generally, each PU includes one chroma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed on a prediction block basis. It is executed. The splitting of CU into PU (or PB of different color channels) is implemented in various spatial patterns. This can be done. The luma PB or chroma PB may contain a matrix of sample values (e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0075] Figure 7 shows a diagram of a video encoder (703) according to another exemplary embodiment of the present disclosure. The O-encoder (703) is used within the current video picture in the video picture sequence. It receives a processing block of sample values (for example, a prediction block), and processes the block. Encoding a coded picture, which is part of a coded video sequence. It is configured to be such that the exemplary video encoder (703) is the example video encoder in Figure 4. (403) can be used instead.
[0076] For example, the video encoder (703) processes blocks such as an 8x8 sample prediction block. The system receives a matrix of sample values for the first sample. The video encoder (703) then processes, for example, rate distortion. RDO optimization is used so that the processing block is coded in the best possible way. This determines whether it is intra-mode, inter-mode, or dual-prediction mode. If it is determined that the block will be coded in intra mode, the video encoder ( 703) uses intra predictive technology to code a picture into a processing block. It is determined that the processing block is coded in intermode or bipredictive mode. If so, the video encoder (703) uses either interpretation technology or biprediction technology, respectively. Using this, processing blocks can be encoded into coded pictures. Several exemplary In this embodiment, as a submode of interpicture prediction, the motion vector is outside the predictor. Without benefiting from the coded motion vector components, one or more motion vector predictions A merge mode derived from the meter may be used. In some other exemplary embodiments, There may be motion vector components applicable to the target block. Therefore, video encoding Da (703) determines the prediction mode of the processing block, such as a mode determination module. This may include components not explicitly shown in Figure 7.
[0077] In the example in Figure 7, the video encoder (703) is connected to each other as shown in the exemplary configuration in Figure 7. The coupled interencoder (730), intraencoder (722), and residual calculator (72 3) Switch (726), residual encoder (724), general-purpose controller (721), and ent Includes a ropi encoder (725).
[0078] The interencoder (730) is a sample of the current block (e.g., a processing block). It receives and references one or more reference blocks in the reference picture (for example, a table Compare the blocks in the previous and subsequent pictures in the displayed order, and predict the interpretation information ( For example, description of redundant information, motion vectors, and merge mode information using intercoding techniques. Generate inter-prediction results based on inter-prediction information using any appropriate technique. For example, it is configured to compute the predicted block. In some examples, the reference pin The Kucha is shown as residual decoder 728 in Figure 7 (as will be explained in more detail below). The encoding is performed using the decoding unit 633 incorporated into the exemplary encoder 620 in Figure 6. This is a decrypted reference picture, decrypted based on the video information.
[0079] The intra encoder (722) is a sample of the current block (e.g., a processing block). It receives the block and compares it to an already coded block in the same picture. Generate the converted quantization coefficients and, if necessary, intra-predictive information (e.g., one or more). It is also configured to generate intra-predictive direction information using intra-coding technology. The tra encoder (722) is based on intra prediction information and reference blocks within the same picture. Therefore, the intra prediction result (for example, the predicted block) can be calculated.
[0080] The general-purpose controller (721) determines general-purpose control data and, based on the general-purpose control data, performs a B It may be configured to control other components of the decoder (703). For example, a general-purpose The controller (721) determines the prediction mode of the block and switches based on the prediction mode. Provides a control signal to (726). For example, if the prediction mode is intra mode, The controller (721) controls the switch (726) to control the residual calculator (723) Select the intra-mode result for and control the entropy encoder (725). Select intra prediction information and include that intra prediction information in the bitstream, block If the description mode of the switch is intermode, the general-purpose controller (721) switches Control (726) to select the interpretation result for use by the residual calculator (723). The entropy encoder (725) is controlled to select the interpretation information and then input Include the ter prediction information in the bitstream.
[0081] The residual calculator (723) takes the received block and the intra encoder (722) or inter - Difference between the prediction result and the result for the block selected from the encoder (730) (residual data) The residual encoder (724) can be configured to calculate the residual data. It may be configured to generate conversion coefficients. For example, the residual encoder (724) generates residual data It can be configured to convert the signal from the spatial domain to the frequency domain and generate conversion coefficients. The conversion coefficients undergo quantization to obtain quantized conversion coefficients. Various exemplary implementations In terms of form, the video encoder (703) also includes the residual decoder (728). 728) is configured to perform the inverse transform and generate the decoded residual data. The residual data is processed by the intra-encoder (722) and inter-encoder (730). It can be used in any way. For example, the interencoder (730) can decode the residual data Based on the data and interpretation information, it is possible to generate decoded blocks, The encoder (722) decodes based on the decoded residual data and intra-predictive information. It is possible to generate a decrypted block. The decrypted block generates the decrypted picture. The picture, which has been properly processed and decoded to achieve this, is then buffered into a memory circuit (not shown). It can also be used as a reference picture.
[0082] The entropy encoder (725) contains a bitstream that is encoded into blocks. It is formatted and configured to perform entropy coding. The rupie encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) uses general-purpose control data and selected predictive information. For example, intra-prediction information and inter-prediction information, residual information, and other appropriate information are included in the BIT. It can be configured to be included in the tostream. Either intermode or bipredictive mode. When coding a block in merge submode, if residual information does not exist There is.
[0083] Figure 8 shows an exemplary video decoder (810) according to another embodiment of the present disclosure. The Odecoder (810) is part of the coded video sequence. The picture is received, the coded picture is decoded and the reconstructed picture is... It is configured to generate a video. For example, the video decoder (810) generates the video in the example in Figure 4. It can be used as a substitute for decoder (410).
[0084] In the example shown in Figure 8, the video decoder (810) is configured as shown in the exemplary configuration of Figure 8. An entropy decoder (871), an inter-decoder (880), a residual decoder ( 873), a reconstruction module (874), and an intra-decoder (872) coupled thereto.
[0085] The entropy decoder (871) is configured to restore from the coded picture specific symbols representing syntax elements from which the coded picture is composed. Such symbols can, for example, be prediction information (e.g., intra-prediction information or inter-prediction information) that can identify a specific sample or metadata used for prediction by an intra-decoder (872) or an inter-decoder (880), such as the mode in which a block is coded ( e.g., intra mode, inter mode, dual prediction mode, merge sub-mode or another sub-mode), for example, the residual information in the form of quantized transform coefficients, etc. In one example, when the prediction mode is inter mode or dual prediction mode, inter-prediction information is provided to the inter-decoder (880), and when the prediction type is intra-prediction type, intra-prediction information is provided to the intra-decoder (872). The residual information can be dequantized and provided to the residual decoder (873). The inter-decoder (880) can be configured to receive inter-prediction information and generate an inter-prediction result based on the inter-prediction information. The intra-decoder (872) can be configured to receive intra-prediction information and generate a prediction result based on the intra-prediction information.
[0086]
[0087]
[0088] The residual decoder (873) performs inverse quantization to extract inverse quantization conversion coefficients, and then performs inverse quantization conversion. It can be configured to process coefficients and convert residuals from the frequency domain to the spatial domain. -da(873) also utilizes specific control information (to include the quantization parameter (QP)) In some cases, this information may be provided by the entropy decoder (871) (this (The data path is not shown because it may only contain a small amount of control information.)
[0089] The reconstruction module (874) uses the output of the residual decoder (873) in the spatial domain. The residuals and (depending on the case, the interpretation module or intrapretation module) (As output) The prediction results are combined and reconstructed as part of the reconstructed video. It can be configured to form reconstructed blocks that form part of the picture. To improve the quality of perception, other appropriate actions, such as deblocking, may be performed. Please take note of this.
[0090] Video encoders (403), (603), and (703), and video decoder (410), ( It should be noted that (510) and (810) can be implemented using any appropriate technique. In some exemplary embodiments, video encoders (403), (603), and (703), and The video decoders (410), (510), and (810) are provided using one or more integrated circuits. It can be implemented as follows. In another embodiment, video encoders (403), (603), and (603), and the video decoders (410), (510), and (810) are, Software instructions It can be implemented using one or more processors.
[0091] Returning to the intra prediction process, there are blocks (for example, luma or chroma prediction blocks, or The sample within the coding block (if it is not further divided into prediction blocks) is To generate a prediction block, use the adjacent line, the next adjacent line, or one other Alternatively, it is predicted by samples of multiple lines, or combinations thereof. Encoded The residual between the actual block and the predicted block is obtained by the transformation and subsequent quantization. It can be processed. Various intra predictive modes may be made available, intra predictive mode Parameters related to selection and other parameters are signaled within the bitstream. It may be done. Various intra prediction modes are, for example, one or for predicting a sample. Prediction samples are selected from predicting multiple line positions, or one or more lines. This may relate to direction and other special intra-predictive modes.
[0092] For example, a set of intra-predictive modes (compatiblely referred to as "intra-mode") is: It may include a predetermined number of directional intra-prediction modes. The above is described in relation to the exemplary implementation shown in Figure 1. Thus, these intra-prediction modes predict the samples to be predicted within a specific block. This allows for the selection of an outside-block sample in a predetermined number of directions. In this exemplary implementation, eight main directional modes corresponding to angles from 45 to 207 degrees with respect to the horizontal axis can be supported and predefined.
[0093] Some other implementations of intra prediction include a wider variety of sky in directional textures. To further leverage inter-mode redundancy, directional intra-mode has finer granularity. It can be further extended to an angle set. For example, the above 8-angle implementation form, as shown in FIG. 9 , may be configured to provide eight nominal angles called V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, and for each nominal angle , a predetermined number (e.g., 7) of finer angles may be added. With such an extension , more total numbers (e.g., 56 in this example) of direction angles can be used for intra prediction corresponding to the same number of predetermined direction intra modes. The prediction angle can be represented by the nominal intra angle degree + angle delta. In the above specific example with 7 finer angle directions for each nominal angle , the angle delta can be -3 to 3 times the step size of 3 degrees .
[0094] In some implementation forms, instead of or in addition to the above direction intra mode, a predetermined number of non-directional intra prediction modes may also be predefined and made available. For example, five non-directional intra modes called sum -of-pixels intra prediction modes may be specified . These non-directional intra mode prediction modes may specifically be called DC, PAETH, SMOOTH , SMOOTH_V, and SMOOTH_H intra modes. The prediction of samples of a specific block under these exemplary non- directional modes is shown in FIG. 10. As an example , FIG. 10 shows a 4×4 block predicted by samples from the upper adjacent line and / or the left adjacent line Block 1002 is shown. A specific sample 1010 within block 1002 may correspond to sample 1004 directly above sample 1010 in the upper adjacent line of block 1002, sample 1006 to the upper left of sample 1010 as the intersection of the upper adjacent line and the left adjacent line, and sample 1008 directly to the left of sample 1010 in the left adjacent line of block 1002. In an exemplary DC intra-prediction mode, the average of the left and upper adjacent samples 1008 and 1004 can be used as the predictor for sample 1010. In an exemplary PAETH intra-prediction mode, the upper, left, and upper-left reference samples 1004, 1008, and 1006 may be fetched, and then any of the values closest to (upper + left - upper left) among these three reference samples may be set as the predictor for sample 1010. In the example of the SMOOTH_V intra prediction mode, sample 1010 can be predicted by vertical quadratic interpolation of the upper-left neighbor sample 1006 and the left neighbor sample 1008. In the exemplary SMOOTH_H intra prediction mode, sample 1010 can be predicted by horizontal quadratic interpolation of the upper-left neighbor sample 1006 and the upper neighbor sample 1004. In the exemplary SMOOTH intra prediction mode, sample 1010 can be predicted by the average of the vertical and horizontal quadratic interpolations. The above implementations of the undirected intra modes are merely shown as non-limiting examples. Adjacent lines, other undirected selections of samples, and specific samples within the prediction block. Another approach is to combine prediction samples to make predictions.
[0095] In various coding levels (picture, slice, block, unit, etc.) Specific intra-predictive mode by encoder from the above-mentioned directional or non-directional mode The selection of the dot can be signaled by a bitstream. In some exemplary implementations... The eight nominal directional modes, along with five non-angle smoothing modes (a total of 13 options), may be signaled first. The signaled modes then correspond to the eight nominal angle If it is one of the degree intra modes, the selected angle delta corresponds to the signaling The index is further signaled to indicate the nominal angle. In this exemplary implementation, all intra-prediction modes are all one for signaling. In addition (for example, adding 5 non-directional modes to 56 directional modes gives 61 intrapredictive modes) (Mode generation) May be indexed.
[0096] In some exemplary implementations, the exemplary 56 or other number of directional intra-prediction modes are: Each sample in the block is projected to the reference subsample position and filtered using a 2-tap bilinear filter. This can be implemented using a unified directionality predictor that interpolates reference samples.
[0097] Some implementations capture the decaying spatial correlation with the reference on the edge. Therefore, an additional filter mode called FILTER INTRA mode can be designed. In these modes, intra predictive reference samples are used for several patches within a block. Therefore, in addition to out-of-block samples, predicted samples within the block may be used. These modes are, for example, predefined and at least luma blocks (or luma blocks only). ) can be made available for intra-prediction. A predefined number (e.g., 5) of filters The input mode can be pre-designed, each of which, for example, is a sample within a 4x2 patch. An n-tap filter (for example, 7 taps) that reflects the correlation between it and the n adjacent elements. It is represented by a set of n-tap filters. In other words, the weight coefficients of an n-tap filter are It can be position-dependent. For example, consider an 8x8 block, 4x2 patch, and 7-tap filter process, as shown in Figure 1. As shown in 1, the 8x8 block 1102 may be divided into 8 4x2 patches. These patches These are shown in Figure 11, B0, B1, B1, B3, B4, B5, B6, and B7. Each patch is shown in Figure 11, R0 to R7. Using those seven adjacent patches, it is possible to predict the samples within the current patch. Yes. Regarding patch B0, all adjacent elements may have already been reconstructed. However, in the case of other patches, some of the adjacent elements are in the current block, and therefore reconstruct It may not be built, in which case the predicted value of the nearest adjacent element will be used as the basis. For example, all adjacent elements of patch B7 as shown in Figure 11 have not been reconstructed. Predicted samples of adjacent elements from patch B7 will be used instead.
[0098] In some implementations of intra prediction, one color component may be one or more other color components. It can be predicted using this. The color components are any of the components in the YCrCb, RGB, XYZ color space, etc. This is also good. For example, a chroma from luma, or CfL) called luma component (e.g., luma group It is possible to predict chromatic components (e.g., chromatic blocks) from a quasi-sample. In some exemplary implementations, many cross-color predictions go from luminous to chroma. or not allowed. For example, a chroma sample within a chroma block is a matching reconstructed It can be modeled as a linear function of the luma sample. CfL prediction is implemented as follows: It can be done. CfL(α) = α × L AC +DC (1)
[0099] In the formula, L AC represents the AC contribution of the Luma component, α represents the parameter of the linear model, and DC is This represents the DC contribution of the chromatic component. For example, the AC component is obtained for each sample in the block. The DC component is obtained for the entire block. Specifically, the reconstructed luma sample is chromium The values can be subsampled to the desired resolution, and then the average luma value (DC of luma) is subtracted from each luma value. Thus, the AC contribution in Luma can be formed. Next, the AC contribution of Luma is used in the linear mode of equation (1) to predict the AC value of the chroma component. To approximate or predict the chroma AC component from the Luma AC contribution, the decoder needs to calculate a scaling parameter. Instead, the exemplary CfL implementation determines parameter α based on the original chroma sample. And they can be signaled within the bitstream. This allows for decoding. The complexity of the chroma component is reduced, and more accurate predictions are obtained. Regarding the DC contribution of the chroma component, In several exemplary implementations, calculations are performed using intraDC mode within chroma content. obtain.
[0100] In some exemplary implementations of the baseline, multi-line intra-prediction may be used. In these implementations, two or more reference lines are available for selection in intra-prediction. The encoder can determine the reference line used to generate the intra predictor. Determine and signal. The reference line index signals before the intra prediction mode. When a non-zero reference line index is signaled, the most likely scenario is that it can be signaled. Only high modes are permitted. Refer to Figure 15, there are four reference lines (from reference line 0 to reference line Examples up to 3 are shown, and each reference line has 6 segments along with the top left reference sample. That is, it consists of segments A to F (shown as 1502 to 1512). Furthermore, segment A F and F may be padded with the nearest samples from segments B and E, respectively.
[0101] Next, perform a transformation of the residuals in either the intra-prediction block or the inter-prediction block. Then, the conversion coefficients can be quantized. For the purpose of performing the conversion, intracode Both the encoded block and the intercoded block were, before the conversion, Multiple transformation blocks (the term "unit transformation" usually refers to a set of three color channels) It is used, but it can also be used interchangeably as a "conversion unit," for example, "Co The "Ding Unit" consists of a chroma coding block and a chroma coding block. It can be further divided into (including) coded blocks (or prediction blocks). In some implementations, coded blocks (or prediction blocks) The maximum division depth of the "encoded block" can be specified (the term "encoded block" is used in the context of coding). (Can be used interchangeably with "block"). For example, such divisions may not exceed two levels. No. The splitting of the prediction block into transformation blocks is done by separating the intra-prediction block and the inter-prediction block. The measurement block may be processed differently. However, in some implementations, Such splits can occur between intra-prediction blocks and inter-prediction blocks. .
[0102] In some exemplary implementations, for intracoding blocks, the conversion party The conversion may be performed so that all conversion blocks have the same size, and the conversion blocks are Code is performed in raster scan order. Changes to such intra coding blocks An example of modular block partitioning is shown in Figure 12. Specifically, Figure 12 shows modular block partitioning. Lock 1202 is the same as shown by 1206, via the intermediate level quadtree partition 1204 This shows that it is divided into 16 transformation blocks of block size. Example for coding. The exemplary raster scan order is indicated by the ordered arrows in Figure 12.
[0103] In some exemplary implementations, and in the case of an interconnecting block, the conversion unit The partitioning is performed recursively up to a predetermined number of levels (e.g., 2 levels) with a partitioning depth. This may also be the case. As shown in Figure 13, the partitioning can be done at any level for any subpartition. It can be recursively stopped or continued. In particular, Figure 13 shows that block 1302 is a quadtree of four It is divided into 1304 subblocks, and one of the subblocks is further divided into 4 second-level transformations While it is divided into blocks, the division of other subblocks stops after the first level, and two This example shows a total of seven transformation blocks of different sizes. An example for coding. The exemplary raster scan order is further indicated by the ordered arrows in Figure 13. Figure 13 shows an exemplary implementation of a quadtree partition with up to two levels of square transformation blocks. However, in some generation implementations, the conversion partitioning is 1:1 (square), 1:2 / The 2:1 and 1:4 / 4:1 conversion block shapes and sizes must be in the range of 4x4 to 64x64. It can support this. In some exemplary implementations, the coding block is For 64x64 or smaller, the conversion block partitioning is (in other words, chroma conversion block The block (which is the same as a coding block under those conditions) can only be applied to the luma component. However, if the width or height of the coding block is greater than 64, then Both the coding block and the chroma coding block are min(W,64) The conversion blocks are implicitly divided into multiples of ×min(H,64) and min(W,32)×min(H,32). It is possible.
[0104] In some exemplary implementations, as shown in Figure 16, coding blocks or This further illustrates another alternative exemplary method for splitting prediction blocks into transformation blocks. As shown in Figure 16, instead of using recursive transformation partitioning, coding blocks A predefined set of partition types is coded according to the conversion type of the block. It can be applied to blocks. In the specific example shown in Figure 16, one of six exemplary partition types is However, this can be applied to divide a coding block into various numbers of transformation blocks. Such a method can be applied to either coding blocks or prediction blocks. In this disclosure, the term “partition type” generally refers to a block (e.g., predictive). This can refer to the way a block (or coding block) is partitioned. "Conversion partition type", "Predictive block partition type", or "Code It can refer to "conversion partition type". For the explanation under "Coding Type," the same concept is referred to as "Coding Block Partition." This also applies to "type," and vice versa.
[0105] More specifically, the partitioning scheme in Figure 16, as shown in Figure 16, applies to any given transformation type. It offers up to six division types. In this method, all coding blocks or A transformation type may be assigned to the prediction block, for example, based on the rate distortion cost. In the example, the split type assigned to the coding block or prediction block is: This can be determined based on the transformation type of the input block or prediction block. As illustrated in Figure 16. As indicated by the four division types, a particular division type is a division of the transformation block. It can accommodate various sizes and patterns (or split types). Various conversion types and various split types The correspondence between the Ip and other factors can be predefined. An exemplary correspondence can be used for rate distortion costs. The large number of transformation types that can be assigned to coding blocks or prediction blocks based on this. The text labels are shown below.
[0106] • PARTITION_NONE: Allocates a conversion size equal to the block size.
[0107] • PARTITION_SPLIT: Converts a block to half its width and half its height. Allocate a size.
[0108] • PARTITION_HORZ: A conversion tool with the same width as the block size and half the height of the block size. Iz layer.
[0109] • PARTITION_VERT: A conversion that is half the width of the block size and the same height as the block size. Iz layer.
[0110] • PARTITION_HORZ4: A conversion tool with the same width as the block size and 1 / 4 the height of the block size. Iz layer.
[0111] • PARTITION_VERT4: A conversion tool with a width of 1 / 4 the block size and the same height as the block size. Iz layer.
[0112] In the example above, all the splitting types shown in Figure 16 apply to the split transformation blocks. Includes a uniform conversion size. This is just an example, not an exhaustive one. Several other implementations In this state, the mixed transformation block size is determined by the division in a specific division type (or pattern). It can be used for the converted block.
[0113] Subsequently, each of the transformation blocks obtained above can undergo a primary transformation. The Mali transform essentially moves the residuals within the transform block from the spatial domain to the frequency domain. In some actual implementations of primary transformations, the above example of extended coding blocks To support block partitioning, multiple conversion sizes (for each dimension of the two dimensions) are supported. (ranging from 4 points to 64 points) and transformed shape (square; width / height ratio of 2:1 / 1:2 , and rectangles with a ratio of 4:1 / 1:4 may be allowed.
[0114] Turning to the actual primary transformation, in some exemplary implementations, the 2D transformation process The hybrid transformation kernel is, for example, the next to the coding residual transformation block. This may include the use of different 1-D transformations for each element. The conversion kernel may include, but is not limited to, the following: a) 4-point, 8-point, 1 DCT-2 with 6 points, 32 points, and 64 points; b) 4 points, 8 points, and 16 points Asymmetrical DSTs (DST-4, DST-7) and their flip versions; c) 4 points, 8 points Integrals, 16-point, and 32-point equivalence transformations. The choice of transformation kernel used for each dimension is: It can be based on rate distortion (RD) criteria. For example, DCT-2 and asymmetric can be implemented. Table 1 lists the basis functions for DST's. [Table 1]
[0115] In some exemplary implementations, the hybrid transformation of a specific primary transformation implementation is used. The availability of the channel can be based on the conversion block size and prediction mode. Dependencies Examples are listed in Table 2. In the case of chroma components, the conversion type selection is performed implicitly. This is also good. In the example, for the intra-predictive residual, the transformation type is as specified in Figure 3. This can be selected according to the intra-prediction mode. For the intra-prediction residual, The conversion type for the Roma block follows the conversion type selection for the Roma block in the same location. You can choose to do so. Therefore, in the case of the chroma component, it can be converted to a bitstream. There is no signaling from the ipu. [Table 2] [Table 3]
[0116] In some implementations, a secondary transformation may be performed on the primary transformation coefficients. For example, as shown in Figure 14, in order to further decorrelate the primary transformation coefficients, reduce LFNST (Low Frequency Inseparable Transform), known as a secondary transform, is performed in the forward direction. Between primary transformation and quantization (in an encoder), and between inverse quantization and inverse primary transformation. It can be applied between the converter and (on the decoder side). Essentially, LFNST is a secondary converter. To proceed, a portion of the primary conversion coefficient, for example, the low-frequency portion (and therefore the conversion b It is possible to take a "reduced" version of the complete set of Locke's primary transformation coefficients. In the example, depending on the transformation block size, a 4x4 inseparable transformation or an 8x8 inseparable transformation is used. Possible transformations can be applied. For example, small transformation blocks (e.g., min(width, height) < 8) A 4x4 LFNST is applied to it, and for larger transformation blocks (e.g., min(width, height)>8), 8 A ×8 LFNST may be applied. For example, if an 8×8 conversion block is affected by a 4×4 LFNST. Only the low-frequency 4x4 portion of the 8x8 primary transformation coefficients undergoes a secondary transformation.
[0117] As specifically shown in Figure 14, the transformation block can be 8x8 (or 16x16). Therefore, the forward primary transform 1402 of the transform block is an 8x8 (or 16x16) primary transform. A re-transformation coefficient matrix 1404 is generated, where each square unit represents a 2×2 (or 4×4) portion. Forward LFNS The input to T does not necessarily have to be the entire set of primary transformation coefficients, for example, 8 × 8 (or 16 × 16). For example, a 4x4 (or 8x8) LFNST can be used for secondary transformation. Therefore, diagonal As shown in section (top left) 1406, the primary transformation coefficient matrix 1404 has a 4×4 (or 8×8) low frequency Only wavenumber primary transform co-effects can be used as input to LFNST. The remaining part of the re-transformation coefficient matrix does not need to undergo a secondary transformation. In this way, After the Ndari transformation, the primary transformation co-effect portion affected by LFNST is secondary This becomes the re-transformation coefficient, but the remaining part that is not affected by LFNST (for example, the shading of matrix 1404) The i-part maintains the corresponding primary transformation coefficient. In some exemplary implementations, The remaining portion that is not subject to secondary transformation may all be set to a coefficient of 0.
[0118] The following describes an example of the application of the inseparable transformation used in LFNST. An example 4x4 LFNST is used. To use this, a 4x4 input block X (for example, the shaded area of the primary transformation matrix 1404 in Figure 14) is used. (Representing the 4x4 low-frequency portion of a primary conversion coefficient block like 1406 min) as follows It can be represented.
number
[0119] This 2D input matrix is first linearized or vectorized in an exemplary order.
number
number
[0120] The inseparable transformation of 4x4 LFNST is then performed,
number
number
number
[0121] The above exemplary LFNST is based on a direct matrix multiplication method for applying a non-separable transform, and as a result, it is implemented in a single pass without multiple iterations. In some further exemplary implementations, in order to minimize the computational complexity and memory space requirements for storing the transform coefficients, the dimension of the non-separable transform matrix (T) of the 4×4 LFNST in the example can be further reduced. Such an implementation can be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (N is 4×4 = 16 in the above example, but may be equal to 64 for an 8×8 block) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) represents the dimension reduction coefficient. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows. The above exemplary LFNST is based on a direct matrix multiplication method for applying a non-separable transform, and as a result, it is implemented in a single pass without multiple iterations. In some further exemplary implementations, in order to minimize the computational complexity and memory space requirements for storing the transform coefficients, the dimension of the non-separable transform matrix (T) of the 4×4 LFNST in the example can be further reduced. Such an implementation can be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (N is 4×4 = 16 in the above example, but may be equal to 64 for an 8×8 block) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) represents the dimension reduction coefficient. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows. The above exemplary LFNST is based on a direct matrix multiplication method for applying a non-separable transform, and as a result, it is implemented in a single pass without multiple iterations. In some further exemplary implementations, in order to minimize the computational complexity and memory space requirements for storing the transform coefficients, the dimension of the non-separable transform matrix (T) of the 4×4 LFNST in the example can be further reduced. Such an implementation can be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (N is 4×4 = 16 in the above example, but may be equal to 64 for an 8×8 block) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) represents the dimension reduction coefficient. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows. The above exemplary LFNST is based on a direct matrix multiplication method for applying a non-separable transform, and as a result, it is implemented in a single pass without multiple iterations. In some further exemplary implementations, in order to minimize the computational complexity and memory space requirements for storing the transform coefficients, the dimension of the non-separable transform matrix (T) of the 4×4 LFNST in the example can be further reduced. Such an implementation can be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (N is 4×4 = 16 in the above example, but may be equal to 64 for an 8×8 block) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) represents the dimension reduction coefficient. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows. The above exemplary LFNST is based on a direct matrix multiplication method for applying a non-separable transform, and as a result, it is implemented in a single pass without multiple iterations. In some further exemplary implementations, in order to minimize the computational complexity and memory space requirements for storing the transform coefficients, the dimension of the non-separable transform matrix (T) of the 4×4 LFNST in the example can be further reduced. Such an implementation can be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (N is 4×4 = 16 in the above example, but may be equal to 64 for an 8×8 block) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) represents the dimension reduction coefficient. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows. The above exemplary LFNST is based on a direct matrix multiplication method for applying a non-separable transform, and as a result, it is implemented in a single pass without multiple iterations. In some further exemplary implementations, in order to minimize the computational complexity and memory space requirements for storing the transform coefficients, the dimension of the non-separable transform matrix (T) of the 4×4 LFNST in the example can be further reduced. Such an implementation can be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (N is 4×4 = 16 in the above example, but may be equal to 64 for an 8×8 block) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) represents the dimension reduction coefficient. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows. The above exemplary LFNST is based on a direct matrix multiplication method for applying a non-separable transform, and as a result, it is implemented in a single pass without multiple iterations. In some further exemplary implementations, in order to minimize the computational complexity and memory space requirements for storing the transform coefficients, the dimension of the non-separable transform matrix (T) of the 4×4 LFNST in the example can be further reduced. Such an implementation can be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (N is 4×4 = 16 in the above example, but may be equal to 64 for an 8×8 block) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) represents the dimension reduction coefficient. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows. The above exemplary LFNST is based on a direct matrix multiplication method for applying a non-separable transform, and as a result, it is implemented in a single pass without multiple iterations. In some further exemplary implementations, in order to minimize the computational complexity and memory space requirements for storing the transform coefficients, the dimension of the non-separable transform matrix (T) of the 4×4 LFNST in the example can be further reduced. Such an implementation can be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (N is 4×4 = 16 in the above example, but may be equal to 64 for an 8×8 block) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) represents the dimension reduction coefficient. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows. The above exemplary LFNST is based on a direct matrix multiplication method for applying a non-separable transform, and as a result, it is implemented in a single pass without multiple iterations. In some further exemplary implementations, in order to minimize the computational complexity and memory space requirements for storing the transform coefficients, the dimension of the non-separable transform matrix (T) of the 4×4 LFNST in the example can be further reduced. Such an implementation can be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (N is 4×4 = 16 in the above example, but may be equal to 64 for an 8×8 block) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) represents the dimension reduction coefficient. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows.
Equation
[0122] In Equation 7, the R rows of the transform matrix are the reduced R basis in the N -dimensional space. Therefore, the transform converts the input vector or the output vector of the reduced R dimension of the N -dimensional space. Therefore, as shown in FIG. 14, the secondary transform coefficients 1408 converted from the primary coefficients 1406 are reduced by a factor of N / R in dimension. The three squares around 1408 in FIG. 14 may be filled with zeros. In Equation 7, the R rows of the transform matrix are the reduced R basis in the N -dimensional space. Therefore, the transform converts the input vector or the output vector of the reduced R dimension of the N -dimensional space. Therefore, as shown in FIG. 14, the secondary transform coefficients 1408 converted from the primary coefficients 1406 are reduced by a factor of N / R in dimension. The three squares around 1408 in FIG. 14 may be filled with zeros. In Equation 7, the R rows of the transform matrix are the reduced R basis in the N -dimensional space. Therefore, the transform converts the input vector or the output vector of the reduced R dimension of the N -dimensional space. Therefore, as shown in FIG. 14, the secondary transform coefficients 1408 converted from the primary coefficients 1406 are reduced by a factor of N / R in dimension. The three squares around 1408 in FIG. 14 may be filled with zeros. In Equation 7, the R rows of the transform matrix are the reduced R basis in the N -dimensional space. Therefore, the transform converts the input vector or the output vector of the reduced R dimension of the N -dimensional space. Therefore, as shown in FIG. 14, the secondary transform coefficients 1408 converted from the primary coefficients 1406 are reduced by a factor of N / R in dimension. The three squares around 1408 in FIG. 14 may be filled with zeros.
[0123] The inverse transform matrix of RTS can be the transpose of its forward transform. An example is an 8x8 LFNST (where For a more diverse explanation, in the case of the above 4×4LFNST (compared to the example reduction factor 4), This can be applied, and therefore a 64x64 directly inseparable transformation matrix is, accordingly It is reduced to a 16x64 direct matrix. Furthermore, in some implementations, the input primary coefficients are... Only a portion, not the whole, may be linearized to the LFNST input vector. For example, the 8x8 example Only a portion of the input primary transformation coefficients may be linearized to the above X vector. Specific example So, of the four 4x4 quadrants of the 8x8 primary transformation coefficient matrix, the bottom right (high frequency coefficients) is It can be excluded, and only the other three quadrants use a predetermined scanning order instead of a 48x1 vector. It is then linearized into a 64x1 vector. In such an implementation, the transformation matrix is inseparable. This can be further reduced from 16x64 to 16x48.
[0124] Therefore, an exemplary reduced 48×16 inverse RST matrix is used on the decoder side, resulting in an 8×8 matrix. It is possible to generate 4x4 quadrants in the upper left, upper right, and lower left of the core (primary) transformation coefficient. Specifically, instead of the 16x64RST which has the same conversion set configuration, a further reduced 16 When the ×48RST matrix is applied, the inseparable secondary transformation is the 4×4 block in the lower right. 48 vectorized from three 4x4 quadrant blocks of an 8x8 primary coefficient block, excluding the 8x8 primary coefficient block. It takes a matrix of a certain number of elements as input. In such an implementation, the omitted 4x4 planar matrix in the lower right corner The Imari transformation coefficients are ignored in the secondary transformation. This further reduced transformation is 48 × 1 The vector is converted into a 16x1 output vector, which is a 4x4 matrix to satisfy 1408 in Figure 14. It is then scanned in reverse. The three squares of secondary transformation coefficients surrounding 1408 are zero-padding. That's fine.
[0125] With the help of such a reduction in the dimensions of RST, notes for storing all LFNST matrices. Memory usage is reduced. In the example above, for example, memory usage is reduced in an implementation without dimensionality reduction. Compared to the current state, it can be reduced from 10KB to 8KB with a reasonably small performance decrease.
[0126] In some implementations, to reduce complexity, LFNST is the target of LFNST. It can be further restricted to be applicable only if all coefficients outside the Imari transformation coefficient portion (e.g., outside the 1404-1406 portion in Figure 14) are not significant. Therefore, LFNST is applicable If used, all primary-only transformation coefficients (e.g., the primary coefficient matrix in Figure 14) are used. The part of 1404 that is not diagonally crossed can be close to 0. Such a restriction applies to the least significant digit. This allows for adjustment of LFNST index signal transmission in the position, and therefore this limitation applies. Additional coefficient scans may be required to check for significant coefficients at specific locations when none are present. To avoid this. In some implementations, the worst-case handling of LFNST (multiplying by the number of pixels) Regarding calculations, the inseparable transformations of 4x4 and 8x8 blocks can be restricted to 8x16 and 8x48 transformations, respectively. In these cases, if LFNST is applied, for other sizes less than 16, the final significant scan position must be less than 8. For blocks with shapes of 4xN, Nx4, and N>8, the above restriction means that LFNST is no longer applied only once to the top-left 4x4 region. When LFNST is applied, all primary-only coefficients are 0, so in such cases, the number of operations required for the primary transformation is reduced. From the encoder's perspective, the coefficient quantization can be simplified when the LFNST transformation is tested. Rate-Distortion Optimized Quantization (RDO) can be applied to the first 16 coefficients at most (in scan order). This can be done, and the remaining coefficients can be set to 0.
[0127] In some exemplary implementations, the available RST kernels are such that each set of translations is several It may be specified as several sets of transformations, including inseparable transformation matrices. For example, A total of four transformation sets and two inseparable transformation lines for each transformation set used in LFNST. A sequence (kernel) may exist. These kernels may be pre-trained offline. Therefore, it is data-driven. The offline-trained transform kernel performs encoding / decoding. Stored in memory or encoded or decoded by an encoding or decoding device for use during the process. It can be hardcoded. The selection of the set of transformations during the encoding or decoding process is in This can be determined by the prediction mode. Mapping from the prediction mode to the transformation set. The 'g' can be predefined. Examples of such predefined mappings are shown in Table 4. As shown in Table 4, there are three cross-component linear model (CCLM) modes (INTRA_LT_CCLM, INT One of RA_T_CCLM or INTRA_L_CCLM is the current block (i.e., 81 <= pred When used with ModeIntra (where ModeIntra <= 83), transformation set 0 is applied to the current chroma block. For each set of transformations, the selected inseparable secondary transformation candidates are: Further specification is possible through explicitly signaled LFNST indices. For example, the index is calculated once per intraCU in the bitstream after the conversion coefficient. It can be gunned. [Table 4]
[0128] In the above implementation, LFNST is such that all coefficients outside the first coefficient subgroup or part thereof are Because it is limited to being applicable only when not statistically significant, the LFNST index The coding may depend on the position of the lowest coefficient. Furthermore, the LFNST index is conte It can be quist coding, but it does not depend on intra prediction mode, and only the first bin is coded It can only be coded in text. Furthermore, LFNST is intra and inter It can be applied to intraCU in both slices, as well as to both luma and chroma. If dual trees are enabled, the LFNST indices for luma and chroma are set separately. Signaling is possible. Inter-slicing (dual tree is disabled) In some cases, a single LFNST index is signaled and used for both luma and chroma. It is used.
[0129] In some exemplary implementations, intra-subpartitioning (ISP) mode is selected. If selected, even if RST is applied to all executable partition blocks Since there is likely a limit to performance improvement, LFNST is disabled and the RST index is signaled It may not be performed. Furthermore, disabling RST for ISP prediction residuals increases the complexity of the coding. It can be reduced. In some further implementations, multiple linear regression intraprediction ( When MIP mode is selected, LFNST is also disabled and the RST index is not signaled. It's not necessary.
[0130] Due to existing maximum conversion size limitations (e.g., 64x64), large CUs exceeding 64x64 ( or any other predetermined size representing the maximum conversion block size) implicitly divides (for example, Considering that TU tiling is performed, LFNST index search is performed on a certain number of decrypted pages. Data buffering for the iPline stage can be increased fourfold. Therefore, in some implementations, the maximum size allowed for LFNST is limited to, for example, 64x64. It may be limited. In some implementations, LFNST is only valid in DCT2 as a primary transformation. It can be done.
[0131] Some other implementations use, for example, three kernels within each set, for example By defining 12 sets of secondary transformations, a secondary internal transformation (IST) can be applied to the Luma component. ) is provided. An intra-mode dependent index is used for conversion set selection. Obtain. Kernel selection within the set may be based on signaled syntactic elements. IST This applies when either DCT2 or ADST is used as both the horizontal and vertical primary transformer. It can be enabled. In some implementations, 4x4 separation according to the block size. You can choose between an impossible transformation or an 8x8 inseparable transformation. min(tx_width,tx If _height) < 8, you can select 4x4IST. For larger blocks, 8 ×8IST can be used. Here, tx_width and tx_height are converted blocks Corresponds to the width and height of the block. The input to IST is a low-frequency primary with a zigzag scanning sequence. It can also be a conversion coefficient.
[0132] Various transformations in the video coding or decoding process, for example, within the residual block The primary transformation of the sample or the secondary transformation of the block of the primary transformation coefficient process. If either conversion method is used, it should be in the 45-degree direction (for example, horizontal direction). Directional texture patterns such as edges (or directions substantially away from the vertical direction) It is not always efficient when capturing data. As mentioned above, some In the exemplary implementation, one or more separations are applied to the secondary transformation of the primary transformation coefficient. It is possible to use impossible conversion designs.
[0133] Transformation block partitioning and transformation types applied to divided transformation blocks These can be interrelated. For example, specific The conversion type depends on the specific partition type. There are cases where this is more suitable. For example, as illustrated in Figure 16, the aforementioned conversion partitioning The method, compared to the recursive divisions such as those described above in Figure 13, uses a non-recursive division type. Presenting all available conversion types, all available partition patterns For conversion blocks divided under n (for example, the conversion partition type in Figure 16) If permitted, the encoder will determine which conversion party to use to obtain the conversion block. Which type of conversion should be used, and for each partitioned conversion block... When deciding which transformation type to use, optimization is performed in a large parameter space. It will be necessary to do so. In fact, a set of conversion types is generally a specific type This conversion partition type may be more suitable than other conversion types. See below for details. In various implementations, the interaction between the conversion partition type and the conversion type is considered. Consider and use to restrict the types of conversions allowed for a particular partition type. Similarly, this leads to a system that restricts the partition types allowed for a particular conversion type. This can be achieved. Such implementation forms are particularly suitable for non-recursive transformation partitioning. In the context in which it is used, the conversion partition pattern and conversion type of each divided conversion block When deciding on a choice, it may be possible to reduce the optimization space required for the encoder.
[0134] These exemplary implementations may be used separately, in any order, or in any way. They may be combined. In the above and below discussions, "coded blocks" Terms such as "coding block" refer to the picture unit where prediction or transformation is performed. It can be used to refer to a coding block. A coding block is a coding block. It may be or may be a chroma coding block. In some situations, coding The coded block may refer to a predicted block. The term refers to the width or height of a coding block, or the maximum value of width and height, or width and Minimum height, or area size (width * height), or aspect ratio (width:height, or height) Used to refer to width.
[0135] Multiple candidate primary conversion types
[0136] In one embodiment, there may be multiple candidate primary transformation types for a block. The block may include the resulting transformation blocks from the partition. Primary transformation block The selection and / or signaling of the type is based on a predefined set of conversion partition types. It may be limited to a predefined set of conversion partition types. A larger set of partition types (for example, the partition types in Figure 16) It can be a subset of a set. In other words, a block transformation partition type If the partition belongs to a predefined set of conversion partition types, then the primary The conversion is selected and signaled. Instead, for example, the conversion partition If the type is of another type, it will not be selected and signaled, but will default. The conversion type of Ruto can be used.
[0137] In one implementation, the primary transformation type is the Discrete Cosine Transform (DCT) type 1 to DCT Type 8 Asymmetric Discrete Sine Transform (ADST); Discrete Sine Transform (DST) Type 1 to DST Type 8 At least one of the following: line graph transformation (LGT) or Carunen-Lobe transformation (KLT) It can contain one.
[0138] In one implementation, the predefined set of conversion partition types is, for example, as shown in Figure 16. It includes only PARTITION_NONE among the various conversion partition types, i.e., conversion The block size is equal to the predicted block (or coding block) size. Therefore, And only if the conversion type belongs to this predefined set, the primary conversion type This may be selected and / or signaled.
[0139] In one implementation, the number of partitions is also predefined by the conversion partition type. This can be considered when determining the set of conversion partitions. For example, a particular conversion partition type The primary transformation type is only used if the number of partitions is below a predefined threshold. A predefined set of conversion partition types that can be selected and / or signaled It is considered part of the implementation. In one implementation, the predefined threshold was an integer from 1 to 16. That's fine.
[0140] Multiple candidate secondary conversion types
[0141] In one embodiment, there may be multiple candidate secondary transformation types for a block. The block may include the resulting converted blocks from the partition. Secondary converted blocks The selection and / or signaling of the type is based on a predefined set of conversion partition types. Applicable only to [specific type]. A predefined set of conversion partition types is available. A larger set of conversion partition types (for example, the conversion partition type shown in Figure 16) It can be a subset of the full set of Ip. In other words, the transformation party of the block Only if the conversion type belongs to a predefined set of conversion partition types, A Kandari conversion is selected and signaled. Instead, for example, the conversion party If the format type is of another type, it will not be selected and signaled. The default secondary conversion type may be used, or a secondary conversion may be performed. It's not necessary.
[0142] In one implementation, secondary translation types may include KLT. KLT is used in different kernels. It may be configured as follows.
[0143] In one implementation, the predefined set of conversion partition types is, for example, as shown in Figure 16. It may only include PARTITION_NONE among various conversion partition types, i.e., The conversion block size is equal to the prediction block (or coding block) size. Therefore, only if the conversion type belongs to this predefined set, the secondary conversion type The type can be selected and / or signaled.
[0144] In one implementation, the number of partitions is also predefined by the conversion partition type. This can be considered when determining the set of conversion partitions. For example, a particular conversion partition type Secondary conversion type only if the number of partitions is below a predefined threshold. A predefined set of conversion partition types that can be selected and / or signaled It is considered part of the implementation. In one implementation, the predefined threshold was an integer from 1 to 16. That's fine.
[0145] In one implementation, the selection and / or signaling of the secondary conversion type is performed by the conversion party. This may be based on a combination of the conversion type and the primary conversion type. A predefined set of conversion partition types, and This may include the primary conversion type in a predefined set of conversion types. The secondary conversion type is when the conversion partition type is PARTITION_NONE, and the block Select and / or only if the primary conversion type used for the buck is DCT or ADST. It may need to be signaled. Otherwise, for example, the conversion partition type If it is of another type, it will be selected and signaled by default instead of being the default type. A ndari conversion type may be used, or a secondary conversion may not be performed.
[0146] Conversion-related signaling
[0147] In this disclosure, the order of transformation-related syntax elements / parameters is taken into consideration when signaling. Various signaling mechanisms aimed at improving efficiency are disclosed.
[0148] In one embodiment, the conversion partition type information is primary / secondary conversion type The selection information may be signaled before the primary / secondary conversion type. The conversion partition is a predefined set of conversion partition types, for example, PARTIT Signaling should only occur when belonging to ION_NONE. Otherwise, the conversion part If the partition does not belong to a predefined set of conversion partition types, Mali / secondary conversion type selection may not need to be signaled. The primary / secondary conversion type is introduced as a predefined default conversion type. It may be released.
[0149] In one embodiment, the primary / secondary conversion type selection information is the conversion partition. Type information may be signaled before it. In this case, the selection of the conversion partition type and The signaling may depend on the primary / secondary conversion type selection information.
[0150] In one implementation, the primary conversion type belongs to a predefined set of conversion types. In only in this case may the conversion partition type information need to be signaled. Otherwise, the conversion partition type information may not need to be signaled. For example, a predefined set of conversion types is DCT type 1 to DCT type 8, ADST, This may include, but is not limited to, DST Type 1 through DST Type 8, LGT, and KLT.
[0151] In one implementation, the secondary conversion type belongs to a predefined set of conversion types. In only certain cases may the conversion partition type information need to be signaled. For example, The predefined set of conversion types is not limited, but includes predefined KLT inputs. It may include a specific KLT that has a kernel associated with the dex. If the secondary conversion type does not belong to a predefined set of conversion types, the conversion type In some cases, the format type information does not need to be signaled. Instead, the conversion package Partition type information is a predefined default conversion part such as PARTITION_NONE. It may be derived as a type.
[0152] Figure 17 shows an exemplary method 1700 for decoding video data. Method 1700 consists of the following steps: step 1710, receiving the coded video bitstream of the data blocks; and step 1720, extracting the transformed partition type associated with the data blocks from the coded video bitstream. Step 1730, the conversion partition type, converts the data block. Pre-configuration of the conversion partition type, specifying the partition pattern for each block to be divided into blocks. In response to belonging to a subset of a defined set, coded video bits Transformation blocks separated from data blocks so that they are signaled in the stream. A step of extracting the conversion type of the conversion associated with the conversion, wherein the conversion type is the conversion type The first predefined set of types, the steps, and the conversion type, according to the conversion type, This includes the step of performing an inverse transformation on the lock.
[0153] In embodiments of this disclosure, any step and / or operation may be performed in any amount as necessary or They can be combined or arranged in sequence. Two or more steps and / or actions These can be executed in parallel.
[0154] The embodiments of this disclosure may be used separately or combined in any order. Furthermore, in each of the methods (or embodiments), the encoder and decoder are processed by a processing circuit (for example). even if implemented by one or more processors or one or more integrated circuits Good. In one example, one or more processors store data in a non-temporary computer-readable medium. The program is executed. Embodiments of this disclosure are suitable for luma blocks or chroma blocks. It may be used.
[0155] The technology described above involves physically storing data on one or more computer-readable media. It can be implemented as computer software using computer-readable instructions. For example, Figure 18 shows a computer system suitable for carrying out a particular embodiment of the disclosed subject matter. This indicates Mu (1800).
[0156] Computer software consists of one or more computer central processing units (CPUs). By using a processing unit (processing unit) and a graphics processing unit (GPU: Graphics Processing Unit), etc. This includes instructions that can be executed directly or through interpretation and microcode execution. To create code, there are assembly, compilation, linking, or similar mechanisms. It may be coded using any suitable machine code or computer language that can accept it. ru.
[0157] The instructions are for, for example, personal computers, tablet computers, servers, and smart computers. Various types of devices including smartphones, game consoles, and Internet of Things devices. It can be executed on a computer or its components.
[0158] The components shown in Figure 18 with respect to the computer system (1800) are essentially illustrative. The scope of use or functionality of the computer software implementing the embodiments of this disclosure is limited to the following: It is not intended to imply any limitations regarding this. Also, the composition of the components is... Any one of the components shown in the exemplary embodiment of the computer system (1800) or It should not be interpreted as having dependencies or requirements related to combinations.
[0159] The computer system (1800) includes a specific human interface input device. It is possible. Such human interface input devices include, for example, tactile input (keys). (strokes, swipes, data globe movements, etc.), voice input (voice, applause, etc.), visual input Through sensory input (such as gestures) and olfactory input (not shown), one or more human users It can respond to input from the user. Using a human interface device, it can respond to voice (speech, (Music, ambient sounds, etc.), images (scanned images, still images, photographic images acquired from cameras, etc.) ), video (including 2D video and 3D video), etc., human consciousness It is also possible to incorporate specific media that are not necessarily directly related to the primary input.
[0160] Input human interface devices include keyboard (1801), mouse (1802), and Rack pad (1803), touchscreen (1810), data glove (not shown), JO One of the following: Istik (1805), Microphone (1806), Scanner (1807), Camera (1808) It may include one or more (each shown individually).
[0161] The computer system (1800) also includes a specific human interface output device. This may include. Such human interface output devices include, for example, haptic output. It can stimulate the senses of one or more human users through sound, light, and smell / taste. Human interface output devices such as touch output devices (e.g., touch) By using a screen (1810), data globe (not shown), or joystick (1805) It is a haptic feedback device, but it does not function as an input device. (Chairs are also possible), audio output devices (speakers (1809), headphones (shown in the diagram) (etc.), visual output devices (CRT screens, LCD screens, plasma screens, OLED screens (regardless of whether they have touchscreen input functionality, each touchscreen Regardless of whether or not they have a visual feedback function, some of these are 2D visual output, or S It is possible to output more than three dimensions using means such as teleographic output. ) including screen (1810), virtual reality glasses (not shown), holographic display (i) and may include a smoke tank (not shown), and a printer (not shown).
[0162] Computer systems (1800) also include human-accessible memory devices and their related Optical media including CD / DVD ROM / RW(1820) which have a media such as CD / DVD. 1821), thumb drive (1822), removable hard drive or solid state drive Eve (1823), legacy magnetic media such as tapes and floppy disks (not shown), This includes dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown). It is possible.
[0163] Those skilled in the art will also know of the "computer-readable media" used in connection with the subject matter currently disclosed. Understand that the term does not include the transmission medium, carrier wave, or other transient signals. It should.
[0164] A computer system (1800) also connects to one or more communication networks (1855). It can include interfaces (1854). The network can be, for example, wireless, wired, or optical. It is possible. Networks can be local, regional, metropolitan, vehicle, and industrial. This can provide real-time and latency-tolerant capabilities. An example of a network is Ethernet. This includes local area networks such as Wi-Fi, GSM, 3G, 4G, 5G, and LTE. Television, including Ruler Network, cable television, satellite television, and terrestrial broadcast television. This includes wired or wireless wide-area digital networks, including CANbus, for vehicles and industrial use. Certain networks are typically (for example, the USB ports of a computer system (1800)) Which of the following is an external network attached to a specific general-purpose data port or peripheral bus (1849) Interface adapters are required, and other networks are typically as described below. By attaching it to the Tembus, it is integrated into the core of the computer system (1800). For example, a recovery interface to a PC computer system or a smartphone Cellular network interface to computer systems. These networks The computer system (1800) can communicate with other entities using one of the following methods. Such communications can be one-way, receive only (e.g., television broadcasting), or one-way transmit only (e.g., For example, CANbus to a specific CANbus device, or bidirectional, for example, local or wide-area This refers to communication with other computer systems using a rear digital network. The protocols and protocol stacks, as described above, are related to their networks and It can be used with each of the network interfaces.
[0165] The aforementioned human interface devices, human-accessible storage devices, and The network interface is connected to the core (1840) of the computer system (1800). It is possible.
[0166] A core (1840) consists of one or more central processing units (CPUs) (1841), graphics processing units. (GPU)(1842), Field-Programmable Gate Area (FPGA)(1843) A dedicated programmable processing unit for a specific state, a hardware accelerator for a specific task (1844 ), and may include graphics adapters (1850), etc. These devices are read ROM (1845), random access memory (1846), user access Internal mass storage devices such as built-in hard drives, SSDs, etc. (1847) cannot be stored in the system. It can be connected via Mubas (1848). In some computer systems, an additional CPU, G To enable expansion by PUs, etc., the system bus is provided in the form of one or more physical plugs. (1848) can be accessed. Peripheral devices can access the core's system bus (1848). It can be connected directly or via a peripheral bus (1849). For example, screen The (1810) can be connected to the graphics adapter (1850). The surrounding bus architecture... Technologies include PCI, USB, etc.
[0167] The CPU (1841), GPU (1842), FPGA (1843), and accelerator (1844) are combined. The computer can then execute specific instructions that can constitute the aforementioned computer code. The code can be stored in ROM (1845) or RAM (1846). The migration data is stored in RAM (1846). While this is also possible, permanent data can be stored, for example, in an internal mass storage device (1847). High-speed storage and retrieval of any memory device is performed using one or more CPUs (1841), GPUs (1842). k This can be made possible through the use of cache memory.
[0168] Computer-readable media are computer codes for performing various computer operations. It may have a medium and computer code specifically for the purposes of this disclosure. It may be designed and configured, or it may be a computer software technology for which one is skilled. It may be of a type that is well known and available to the public.
[0169] As a non-limiting example, computer systems with architecture (1800), particularly A(1840) is (one or more) processors (including CPUs, GPUs, FPGAs, accelerators, etc.). A software that is embodied in one or more tangible computer-readable media. As a result of performing A, functionality can be provided. Such computer-readable media This includes the user-accessible high-capacity storage device described above, as well as the core internal high-capacity storage device. Related to specific memory devices of the core (1840) with a non-transient nature, such as storage (1847) and ROM (1845). It may be attached to a medium. Software that implements various embodiments of this disclosure may be It can be stored in a device and executed by the core (1840). The computer-readable medium is Depending on the specific needs, it may include one or more memory devices or chips. The software is applied to the core (1840), specifically to the processors within it (CPU, GPU, and FP). Defining data structures stored in RAM (1846), including GA, and software This includes modifying such data structures according to processes defined by the client. To perform a specific process or a specific part of a specific process as described herein. It is possible. In addition, or alternatively, computer systems are described herein. To execute a specific process or a specific part of a specific process, the software A circuit that is hardwired or can work with software instead of or with other The logic embodied in this method (e.g., accelerator (1844)) provides functionality as a result. It is possible. References to software should, where appropriate, include the logic. This is possible, and vice versa. References to computer-readable media are permitted as needed. A circuit (such as an integrated circuit (IC)) that stores the software for execution. This disclosure may include circuits that embody the logic for, or both. This encompasses any suitable combination of hardware and software.
[0170] While this disclosure has described several exemplary embodiments, modifications within the scope of this disclosure are also possible. There are substitution examples and various alternative equivalents. Therefore, those skilled in the art will not explicitly state otherwise. Although not shown or explained, the principles of this disclosure are embodied and therefore It will be understood that a great many systems and methods can be devised within the realm of the mind and its scope. Note A: Acronym JEM: Collaborative Search Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High-Efficiency Video Coding SEI: Supplemental Enhancement Information VUI: Video Usability Information From GOP:Group to Pictures TU: Conversion Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Virtual Reference Decoder SNR: Signal-to-Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid crystal display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-only memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logical Device LAN: Local Area Network GSM: Global Mobile Communications System LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral component interconnection FPGA: Field-Programmable Gate Area SSD: Solid State Drive IC: Integrated Circuit HDR: High Dynamic Range SDR: Standard Dynamic Range JVET: Joint Video Exploration Team MPM: Most Probable Mode WAIP: Wide-angle intra-prediction CU: Coding Unit PU: Prediction Unit TU: Conversion Unit CTU: Coding Tree Unit PDPC: Location-dependent predictive combination ISP: Intra Subpartition SPS: Sequence Parameter Settings PPS: Picture Parameter Set APS: Adaptive Parameter Set VPS: Video Parameter Set DPS: Decryption parameter set ALF: Adaptive Loop Filter SAO: Sample Adaptive Offset CC-ALF: Cross-Component Adaptive Loop Filter CDEF: Constrained Directional Enhancement Filter CCSO: Cross-component sample offset LSO: Local Sample Offset LR: Loop Restoration Filter AV1: AOMedia Video1 AV2: AOMedia Video2 [Explanation of symbols]
[0171] 101 samples 102 Arrow 103 Arrow 104 square blocks 201 blocks 300 Communication Systems 310 Terminal device 320 Terminal devices 330 Terminal devices 350 Networks 400 Communication Systems 401 Video Source 402 Stream 403 Video Encoder 404 Video data, video bitstream 405 Streaming Server 406 Client Subsystem 407 Input copy, video data 410 Video Decoder 411 Output Stream 412 displays 413 Video Acquisition Subsystem 420 Electronic equipment 430 Electronic equipment 501 Channel 510 Video Decoder 512 Rendering equipment, displays 515 buffer memory 520 Parser 521 Symbols 530 Electronic equipment 531 Receiver 551 Reverse Conversion Unit 552 Intrapicture Prediction Units 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Picture Memory 558 Picture buffer 601 Video Sources 603 Video Coder, Video Encoder 620 Electronic equipment 630 Source Coder 632 Coding Engine 633 Local video decoder, decoding unit 634 Reference picture cache, reference picture memory 635 Predictor 640 Transmitter 643 video sequences 645 Entropy Coder 650 Controller 660 communication channels 703 Video Encoder 721 General-purpose controller 722 Intra Encoders 723 Residual Calculator 724 Residual Encoder 725 Entropy Encoder 726 switches 728 Residual Decoder 730 Interencoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Reconfiguration Module 880 Interdecoder 1002 blocks 1004 Upper sample 1006 Top left sample 1008 Left sample 1010 samples 1102 blocks 1202 coded blocks 1204 Quadtree splitting 1302 blocks 1304 Quadrutree subblock 1402 Forward Primary Transform 1404 Primary Transformation Coefficient Matrix 1406 Primary coefficient, shaded area 1408 Secondary conversion coefficient 1800 Computer System 1801 Keyboard 1802 Mouse 1803 Trackpad 1805 Joystick 1806 Mike 1807 Scanner 1808 Camera 1809 Audio Output Device Speaker 1810 Touchscreen 1821 Optical media 1822 Sam Drive 1823 Solid State Drive 1840 cores 1843 Field-Programmable Gate Area (FPGA) 1844 Hardware Accelerator 1845 Read-only memory (ROM) 1846 random access memory 1847 Core internal mass storage 1848 System Bus 1849 Local buses 1850 Graphics Adapter 1854 Interface 1855 Communication Network
Claims
1. A method for video processing, The steps include determining the conversion partition type associated with the data block, The steps include encoding a first syntax element indicating the conversion partition type into a video bitstream, The conversion partition type belongs to a subset of a predefined set of conversion partition types, each specifying a partition pattern for dividing the data block into subblocks, and in response that the number of conversion partitions associated with each conversion partition type in the subset of the predefined set of conversion partition types is less than or equal to a predefined threshold, A step of determining the transformation type of a transformation associated with a subblock separated from the data block, wherein the transformation type belongs to a first predefined set of transformation types. The steps include encoding a second syntax element indicating the conversion type into the video bitstream, In response that the conversion partition type does not belong to the subset of the predefined set of conversion partition types, A step of determining the transformation type of the transformation associated with the subblock separated from the data block as the default transformation type, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream and Methods that include...
2. The transformation is a primary transformation, and the first predefined set of transformation types is: Discrete cosine transform (DCT) type 1 to DCT type 8, Asymmetric Discrete Sine Transform (ADST), Discrete sine transform (DST) type 1 to DST type 8, Line graph transformation (LGT), and The method according to claim 1, comprising the Karunen-Löwe transformation (KLT).
3. The method according to claim 1 or 2, wherein the subset of the predefined set of conversion partition types consists of PARTITION_NONE for no conversion block partitioning.
4. The method according to any one of claims 1 to 3, wherein the predefined threshold includes an integer between 1 and 16.
5. The method according to any one of claims 1 to 4, wherein the transformation is a secondary transformation, and the first predefined set of transformation types includes KLT.
6. The method according to any one of claims 1 to 5, wherein the subset of the predefined set of conversion partition types consists of PARTITION_NONE for no conversion block partitioning.
7. The method according to any one of claims 1 to 6, wherein the number of conversion partitions associated with each conversion partition type in the subset of the predefined set of conversion partition types is less than or equal to a predefined threshold.
8. The method according to any one of claims 1 to 7, wherein the conversion type further indicates that the type of primary conversion associated with the conversion block belongs to a second predefined set of conversion types.
9. The method according to claim 8, wherein the subset of the predefined set of conversion partition types consists of PARTITION_NONE for no conversion block partitioning, and the second predefined set of conversion types consists of DCT and ADST.
10. A method for video processing, The steps include determining the transformation type associated with a subblock of a data block, The steps include encoding a first syntax element indicating the conversion type into a video bitstream, If the aforementioned conversion type belongs to a subset of a predefined set of conversion types, The steps include determining the conversion partition type associated with the data block, The steps include encoding a second syntax element indicating the conversion partition type into the video bitstream, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream, Methods that include...
11. If the conversion type does not belong to the predefined set of conversion partition types, The method of claim 10, further comprising the step of determining that the conversion partition type associated with the data block is a predefined default conversion partition type, wherein the predefined default conversion partition type includes PARTITION_NONE.
12. The aforementioned transformation is a primary transformation, and the predefined set of transformation types is: Discrete cosine transform (DCT) type 2, Asymmetric Discrete Sine Transform (ADST), DCT Type 1 to DCT Type 8, Discrete sine transform (DST) type 1 to DST type 8, Line graph transformation (LGT), or The method according to claim 10 or 11, comprising the Karunen-Löwe transformation (KLT).
13. The method according to any one of claims 10 to 12, wherein the transformation is a secondary transformation, and the predefined set of transformation types includes a KLT having a kernel associated with a predefined KLT index.
14. A method for transmitting a video bitstream generated by a method of encoding a conversion block into a video bitstream, The method for encoding the conversion block into a video bitstream is: The steps include determining the conversion partition type associated with the data block, The steps include encoding a first syntax element indicating the conversion partition type into a video bitstream, The conversion partition type belongs to a subset of a predefined set of conversion partition types, each specifying a partition pattern for dividing the data block into subblocks, and in response that the number of conversion partitions associated with each conversion partition type in the subset of the predefined set of conversion partition types is less than or equal to a predefined threshold, A step of determining the transformation type of a transformation associated with a subblock separated from the data block, wherein the transformation type belongs to a first predefined set of transformation types. The steps include encoding a second syntax element indicating the conversion type into the video bitstream, In response that the conversion partition type does not belong to the subset of the predefined set of conversion partition types, A step of determining the transformation type of the transformation associated with the subblock separated from the data block as the default transformation type, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream and Methods that include...
15. A method for storing a video bitstream generated by a method of encoding a conversion block into a video bitstream, The method for encoding the conversion block into a video bitstream is: The steps include determining the conversion partition type associated with the data block, The steps include encoding a first syntax element indicating the conversion partition type into a video bitstream, The conversion partition type belongs to a subset of a predefined set of conversion partition types, each specifying a partition pattern for dividing the data block into subblocks, and in response that the number of conversion partitions associated with each conversion partition type in the subset of the predefined set of conversion partition types is less than or equal to a predefined threshold, A step of determining the transformation type of a transformation associated with a subblock separated from the data block, wherein the transformation type belongs to a first predefined set of transformation types. The steps include encoding a second syntax element indicating the conversion type into the video bitstream, In response that the conversion partition type does not belong to the subset of the predefined set of conversion partition types, A step of determining the transformation type of the transformation associated with the subblock separated from the data block as the default transformation type, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream and Methods that include...
16. A method for transmitting a video bitstream generated by a method of encoding a conversion block into a video bitstream, The method for encoding the conversion block into a video bitstream is: The steps include determining the transformation type associated with a subblock of a data block, The steps include encoding a first syntax element indicating the conversion type into a video bitstream, If the aforementioned conversion type belongs to a subset of a predefined set of conversion types, The steps include determining the conversion partition type associated with the data block, The steps include encoding a second syntax element indicating the conversion partition type into the video bitstream, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream, Methods that include...
17. A method for storing a video bitstream generated by a method for encoding a conversion block into a video bitstream, The method for encoding the conversion block into a video bitstream is: The steps include determining the transformation type associated with a subblock of a data block, The steps include encoding a first syntax element indicating the conversion type into a video bitstream, If the aforementioned conversion type belongs to a subset of a predefined set of conversion types, The steps include determining the conversion partition type associated with the data block, The steps include encoding a second syntax element indicating the conversion partition type into the video bitstream, The steps include: performing a transformation on the subblock using the transformation type to obtain the corresponding transformation block; The steps include encoding the conversion block into the video bitstream, Methods that include...