Video coding and decoding method and device
By introducing a multi-reference row (MRL) intra prediction method based on template matching in video encoding technology, the problem of intra prediction in the prior art is solved, and more efficient encoding and decoding performance is achieved.
Patent Information
- Application Number
- CN202510357549.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-04-22
- Filing Date
- 2022-04-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing video encoding technology has problems with inefficiency in intra prediction, especially when dealing with multi-direction prediction, and the prior art is difficult to effectively utilize multiple reference areas to improve encoding efficiency.
A multi-reference row (MRL) intra prediction method based on template matching is proposed. By determining the predicted samples of the template area and the corresponding reconstructed samples from multiple reference areas, the optimal reference area is selected for intra prediction.
This method improves the encoding efficiency of intra prediction through more refined reference area selection and template matching, reduces the size of the bitstream, and improves the performance of video decoding.
Smart Images

Figure CN120223902A_ABST
Abstract
Description
Incorporation by Reference
[0001] This application claims priority to U.S. Patent Application No. 17 / 727,570, filed on April 22, 2022, entitled "Template Matching Based Intra Prediction", which claims priority to U.S. Provisional Application No. 63 / 179,891, filed on April 26, 2021, entitled "Template Matching Based Intra Prediction". The entire disclosure of the prior applications is incorporated herein by reference. Technical Field
[0002] The embodiments described in this application generally relate to video coding and decoding. Background Art
[0003] The background description provided herein is for the purpose of presenting the background of the present application. The work of the named inventors, the work described in this background section, and the content within the scope of the embodiments of this specification may not, as of the time of filing, constitute prior art and are not expressly or implicitly admitted as prior art against the present application.
[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video may include a sequence of images, each image having a certain spatial dimension, such as 1920x1080 luminance samples and associated chrominance samples. The image sequence may have a fixed or variable image rate (commonly known as the frame rate), for example, 60 images per second or 60 Hz. Uncompressed video requires a specific bit rate. For example, a 1080p60 4:2:0 (1920x1080 luminance sample resolution at 60 Hz frame rate) video with 8 bits per sample requires a bandwidth of nearly 1.5 Gbps. A one-hour-long video of this type requires more than 600 GB of storage space.
[0005] One purpose of video encoding and decoding is to reduce the redundancy of the input video signal through compression. In some cases, compression can reduce the bandwidth and / or memory requirements by at least two orders of magnitude. Lossless compression, lossy compression, or a combination thereof can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal. When lossy compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to serve the desired purpose. Lossy compression is widely adopted in the video field. The amount of allowable distortion depends on the application. For example, users of some consumer live applications can tolerate more distortion than users of television program applications. The achievable compression ratio can reflect that the greater the allowable / tolerable distortion, the higher the compression ratio that can be produced.
[0006] Video encoders and decoders can utilize several broad categories of techniques, such as including motion compensation, transformation, quantization, and entropy coding.
[0007] Video coding and decoding techniques can include techniques known as intra-frame coding. In intra-frame coding, the representation of sample values does not require reference to samples or other data in a previously reconstructed reference image. In some video codecs, an image is spatially subdivided into sample blocks. When all sample blocks are encoded in the intra-frame mode, the image can be an intra-frame image. Intra-frame images and their derivatives (such as independent decoder refresh images) can be used to reset the decoder state and can thus be used as the first image in an encoded video bitstream and a video session, or as a still image. Samples of intra-frame blocks can undergo transformation, and the transformation coefficients can be quantized before entropy coding. Intra-frame prediction can be a technique that minimizes the sample values in the pre-transform domain. In some cases, the smaller the transformed DC value, the smaller the AC coefficients, and the fewer bits required to represent the block with a given quantization step after entropy coding.
[0008] (For example, as known from encoding techniques of the MPEG-2 generation) Traditional intra-frame coding does not use intra-frame prediction. However, some newer video compression techniques include attempts, such as techniques for surrounding sample data and / or metadata, which can obtain the above-mentioned surrounding sample data and / or metadata during the encoding and / or decoding of spatially adjacent and earlier decoded-order block data. Such techniques are henceforth referred to as "intra-frame prediction" techniques. Note that in at least some cases, intra-frame prediction uses only reference data from the currently being reconstructed image (rather than a reference image).
[0009] There are many different forms of intra prediction. When more than one such technique can be used in a given video coding technology, the technique used can be coded in an intra prediction mode. In some cases, the mode can have sub - modes and / or parameters, which can be coded separately or included in the mode codeword. Which codeword is used for a given combination of mode, sub - mode, and / or parameter can affect the coding efficiency gain through intra prediction, and so can the entropy coding technique used to convert the codeword into a bitstream.
[0010] A certain intra prediction mode was introduced with H.264, refined in H.265, and further refined in new coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Adjacent sample values can be used to form a prediction block, and the adjacent sample values belong to samples that are already available. The sample values of the adjacent samples are copied into the prediction block according to a direction. The information about the direction used can be coded in the bitstream or can be predicted itself.
[0011] See Figure 1 , in the lower - right, a subset of nine known prediction directions out of 33 possible prediction directions of H.265 (corresponding to 33 angular modes of 35 intra modes) is depicted. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction of the sample being predicted. For example, arrow (102) indicates that the prediction direction of sample (101) is from one sample or multiple samples to the upper - right corner, at a 45 - degree angle to the horizontal direction. Similarly, arrow (103) indicates that the prediction direction of sample (101) is from one sample or multiple samples to the lower - left of sample (101), at a 22.5 - degree angle to the horizontal direction.
[0012] Still referring to Figure 1 , in the upper - left, a 4×4 sampled square block (104) (represented by the bold dashed line) is shown. The square block (104) includes 16 samples, and each sample is marked with 'S' for its position in the Y - dimension (e.g., row index) and its position in the X - dimension (e.g., column index). For example, sample S21 is the second sample in the Y - dimension (starting from the top) and the first sample in the X - dimension (starting from the left). Similarly, sample S44 in block (104) is the fourth sample in both the Y and X dimensions. Since the size of the block is 4×4 samples, S44 is located in the lower - right corner. Reference samples following a similar numbering scheme are also shown. The reference samples are marked with 'R' and their Y - position (e.g., row index) and X - position (column index) relative to block (104). In H.264 and H.265, the predicted samples are adjacent to the block being reconstructed; thus, negative values do not need to be used.
[0013] Intra picture prediction works by copying reference sample values from neighboring samples covered by the predicted direction indicated by the signal. For example, assume that the encoded video bitstream includes signaling indicating that the prediction direction of the block is consistent with arrow (102) - that is, from one or more prediction samples to the upper right corner, at a 45-degree angle to the horizontal plane, to predict the samples. In this case, samples S41, S32, S23, and S14 are predicted using the same reference sample R05. Then, sample S44 is predicted using reference sample R08.
[0014] In some cases, to calculate the reference sample, the values of multiple reference samples can be combined, for example, by interpolation; especially when the direction is not divisible by 45 degrees.
[0015] As video coding technology develops, the number of possible directions is increasing. In H.264 (in 2003), nine different directions can be represented. This number increased to 33 in H.265 (in 2013), and JEM / VVC / BMS can support up to 65 directions at the time of publication. Some experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent those possible directions with a small number of bits, while bearing the adverse results brought by the less likely directions. In addition, these directions themselves can sometimes be predicted from the neighboring directions used by adjacent decoded blocks.
[0016] Figure 2 A schematic diagram (201) showing the 65 intra prediction directions of JEM is presented to show the increasing number of prediction directions over time.
[0017] The mapping method of the intra prediction direction bits representing the direction in the encoded video bitstream can be different in different video coding technologies; it can cover, for example, a simple direct mapping from the prediction direction to the intra prediction mode or to the codeword, to a complex adaptive scheme involving most possible modes, and similar techniques. However, in all cases, there may be certain directions that are statistically less likely to appear in the video content compared to other directions. Since the goal of video compression is to reduce redundancy, in a well-functioning video coding technology, those less likely directions will be represented by more bits compared to the more likely directions. Summary of the Invention
[0018] Aspects of the present application provide methods and apparatuses for video coding / decoding. In some examples, a video decoding apparatus includes a receiving circuit and a processing circuit.
[0019] According to one aspect of the present application, a decoding method performed in a decoder is provided. In this method, encoded information about a coding unit (CU), a template region, and a plurality of reference regions can be received from an encoded video bitstream. The encoded information includes a first syntax element that indicates whether the CU is predicted according to a multi-reference line (MRL) intra prediction mode based on template matching. The template region can be adjacent to the CU, and the plurality of reference regions can be adjacent to the template region. In response to the first syntax element indicating that the CU is predicted according to the MRL intra prediction mode based on template matching, a plurality of cost values can be determined between (i) respective predicted samples of the template region determined based on samples in each of the plurality of reference regions and (ii) the reconstructed samples of the template region corresponding to the respective predicted samples. One reference region can be determined from the plurality of reference regions based on the plurality of cost values. Samples of the CU can be reconstructed based on samples in the determined reference region.
[0020] In some embodiments, each of the plurality of cost values is determined according to the difference between (i) respective predicted samples of the template region determined based on samples in a corresponding one of the plurality of reference regions and (ii) the reconstructed samples of the template region corresponding to the respective predicted samples.
[0021] In some embodiments, the template region can further include one or a combination of the following: (i) top samples in a row above the top edge of the CU and (ii) side samples in a column adjacent to the left edge of the CU.
[0022] In some embodiments, the template region can further include junction samples located above the side samples and adjacent to the top samples.
[0023] In some embodiments, the plurality of reference regions can include: (i) a first reference region that includes a row segment above the top edge of the template region and a column segment adjacent to the left edge of the template region; (ii) a second reference region that includes a row segment above the row segment of the first reference region and a column segment adjacent to the column segment of the first reference region; and (iii) a third reference region that includes a row segment above the row segment of the second reference region and a column segment adjacent to the column segment of the second reference region.
[0024] In some embodiments, the template region may further include one or a combination of the following: (i) top samples in the first row above the top edge of the CU and top samples in the second row above the first row, and (ii) side samples in the first column adjacent to the left edge of the CU and side samples in the second column next to the first column.
[0025] In some embodiments, the template region may further include junction samples located above the side samples in the first and second columns and to the left of the top samples in the first and second rows.
[0026] In some embodiments, the plurality of reference regions may include: a first reference region including a row segment above the top edge of the template region and a column segment adjacent to the left edge of the template region; and a second reference region including a row segment above the row segment of the first reference region and a column segment adjacent to the column segment of the first reference region.
[0027] In this method, a second syntax element may be further decoded from the encoded information, the second syntax element indicating whether to perform intra prediction on the CU according to a template-based intra mode derivation (TIMD) mode, the TIMD mode including a set of candidate intra prediction modes; in response to the first syntax element indicating to perform prediction on the CU according to the template matching-based MRL intra prediction mode and the second syntax element indicating to perform intra prediction on the CU according to the TIMD mode, determining the corresponding prediction samples of the template region based on the following information: (i) the samples in the corresponding one of the plurality of reference regions, and (ii) the corresponding candidate intra prediction mode in the set of candidate intra prediction modes; determining the plurality of cost values, each of the plurality of cost values being determined according to the sum of the absolute transform differences between the respective prediction samples of the template region and the reconstructed samples of the template region corresponding to the respective prediction samples; determining a pair of a reference region from the plurality of reference regions and an intra prediction mode from the set of candidate intra prediction modes, the pair of the reference region and the intra prediction mode being associated with the lowest cost value among the plurality of cost values; and reconstructing the samples of the CU based on the samples in the determined pair of the reference region and the intra prediction mode.
[0028] In some embodiments, in response to a second syntax element indicating that the CU is not intra-predicted according to the TIMD mode, syntax elements associated with another intra-coding mode are decoded, where the another intra-coding mode may include one of matrix-based intra prediction (MIP) and most probable mode (MPM).
[0029] According to another aspect of the present application, there is provided an apparatus. The apparatus has a processing circuit. The processing circuit is configured to execute the video encoding method of the present application.
[0030] Aspects of the present application also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to execute a method of video decoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Additional features, properties, and various advantages of the subject matter of the present application will become more apparent from the following detailed description and the accompanying drawings, in which:
[0032] Figure 1 is a schematic diagram of an exemplary subset of intra-prediction modes.
[0033] Figure 2 is a schematic diagram of exemplary intra-prediction directions.
[0034] Figure 3 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.
[0035] Figure 4 is a schematic diagram of a simplified block diagram of another communication system according to an embodiment.
[0036] Figure 5 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.
[0037] Figure 6 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.
[0038] Figure 7 shows a block diagram of an encoder according to another embodiment.
[0039] Figure 8 shows a block diagram of a decoder according to another embodiment.
[0040] Figure 9 is a schematic diagram of multi-reference row (MRL) intra-prediction according to an embodiment.
[0041] Figure 10 is a schematic diagram of template-based intra-mode derivation (TIMD) according to an embodiment.
[0042] Figure 11A Shows a first exemplary template of the template matching-based MRL of an embodiment.
[0043] Figure 11B Shows a second exemplary template of the template matching-based MRL of an embodiment.
[0044] Figure 11C Shows a third exemplary template of the template matching-based MRL of an embodiment.
[0045] Figure 11D Shows a fourth exemplary template of the template matching-based MRL of an embodiment.
[0046] Figure 12 Shows an overview flowchart of an exemplary decoding process of some embodiments of the present application.
[0047] Figure 13 Shows an overview flowchart of an exemplary encoding process of some embodiments of the present application.
[0048] Figure 14 Is a schematic diagram of a computer system of an embodiment. Detailed Description
[0049] Figure 3 Is a simplified block diagram of a communication system (300) according to an embodiment disclosed in the present application. The communication system (300) includes a plurality of terminal devices, and the terminal devices can communicate with each other through, for example, a network (350). For example, the communication system (300) includes a first terminal device (310) and a second terminal device (320) interconnected through a network (350). In Figure 3 In the embodiment, the first terminal device (310) and the second terminal device (320) perform unidirectional data transmission. For example, the first terminal device (310) can encode video data (such as a video image stream collected by the terminal device (310)) for transmission through the network (350) to another second terminal device (320). The encoded video data is transmitted in the form of one or more encoded video bitstreams. The second terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video data, and display a video image according to the recovered video data. Unidirectional data transmission is relatively common in applications such as media services.
[0050] In another embodiment, a communication system (300) includes a third terminal device (330) and a fourth terminal device (340) that perform a two-way transmission of encoded video data, which may occur, for example, during a video conference. For two-way data transmission, each of the third terminal device (330) and the fourth terminal device (340) may encode video data (such as a video image stream captured by the terminal device) for transmission over a network (350) to the other of the third terminal device (330) and the fourth terminal device (340). Each of the third terminal device (330) and the fourth terminal device (340) may also receive the encoded video data transmitted by the other of the third terminal device (330) and the fourth terminal device (340), may decode the encoded video data to recover the video data, and may display video images on an accessible display device based on the recovered video data.
[0051] In Figure 2 an embodiment, the first terminal device (310), the second terminal device (320), the third terminal device (330), and the fourth terminal device (340) may be servers, personal computers, and smart phones, but the principles disclosed in this application are not limited thereto. The embodiments disclosed in this application are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (350) represents any number of networks that transfer encoded video data between the first terminal device (310), the second terminal device (320), the third terminal device (330), and the fourth terminal device (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) may exchange data in circuit-switched and / or packet-switched channels. The network may include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this discussion of the application, unless otherwise explained below, the architecture and topology of the network (350) may be immaterial to the operation disclosed in this application.
[0052] As an example, Figure 4 shows how a video encoder and a video decoder are placed in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, and so on.
[0053] A streaming system may include an acquisition subsystem (413), which may include a video source (401) such as a digital camera, and the video source creates an uncompressed video image stream (402). In an embodiment, the video image stream (402) includes samples captured by the digital camera. Compared with the encoded video data (404) (or encoded video stream), the video image stream (402) is depicted as a thick line to emphasize the high data volume of the video image stream. The video image stream (402) may be processed by an electronic device (420), which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the video image stream (402), the encoded video data (404) (or encoded video stream (404)) is depicted as a thin line to emphasize the lower data volume of the encoded video data (404) (or encoded video stream (404)), which may be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as Figure 4 the client subsystem (406) and the client subsystem (408) in
[0054]
[0055] Figure 5 may access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and generates an output video image stream (411) that can be presented on a display (412) (such as a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (404), video data (407), and video data (409) (such as video streams) may be encoded according to certain video coding / compression standards. Embodiments of such standards include ITU-T H.265. In an embodiment, a video coding standard under development is informally referred to as Versatile Video Coding (VVC), and this application may be used in the context of the VVC standard.
[0054] It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).
[0055] Figure 5is a block diagram of a video decoder (510) according to an embodiment disclosed in the present application. The video decoder (510) may be provided in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 3 the video decoder (410) in the embodiment.
[0056] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510); in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data and other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not labeled). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and an entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) may be provided external to the video decoder (510) (not labeled). In other cases, a buffer memory (not labeled) is provided externally to the video decoder (510) to, for example, prevent network jitter, and another buffer memory (515) may be configured inside the video decoder (510) to, for example, handle the playback timing. And when the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (515), or the buffer memory may be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (515) may also be required, which may be relatively large and may have an adaptive size, and may be at least partially implemented in an operating system or a similar element (not labeled) external to the video decoder (510).
[0057] The video decoder (510) may include a parser (520) to reconstruct symbols (521) according to the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510), and potential information for controlling a display device (512) (e.g., a display screen), etc., which is not a component of the electronic device (530) but may be coupled to the electronic device (530), such as Figure 5As shown. The control information for the display device may be a parameter set segment (not labeled) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (520) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be carried out according to video coding techniques or standards and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (520) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroup may include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (420) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0058] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).
[0059] Depending on the type of the encoded video image or a part of the encoded video image (e.g., inter-frame image and intra-frame image, inter-frame block and intra-frame block) and other factors, the reconstruction of the symbols (521) may involve multiple different units. Which units are involved and the way they are involved may be controlled by the subgroup control information parsed by the parser (520) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (520) and multiple units below are not described.
[0060] In addition to the functional blocks already mentioned, the video decoder (510) may be conceptually divided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and may be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide them into the functional units below.
[0061] The first unit is a scaler / inverse transformer (551). The scaler / inverse transformer (551) receives the quantized transform coefficients as symbols (521) and control information from the parser (520), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transformer (551) can output blocks including sample values, and the sample values can be input into the aggregator (555).
[0062] In some cases, the output samples of the scaler / inverse transformer (551) can belong to intra-coded blocks; that is: blocks that do not use predictive information from a previously reconstructed image, but can use predictive information from a previously reconstructed part of the current image. Such predictive information can be provided by the intra-image predictor (552). In some cases, the intra-image predictor (552) generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed information extracted from the current image buffer (558). For example, the current image buffer (558) buffers the partially reconstructed current image and / or the fully reconstructed current image. In some cases, based on each sample, the aggregator (555) adds the predictive information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transformer (551).
[0063] In other cases, the output samples of the scaler / inverse transformer (551) can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation predictor (553) can access the reference image memory (557) to extract samples for prediction. After motion compensation of the extracted samples according to the symbol (521), these samples can be added by the aggregator (555) to the output of the scaler / inverse transformer (551) (which is called the residual sample or residual signal in this case), thereby generating output sample information. The motion compensation predictor (553) obtaining the prediction samples from the address in the reference image memory (557) can be controlled by a motion vector, and the motion vector is in the form of the symbol (521) for use by the motion compensation predictor (553), and the symbol (521) includes, for example, X, Y, and reference image components. Motion compensation can also include interpolation of the sample values extracted from the reference image memory (557) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.
[0064] The output samples of the aggregator (555) can be employed by various loop filtering techniques in the loop filter (556). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and the parameters can be available for the loop filter (556) as symbols (521) from the parser (520). However, in other embodiments, video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) portions of the encoded image or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0065] The output of the loop filter (556) can be a sample stream that can be output to the display device (512) and stored in the reference image memory (557) for subsequent inter-frame image prediction.
[0066] Once fully reconstructed, some encoded images can be used as reference images for future prediction. For example, once the encoded image corresponding to the current image is fully reconstructed and the encoded image is identified as a reference image (by, for example, the parser (520)), the current image buffer (558) can become part of the reference image memory (557), and a new current image buffer can be reallocated before starting to reconstruct subsequent encoded images.
[0067] The video decoder (510) can perform decoding operations according to, for example, predetermined video compression techniques in the ITU-T H.265 standard. The encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference image size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.
[0068] In an embodiment, a receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be part of an encoded video sequence. The additional data may be used by a video decoder (510) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.
[0069] Figure 6 is a block diagram of a video encoder (603) according to an embodiment disclosed in the present application. The video encoder (603) is disposed in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used to replace Figure 4 the video encoder (403) in the embodiment.
[0070] The video encoder (603) may receive video samples from a video source (601) (which is not Figure 5 part of the electronic device (620) in the embodiment), and the video source may capture video images to be encoded by the video encoder (603). In another embodiment, the video source (601) is part of the electronic device (620).
[0071] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 Y CrCb, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0072] According to an embodiment, the video encoder (603) may encode and compress images of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by an application. Enforcing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units as described below. For the sake of brevity, the couplings are not labeled in the figures. Parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be used for other suitable functions that relate to optimizing the video encoder (603) for a particular system design.
[0073] In some embodiments, the video encoder (603) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input image to be encoded and reference images) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder creates sample data (since in the video compression techniques contemplated in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference image memory (634). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the contents in the reference image memory (634) are also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference image samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference image synchronization principle (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is also used in some related techniques.
[0074] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder that has been described in detail above in connection with Figure 4 However, briefly referring additionally to Figure 5 , when symbols are available and the entropy encoder (645) and the parser (520) can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding part of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633).
[0075] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding existing in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, the present application focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is reciprocal to the decoder technology described comprehensively. A more detailed description is only needed in certain areas and is provided below.
[0076] During operation, in some embodiments, the source encoder (630) may perform motion compensation predictive coding. Referring to one or more previously encoded images designated as "reference images" in the video sequence, the motion compensation predictive coding performs predictive coding on the input image. In this way, the coding engine (632) encodes the difference between the pixel blocks of the input image and the pixel blocks of the reference image, and the reference image can be selected as the prediction reference for the input image.
[0077] The local video decoder (633) may decode the encoded video data of the image that can be designated as a reference image based on the symbols created by the source encoder (630). The operation of the coding engine (632) may be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 6 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that can be performed by the video decoder on the reference image and can store the reconstructed reference image in the reference image cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference image, which has the same content (without transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.
[0078] The predictor (635) may perform a prediction search for the coding engine (632). That is, for a new image to be encoded, the predictor (635) may search in the reference image memory (634) for sample data (as candidate reference pixel blocks) or some metadata, such as reference image motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new image. The predictor (635) may operate on a per-pixel block basis of the sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (635), it can be determined that the input image may have a prediction reference obtained from multiple reference images stored in the reference image memory (634).
[0079] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding the video data.
[0080] The outputs of all the above functional units can be entropy - encoded in an entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable - length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0081] The transmitter (640) can buffer one or more encoded video sequences created by the entropy encoder (645) to prepare for transmission over a communication channel (660), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) can merge the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).
[0082] The controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a certain encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding image. For example, an image can typically be assigned to any of the following image types:
[0083] An intra - frame image (I - image), which can be an image that can be encoded and decoded without using any other image in the sequence as a prediction source. Some video codecs allow different types of intra - frame images, including, for example, Independent Decoder Refresh (IDR) images. Those skilled in the art are aware of the variants of I - images and their corresponding applications and characteristics.
[0084] A predictive image (P - image), which can be an image that can be encoded and decoded using intra - frame prediction or inter - frame prediction, where the intra - frame prediction or inter - frame prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0085] A bi - directional predictive image (B - image), which can be an image that can be encoded and decoded using intra - frame prediction or inter - frame prediction, where the intra - frame prediction or inter - frame prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive images can use more than two reference images and associated metadata for reconstructing a single block.
[0086] Source images can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined according to the coding assignment of the corresponding image applied to the block. For example, blocks of an I image can be non-predictively encoded, or the blocks can be predictively encoded with reference to already encoded blocks of the same image (spatial prediction or intra-frame prediction). Pixel blocks of a P image can be predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference image. Blocks of a B image can be predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference images.
[0087] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (603) can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0088] In an embodiment, the transmitter (640) can transmit additional data when transmitting the encoded video. The source encoder (630) can include such data as part of what can be an encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0089] The captured video can be a plurality of source images (video pictures) in a time sequence. Intra-image prediction (often abbreviated to intra-frame prediction) exploits the spatial correlation within a given image, while inter-image prediction exploits the (temporal or other) correlation between images. In an embodiment, the particular image being encoded / decoded is segmented into blocks, and the particular image being encoded / decoded is referred to as the current image. When a block in the current image is similar to a reference block in a reference image that has been previously encoded and is still buffered in the video, the block in the current image can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference image, and in the case of using multiple reference images, the motion vector can have a third dimension identifying the reference image.
[0090] In some embodiments, bidirectional prediction techniques can be used for inter-frame image prediction. According to the bidirectional prediction technique, two reference images are used, such as a first reference image and a second reference image that are both before the current image in the video in decoding order (but may be past and future respectively in display order). A block in the current image can be encoded by a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. Specifically, the block can be predicted by a combination of the first reference block and the second reference block.
[0091] In addition, merge mode techniques can be used for inter-frame image prediction to improve coding efficiency.
[0092] According to some embodiments disclosed in the present application, predictions such as inter-frame image prediction and intra-frame image prediction are performed on a block-by-block basis. For example, according to the HEVC standard, an image in a video image sequence is segmented into coding tree units (CTUs) for compression, and the CTUs in the image have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be split into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an embodiment, each CU is analyzed to determine the prediction type for the CU, such as an inter-frame prediction type or an intra-frame prediction type. In addition, depending on the temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed on a prediction block basis. Taking the luminance prediction block as an example of the prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0093] Figure 7 is a diagram of a video encoder (703) according to another embodiment disclosed in the present application. The video encoder (703) is configured to receive sample values in a processing block (e.g., a prediction block) within a current video image in a video image sequence and encode the processing block into an encoded image that is part of an encoded video sequence. In this embodiment, the video encoder (703) is used to replace Figure 4The video encoder (403) in the embodiment.
[0094] In HEVC embodiments, the video encoder (703) receives a matrix of sample values for processing blocks, such as a prediction block of 8×8 samples. The video encoder (703) uses, for example, rate-distortion (RD) optimization to determine whether to use an intra mode, an inter mode, or a bi-prediction mode to encode the processing block. When encoding a processing block in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into an encoded image; and when encoding a processing block in the inter mode or the bi-prediction mode, the video encoder (703) may use inter prediction or bi-prediction techniques respectively to encode the processing block into an encoded image. In some video coding techniques, the merge mode may be an inter-picture prediction sub-mode, in which a motion vector is derived from one or more motion vector predictors without resorting to encoded motion vector components external to the predictors. In some other video coding techniques, there may be motion vector components applicable to the subject block. In an embodiment, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the processing block mode.
[0095] In Figure 7 the embodiments of, the video encoder (703) includes an inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in Figure 7
[0096] The inter-frame encoder (730) is configured to receive samples of a current block (such as a processing block), compare the block with one or more reference blocks in a reference image (such as blocks in a previous image and a subsequent image), generate inter-frame prediction information (such as redundancy information description, motion vectors, merge mode information according to inter-frame coding techniques), and calculate an inter-frame prediction result (such as a predicted block) based on the inter-frame prediction information using any suitable technique. In some embodiments, the reference image is a decoded reference image decoded based on the encoded video information.
[0097] The intra-frame encoder (722) is configured to receive samples of a current block (such as a processing block), compare the block with encoded blocks in the same image in some cases, generate quantization coefficients after transformation, and also generate intra-frame prediction information in some cases (such as intra-frame prediction direction information according to one or more intra-frame coding techniques). In an embodiment, the intra-frame encoder (722) also calculates an intra-frame prediction result (such as a predicted block) based on the intra-frame prediction information and reference blocks in the same image.
[0098] The general controller (721) is used to determine general control data and control other components of the video encoder (703) based on the general control data. In an embodiment, the general controller (721) determines the mode of a block and provides a control signal to the switch (726) based on the mode. For example, when the mode is the intra mode, the general controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select the intra prediction information and add the intra prediction information to the bitstream; and when the mode is the inter mode, the general controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select the inter prediction information and add the inter prediction information to the bitstream.
[0099] The residual calculator (723) is used to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is used to operate based on the residual data to encode the residual data to generate transform coefficients. In an embodiment, the residual encoder (724) is used to convert the residual data from the time domain to the frequency domain and generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is used to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded image, and in some embodiments, the decoded image can be buffered in a memory circuit (not shown) and used as a reference image.
[0100] The entropy encoder (725) is used to format the bitstream to produce an encoded block. The entropy encoder (725) generates various information according to a suitable standard such as the HEVC standard. In an embodiment, the entropy encoder (725) is used to obtain general control data, the selected prediction information (such as intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. It should be noted that according to the disclosed subject matter, there is no residual information when encoding a block in the merge submode of the inter mode or the bi - directional prediction mode.
[0101] Figure 8FIG. is a diagram of a video decoder (810) according to another embodiment disclosed in the present application. The video decoder (810) is configured to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed image. In an embodiment, the video decoder (810) is configured to replace Figure 4 the video decoder (410) in the embodiment.
[0102] In Figure 8 an embodiment, the video decoder (810) includes an entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) coupled together as shown in Figure 7 .
[0103] The entropy decoder (871) may be configured to reconstruct certain symbols based on the encoded image, where the symbols represent syntax elements that make up the encoded image. Such symbols may include, for example, a mode used to encode the block (e.g., an intra mode, an inter mode, a bidirectional prediction mode, a merge sub-mode of the latter two, or another sub-mode), prediction information that can separately identify certain samples or metadata for use by the intra-frame decoder (872) or the inter-frame decoder (880) for prediction (e.g., intra prediction information or inter prediction information), residual information in the form of, for example, quantized transform coefficients, and so on. In an embodiment, when the prediction mode is an inter or bidirectional prediction mode, the inter prediction information is provided to the inter-frame decoder (880); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra-frame decoder (872). The residual information may be dequantized and provided to the residual decoder (873).
[0104] The inter-frame decoder (880) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.
[0105] The intra-frame decoder (872) is configured to receive the intra prediction information and generate a prediction result based on the intra prediction information.
[0106] The residual decoder (873) is configured to perform dequantization to extract dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to obtain quantization parameter QP), and the information may be provided by the entropy decoder (871) (data paths are not labeled as this is only low-volume control information).
[0107] The reconstruction module (874) is used to merge, in the spatial domain, the residual output by the residual decoder (873) with the prediction result (which can be output by an inter-frame prediction module or an intra-frame prediction module) to form a reconstructed block, and the reconstructed block can be part of a reconstructed image, and the reconstructed image can in turn be part of a reconstructed video. It should be noted that other suitable operations such as deblocking operations can be performed to improve the visual quality.
[0108] It should be noted that any suitable technology can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In an embodiment, one or more integrated circuits can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810).
[0109] This application includes improvements to intra-frame prediction based on template matching.
[0110] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, these two standardization organizations jointly formed the JVET (Joint Video Exploration Team) to explore the possibility of developing the next video coding standard after HEVC. In April 2018, the JVET officially launched the standardization process for the next-generation video coding after HEVC. The new standard is named Versatile Video Coding (VVC), and the JVET was renamed the Joint Video Expert Team. In July 2020, the draft of H.266 / VVC version 1 was completed. In January 2021, an ad hoc group was established to study enhanced compression beyond the capabilities of VVC.
[0111] In intra mode acquisition at the decoder side, the intra mode can be obtained using relevant syntax elements signaled in the bitstream. Alternatively, the intra mode can be obtained at the decoder side without using the relevant syntax elements signaled in the bitstream. There are multiple methods for obtaining the intra mode at the decoder side, and "intra mode acquisition at the decoder side" is not limited to the exemplary methods described in this application.
[0112] Multiple reference line (MRL) intra prediction can use more reference lines (or reference regions) for intra prediction. As Figure 9 shown, an example of 4 reference lines is depicted, where the samples in segments A and F are not extracted from the reconstructed neighboring samples, but are padded (or filled) with the nearest samples in segments B and E respectively. HEVC intra prediction uses the nearest reference line (e.g., reference line 0). In MRL, two additional lines (e.g., reference line 1 and reference line 3) can be used. Thus, the samples in these two additional lines can be used for intra prediction of the block unit (902).
[0113] The index of the selected reference line (e.g., mrl_idx) can be signaled and used to generate the intra predictor for the block unit (902). For reference line indices greater than 0, only additional reference line modes can be included in the most probable mode (MPM) list, and only the MPM indices, excluding the remaining modes (e.g., intra prediction modes not included in the MPM list), can be signaled. The reference line index can be signaled before the intra prediction mode, and the planar mode and DC mode can be excluded from the intra prediction mode when a non-zero reference line index is signaled.
[0114] Template-based intra mode derivation (TIMD) can use the reference samples of the current CU as a template and select the best intra prediction mode from a set of candidate intra prediction modes associated with TIMD. As Figure 10As shown, the neighboring reconstructed samples of the current CU (1002) can be used as a template (1004). The reconstructed samples in the template (1004) can be compared with the predicted samples of the template (1004). The reference samples (1006) of the template (1004) can be used to generate predicted samples. The reference samples (1006) can be neighboring reconstructed samples around the template (1004). A cost function can be used to calculate the cost (or distortion) between the predicted samples and the reconstructed samples in the template (1004) based on a corresponding one of a set of candidate intra prediction modes. The intra prediction mode with the minimum cost (or distortion) can be selected as the best intra prediction mode for inter prediction of the current CU (1002).
[0115] Table 1 shows an exemplary encoding process associated with TIMD. As shown in Table 1, when the decoder-side intra mode derivation (DIMD) flag (e.g., DIMD_flag) is not 1 (or not true), the TIMD flag (e.g., TIMD_flag) can be signaled. When DIMD_flag is 1, the current CU / PU is using DIMD, and the ISP flag (e.g., ISP_flag) can be parsed to see if ISP is used for the current CU / PU. When DIMD_flag is not 1, the TIMD_flag is parsed. When TIMD_flag is 1, TIMD can be applied to the current CU / PU without applying other intra coding tools (e.g., ISP is not allowed when using TIMD). When TIMD_flag is not 1, other intra coding tools such as (MIP, MRL, MPM, etc.) associated syntax elements can be parsed in the decoder. Table 1: Pseudocode for TIMD Signaling
[0116] In the present application, MRL based on decoder-side template matching can be applied. MRL based on template matching can use the template to find a reference row (e.g., the best reference row) in the candidate reference rows of the current CU / PU, and this reference row can be used for intra prediction of the current CU / PU. The reference index of the current CU can be obtained by a method based on template matching instead of being signaled in the bitstream. Examples of templates can be seen, but are not limited to, Figures 11A to 11D .
[0117] In one embodiment, the template can include only one column and / or one row, as Figure 11A and Figure 11B shown. Correspondingly, the candidate reference rows can include reference rows 1, 2, and 3. For example, in Figure 11AIn, the current CU (1102) may have the following templates, which may include one or a combination of the following: (i) the top samples (1104A) in a row above the top edge of the current CU (1102), and (ii) the side samples (1104B) in a column adjacent to the left edge of the current CU (1102). The samples (1106) in reference lines 1, 2, and 3 may be applied to generate the predicted samples of the template. In Figure 11B In, the template of the current CU (1102) may further include a joint sample (1104C) located above the side sample 1104B and adjacent to the top sample 1104A.
[0118] In Figure 11C In, the template of the current CU (1102) may include one or a combination of the top samples in multiple rows and the side samples in multiple columns. For example, (i) the top sample (1104A) in the first row above the top edge of the current CU (1102) and the top sample (1104D) in the second row above the first row, and (ii) the side sample (1104B) in the first column adjacent to the left edge of the current CU (1102), and the side sample (1104E) in the second column next to the first column, one or a combination of them. Accordingly, the candidate reference lines may be reference line 2 and reference line 3. The samples (1106) in reference lines 2 and 3 may be applied to generate the predicted samples of the template. In Figure 11D In, the template may further include a joint sample (1104C), which is located above or over the side samples (1104B) and (1104E) of the first column and the second column and to the left of the top samples (1104A) and (1104D) of the first row and the second row.
[0119] For each candidate reference line, the predicted samples of the template may be generated. A cost function may be used to calculate the cost between the predicted samples of the template and the reconstructed samples of the template corresponding to the predicted samples of the template. For example, the sum of absolute differences (SAD) or the sum of absolute transformed differences (SATD) between the predicted samples of the template and the reconstructed samples of the template corresponding to the predicted samples of the template may be calculated. The reference line with the minimum cost (e.g., the minimum SAD value or the minimum SATD value) may be selected as the reference line or the best reference line of the current CU / PU. Accordingly, the samples in the reference line may be applied to perform intra prediction on the current CU / PU.
[0120] The application of the MRL based on template matching can be indicated by the MRL information based on template matching, such as the MRL flag based on template matching (e.g., MRL_flag). For example, when the value of the MRL flag based on template matching is 1, it indicates that the current CU / PU uses the MRL based on template matching. Otherwise, when the value of MRL_flag is not 1, the MRL based on template matching is not used.
[0121] In this application, the MRL based on template matching can be combined with another template-based pattern. For example, the combination of the MRL based on template matching and TIMD described in Figure 10 can be applied. An exemplary pseudocode for the combination of the MRL based on template matching and TIMD can be shown in Table 2. Table 2: Pseudocode for the combination of the MRL based on template matching and TIMD
[0122] As shown in Table 2, the TIMD flag (e.g., TIMD_flag) can be represented by a signal. When TIMD_flag is 1, it indicates that the current CU / PU is using TIMD. Further, the MRL flag based on template matching (e.g., MRL_flag) can be parsed to determine whether the MRL based on template matching is used for the current CU / PU. When TIMD_flag is not 1, the syntax elements related to other intra-frame coding tools (e.g., MIP, MPM, etc.) can be parsed in the decoder.
[0123] When both TIMD_flag and MRL_flag are equal to 1, the same template (e.g., any template shown in Figures 11A to 11D ) can be used to determine the selected intra-mode (e.g., the best intra-mode) and the selected reference line (e.g., the best reference line). Thus, a pair of selectedIntraMode (e.g., bestIntraMode) and selectedRefereneLine (e.g., bestReferenceLine) can be obtained. To determine the intra-mode and reference line to be used, a search process can be performed through a set of candidate intra-modes and candidate reference lines associated with TIMD (e.g., reference lines 1-3 in Figures 11A to 11D ). For each candidate intra-mode in the set of candidate intra-modes and each candidate reference line in the set of candidate reference lines, the SATD or SAD between the predicted samples of the template and the reconstructed samples of the template can be calculated. Thus, multiple pairs of intra-modes and reference lines can be obtained. A pair of intra-mode and reference line with the minimum cost can be selected as the intra-mode and reference line of the current CU / PU. Thereafter, the selected intra-mode and the selected reference line can be used to perform intra-prediction.
[0124] Figure 12 FIG. shows an overview flowchart of an exemplary decoding process (1200) of some embodiments of the present application. Figure 13 FIG. shows an overview flowchart of an exemplary encoding process (1300) of some embodiments of the present application. The proposed methods can be used alone or in any order combination. Further, each of the methods (or embodiments), encoders, and decoders can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium.
[0125] In each embodiment, any operations of the processes (e.g., (1200) and (1300)) can be combined or arranged in any number or order as needed. In each embodiment, two or more operations of the operations of the processes (e.g., (1200) and (1300)) can be executed in parallel.
[0126] The processes (e.g., (1200) and (1300)) can be used for reconstruction and / or encoding of blocks to generate prediction blocks for the blocks in reconstruction. In each embodiment, the processes (e.g., (1200) and (1300)) are executed by a processing circuit, such as the processing circuits in terminal devices (210), (220), (230), and (240), the processing circuit that executes the function of video encoder (303), the processing circuit that executes the function of video decoder (310), the processing circuit that executes the function of video decoder (410), and the processing circuit that executes the function of video encoder (503), etc. In some embodiments, the processes (e.g., (1200) and (1300)) are implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the processes (e.g., (1200) and (1300)).
[0127] As Figure 12 shown, the process (1200) can start from (S1201) and then proceed to (S1210). At (S1210), the encoded information about the coding unit (CU), the template region, and the multiple reference regions can be received from the encoded video bitstream. The encoded information can include a first syntax element that indicates whether the CU is predicted according to the multi-reference line (MRL) intra prediction mode based on template matching. The template region can be adjacent to the CU, and the multiple reference regions can be adjacent to the template region.
[0128] At (S1220), in response to the first syntax element indicating that the CU can be predicted according to the MRL intra prediction mode based on template matching, a plurality of cost values between (i) respective predicted samples of the template region obtained from samples in each of the plurality of reference regions and (ii) the reconstructed samples of the template region corresponding to the respective predicted samples can be determined.
[0129] At (S1230), a reference region can be determined from the plurality of reference regions based on the plurality of cost values.
[0130] At (S1240), respective samples of the CU can be reconstructed based on the samples in the determined reference region.
[0131] In some embodiments, each of the plurality of cost values can be determined according to the difference between (i) respective predicted samples of the template region determined based on samples in a corresponding one of the plurality of reference regions and (ii) the reconstructed samples of the template region corresponding to the respective predicted samples.
[0132] In some embodiments, the template region can further include one or a combination of the following: (i) top samples in a row above the top edge of the CU and (ii) side samples in a column adjacent to the left edge of the CU.
[0133] In some embodiments, the template region can further include junction samples located above the side samples and adjacent to the top samples.
[0134] In some embodiments, the plurality of reference regions can include: (i) a first reference region including a row segment above the top edge of the template region and a column segment adjacent to the left edge of the template region; (ii) a second reference region including a row segment above the row segment of the first reference region and a column segment adjacent to the column segment of the first reference region; and (iii) a third reference region including a row segment above the row segment of the second reference region and a column segment adjacent to the column segment of the second reference region.
[0135] In some embodiments, the template region can further include one or a combination of the following: (i) top samples in the first row above the top edge of the CU and (ii) side samples in the first column adjacent to the left edge of the CU. The template region can also include (i) top samples in the second row above the first row and (ii) side samples in the second column next to the first column.
[0136] In some embodiments, the template region can further include junction samples located above the side samples in the first and second columns and to the left of the top samples in the first and second rows.
[0137] In some embodiments, the plurality of reference regions may include: a first reference region including a row segment located above a top-side edge of the template region and a column segment adjacent to a left-side edge of the template region; and a second reference region including a row segment above the row segment of the first reference region and a column segment adjacent to the column segment of the first reference region.
[0138] In process (1200), a second syntax element may be further decoded from the encoded information, and the second syntax element may indicate whether to perform intra prediction on the CU according to a template-based intra mode derivation (TIMD) mode. The TIMD mode may include a set of candidate intra prediction modes. In response to the first syntax element indicating that the CU is predicted according to a template matching-based MRL intra prediction mode and the second syntax element indicating that intra prediction on the CU is performed based on the TIMD mode, the respective prediction samples of the template region may be determined based on the following information: (i) samples in each of the plurality of reference regions among the plurality of reference regions, and (ii) each candidate intra prediction mode in the set of candidate intra prediction modes. A plurality of cost values may be determined. Each of the plurality of cost values may be determined according to the sum of the absolute transform differences between the respective prediction samples of the template region and the reconstructed samples of the template region corresponding to the respective prediction samples. A pair may be determined from one of the plurality of reference regions and one intra prediction mode from the set of candidate intra prediction modes, and the pair of the reference region and the intra prediction mode is associated with the lowest cost value among the plurality of cost values. The samples of the CU may be reconstructed based on the samples in the determined pair of the reference region and the intra prediction mode.
[0139] In some embodiments, in response to the second syntax element indicating that intra prediction on the CU is not performed based on the TIMD mode, the syntax element associated with another intra coding mode may be decoded, where the another intra coding mode may include one of matrix-based intra prediction (MIP) and MPM.
[0140] As Figure 13 shown, process (1300) may start from (S1301) and then proceed to (S1310). At (S1310), a plurality of cost values may be determined between (i) the respective prediction samples of a template region of the CU obtained based on the samples in each of the plurality of reference regions of the coding unit (CU) in the image and (ii) the reconstructed samples of the template region corresponding to the respective prediction samples. The template region may be adjacent to the CU, and the plurality of reference regions may be adjacent to the template region.
[0141] For example, each of the plurality of cost values may be determined based on the difference between each predicted sample of the template region determined based on the samples in each of the plurality of reference regions and the reconstructed samples of the template region corresponding to the respective predicted samples.
[0142] At (S1320), the reference region may be determined from the plurality of reference regions based on the plurality of cost values. For example, the determined reference region may be associated with the lowest cost value among the plurality of cost values.
[0143] At (S1330), the CU may be predicted based on the determined reference region and the template region according to a multi-reference line (MRL) intra prediction mode based on template matching.
[0144] At (S1340), a first syntax element may be generated that indicates whether intra prediction of the CU is performed according to the MRL intra prediction mode based on template matching.
[0145] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 14 FIG. shows a computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter.
[0146] The computer software may be encoded using any suitable machine code or computer language that may be assembled, compiled, linked, or the like to create code that includes instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.
[0147] These instructions may be executed on various types of computers or their components, including (for example) personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0148] Figure 14 The components shown in FIG. for the computer system (1400) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. The configuration of the components should also not be construed as having any dependence on or requirement for any one component or combination thereof shown in the exemplary embodiments of the computer system (1400).
[0149] A computer system (1400) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to one or more human users through, for example, tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, taps), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices may also be used to collect certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0150] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1401), mouse (1402), touchpad (1403), touch screen (1410), data glove (not shown), joystick (1405), microphone (1406), scanner (1407), camera (1408).
[0151] The computer system (1400) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through the screen (1410), data glove (not shown), or joystick (1405), but there may also be tactile feedback devices that do not serve as input devices), audio output devices (e.g., speakers (1409), headphones (not depicted)), visual output devices (e.g., screens (1410) for CRT screens, LCD screens, plasma screens, OLED screens, each screen having or not having touch screen input capabilities, each screen having or not having tactile feedback capabilities, where some screens are capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output; virtual reality glasses (not depicted), holographic displays, and fog machines (not depicted), and printers (not depicted)).
[0152] The computer system (1400) may also include human-accessible storage devices and their associated media, such as optical media including media (1421) such as CD / DVD ROM / RW (1420) with CD / DVDs, thumb drives (1422), removable hard disk drives or solid state drives (1423), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0153] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not encompass a transmission medium, a carrier wave, or other transitory signals.
[0154] The computer system (1400) may also include an interface (1454) to one or more communication networks (1455). The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan area, vehicular, and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks (such as Ethernet), wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), television cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANBus), etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1449) (e.g., the USB port of the computer system (1400)); other networks are typically integrated into the core of the computer system (1400) by attaching to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). Using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be one-way, receive-only (e.g., broadcast TV), send-only one-way (e.g., CANbus to some CANbus devices), or two-way (e.g., to other computer systems using local or wide area digital networks). Certain protocols and protocol stacks can be used on each of the networks and network interfaces as described above.
[0155] The above-described human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the core (1440) of the computer system (1400).
[0156] The kernel (1440) may include one or more central processing units (CPUs) (1441), a graphics processing unit (GPU) (1442), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1443), a hardware accelerator for certain tasks (1444), a graphics adapter (1450), etc. These devices, together with a read-only memory (ROM) (1445), a random access memory (1446), an internal mass storage device such as an internal non-user-accessible hard disk drive, SSD, etc. (1447), may be connected via a system bus (1448). In some computer systems, the system bus (1448) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the system bus (1448) of the kernel or attached to the system bus (1448) via a peripheral bus (1449). In one example, a screen (1410) may be connected to the graphics adapter (1450). The architecture of the peripheral bus includes PCI, USB, etc.
[0157] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) may execute certain instructions, and the combination of these instructions may constitute the aforementioned computer code. The computer code may be stored in the ROM (1445) or the RAM (1446). Transitional data may also be stored in the RAM (1446), while permanent data may be stored in, for example, the internal mass storage device (1447). Fast storage and retrieval of any memory device may be enabled by using a cache memory that may be closely associated with one or more CPUs (1441), GPUs (1442), the mass storage device (1447), the ROM (1445), the RAM (1446), etc.
[0158] Computer code for performing various computer-implemented operations may be present on a computer-readable medium. The medium and the computer code may be those specially designed and constructed for the purposes of this application, or they may be of the type well-known and available to those skilled in the computer software art.
[0159] By way of example and not limitation, a computer system having an architecture (1400) and in particular a core (1440) can provide functionality as a result of one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software implemented in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as introduced above, as well as certain storage devices of the core (1440) having a non-volatile nature (such as the on-core mass storage device (1447) or ROM (1445)). The software implementing various embodiments of the present application can be stored in such devices and executed by the core (1440). Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core (1440) and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.) to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (1446) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality as a result of being logically hardwired or otherwise implemented in circuitry (e.g., accelerator (1444)), which can operate in place of or in conjunction with the software to execute specific processes or specific parts of specific processes described herein. Where appropriate, references to software can include logic and vice versa. Where appropriate, references to computer-readable media can include circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry containing logic for execution, or both. The present application encompasses any suitable combination of hardware and software. Appendix A: Acronyms JEM: joint exploration model VVC: versatile video coding BMS: benchmark set MV: motion vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOP: Groups of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communication LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: solid - state drive IC: Integrated Circuit CU: Coding Unit
[0160] Although this application has described multiple exemplary embodiments, various changes, permutations, and various alternatives of the embodiments are within the scope of this application. Therefore, it should be understood that those skilled in the art can design multiple systems and methods that, although not explicitly shown or described herein, embody the principles of this application and are thus within the spirit and scope of this application.
Claims
1. A video decoding method, characterized in that, Comprising: Receiving, from an encoded video bitstream, encoding information about a coding unit (CU), the encoding information including a first syntax element indicating prediction of the CU using a multiple reference line (MRL) intra prediction mode based on template matching; Determining at least two cost values associated with a template region adjacent to the CU, each of the at least two cost values corresponding to each of at least two reference regions, the at least two reference regions including at least one reference line; wherein, Each of the at least two cost values is determined based on a difference between: (i) respective predicted samples of the template region determined based on samples in one of the at least two reference regions, and (ii) reconstructed samples of the template region corresponding to the respective predicted samples; Determining, based on the at least two cost values, a reference region having a lowest cost value from the at least two reference regions; Predicting the CU based on the determined reference region.
2. The method according to claim 1, characterized in that, The template region includes: One or a combination of the following: (i) top samples in a row above the top edge of the CU and (ii) side samples in a column adjacent to the left edge of the CU.
3. The method according to claim 2, wherein The template region further includes: The top samples and the side samples; and Junction samples located above the side samples and adjacent to the top samples.
4. The method according to claim 3, characterized in that The multiple reference regions include: A first reference region including a row segment above the top edge of the template region and a column segment adjacent to the left edge of the template region; A second reference region including a row segment above the row segment of the first reference region and a column segment adjacent to the column segment of the first reference region; and A third reference region including a row segment above the row segment of the second reference region and a column segment adjacent to the column segment of the second reference region.
5. The method according to claim 3, characterized in that, The multiple reference regions at least include: A first reference region including a row segment above the top edge of the template region and a column segment adjacent to the left edge of the template region; A second reference region including a row segment above the row segment of the first reference region and a column segment adjacent to the column segment of the first reference region.
6. The method according to claim 3, wherein The multiple reference regions at least include: A first reference region including a row segment above the top edge of the template region and a column segment adjacent to the left edge of the template region; A third reference region including a row segment above the row segment of the second reference region and a column segment adjacent to the column segment of the second reference region.
7. The method according to claim 1, characterized in that, The template region includes: One or a combination of the following: (i) top samples in a first row above the top edge of the CU and top samples in a second row above the first row, and (ii) side samples in a first column adjacent to the left edge of the CU and side samples in a second column next to the first column.
8. The method according to claim 7, wherein In the template region further includes: The top samples in the first row and the second row; The side samples in the first column and the second column; and The joint samples located above the side samples in the first column and the second column and to the left of the top samples in the first row and the second row.
9. The method according to claim 8, wherein The plurality of reference regions include: A first reference region including a row segment above the top edge of the template region and a column segment adjacent to the left edge of the template region; and A second reference region including a row segment above the row segment of the first reference region and a column segment adjacent to the column segment of the first reference region.
10. The method according to claim 1, characterized in that The coded information further includes a second syntax element that indicates whether the CU is intra-predicted according to a template based intra mode derivation (TIMD) mode, the TIMD mode being associated with a set of candidate intra-prediction modes; Based on the second syntax element indicating that the CU is not intra-predicted according to the TIMD mode, a third syntax element associated with another intra-coding mode is decoded, the another intra-coding mode including one of matrix-based intra-prediction (MIP) and most probable mode (MPM); and Based on at least one of the second syntax element and the third syntax element, one of the TIMD mode, the MIP, and the MPM is used to reconstruct the samples of the CU.
11. The method according to claim 9, wherein Further includes: Based on the first syntax element indicating that the CU is predicted according to a multiple reference line (MRL) intra-prediction mode of template matching and the second syntax element indicating that the CU is intra-predicted according to the TIMD mode, Based on the following information to determine the corresponding prediction samples of the template region adjacent to the CU: (i) the samples in a corresponding one of at least two reference regions, and (ii) the corresponding candidate intra-prediction mode in the set of candidate intra-prediction modes; Determine the cost value associated with the template region according to the sum of the absolute transform differences between the respective prediction samples of the template region and the reconstructed samples of the template region corresponding to the respective prediction samples, each cost value in the cost values corresponding to a combination of one of the at least two reference regions and an intra-prediction mode; Determine a combination of one of the at least two reference regions and one intra-prediction mode in the set of candidate intra-prediction modes, the combination of the reference region and the intra-prediction mode being associated with the lowest cost value among the at least two cost values; And Based on the samples in the determined combination of the reference region and the intra-prediction mode to reconstruct the samples of the CU.
12. The method according to claim 1, wherein When a fourth syntax element indicates that decoder-side intra mode derivation (DIMD) is used to decode the CU, the coding information includes the first syntax element.
13. A video encoding method, characterized in that, Comprising: setting a value of a first syntax element that indicates prediction of the CU based on a multiple reference line (MRL) intra prediction mode based on template matching; determining at least two cost values associated with a template region adjacent to the CU, each of the at least two cost values corresponding to each of at least two reference regions, the at least two reference regions including at least one reference line; wherein each of the at least two cost values is determined based on a difference between: (i) respective predicted samples of the template region determined based on samples in one of the at least two reference regions, and (ii) reconstructed samples of the template region corresponding to the respective predicted samples; determining, based on the at least two cost values, a reference region with a lowest cost value from the at least two reference regions; predicting the CU based on the determined reference region.
14. A non-volatile computer-readable storage medium storing an encoded video bitstream, characterized in that, The encoded video bitstream is decoded by the method according to any one of claims 1-12, or is generated by the method according to claim 13.
15. A computer device, characterized in that, Comprising a processor and a memory, wherein computer-readable instructions are stored in the memory, and when the instructions are executed by the processor, the processor is caused to implement the method according to any one of claims 1-13.