Video coding and decoding method and device
By introducing the decoding-end intra-mode derivation (DIMD) method in video encoding technology, combined with ISP and TIMD technology, the problem of intra-prediction mode derivation in the prior art is solved, and more efficient video compression is achieved.
Patent Information
- Application Number
- CN202510540703.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-04-22
- Filing Date
- 2022-04-26
- Publication Date
- 2025-06-06
AI Technical Summary
The existing video encoding technology has problems of inefficiency in intra prediction mode derivation, especially when processing complex video content, it is difficult to effectively utilize the potential of intra mode.
A decoding-end intra-mode derivation (DIMD) method is proposed. By deriving the intra-prediction mode of the current block based on adjacent blocks at the decoding end, and combining intra-subregion division (ISP) and template-based intra-mode derivation (TIMD) technology, the selection of intra-prediction mode is optimized.
Through the DIMD method, the efficiency and accuracy of intra prediction are improved, the bit requirements during encoding and decoding are reduced, and the overall performance of video compression is improved.
Smart Images

Figure CN120111228A_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application with an application date of April 26, 2022, Chinese patent application number 202280003429.9, and invention name “Intra-frame mode derivation at the decoding end”. Technical Field
[0002] The present disclosure describes embodiments generally related to video encoding / decoding, and more particularly to a method and apparatus for video decoding. Background Art
[0003] The background description provided herein is for the purpose of presenting the content of the present disclosure in general. The extent to which the work of the inventors currently named is described in the background section and various aspects of this specification is not intended to be prior art at the time of filing this application, and it is never explicitly or implicitly admitted to be prior art of this application.
[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. An uncompressed digital video may include a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. The series of pictures may have a fixed or variable picture rate (also informally referred to as a frame rate) of, for example, 60 pictures per second or 60 Hz per second. Uncompressed video has specific bit rate requirements. For example, 1080P604:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at 60 Hz frame rate) requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires more than 600 GB of storage space.
[0005] One purpose of video encoding and decoding can be to reduce redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by two orders of magnitude or more. Lossless compression and lossy compression and their combinations can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal. When lossy compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. Taking video as an example, lossy compression is widely used. The amount of tolerable distortion depends on the application, for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression rate can reflect that higher allowable / tolerable distortion can produce higher compression rates.
[0006] Video encoders and decoders may utilize several broad categories of techniques including, for example, motion compensation, transforms, quantization, and entropy coding.
[0007] Video codec techniques may include a technique known as intra-coding. In intra-coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. When all blocks of samples are encoded in intra-mode, the picture may be an intra-picture. Intra-pictures and their derivatives (such as independent decoder refresh pictures) can be used to reset the decoder state and can therefore be used as the first picture in an encoded video stream and video session, or as a still image. The samples of the intra-block may be transformed, and the transform coefficients may be quantized prior to entropy coding. Intra-prediction may be a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficient, the fewer bits are required to represent the entropy-coded block at a given quantization step size.
[0008] For example, conventional intra-frame coding known from MPEG-2 generation coding techniques does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to use, for example, surrounding sample data and / or metadata obtained during encoding and / or decoding of spatially adjacent and preceding data blocks in decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. It should be noted that, at least in some cases, intra-frame prediction uses only reference data from the current picture being reconstructed, and not reference data from reference pictures.
[0009] Intra-frame prediction can take many different forms. When more than one such technique can be used in a given video coding technique, the technique used can be encoded in the intra-frame prediction mode. In some cases, a mode can have sub-modes and / or parameters, and these sub-modes and / or parameters can be encoded separately or included in the mode codeword. Which codeword is used for a given mode, sub-mode and / or parameter combination may have an impact on the coding efficiency gain through intra-frame prediction, and the entropy coding technique used to convert the codeword into a bitstream may also have an impact on it.
[0010] A certain intra-frame prediction mode was introduced in H.264, which was improved in H.265 and further improved in new coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC) and Benchmark Set (BMS). The prediction block can be formed using the values of neighboring samples belonging to already available samples. The sample values of neighboring samples are copied to the prediction block according to the direction. The sample values of neighboring samples are copied to the prediction block according to the direction. The reference to the direction used can be encoded in the bitstream, or it can be predicted itself.
[0011] refer to Figure 1 , depicted in the lower right is a subset of 9 known prediction directions from the 33 possible prediction directions of H.265 (corresponding to the 33 angular modes in the 35 intra-frame modes). The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction of the sample being predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples at the upper right and at a 45 degree angle to the horizontal line. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples at a 22.5 degree angle to the horizontal line at the lower left of sample (101).
[0012] Still reference Figure 1 , a square block (104) of 4×4 samples is depicted in the upper left (indicated by bold dashed lines). The square block (104) includes 16 samples, each of which is labeled with an "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in both the Y and X dimensions in the block (104). Since the size of the block is 4×4 samples, S44 is in the lower right corner. Figure 1 Reference samples are further shown in , which follow a similar numbering scheme. Reference samples are labeled with R and their Y position (e.g., row index) and X position (column index) relative to the block (104). In H.264 and H.265, the prediction samples are adjacent to the block being reconstructed, so there is no need to use negative values.
[0013] Intra-picture prediction can work by copying reference sample values from neighboring samples occupied by the signaled prediction direction. For example, assume that the encoded video bitstream includes signaling that, for this block, indicates a prediction direction consistent with arrow (102), i.e., the sample is predicted from one or more prediction samples at the top right and at a 45 degree angle to the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Then, sample S44 is predicted based on reference sample R08.
[0014] In some cases, the values of multiple reference samples may be combined, such as by interpolation, in order to calculate the reference sample, especially when the direction is not divisible by 45 degrees.
[0015] As video coding technology develops, the possible directions are increasing. In H.264 (2003), nine different directions can be represented. In H.265 (2013), this increased to 33 directions, and at the time of this disclosure, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions and use certain techniques in entropy coding to represent these possible directions with a small number of bits, accepting a certain cost for less likely directions. In addition, the direction itself can sometimes be predicted based on the adjacent directions used in adjacent, already decoded adjacent blocks.
[0016] Figure 2 A schematic diagram (201) of 65 intra prediction directions according to JEM is shown to illustrate the increasing number of prediction directions over time.
[0017] The mapping of the intra prediction direction bits representing the direction in the coded video bitstream can vary from one video coding technique to another, e.g. it can range from a simple direct mapping of the prediction direction to intra prediction modes to codewords to complex adaptive schemes involving the most probable mode and similar techniques. However, in all cases there may be some directions that are less likely to occur in the video content statistics than some other directions. Since the goal of video compression is to reduce redundancy, in a well-functioning video coding technique those less likely directions will be represented by more bits than more probable directions. Summary of the invention
[0018] Aspects of the present disclosure provide methods and apparatus for video encoding / decoding.In some examples, an apparatus for video decoding includes a receiving circuit and a processing circuit.
[0019] The present application provides a method for video decoding, which is performed in a video decoder. Coding information of a current block and neighboring blocks of the current block can be received from an encoded video code stream. First information associated with the current block can be obtained in the coding information. The first information can indicate whether the current block is intra-predicted based on decoderside intra mode derivation (DIMD), in which the intra prediction mode of the current block is derived based on neighboring blocks. Second information associated with the current block can be obtained in the coding information. The second information can indicate whether the current block is divided based on an intra sub-partition (ISP) mode. A context model index can be determined in response to (i) first information indicating that the current block is intra-predicted based on DIMD, or (ii) second information indicating that an upper neighboring block or a left neighboring block of the current block in the neighboring blocks is divided based on an ISP mode. The current block can be decoded from the encoded video code stream based on at least the context model index.
[0020] In some embodiments, the second information may be decoded in response to first information indicating that the current block is intra-predicted based on DIMD. In addition, the syntax element associated with the ISP mode may be decoded in response to second information indicating that the current block is divided based on the ISP mode.
[0021] In the method, determining the context model in response to (i) first information or (ii) second information is performed in response to first information indicating that the current block is intra-predicted based on DIMD. Third information associated with the current block can be obtained in the encoding information in response to the first information indicating that the current block is not intra-predicted based on DIMD. The third information can indicate whether the current block is intra-predicted based on a template based intra mode derivation (TIMD) including a set of candidate intra prediction modes. The context model index can be determined in response to (i) third information indicating that the current block is intra-predicted based on the TIMD, or (ii) second information indicating that the upper neighboring block or the left neighboring block of the current block in the neighboring blocks is divided based on the ISP mode.
[0022] In some embodiments, the second information may be decoded in response to third information indicating that the current block is intra-predicted based on TIMD. Accordingly, the syntax element associated with the ISP mode may be decoded in response to second information indicating that the current block is divided based on the ISP mode.
[0023] In the method, in response to second information indicating that the current block is not divided based on the ISP mode, a syntax element associated with another intra-frame coding mode can be decoded. The other intra-frame coding mode includes one of matrix-based intra prediction (MIP), multiple reference line intra prediction (MRL) and most probable mode (MPM).
[0024] In some embodiments, in response to first information indicating that the current block is intra-predicted based on DIMD, the context model index may be determined to be 1. In response to first information indicating that the current block is not intra-predicted based on DIMD, the context model index may be determined to be 0.
[0025] In some embodiments, the context model index may be used in Context-based Adaptive Binary Arithmetic Coding (CABAC).
[0026] In some embodiments, in response to third information indicating that the current block is intra-predicted based on TIMD, the context model index may be determined to be 1. In response to third information indicating that the current block is not intra-predicted based on TIMD, the context model index may be determined to be 0.
[0027] According to another aspect of the present disclosure, a method for video decoding is provided, which is performed in a video decoder. In the method, encoding information of a current block and adjacent blocks of the current block can be received from an encoded video code stream. The first information associated with the current block in the encoding information can be decoded. The first information can indicate whether the current block is intra-predicted based on a decoding-end intra-frame mode derivation (DIMD), in which the intra-frame prediction mode of the current block is derived based on adjacent blocks. In response to the first information indicating that the current block is intra-predicted based on DIMD, a first intra-frame prediction mode can be determined based on DIMD. In addition, a second intra-frame prediction mode can be determined based on a syntax element included in the encoding information and associated with a most probable mode (MPM) and an MPM remainder.
[0028] In the method, second information associated with the current block in the encoding information may be decoded in response to first information indicating that the current block is intra-predicted based on DIMD. The second information may indicate whether the current block is divided based on an intra sub-region partitioning (ISP) mode. In response to the second information indicating that the current block is divided based on the ISP mode, a syntax element associated with the ISP mode may be decoded.
[0029] In the method, in response to second information indicating that the current block is not split based on the ISP mode, a syntax element associated with another intra-frame coding mode may be decoded. The other intra-frame coding mode may include one of matrix-based intra-frame prediction (MIP), multiple reference line intra-frame prediction (MRL), and most probable mode (MPM).
[0030] In the method, a first intra-frame predictor may be determined based on a first intra-frame prediction mode, a second intra-frame predictor may be determined based on a second intra-frame prediction mode, and a final intra-frame predictor may be determined based on the first intra-frame predictor and the second intra-frame predictor.
[0031] In some embodiments, the final intra-frame predictor may be determined to be equal to the sum of: (i) a product of a first weight and a first intra-frame predictor, and (ii) a product of a second weight and the second intra-frame predictor. The first weight may be equal to or greater than the second weight, and the sum of the first weight and the second weight may be equal to 1.
[0032] In some embodiments, the final intra-frame predictor may be determined to be equal to the sum of: (i) a product of the first weight and the first intra-frame predictor, (ii) a product of the second weight and the second intra-frame predictor, and (iii) a product of the third weight and the planar mode-based intra-frame predictor. The first weight may be equal to or greater than the second weight, and the sum of the first weight, the second weight, and the third weight may be equal to 1.
[0033] According to another aspect of the present disclosure, a device is provided. The device has a processing circuit. The processing circuit can be configured to perform any method of video encoding.
[0034] Some aspects of the present invention also provide a non-transitory computer-readable medium storing instructions which, when executed by a computer for video decoding, cause the computer to perform any method for video decoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0036] Figure 1 is a schematic diagram of an exemplary subset of intra prediction modes.
[0037] Figure 2 is a diagram of exemplary intra prediction directions.
[0038] Figure 3 is a schematic diagram of a simplified block diagram of a communication system according to one embodiment.
[0039] Figure 4 is a schematic diagram of a simplified block diagram of a communication system according to another embodiment.
[0040] Figure 5 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment.
[0041] Figure 6 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment.
[0042] Figure 7 A block diagram of an encoder according to another embodiment is shown.
[0043] Figure 8 A block diagram of a decoder according to another embodiment is shown.
[0044] Fig. 9 is a schematic diagram of template-based intra mode derivation (TIMD) according to one embodiment.
[0045] Fig.10 A flow chart outlining a first exemplary decoding process according to some embodiments of the present disclosure is shown.
[0046] Fig.11 A flow chart outlining a second exemplary decoding process according to some embodiments of the present disclosure is shown.
[0047] Fig.12 A flowchart outlining a first exemplary encoding process according to some embodiments of the present disclosure is shown.
[0048] Fig.13 A flowchart outlining a second exemplary encoding process according to some embodiments of the present disclosure is shown.
[0049] Fig.14 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION
[0050] Figure 3 A simplified block diagram of a communication system (300) according to one embodiment of the present disclosure is shown. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). For example, in Figure 3 In the embodiment, the first terminal device performs unidirectional transmission of data to (310) and (320). For example, the terminal device (310) can encode video data (e.g., a video picture stream collected by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video code streams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to restore the video picture, and display the video picture based on the restored video data. Unidirectional data transmission may be common in media service applications, etc.
[0051] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data that may occur, for example, during a video conference. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., a video picture stream captured by the terminal device) for transmission to the other terminal device of the terminal devices (330) and (340) via a network (350). Each of the terminal devices (330) and (340) can also receive encoded video data sent by the other terminal device of the terminal devices (330) and (340), and can decode the encoded video data to restore the video picture, and can display the video picture on an accessible display device based on the restored video data.
[0052] exist Figure 3 In the example of, terminal device (310), terminal device (320), terminal device (330) and terminal device (340) may be shown as a server, a personal computer and a smart phone, but the principles of the present disclosure may not be limited thereto. Embodiments of the present disclosure are applicable to applications of laptop computers, tablet computers, media players and / or dedicated video conferencing devices. Network (350) represents any number of networks that transmit encoded video data between terminal devices (310), terminal devices (320), terminal devices (330) and terminal devices (340), including, for example, wired (wired) and / or wireless communication networks. Communication network (350) can exchange data in circuit switching channels and / or packet switching channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purposes of this application, the architecture and topology of network (350) may be unimportant to the operation of the present disclosure unless explained below.
[0053] As examples of applications of the disclosed subject matter, Figure 4The placement of the video encoder and video decoder in a streaming environment is shown. The disclosed subject matter can be equally applicable to other video-enabled applications, including, for example, video conferencing, digital television, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0054] The streaming system may include an acquisition subsystem (413), which may include a video source (401), such as a digital camera, which creates, for example, an uncompressed video picture stream (402). In one example, the video picture stream (402) includes samples captured by a digital camera. The video picture stream (402), depicted as thick lines to emphasize the high amount of data compared to the encoded video data (404) (or the encoded video bitstream), can be processed by an electronic device (420), which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or the encoded video bitstream (404)), depicted as thin lines to emphasize the lower amount of data compared to the video picture stream (402), can be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as Figure 4 The client subsystem (406) and the client subsystem (408) in the embodiment of the present invention can access the streaming server (405) to retrieve the copy (407) and the copy (409) of the encoded video data (404). The client subsystem (406) can include a video decoder (410), for example, in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and creates an output video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or another presentation device (not shown). In some streaming systems, the encoded video data (404), the video data (407), and the video data (409) (e.g., a video bitstream) can be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC.
[0055] It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).
[0056] Figure 5 A block diagram of a video decoder (510) according to an embodiment of the present disclosure is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 A video decoder (410) is shown in an example.
[0057] The receiver (531) can receive one or more encoded video sequences to be decoded by the video decoder (510); in the same embodiment or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence can be received from a channel (501), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (531) can receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which can be forwarded to their respective use entities (not shown). The receiver (531) can separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) can be set outside the video decoder (510) (not shown). In other cases, a buffer memory (not shown) may be provided external to the video decoder (510), for example, to prevent network jitter, and another buffer memory (515) may be provided internal to the video decoder (510), for example, to handle playback timing. When the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be required, or the buffer memory may be made smaller. For use on a packet network such as the Internet, a buffer memory (515) may also be required, which may be relatively large and may advantageously have an adaptive size, and may be implemented at least partially in an operating system or similar element (not shown) outside the video decoder (510).
[0058] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510) and potential information for controlling a display device such as a display device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but can be coupled to the electronic device (530), such as Figure 5. The control information for the display device may be in the form of a parameter set fragment (not shown) of Supplemental Enhancement Information (SEI) or Video Usability Information (VUI). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be based on a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a subgroup parameter set for at least one subgroup of the pixel subgroups in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (Group of Pictures, GOP), a picture, a block, a slice, a macroblock, a coding unit (CodingUnit, CUs, a block, a transform unit (Transform Unit, TU), a prediction unit (Prediction Unit, PU), etc. The parser (520) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.
[0059] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).
[0060] Depending on the type of the coded video picture or part of the coded video picture (e.g., inter-frame and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of the symbol (521) can involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). For clarity, this subgroup control information flow between the parser (520) and the multiple units below is not described.
[0061] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into a number of functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.
[0062] The first unit is a sealer / inverse transform unit (551). The sealer / inverse transform unit (551) receives quantized transform coefficients as symbols (521) from the parser (520) and control information, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (551) can output a block including sample values, which can be input into an aggregator (555).
[0063] In some cases, the output samples of the scaler / inverse transform (551) may belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information extracted from the current picture buffer (558). For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (555) adds the prediction information already generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.
[0064] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit (553) may access the reference picture memory (557) to obtain samples for prediction. After the extracted samples are motion compensated according to the symbols (521) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (551) (in this case referred to as residual samples or residual signals) by the aggregator (555), thereby generating output sample information. The address at which the motion compensated prediction unit (553) extracts the predicted samples from the reference picture memory (557) may be controlled by a motion vector, and the motion vector is provided to the motion compensated prediction unit (553) in the form of a symbol (521), which may have, for example, X, Y and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (557) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.
[0065] The output samples of the aggregator (555) may be controlled by various loop filtering techniques in a loop filter unit (556). The video compression techniques may include loop filter techniques controlled by parameters contained in the coded video sequence (also referred to as the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), however, the video compression techniques may also be responsive to meta information obtained during decoding of the coded video sequence or a previous (in decoding order) portion of the coded video sequence, and to previously reconstructed and loop filtered sample values.
[0066] The output of the loop filter unit (556) may be a sample stream that may be output to a display device (512) and stored in a reference picture memory (557) for subsequent inter-picture prediction.
[0067] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified (e.g., by the parser (520)) as a reference picture, the current picture buffer (558) may become part of the reference picture memory (557) and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.
[0068] The video decoder (510) may perform decoding operations according to a predetermined video compression technology in a standard such as ITU-T H.265. The coded video sequence may follow the syntax specified by the video compression technology or standard used, in the sense that the coded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available under the profile. For compliance, the complexity of the coded video sequence is also required to be within the range defined by the video compression technology or standard level. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (e.g., measured in mega samples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further constrained by the metadata of the HRD buffer management signaled in the coded video sequence and the HRD buffer management signaled in the coded video sequence.
[0069] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The video decoder (510) may use the additional data to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0070] Figure 6 A block diagram of a video encoder (603) according to an embodiment of the present disclosure is shown. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) can be used to replace Figure 4 The video encoder (403) in the example.
[0071] The video encoder (603) may receive a video from a video source (601) (not Figure 6 In another example, the video source (601) is a part of the electronic device (620).
[0072] The video source (601) can provide a source video sequence to be encoded by the video encoder (603) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CRCB4:2:0, YCRCB4:4:4). In a media service system, the video source (601) can be a storage device that stores previously prepared videos. In a video conferencing system, the video source (601) can be a camera that collects local image information as a video sequence. The video data can be provided as multiple individual pictures that are given motion when viewed in sequence. The picture itself can be constructed as a spatial array of pixels, where each pixel can contain one or more samples, depending on the sampling structure, color space, etc. used. The relationship between pixels and samples can be easily understood by those skilled in the art. The following description focuses on samples.
[0073] According to one embodiment, the video encoder (603) can encode and compress the pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Performing the appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to the other functional units. For clarity, the coupling is not shown here. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, ...), picture size, picture group (group of pictures, GOPs) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a certain system design.
[0074] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an oversimplified description, in one example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and a reference picture) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video code stream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same sample values as the samples "seen" by the decoder when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that can occur if synchronization cannot be maintained, eg due to channel errors) is also used in some related techniques.
[0075] The operation of the "local" decoder (633) can be combined with, for example, Figure 5 The operation of the "remote" decoder of the video decoder (510) described in detail is identical. Figure 5 However, when the blocking symbols are available and the symbols can be losslessly encoded / decoded into the encoded video sequence by the entropy encoder (645) and the parser (520), the entropy decoding portion of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).
[0076] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also be present in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on the decoder operation. The description of encoder technology can be brief because encoder technology is contrary to the decoder technology described comprehensively. A more detailed description is only required in certain parts and is provided below.
[0077] During operation, in some examples, the source encoder (630) may perform motion compensated predictive encoding, which predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (632) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture, which pixel blocks may be selected as a prediction reference for the input picture.
[0078] The local video decoder (633) can decode the encoded video data of the picture that can be designated as the reference picture based on the symbols created by the source encoder (630). The operation of the encoding engine (632) can advantageously be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 6 When the video encoder (603) is decoded on a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that can be performed by the video decoder on the reference picture and can cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) can store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture that will be obtained by the remote video decoder.
[0079] The predictor (635) may perform a prediction search on the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor (635) may operate on a pixel block by pixel block basis to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).
[0080] The controller (650) can manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.
[0081] The outputs of all the above functional units may be entropy encoded in an entropy encoder (645). The entropy encoder (645) converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0082] The transmitter (640) can buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission over a communication channel (660), which can be a hardware / software link to a storage device storing the encoded video data. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0083] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned any of the following picture types:
[0084] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of those variants of I pictures and their corresponding applications and features.
[0085] A predictive picture (P picture) may be a picture that may be encoded and decoded using intra prediction or inter prediction, which predicts sample values of each block using at most one motion vector and at most one reference index.
[0086] Bidirectional predictive pictures (B pictures) can be pictures that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses up to two motion vectors and up to two reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.
[0087] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and coded block by block. The blocks may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to the block's corresponding picture. For example, a block of an I picture may be non-predictively coded, or the block may be predictively coded (spatial prediction or intra-frame prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction or by temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be predictively coded by spatial prediction or by temporal prediction with reference to one or two previously coded reference pictures.
[0088] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In its operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0089] In one embodiment, the transmitter (640) may send additional data with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0090] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded (referred to as the current picture) is divided into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case of using multiple reference pictures, the motion vector may have a third dimension that identifies the reference picture.
[0091] In some embodiments, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture, which are both before the current picture in the video in decoding order (but may be in the past and future in display order, respectively). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block may be predicted by a combination of the first reference block and the second reference block.
[0092] In addition, merge mode technology can be used in inter-picture prediction to improve coding efficiency.
[0093] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the High Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Each CTU can be recursively partitioned into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In the example, each CU is analyzed to determine the prediction type for the CU, such as an inter-prediction type or an intra-prediction type. According to temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0094] Figure 7A schematic diagram of a video encoder (703) according to another embodiment of the present disclosure is shown. The video encoder (703) is configured to receive a processed block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence and encode the processed block into an encoded picture as part of an encoded video sequence. In the example, the video encoder (703) is used instead of Figure 4 The video encoder (403) in the example.
[0095] In the HEVC example, the video encoder (703) receives a matrix of sample values for a processing block, which is, for example, a prediction block of 8×8 samples. The video encoder (703) uses, for example, rate-distortion optimization to determine whether to best encode the processing block using intra mode, inter mode, or bidirectional prediction mode. When the processing block is encoded in intra mode, the video encoder (703) can encode the processing block into a coded picture using intra prediction techniques; when the processing block is encoded in inter mode or bidirectional prediction mode, the video encoder (703) can encode the processing block into a coded picture using inter prediction or bidirectional prediction techniques, respectively. In some video coding techniques, the merge mode can be an inter-picture prediction submode, in which the motion vector is derived from one or more motion vector predictors without the aid of encoded motion vector components outside the predictors. In some other video coding techniques, there may be a motion vector component applicable to the main block. In the example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.
[0096] exist Figure 7 In the example of FIG. 7 , the video encoder ( 703 ) includes Figure 7 An inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) are shown coupled together.
[0097] The inter-frame encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., a block in a previous picture and a block in a subsequent picture), generate inter-frame prediction information (e.g., a description of redundant information according to an inter-frame coding technique, motion vectors, merge mode information), and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the encoded video information.
[0098] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with blocks already encoded in the same picture in some cases, generate quantization coefficients after transformation, and in some cases (e.g., intra prediction direction information according to one or more intra coding techniques) also generate intra prediction information. In one example, the intra encoder (722) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.
[0099] The general controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, when the mode is intra-frame mode, the general controller (721) controls the switch (726) to select the intra-frame mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select the intra-frame prediction information and include the intra-frame prediction information in the bitstream; when the mode is inter-frame mode, the general controller (721) controls the switch (726) to select the inter-frame prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select the inter-frame prediction information and include the inter-frame prediction information in the bitstream.
[0100] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients are then quantized to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded blocks are appropriately processed to generate a decoded picture, and in some examples, the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.
[0101] The entropy encoder (725) is configured to format the bitstream to include the coded blocks. The entropy encoder (725) is configured to include various information according to an appropriate standard (e.g., the HEVC standard). In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. It should be noted that according to the disclosed subject matter, when the block is encoded in the inter-frame mode or the merge sub-mode of the bidirectional prediction mode, there is no residual information.
[0102] Figure 8 A schematic diagram of a video decoder (810) according to another embodiment of the present disclosure is shown. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In one example, the video decoder (810) is used to replace Figure 4 A video decoder (410) is shown in an example.
[0103] exist Figure 8 In the example of FIG. 8 , the video decoder (810) includes an entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) coupled together. Figure 8 shown.
[0104] The entropy decoder (871) can be configured to reconstruct certain symbols from the encoded picture, which represent syntax elements constituting the encoded picture. Such symbols may include, for example, a mode for encoding the block (e.g., intra mode, inter mode, bidirectional prediction mode, a merged submode of the latter two, or another submode), prediction information (e.g., intra prediction information or inter prediction information) that can identify certain samples or metadata used by the intra decoder (872) or the inter decoder (880) for prediction, and residual information represented in the form of, for example, quantized transform coefficients. In one example, when the prediction mode is an inter prediction mode or a bidirectional prediction mode, the inter prediction information is provided to the inter decoder (880); when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can be inverse quantized and provided to the residual decoder (873).
[0105] The inter-frame decoder (880) is configured to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.
[0106] The intra decoder (872) is configured to receive intra prediction information and generate an intra prediction result based on the intra prediction information.
[0107] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (Quantizer Parameter, QP), and this information can be provided by the entropy decoder (871) (the data path is not shown because this may only be a low amount of control information).
[0108] The reconstruction module (874) is configured to combine the residual output by the residual decoder (873) and the prediction result (output by the inter-frame or intra-frame prediction module as appropriate) in the spatial domain to form a reconstructed block, which can be part of a reconstructed picture, which in turn can be part of a reconstructed video. It should be noted that other suitable operations such as deblocking operations can be performed to improve visual quality.
[0109] It should be noted that the video encoder (403), video encoder (603) and video encoder (703) and video decoder (410), video decoder (510) and video decoder (810) can be implemented using any suitable technology. In one embodiment, the video encoder (403), video encoder (603) and video encoder (703) and video decoder (410), video decoder (510) and video decoder (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoder (403), video encoder (603) and video encoder (603) and video decoder (410), video decoder (510) and video decoder (810) can be implemented using one or more processors that execute software instructions.
[0110] The present disclosure includes improvements to decoding-side intra-mode derivation.
[0111] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) released the H.265 / HEVC standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4), respectively. In 2015, the two standards organizations jointly formed the Joint Video Exploration Team (JVET) to explore the potential of developing the next video coding standard beyond HEVC. In April 2018, JVET officially launched the standardization process for the next generation of video coding beyond HEVC. The new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Expert Team. In July 2020, H.266 / VVC version 1 was finalized. In January 2021, an ad hoc group was established to investigate enhanced compression beyond the capabilities of VVC.
[0112] In decoding-side intra mode derivation (DIMD), the intra mode may be derived using relevant syntax elements signaled in the bitstream, or the intra mode may be derived at the decoding side without using relevant syntax elements signaled in the bitstream. Many methods may be used to derive the intra mode at the decoding side, and the expression "decoding-side intra mode derivation" is not limited to the methods described in the present disclosure.
[0113] In DIMD, two intra modes can be derived from multiple candidate intra modes of the current CU / PU based on the reconstructed adjacent samples of the current CU / PU. In DIMD, texture gradient analysis can be performed at the encoding end and the decoding end to generate multiple candidate intra modes based on the reconstructed adjacent samples. Each of the multiple candidate intra modes can be associated with a respective gradient history (or respective gradient). Two intra modes (e.g., intraMode1 and intraMode2) with the highest gradient history (or with the highest gradient in the histogram) can be selected. The intra mode value predictions of the two selected intra modes (e.g., intraMode1 and intraMode2) can be combined with the planar mode prediction value using a weighted sum. Based on the combination of intraMode1, intraMode2 and planar modes, the final intra mode prediction value for the current CU / PU can be formed.
[0114] Table 1 shows an exemplary DIMD signaling process. As shown in Table 1, a DIMD flag (e.g., DIMD_flag) may be signaled before an ISP flag (e.g., ISP_flag). When DIMD_flag is 1 (or true), it may indicate that the current CU / PU uses DIMD, and then the ISP_flag may be further parsed to verify whether the current CU / PU applies ISP. When DIMD_flag is not 1 (or false), it may indicate that the current CU / PU does not use DIMD. Therefore, syntax elements related to other intra-coding tools (e.g., MIP, MRL, MPM, etc.) may be parsed in the decoder.
[0115] The context modeling of the DIMD flag may depend on the neighboring CU / PU. For example, the context modeling of the DIMD flag may depend on (i) the availability of the left neighboring CU / PU or the upper neighboring CU / PU, and (ii) whether the left neighboring CU / PU or the upper neighboring CU / PU also uses DIMD. If the left neighboring CU / PU or the upper neighboring CU / PU exists and uses DIMD, the context index (e.g., ctxIdx) may be 1. When the left neighboring CU / PU and the upper neighboring CU / PU exist at the same time and use DIMD, ctxIdx may be 2. Otherwise, ctxIdx may be 0. Table 1 Pseudo code of DIMD signaling notification
[0116] Template-based intra mode derivation (TIMD) can use the reference sample of the current CU as a template and select an intra mode from a set of candidate intra prediction modes associated with TIMD. For example, the selected intra mode can be determined as the best intra mode based on a cost function. Fig. 9 As shown, the neighboring reconstructed samples of the current CU (902) can be used as a template (904). The reconstructed samples in the template (904) can be compared with the predicted samples of the template (904). The predicted samples can be generated using the reference samples (906) of the template (904). The reference samples (906) can be neighboring reconstructed samples around the template (904). The cost function can be used to calculate the cost (or distortion) between the predicted samples and the reconstructed samples in the template (904) based on a corresponding one of the set of candidate intra-frame prediction modes. The intra-frame prediction mode with the minimum cost (or distortion) can be selected as the intra-frame prediction mode (e.g., the best intra-frame prediction mode) to perform inter-frame prediction on the current CU (902).
[0117] Table 2 shows an exemplary encoding process associated with TIMD. As shown in Table 1, a TIMD flag (e.g., TIMD_flag) can be signaled when the DIMD flag (e.g., DIMD_flag) is not 1 (or not true). When DIMD_flag is 1, the current CU / PU uses DIMD, and the ISP flag (e.g., ISP_flag) can be parsed to see whether the current CU / PU uses ISP. When DIMD_flag is not 1, TIMD_flag is parsed. When TIMD_flag is 1, TIMD can be applied to the current CU / PU without applying other intra-coding tools (e.g., ISP is not allowed when TIMD is used). When TIMD_flag is not 1 (or is false), syntax elements related to other intra-coding tools (e.g., MIP, MRL, MPM, etc.) can be parsed in the decoder. Table 2 Pseudo code of TIMD signaling notification
[0118] In the present disclosure, a combination of ISP and TIMD may be provided. Table 3 shows an exemplary pseudo code of the combination of ISP and TIMD. Table 3 Pseudo code of ISP and TIMD combination
[0119] As shown in Table 3, a TIMD flag (e.g., TIMD_flag) can be signaled when a DIMD flag (e.g., DIMD_flag) is not 1 (or false). When the DIMD flag is 1 (or true), it indicates that the current CU / PU uses DIMD. The ISP flag (e.g., ISP_flag) can be further parsed to verify whether the ISP is used for the current CU / PU. When the DIMD flag is not 1, the TIMD flag can be parsed. When the TIMD flag is 1, it can indicate that TIMD is applied to the current CU / PU. Accordingly, the ISP flag can be parsed to see if the ISP is applied to the current CU / PU. When the TIMD flag is not 1, syntax elements related to other intra-coding tools (e.g., MIP, MRL, MPM, etc.) can be parsed in the decoder.
[0120] In some embodiments, for the ISP flags provided in Table 3, context-adaptive binary arithmetic coding (CABAC) context modeling may depend on the value of the DIMD flag. In one example, the allocation of context indices (e.g., ctxIdx) associated with CABAC may depend only on the DIMD flag. For example, when the DIMD flag is equal to 1, ctxIdx may be set equal to 1. Otherwise, ctxIdx may be set equal to 0.
[0121] In another example, allocation of a context index (e.g., ctxIdx) depends not only on the DIMD flag, but also on one or more other factors, such as whether the adjacent CUs are split based on the ISP (e.g., the upper adjacent CU is split based on the ISP, or the left adjacent CU is split based on the ISP).
[0122] In some embodiments, for the ISP flags described in Table 3, CABAC context modeling may depend on the value of the TIMD flag. In one example, the allocation of a context index (e.g., ctxIdx) associated with CABAC may depend only on the TIMD flag. For example, when the TIMD flag is equal to 1, ctxIdx may be set equal to 1. Otherwise, ctxIdx may be set equal to 0.
[0123] In another example, allocation of a context index (e.g., ctxIdx) depends not only on the TIMD flag, but also on one or more other factors, such as whether the adjacent CU is split based on the ISP (e.g., the upper adjacent CU is split based on the ISP, or the left adjacent CU is split based on the ISP).
[0124] In the present disclosure, instead of deriving two intra-frame modes (e.g., intraMode1 and intraMode2) at the decoding end, in some embodiments, one intra-frame mode can be signaled in the codestream and the other intra-frame mode can be derived at the decoding end based on DIMD.
[0125] In one embodiment, an intra-mode (e.g., intraMode1) can be signaled in the bitstream using the most probable mode (MPM) MPM and MPM remainder related syntax elements. Table 4 shows an exemplary pseudo code for signaling an intra-mode (e.g., intraMode1) using MPM and MPM remainder related syntax elements. Table 4 Pseudo code of the proposed process
[0126] As shown in Table 4, when DIMD_flag is 1, it means that the current CU / PU uses DIMD, and ISP_flag can be further parsed to check whether the current CU / PU uses ISP. After checking whether ISP is used for the current CU / PU, the first intra mode (e.g., intraMode1) can be derived by parsing the syntax elements related to MPM and MPM remainder. For example, the first intra mode (e.g., intraMode1) can be determined based on the candidate intra modes in the MPM list. When DIMD_flag is not 1, syntax elements related to other intra coding tools (e.g., MIP, MRL, etc.) can be parsed in the decoder.
[0127] Therefore, the first intra-frame mode (e.g., intraMode1) can be signaled in the bitstream and defined by parsing the syntax elements related to the MPM and the MPM remainder, and the second intra-frame mode (e.g., intraMode2) can be derived at the decoding end based on the DIMD without signaling in the bitstream. It should be noted that intraMode1 and intraMode2 cannot be the same.
[0128] In the present disclosure, the intra predictor obtained based on intraMode1 and the intra predictor obtained based on intraMode2 may be used to generate a final intra predictor (eg, finalIntraPredictor) for the current CU / PU. Equation (1) is an example of such a final intra predictor. finalIntraPredictor=weight1*(intra predictor of intraMode1)+weight2*(intra predictor of intraMode2) equation (1) Among them, weight1 is not less than weight2, and weight1+weight2=1.
[0129] In some embodiments, (weight1, weight2) may have four available values, such as (7 / 8, 1 / 8), (6 / 8, 2 / 8), (5 / 8, 3 / 8), and (4 / 8, 4 / 8). The value of (weight1, weight2) may depend on the width or height of the adjacent intra mode or the current CU / PU.
[0130] In another embodiment, the final intra predictor of the current CU / PU may be a weighted sum of the intra predictor of intraMode1, the intra predictor of intraMode2, and the intra predictor of planar. Equation (2) is an example of such a final intra predictor. It should be noted that intraMode1, intraMode2, and planar cannot be the same. finalIntraPredictor=weight1*(intra predictor of intraMode1)+weight2*(intra predictor of intraMode2)+weight3*(intra predictor of planar) Equation (2) Among them, weight1 is not less than weight2, and weight1+weight2+weight3=1.
[0131] (weight1, weight2, weight3) can have 6 available values, such as (6 / 8, 1 / 8, 1 / 8), (5 / 8, 2 / 8, 1 / 8), (4 / 8, 3 / 8, 1 / 8), (5 / 8, 1 / 8, 2 / 8), (4 / 8, 2 / 8, 2 / 8) and (3 / 8, 3 / 8, 2 / 8). The value of (weight1, weight2, weight3) can depend on the adjacent intra mode or the width or height of the current CU / PU.
[0132] Fig.10 A flow chart outlining a first exemplary decoding process (1000) according to some embodiments of the present disclosure is shown. Fig.11 A flow chart outlining a second exemplary decoding process (1100) according to some embodiments of the present disclosure is shown. Fig.12 A flow chart outlining a first exemplary encoding process (1200) according to some embodiments of the present disclosure is shown. Fig.13 A flow chart outlining a second exemplary encoding process (1300) according to some embodiments of the present disclosure is shown. The proposed processes can be used alone or in combination in any order. In addition, each process (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0133] In an embodiment, any operation of the process (e.g., (1000), (1100), (1200), and (1300)) can be combined or arranged in any number or in any order as needed. In an embodiment, two or more operations of the process (e.g., (1000), (1100), (1200), and (1300)) can be performed in parallel.
[0134] The processes (e.g., (1000), (1100), (1200), and (1300)) may be used for reconstruction and / or encoding of a block to generate a prediction block for the block being reconstructed. In various embodiments, the processes (e.g., (1000), (1100), (1200), and (1300)) are performed by processing circuits such as: processing circuits in terminal device (210), terminal device (220), terminal device (230), and terminal device (240), processing circuits that perform the functions of a video encoder (303), processing circuits that perform the functions of a video decoder (310), processing circuits that perform the functions of a video decoder (410), processing circuits that perform the functions of a video encoder (503), and the like. In some embodiments, the processes (e.g., (1000), (1100), (1200), and (1300)) are implemented as software instructions, so that when the processing circuit executes the software instructions, the processing circuit performs the processes (e.g., (1000), (1100), (1200), and (1300)).
[0135] like Fig.10 The process (1000) may start from (S1001) to (S1010). In (S1010), encoding information of a current block and adjacent blocks of the current block may be received from an encoded video bitstream.
[0136] In (S1020), first information associated with the current block in the encoding information may be decoded. The first information may indicate whether the current block is intra-predicted based on DIMD, in which the intra-prediction mode of the current block is derived based on neighboring blocks.
[0137] In (S1030), second information associated with the current block may be obtained in the encoding information. The second information may indicate whether the current block is divided based on an intra sub-region partitioning (ISP) mode.
[0138] In (S1040), the context model index can be determined in response to (i) first information indicating that the current block is intra-predicted based on DIMD, or (ii) second information indicating that the upper neighboring block or the left neighboring block of the current block among the neighboring blocks is divided based on the ISP mode.
[0139] In (S1050), the current block may be decoded from the encoded video stream based at least on the context model index. Of course, the current block may be reconstructed based on the decoded first information and / or the second information. For example, when the first information indicates that the current block is intra-predicted based on DIMD, the current block may be reconstructed based on DIMD; when the second information indicates that the current block is divided based on the ISP mode, the current block may be reconstructed based on the ISP mode.
[0140] In some embodiments, the second information may be decoded in response to first information indicating that the current block is intra-predicted based on DIMD. In addition, the syntax element associated with the ISP mode may be decoded in response to second information indicating that the current block is divided based on the ISP mode.
[0141] In process (1000), determining the context model in response to (i) the first information or (ii) the second information may be performed in response to the first information indicating that the current block is intra-predicted based on DIMD. In response to the first information indicating that the current block is not intra-predicted based on DIMD, third information associated with the current block may be obtained in the encoding information. The third information may indicate whether the current block is intra-predicted based on a template-based intra mode derivation (TIMD) including a set of candidate intra-prediction modes. The context model index may be determined in response to (i) the third information indicating that the current block is intra-predicted based on TIMD, or (ii) the second information indicating that the upper neighboring block or the left neighboring block of the current block in the neighboring blocks is divided based on the ISP mode.
[0142] In some embodiments, the second information may be decoded in response to third information indicating that the current block is intra-predicted based on TIMD. Thus, the syntax elements associated with the ISP mode may be decoded in response to second information indicating that the current block is divided based on the ISP mode.
[0143] In the method, in response to second information indicating that the current block is not divided based on the ISP mode, a syntax element associated with another intra-frame coding mode may be decoded. The other intra-frame coding mode may include one of the following: matrix-based intra prediction (MIP), multiple reference line intra prediction (MRL), and most probable mode (MPM).
[0144] In some embodiments, in response to first information indicating that the current block is intra-predicted based on DIMD, the context model index may be determined to be 1. In response to first information indicating that the current block is not intra-predicted based on DIMD, the context model index may be determined to be 0.
[0145] In some embodiments, the context model index may be used in context-based adaptive binary arithmetic coding (CABAC).
[0146] In some embodiments, in response to third information indicating that the current block is intra-predicted based on TIMD, the context model index may be determined to be 1. In response to third information indicating that the current block is not intra-predicted based on TIMD, the context model index may be determined to be 0.
[0147] like Fig.11 As shown, the process (1100) may start from (S1101) to (S1110). In (S1110), encoding information of a current block and adjacent blocks of the current block may be received from an encoded video code stream.
[0148] In (S1120), first information associated with the current block in the encoding information may be decoded. The first information may indicate whether the current block is intra-predicted based on DIMD, where the intra-prediction mode of the current block is derived based on neighboring blocks.
[0149] In (S1130), in response to first information indicating that the current block is intra predicted based on DIMD, a first intra prediction mode may be determined based on DIMD. In addition, a second intra prediction mode may be determined based on a syntax element included in the encoding information and associated with the MPM and the MPM remainder.
[0150] In (S1140), the current block may be reconstructed based on the first intra prediction mode and the second intra prediction mode.
[0151] In the process (1100), second information associated with the current block in the encoding information may be decoded in response to first information indicating that the current block is intra-predicted based on DIMD. The second information may indicate whether the current block is divided based on the ISP mode. In response to the second information indicating that the current block is divided based on the ISP mode, a syntax element associated with the ISP mode may be decoded.
[0152] In the process (1100), in response to second information indicating that the current block is not divided based on the ISP mode, a syntax element associated with another intra-frame coding mode may be decoded. The other intra-frame coding mode includes one of MIP, MRL and MPM.
[0153] In the process (1100), a first intra predictor may be determined based on a first intra prediction mode. A second intra predictor may be determined based on a second intra prediction mode. A final intra predictor may be determined based on the first intra predictor and the second intra predictor.
[0154] In some embodiments, the final intra-frame predictor may be determined to be equal to the sum of: (i) the product of the first weight and the first intra-frame predictor, and (ii) the product of the second weight and the second intra-frame predictor. The first weight may be equal to or greater than the second weight, and the sum of the first weight and the second weight may be equal to 1.
[0155] In some embodiments, the final intra-frame predictor may be determined to be equal to the sum of: (i) the product of the first weight and the first intra-frame predictor, (ii) the product of the second weight and the second intra-frame predictor, and (iii) the product of the third weight and the intra-frame predictor based on a planar mode. The first weight may be equal to or greater than the second weight, and the sum of the first weight, the second weight, and the third weight may be equal to 1.
[0156] like Fig.12 , the process (1200) may start from (S1201) to (S1210). In (S1210), first information associated with a current block may be generated. The first information may indicate whether the current block in the video picture is intra-predicted based on a decoding-side intra mode derivation (DIMD), in which the intra prediction mode of the current block is derived based on neighboring blocks of the current block.
[0157] In (S1220), second information associated with the current block may be generated. The second information may indicate whether the current block is divided based on an intra sub-region partitioning (ISP) mode.
[0158] In (S1230), the context model index can be determined in response to (i) first information indicating that the current block is intra-predicted based on DIMD, or (ii) second information indicating that the upper neighboring block or the left neighboring block of the current block among the neighboring blocks is divided based on the ISP mode.
[0159] In (S1240), the current block may be encoded based on at least the context model indicated by the context model index.
[0160] like Fig.13 As shown, the process (1300) can start from (S1301) to (S1310), wherein in (S1310), the first intra-frame prediction mode of the current block can be determined based on DIMD, in which the first intra-frame prediction mode of the current block is derived based on the neighboring blocks of the current block.
[0161] In (S1320), a second intra prediction mode of the current block may be determined based on a most probable mode (MPM) list and an MPM remainder list associated with the current block.
[0162] In (S1330), the current block may be intra predicted based on the first intra prediction mode and the second intra prediction mode.
[0163] In (S1340), first information and syntax elements may be generated. The first information may indicate whether the current block is intra-predicted based on DIMD. The syntax element may be associated with the MPM list and the MPM remainder list and indicate a second intra-prediction mode of the current block.
[0164] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.14 A computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0165] Computer software may be encoded using any suitable machine code or computer language that may be assembled, compiled, linked, or similarly structured to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or by way of interpreted code, microcode execution, etc.
[0166] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, IoT devices, etc.
[0167] Fig.14 The components for the computer system (1400) shown in the example are exemplary in nature and are not intended to impose any limitation on the use or functionality of the computer software implementing the disclosed embodiments. The configuration of the components should not be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary embodiment of the computer system (1400).
[0168] The computer system (1400) may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input by one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to a person's conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0169] The human-machine interface input device may include one or more of the following (only one of each is shown): keyboard (1401), mouse (1402), touchpad (1403), touch screen (1410), data gloves (not shown), joystick (1405), microphone (1406), scanner (1407), camera (1408).
[0170] The computer system (1400) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (1410), a data glove (not shown), or a joystick (1405), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (1409), headphones (not shown)), visual output devices (e.g., screens (1410), to include CRT screens, LCD screens, plasma screens, OLED screens, each screen with or without touch screen input capabilities, each screen with or without tactile feedback capabilities, some of which may be capable of outputting two-dimensional visual outputs or more than three-dimensional outputs through devices such as stereographic output, virtual reality glasses (not shown), holographic displays and smoke boxes (not shown), and printers (not shown).
[0171] The computer system (1400) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1420) with CD / DVD or similar media (1421), thumb drives (1422), removable hard drives or solid-state drives (1423), traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.
[0172] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0173] The computer system (1400) may also include an interface (1454) to one or more communication networks (1455). The network may be, for example, a wireless network, a wired network, an optical network. The network may also be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area network digital networks including cable TV, satellite TV and terrestrial broadcast television, in-vehicle and industrial television including CANBus, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1449) (e.g., a USB port of the computer system (1400)); as described below, other network interfaces are typically integrated into the kernel of the computer system (1400) by connecting to the system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). The computer system (1400) can use any of these networks to communicate with other entities. This communication can be one-way receive only (e.g., broadcast television), one-way send only (e.g., CANbus connected to certain CANbus devices), or two-way, such as connecting to other computer systems using a local area network or wide area digital network. As described above, certain protocols and protocol stacks can be used on each of these networks and network interfaces.
[0174] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be connected to the kernel (1440) of the computer system (1400).
[0175] The kernel (1440) may include one or more central processing units (CPUs) (1441), graphics processing units (GPUs) (1442), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1443), hardware accelerators for certain tasks (1444), graphics adapters (1450), etc. These devices, along with read-only memory (ROM) (1445), random access memory (1446), internal mass storage (1447) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (1448). In some computer systems, the system bus (1448) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the kernel's system bus (1448) or to the kernel's system bus (1448) via a peripheral device bus (1449). In one example, a screen (1410) may be connected to a graphics adapter (1450). The architecture of the peripheral bus includes PCI, USB, etc.
[0176] The CPU (1441), GPU (1442), FPGA (1443) and accelerator (1444) can execute certain instructions, which in combination can constitute the above-mentioned computer code. The computer code can be stored in ROM (1445) or RAM (1446). Transition data can also be stored in RAM (1446), while permanent data can be stored in, for example, internal mass storage (1447). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely related to one or more CPUs (1441), GPUs (1442), mass storage (1447), ROM (1445), RAM (1446), etc.
[0177] The computer readable medium may have thereon computer codes for performing various computer-implemented operations. The media and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or they may be of a type well known and available to those skilled in the art of computer software.
[0178] As a non-limiting example, a computer system having an architecture (1400), particularly a kernel (1440), can provide functionality by executing software contained in one or more tangible computer-readable media through one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.). Such computer-readable media can be media associated with a user-accessible mass storage as described above, as well as certain memories of a kernel (1440) that are non-transitory, such as a kernel internal mass storage (1447) or ROM (1445). Software that implements various embodiments of the present disclosure can be stored in such a device and executed by the kernel (1440). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the kernel (1440), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific parts of specific processes or specific procedures described herein, including defining data structures stored in RAM (1446) and modifying these data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1544)) that may operate in place of or in conjunction with software to perform specific processes or specific portions of specific processes described herein. Where appropriate, references to portions of software may include logic and vice versa. Where appropriate, references to portions of computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software. Appendix A: Abbreviations JEM: Joint Exploration Mode VVC: Next Generation Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplemental Enhancement Information VUI: Video Availability Information GOPs: Group of Pictures TUs: Transformation Units PUs: Prediction Units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: prediction blocks HRD: Hypothesized Reference Decoder SNR: Signal to Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT: cathode ray tube LCD: Liquid Crystal Display OLED: Organic Light Emitting Diode CD: compact disc DVD: Digital Video Disc ROM: Read Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: Solid State Drive IC: Integrated Circuit CU: Coding Unit
[0179] Although the present disclosure has described a number of exemplary embodiments, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.
Claims
1. A method for video decoding, the method include: Receiving encoding information of a current block and adjacent blocks of the current block from an encoded video bitstream; Acquire first information from the encoding information, the first information indicating whether the current block is intra-predicted based on a template-based intra-mode derivation TIMD including a set of candidate intra-prediction modes; Acquire second information from the encoding information, where the second information indicates whether the current block is divided based on an intra-frame sub-region division ISP mode; Determine a context model index based on (i) the first information indicating that the current block is intra-predicted based on TIMD, and / or (ii) the second information indicating that the current block is divided based on an intra-sub-region division ISP mode; as well as The current block is decoded from the encoded video stream based at least on the context model index.
2. The method according to claim 1, in, Determining the context model index comprises: The second information indicates that the upper neighboring block or the left neighboring block of the current block among the neighboring blocks is divided based on the ISP mode.
3. The method according to claim 1, in, The method further comprises: decoding the second information based on the first information indicating that the current block is intra-predicted based on TIMD; and Based on the second information indicating that the current block is divided based on the ISP mode, a syntax element associated with the ISP mode is decoded.
4. The method according to claim 3, in, The method further comprises: Based on the second information indicating that the current block is not divided based on the ISP mode, a syntax element associated with another intra-frame coding mode is decoded, and the other intra-frame coding mode includes one of matrix-based intra-frame prediction MIP, multi-reference row intra-frame prediction MRL and most likely mode MPM.
5. The method according to claim 1, in, The obtaining the first information from the coded information comprises: Acquire third information associated with the current block in the encoding information, the third information indicating whether the current block performs intra-frame prediction based on a decoding-side intra-frame mode derivation DIMD, in which the intra-frame prediction mode of the current block is derived based on the neighboring blocks; Based on the third information indicating that the current block is not intra-predicted based on DIMD, first information associated with the current block is obtained in the encoding information.
6. The method according to claim 1, in, The context model index is used in context-based adaptive binary arithmetic coding (CABAC).
7. The method according to claim 1, in, The method further comprises one of the following: Based on the first information indicating that the current block is intra-predicted based on TIMD, determining that the context model index is 1; as well as Based on the first information indicating that the current block is not intra-predicted based on TIMD, the context model index is determined to be 0.
8. A method for video encoding, the method include: Generate first information indicating whether the current block is intra-predicted based on a template-based intra-mode derivation TIMD including a set of candidate intra-prediction modes; Generate second information, wherein the second information indicates whether the current block is divided based on the intra-frame sub-region division ISP mode; Determine a context model index based on (i) the first information indicating that the current block is intra-predicted based on TIMD, and / or (ii) the second information indicating that the current block is divided based on an intra-sub-region division ISP mode or the second information indicating that an upper neighboring block or a left neighboring block of the current block among the neighboring blocks is divided based on an ISP mode; as well as The current block is encoded based on at least a context model indicated by the context model index.
9. A video decoding device, include: A processing circuit configured to: execute the method according to any one of claims 1-7.
10. A video encoding device, include: A processing circuit configured to: execute the method according to claim 8.
11. A non-transitory computer-readable medium storing instructions, which, when executed by a computer for video decoding, cause the computer to perform the method according to any one of claims 1 to 7.
12. A method for storing a video stream, It is characterized in that The video code stream is generated according to the encoding method according to claim 8, or is decoded based on the decoding method according to any one of claims 1-7.