Method for video decoding and related apparatus

By determining the applicable intra prediction mode in video decoding, combining bidirectional intra prediction and multi-reference line selection, the problem of low processing efficiency of non-adjacent reference line in the prior art is solved, and the encoding efficiency is improved.

CN120111214APending Publication Date: 2025-06-06TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510279277.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-01-07
Filing Date
2022-01-21
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When the prior art combines bidirectional intra prediction and multi-reference line selection, it is difficult to effectively process reference samples in non-adjacent reference lines, resulting in a decrease in encoding efficiency.

Method used

By receiving the encoded video bitstream of the block, it is determined whether one-way intra prediction or two-way intra prediction is applicable to the block. The block-based pattern information includes reference line index, intra prediction mode and block size. If it is determined that the bi-way intra prediction is applicable to the block, bi-way intra prediction is performed on the block.

Benefits of technology

Compatible with bidirectional intra prediction and multi-reference line selection in video decoding is realized, and encoding efficiency is improved, especially when processing non-adjacent reference lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111214A_ABST
    Figure CN120111214A_ABST
Patent Text Reader

Abstract

A method, apparatus, and computer readable storage medium for video decoding. The method includes receiving an encoded video bitstream of a block, the encoded video bitstream including mode information of the block. The method further includes determining whether unidirectional intra prediction or bidirectional intra prediction is applicable to the block based on mode information of the block, the mode information of the block including at least one of a reference cue index of the block, an intra prediction mode of the block, and a size of the block; if it is determined that the one-way intra prediction is applicable to the block, performing one-way intra prediction on the block; and performing bidirectional intra prediction on the block if it is determined that the bidirectional intra prediction is applicable to the block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application is based on and claims the benefit of priority of U.S. Provisional Application No. 63 / 215,888 filed on June 28, 2021 and U.S. Non-Provisional Application No. 17 / 570,603 filed on January 7, 2022, both of which are hereby incorporated by reference in their entirety.

[0003] This application is a divisional application for the Chinese patent application with application number 202280005806.2, application date January 21, 2022, and invention name “Coordinated design of bidirectional intra-frame prediction and multi-reference line selection”. Technical Field

[0004] The present disclosure relates to video encoding and / or decoding techniques, and in particular, to methods and related devices for video decoding. Background Art

[0005] The background description provided herein is for the purpose of generally introducing the background of the present disclosure. The content described in this background section is not necessarily prior art to the present disclosure.

[0006] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. An uncompressed digital video may include a series of pictures, each picture having a spatial dimension of, for example, 1920×1080 luminance samples and associated full chroma samples or sub-sampled chroma samples. A series of pictures may have a fixed or variable picture rate (alternatively referred to as a frame rate), such as 60 pictures per second or 60 frames per second. Uncompressed video has specific bit rate requirements for streaming or data processing. For example, at 8 bits per pixel per color channel, a video with a pixel resolution of 1920×1080, a frame rate of 60 frames per second, and a chroma subsampling of 4:2:0 requires a bandwidth of nearly 1.5 Gbit / s. One hour of such a video requires more than 600 Gbytes of storage space.

[0007] One purpose of video encoding and decoding can be to reduce redundancy in an uncompressed input video signal by compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements by two orders of magnitude or more in some cases. Both lossless compression and lossy compression and their combinations can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal via a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not fully retained during the encoding process and cannot be fully restored during the decoding process. When lossy compression is used, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal is useful for the intended application despite some information loss. In the case of video, lossy compression is widely used in many applications. The amount of distortion allowed depends on the application. For example, users of certain consumer streaming applications can tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a specific encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows encoding algorithms that produce higher losses and higher compression ratios.

[0008] Video encoders and decoders may utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transforms, quantization, and entropy coding.

[0009] Video codec techniques may include techniques known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. When all sample blocks are encoded in intra-mode, the picture may be referred to as an intra-picture. Intra-pictures and their derivatives (e.g., independent decoder refresh pictures) may be used to reset decoder states, and may therefore be used as the first picture in a coded video bitstream and video session or as a still image. The samples of the intra-predicted blocks may be transformed into the frequency domain, and the generated transform coefficients may be quantized before entropy coding. Intra-prediction represents a technique for minimizing sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after the transform, and the smaller the AC coefficient, the fewer bits are required to represent the block after entropy coding at a given quantization step size.

[0010] Traditional intra-frame coding, such as known from, for example, MPEG-2 generation coding techniques, does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to encode / decode blocks based on, for example, metadata and / or surrounding sample data obtained during encoding / decoding that is spatially adjacent and precedes the data block being intra-coded or decoded in decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. Note that, in at least some cases, intra-frame prediction uses only reference data from the current picture being reconstructed, and does not use reference data from other reference pictures.

[0011] There can be many different forms of intra-frame prediction. When more than one such technique can be used in a given video coding technique, the technique used can be referred to as an intra-frame prediction mode. One or more intra-frame prediction modes can be provided in a particular codec. In some cases, a mode can have sub-modes and / or can be associated with various parameters, and the mode / sub-mode information of a video block and the intra-frame coding parameters can be encoded separately or collectively included in a mode codeword. The codeword to be used for a given mode, sub-mode, and / or parameter combination can have an impact on the coding efficiency gain through intra-frame prediction, and therefore can have an impact on the entropy coding technique used to convert the codeword into a bitstream.

[0012] A certain mode of intra prediction was introduced by H.264, refined in H.265, and further refined in newer coding techniques such as the Joint Exploitation Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Typically, for intra prediction, the values ​​of neighboring samples that have become available can be used to form a block of prediction values. For example, the available values ​​of a particular set of neighboring samples along a particular direction and / or line can be copied into a block of prediction values. A reference to the direction used can be encoded in the bitstream, or it can itself be predicted.

[0013] Reference Figure 1A , depicted at the bottom right is a subset of 9 prediction value directions specified in the 33 possible intra-frame prediction directions of H.265 (corresponding to the 33 angular modes of the 35 intra-frame modes specified in H.265). The point (101) where the arrows converge represents the sample being predicted. The arrows indicate the direction in which neighboring samples are used to predict the sample at 101. For example, arrow (102) indicates that sample (101) is predicted from one or more neighboring samples at a 45 degree angle to the horizontal direction at the upper right. Similarly, arrow (103) indicates that sample (101) is predicted from one or more neighboring samples at a 22.5 degree angle to the horizontal direction at the lower left of sample (101).

[0014] Still refer to Figure 1A, depicted at the top right is a square block (104) of 4×4 samples (indicated by the dashed bold line). The square block (104) includes 16 samples, each of which is labeled by "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in block (104) in both the Y dimension and the X dimension. Since the size of the block is 4×4 samples, S44 is at the bottom right. Reference samples are further shown, which follow a similar numbering scheme. Reference samples are labeled by R, their Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, prediction samples that are adjacent to the block being reconstructed are used.

[0015] The intra picture prediction of block 104 can start by copying reference sample values ​​from neighboring samples according to the prediction direction signaled by the signal. For example, assume that the encoded video bitstream includes the following signaling, which indicates the prediction direction of arrow (102) for this block 104, that is, the samples are predicted according to one or more prediction samples at a 45 degree angle to the horizontal direction in the upper right direction. In this case, samples S41, S32, S23 and S14 are predicted according to the same reference sample R05. Then, sample S44 is predicted according to R08.

[0016] In some cases, the values ​​of multiple reference samples may be combined, such as by interpolation, in order to calculate the reference sample; in particular when the direction is not divisible by 45 degrees.

[0017] As video coding technology continues to develop, the number of possible directions is also increasing. For example, in H.264 (2003), nine different directions were available for intra prediction. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions when it is disclosed. Experimental studies have been conducted to help identify the most suitable intra prediction directions, and certain techniques in entropy coding are used to encode those most suitable directions with a small number of bits, accepting some bit penalty for the direction. In addition, the direction itself can sometimes be predicted based on the neighboring directions used in the intra prediction of decoded neighboring blocks.

[0018] Figure 1B A schematic diagram (180) depicting 65 intra prediction directions according to the JEM is shown to illustrate the increasing number of prediction directions in various encoding techniques developed over time.

[0019] The manner in which the bits representing the intra prediction direction are mapped to the prediction direction in the coded video bitstream may vary from one video coding technique to another; and the mapping may vary, for example, from a simple direct mapping of prediction direction to intra prediction mode to codeword to complex adaptive schemes involving most probable modes and similar techniques. However, in all cases, there may be some intra prediction directions that are statistically less likely to occur in the video content than some other directions. Since the goal of video compression is to reduce redundancy, in a well-functioning video coding technique, those less likely directions will be represented by a larger number of bits than more likely directions.

[0020] Inter-picture prediction or inter-prediction may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or part thereof (reference picture) may be used to predict a newly reconstructed picture or picture part (e.g., block) after spatial shifting in the direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. An MV may have two dimensions, X and Y, or three dimensions, the third dimension being an indication of the reference picture in use (similar to the time dimension).

[0021] In some video compression techniques, the current MV applicable to certain areas of sample data can be predicted from other MVs, such as the following other MVs, which are related to other areas of sample data that are spatially adjacent to the area being reconstructed and precede the current MV in decoding order. Doing so can significantly reduce the amount of data required to encode the MV by relying on removing redundancy in the relevant MV, thereby increasing compression efficiency. MV prediction can work effectively, for example, because when encoding an input video signal derived from a camera (called natural video), there is a statistical possibility that areas larger than the area applicable to a single MV in the video sequence move in similar directions, and therefore in some cases similar motion vectors derived from MVs of neighboring areas can be used for prediction. This results in the actual MV for a given area being similar or identical to the MV predicted from the surrounding MVs. Such an MV can be represented after entropy coding with less than the number of bits used when directly encoding the neighboring MVs instead of predicting from the neighboring MVs. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors when calculating predicted values ​​from several surrounding MVs.

[0022] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). Among the various MV prediction mechanisms specified by H.265, the following is a technique referred to as "spatial merging" hereinafter.

[0023] Specifically, refer to Figure 2 , the current block (201) includes samples that have been discovered by the encoder during the motion search process, which are predictable from a previous block of the same size that has been spatially offset. Instead of encoding the MV directly, the MV can be derived from metadata associated with one or more reference pictures, for example, using the MV associated with any of the five surrounding samples denoted A0, A1 and B0, B1, B2 (corresponding to 202 to 206, respectively), and the MV is derived from metadata of the nearest reference picture (in decoding order). In H.265, MV prediction can use the prediction value of the same reference picture that the neighboring block is also using.

[0024] In addition to unidirectional intra prediction, reference pixels along two prediction directions can be combined to obtain directional prediction values, which can be called bidirectional intra prediction or intra bidirectional prediction (IBP). In some implementations of applying IBP, two reference pixels located along the adjacent upper reference line and the adjacent left reference are weighted to achieve the prediction value. When both IBP and multi-reference line selection (MRLS) are applied to the coding block at the same time, some problems may arise. For example, how to apply IBP to the current block when selecting reference samples in one or more non-adjacent reference lines. Summary of the invention

[0025] The present disclosure describes various embodiments of methods, apparatus, and computer-readable storage media for video encoding and / or decoding.

[0026] According to one aspect, an embodiment of the present disclosure provides a method for video decoding. The method includes receiving a coded video bitstream of a block, wherein the coded video bitstream includes mode information of the block. The method also includes determining whether unidirectional intra prediction or bidirectional intra prediction is applicable to the block based on the mode information of the block, the mode information of the block including at least one of a reference line index of the block, an intra prediction mode of the block, and a size of the block; if it is determined that unidirectional intra prediction is applicable to the block, unidirectional intra prediction is performed on the block; and if it is determined that bidirectional intra prediction is applicable to the block, bidirectional intra prediction is performed on the block.

[0027] According to one aspect, an embodiment of the present disclosure provides a device for video decoding. The device includes a receiving module, which is configured to receive a coded video bitstream of a block, wherein the coded video bitstream includes mode information of the block. The device also includes a determining module, which is configured to determine whether unidirectional intra prediction or bidirectional intra prediction is applicable to the block based on the mode information of the block, the mode information of the block includes at least one of the reference line index of the block, the intra prediction mode of the block and the size of the block; a unidirectional intra prediction module, which is configured to perform unidirectional intra prediction on the block if it is determined that unidirectional intra prediction is applicable to the block; and a bidirectional intra prediction module, which is configured to perform bidirectional intra prediction on the block if it is determined that bidirectional intra prediction is applicable to the block.

[0028] According to another aspect, an embodiment of the present disclosure provides a device for video decoding. The device includes a memory storing instructions; and a processor communicating with the memory. When the processor executes the instructions, the processor performs the above method for video decoding.

[0029] In another aspect, an embodiment of the present disclosure provides a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform the above-described method for video decoding.

[0030] The above aspects and other aspects and their implementation methods are described in more detail in the drawings, the specification and the claims. Through the technical solution of the present disclosure, it is determined whether to use unidirectional intra-frame prediction or bidirectional intra-frame prediction for the coding block based on the mode information of the received coding block, and the mode information includes the reference line index, intra-frame prediction mode and / or size of the coding block. That is to say, it is determined whether to use unidirectional intra-frame prediction or bidirectional intra-frame prediction for the coding block according to the reference line index, intra-frame prediction mode and / or size of the coding block, thereby providing a coding method compatible with IBP and MRLS. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings.

[0032] Figure 1A A schematic diagram showing an exemplary subset of intra prediction direction modes;

[0033] Figure 1B A diagram showing exemplary intra prediction directions;

[0034] Figure 2 A schematic diagram showing spatial merging candidates of a current block and its surroundings for motion vector prediction in one example is shown;

[0035] Figure 3A schematic diagram showing a simplified block diagram of a communication system (300) according to an example implementation;

[0036] Figure 4 A schematic diagram showing a simplified block diagram of a communication system (400) according to an example implementation;

[0037] Figure 5 A schematic diagram showing a simplified block diagram of a video decoder according to an example implementation;

[0038] Figure 6 A schematic diagram showing a simplified block diagram of a video encoder according to an example implementation;

[0039] Figure 7 shows a block diagram of a video encoder according to another example embodiment;

[0040] Figure 8 shows a block diagram of a video decoder according to another example embodiment;

[0041] Fig. 9 A directional intra prediction mode according to an example implementation of the present disclosure is shown;

[0042] Fig.10 A non-directional intra prediction mode according to an example implementation of the present disclosure is shown;

[0043] Fig.11 A recursive intra prediction mode according to an example implementation of the present disclosure is shown;

[0044] Fig.12 An intra prediction scheme based on various reference lines according to an example implementation of the present disclosure is shown;

[0045] Fig.13A Bidirectional intra prediction according to an example implementation of the present disclosure is shown;

[0046] Fig. 13B Another bidirectional intra prediction according to an example implementation of the present disclosure is shown;

[0047] Fig.14 A flowchart showing a method according to an example implementation of the present disclosure is shown;

[0048] Fig.15 Bidirectional intra prediction based on multiple reference lines according to an example implementation of the present disclosure is shown;

[0049] Fig.16 Another bidirectional intra prediction based on multiple reference lines according to an example implementation of the present disclosure is shown;

[0050] Fig.17 A schematic diagram of a computer system according to an example implementation of the present disclosure is shown. DETAILED DESCRIPTION

[0051] The present invention will now be described in detail hereinafter with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of implementation by way of illustration. However, it is noted that the present invention may be implemented in a variety of different forms, and therefore, the subject matter covered or claimed is intended to be construed as not being limited to any of the implementations to be set forth below. It is also noted that the present invention may be implemented as a method, apparatus, component, or system. Thus, embodiments of the present invention may, for example, take the form of hardware, software, firmware, or any combination thereof.

[0052] Throughout the specification and claims, terms may have slightly different meanings that are suggested or implied in the context beyond the explicitly stated meaning. The phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" as used herein do not necessarily refer to different embodiments. Similarly, the phrases "in one implementation" or "in some implementations" as used herein do not necessarily refer to the same implementation, and the phrases "in another implementation" or "in other implementations" as used herein do not necessarily refer to different implementations. For example, it is intended that the claimed subject matter includes combinations of the whole or part of the exemplary embodiments / implementations.

[0053] Typically, terms can be understood at least in part from usage in context. For example, as used herein, terms such as "and", "or" or "and / or" can include various meanings, which can depend at least in part on the context in which such terms are used. Typically, if "or" is used for an association list, such as A, B or C, then "or" is intended to represent A, B and C, which are used in an inclusive sense herein, and A, B or C, which are used in an exclusive sense herein. In addition, at least in part depending on the context, as used herein, the terms "one or more" or "at least one" can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Similarly, terms such as "one", "an" or "the" can also be understood to convey singular usage or to convey plural usage, which depends at least in part on the context. In addition, the term "based on" or "determined by..." can be understood to not necessarily be intended to convey an exclusive set of factors, and can alternatively allow the presence of additional factors that are not necessarily explicitly described, which also depends at least in part on the context.

[0054] Figure 3 A simplified block diagram of a communication system (300) according to an embodiment of the present disclosure is shown. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and a terminal device (320) interconnected via the network (350). Figure 3 In an embodiment, a first pair of terminal devices (310) and a terminal device (320) perform unidirectional data transmission. For example, the terminal device (310) may encode video data (e.g., a video picture stream captured by the terminal device (310)) for transmission to another terminal device (320) via a network (350). The encoded video data is transmitted in the form of one or more encoded video bitstreams. The terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to restore the video picture, and display the video picture based on the restored video data. The unidirectional data transmission may be implemented in a media service application, etc.

[0055] In another embodiment, the communication system (300) includes a second pair of a third terminal device (330) and a fourth terminal device (340) for performing bidirectional transmission of encoded video data, which can be implemented, for example, during a video conferencing application. For the bidirectional transmission of data, each of the third terminal device (330) and the fourth terminal device (340) can encode video data (e.g., a video picture stream captured by the terminal device) for transmission to the other of the third terminal device (330) and the fourth terminal device (340) through a network (350). Each of the third terminal device (330) and the fourth terminal device (340) can also receive the encoded video data transmitted by the other of the third terminal device (330) and the fourth terminal device (340), and can decode the encoded video data to restore the video picture, and can display the video picture on an accessible display device based on the restored video data.

[0056] exist Figure 3In the example of , terminal devices (310), (320), (330) and (340) can be implemented as servers, personal computers and smart phones, but the applicability of the basic principles of the present disclosure may not be limited to this. Implementations of the present disclosure can be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, etc. Network (350) represents any number or type of network that transmits encoded video data between terminal devices (310), (320), (330) and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) can exchange data in circuit switching, packet switching and / or other types of channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purposes of this discussion, unless otherwise explicitly explained in this article, the architecture and topology of the network (350) may be irrelevant to the operation of the present disclosure.

[0057] As examples of applications of the disclosed subject matter, Figure 4 The placement of the video encoder and video decoder in a video streaming environment is shown. The disclosed subject matter may be equally applicable to other video applications, including, for example: video conferencing, digital television broadcasting, gaming, virtual reality, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0058] The video streaming system may include a video capture subsystem (413), which may include a video source (401), such as a digital camera, for creating an uncompressed video picture or image stream (402). In an example, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. The video picture stream (402) is depicted as a thick line to emphasize the high amount of data when compared to the encoded video data (404) (or encoded video bitstream), and the video picture stream (402) may be processed by an electronic device (420) coupled to the video source (401) including a video encoder (403). The video encoder (403) may include hardware, software, or a combination thereof to implement or implement aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)) is depicted as a thin line to emphasize the lower data volume when compared to the uncompressed video picture stream (402), and the encoded video data (404) can be stored on the streaming server (405) for future use or directly to a downstream video device (not shown). One or more streaming client subsystems, such as Figure 4The client subsystems (406) and (408) in the video transmission server (405) can access the streaming server (405) to read the copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in the electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and creates an outgoing video picture stream (411) that is uncompressed and can be presented on a display (412) (e.g., a display screen) or other rendering device (not depicted). The video decoder 410 can be configured to perform some or all of the various functions described in the present disclosure. In some streaming transmission systems, the encoded video data (404), (407) and (409) (e.g., video bitstreams) can be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T H.265 Recommendation. In the example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC and other video coding standards.

[0059] Note that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0060] Figure 5 A block diagram of a video decoder (510) according to any embodiment of the present disclosure below is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., receiving circuitry). The video decoder (510) may be used instead of Figure 4 A video decoder (410) is shown in an example.

[0061] The receiver (531) can receive one or more encoded video sequences to be decoded by the video decoder (510). In the same embodiment or another embodiment, one encoded video sequence can be decoded at a time, wherein the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence can be associated with multiple video frames or images. The encoded video sequence can be received from a channel (501), which can be a hardware / software link to a storage device storing the encoded video data or a streaming source sending the encoded video data. The receiver (531) can receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams that can be forwarded to their respective processing circuit systems (not depicted). The receiver (531) can separate the encoded video sequence from the other data. In order to combat network jitter, a buffer memory (515) can be set between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) can be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be external to the video decoder (510) and separate from the video decoder (510) (not depicted). In still other applications, there may be a buffer memory (not depicted) external to the video decoder (510) to, for example, combat network jitter, and additional buffer memory (515) may be present within the video decoder (510) to, for example, handle playback timing. When the receiver (531) is receiving data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be required, or the buffer memory (515) may be small. For use over a best effort packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and the size of the buffer memory (515) may be relatively large. Such a buffer memory may be implemented to have an adaptive size, and may be implemented at least in part in an operating system or similar element (not depicted) external to the video decoder (510).

[0062] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510) and may include information for controlling a presentation device such as a display (512) (e.g., a display screen), which may or may not be part of the electronic device (530) but may be coupled to the electronic device (530), such as Figure 5. The control information for the rendering device may take the form of a supplementary enhancement information (SEI (Supplemental Enhancement Information) message) or a video usability information (VUI) parameter set fragment (not depicted). The parser (520) may parse / entropy decode the coded video sequence received by the parser (520). The entropy coding of the coded video sequence may conform to a video coding technique or a video coding standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a subgroup parameter set for at least one subgroup of the pixel subgroups in the video decoder from the coded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a group of pictures (Group of Pictures, GOP), a picture, a tile, a slice, a macroblock, a coding unit (Coding Unit, CU), a block, a transform unit (TransformUnit, TU), a prediction unit (Prediction Unit, PU), etc. The parser (520) may also extract information such as transform coefficients (eg, Fourier transform coefficients), quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0063] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515), thereby creating symbols (521).

[0064] Depending on the type of coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbol (521) may involve a number of different processing or functional units. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by the parser (520). For clarity, the flow of such subgroup control information between the parser (520) and the following multiple processing or functional units is not depicted.

[0065] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into a number of functional units as described below. In a practical implementation operating under commercial constraints, many of these functional units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, the conceptual subdivision into functional units is adopted in the following disclosure.

[0066] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives quantized transform coefficients as symbols (521) and control information from the parser (520), including information indicating which inverse transform to use, block size, quantization factors / parameters, quantization scaling matrix, etc. The scaler / inverse transform unit (551) may output blocks including sample values, which may be input into an aggregator (555).

[0067] In some cases, the output samples of the scaler / inverse transform unit (551) may be from an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture, but instead uses predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may use surrounding block information that has already been reconstructed and stored in a current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (555) may add the prediction information that has been generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.

[0068] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-coded and possibly motion compensated block. In such a case, the motion compensated prediction unit (553) may access the reference picture memory (557) to obtain samples for inter-picture prediction. After motion compensation of the obtained samples according to the symbol (521) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (551) (the output of unit 551 may be referred to as residual samples or residual signal) by the aggregator (555) to generate output sample information. The address within the reference picture memory (557) from which the motion compensated prediction unit (553) obtains the predicted samples may be controlled by a motion vector, which may be obtained by the motion compensated prediction unit (553) in the form of a symbol (521) that may have, for example, an X component, a Y component (offset) and a reference picture component (time). Motion compensation may also include interpolation of sample values ​​retrieved from a reference picture memory (557) when using sub-sample accurate motion vectors, and may also be associated with motion vector prediction mechanisms, etc.

[0069] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream) and obtained by the loop filter unit (556) as symbols (521) from the parser (520), but may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of the coded picture or coded video sequence, as well as to previously reconstructed and loop filtered sample values. Several types of loop filters may be included as part of the loop filter unit 556 in various orders, as will be described in further detail below.

[0070] The output of the loop filter unit (556) may be a sample stream that may be output to a rendering device (512) and stored in a reference picture memory (557) for future inter-picture prediction.

[0071] Certain coded pictures, once fully reconstructed, can be used as reference pictures for future inter-picture prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by a parser (520)), the current picture buffer (558) can become part of the reference picture memory (557) and a new current picture buffer can be reallocated before starting to reconstruct a subsequent coded picture.

[0072] The video decoder (510) may perform decoding operations according to a predetermined video compression technique adopted in a standard such as ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence follows both the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence is within the limits defined by the level of the video compression technology or standard. In some cases, the level may limit the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.

[0073] In some example embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The video decoder (510) may use the additional data to correctly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0074] Figure 6 A block diagram of a video encoder (603) according to an example implementation of the present disclosure is shown. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may also include a transmitter (640) (e.g., a transmission circuit system). The video encoder (603) may be used instead of Figure 4 The video encoder (403) in the example.

[0075] The video encoder (603) can capture video images to be encoded by the video encoder (603) from a video source (601) (in Figure 6 In one example, the video source (601) is not a part of the electronic device (620) and receives the video samples. In another example, the video source (601) can be implemented as a part of the electronic device (620).

[0076] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603), the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) may be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures or images that are given motion when viewed sequentially. The picture itself may be organized as a spatial pixel array, wherein each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by those skilled in the art. The following description focuses on samples.

[0077] According to some example embodiments, the video encoder (603) may encode and compress pictures of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed constitutes a function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to and control other functional units as described below. For clarity, the coupling is not depicted. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technique, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other suitable functions, which are related to the video encoder (603) optimized for a specific system design.

[0078] In some example embodiments, the video encoder (603) may be configured to operate in an encoding loop. As an extremely simplified description, in an example, the encoding loop may include: a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture); and a (local) decoder (633), which is embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder would create sample data, although the embedded decoder 633 processes the encoded video stream of the source encoder 630 without entropy coding (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video bitstream in entropy coding is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream results in a bit-exact result that is independent of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" as reference picture samples exactly the same sample values ​​as the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the creation of offsets in cases where synchronization cannot be maintained, e.g. due to channel errors) is also used to improve encoding quality.

[0079] The operation of the "local" decoder (633) may be the same as the operation of a "remote" decoder, such as the video decoder (510), described above in conjunction with Figure 5 The video decoder (510) is described in detail. However, reference is also briefly made to Figure 5, since the symbols are available and encoding of the symbols into a coded video sequence by the entropy encoder (645) and decoding of the symbols by the parser (520) can be lossless, the entropy decoding portion of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633) in the encoder.

[0080] At this point it can be observed that any decoder technology other than parsing / entropy decoding present in the decoder must also be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter sometimes focuses on the decoder operation, which is related to the decoding portion of the encoder. Since the encoder technology is opposite to the decoder technology that has been fully described, the description of the encoder technology can be simplified. A more detailed description of the encoder is provided below only in certain areas or aspects.

[0081] In some example implementations, during operation, the source encoder (630) may perform motion compensated predictive coding, which predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (632) encodes the differences (or residuals) of the color channels between pixel blocks of the input picture and pixel blocks of a reference picture, which may be selected as a prediction reference for the input picture.

[0082] The local video decoder (633) can decode the encoded video data of the picture that can be designated as the reference picture based on the symbols created by the source encoder (630). The operation of the encoding engine (632) can advantageously be lossy processing. When it is possible to Figure 6 When the encoded video data is decoded at a local video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors therein. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference picture, and may cause the reconstructed reference picture to be stored in the reference picture memory (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture that has common content (absent transmission errors) with the reconstructed reference picture that will be obtained by the far-end (remote) video decoder.

[0083] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that may be used as suitable prediction references for the new picture. The predictor (635) may operate on a sample block-by-pixel block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references obtained from multiple reference pictures stored in the reference picture memory (634).

[0084] The controller (650) may manage encoding operations of the source encoder (630), including, for example, setting of parameters and sub-group parameters for encoding video data.

[0085] The outputs of all the above functional units may be entropy encoded in an entropy encoder (645). The entropy encoder (645) converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing them according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0086] The transmitter (640) can buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission via the communication channel (660), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0087] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain encoding picture type to each encoded picture, which may affect the encoding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned one of the following picture types:

[0088] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of those variations of I pictures and their respective applications and features.

[0089] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, where inter prediction uses at most one motion vector and a reference index to predict sample values ​​for each block.

[0090] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, the inter prediction uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0091] The source picture can typically be spatially subdivided into a plurality of sample coding blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and encoded block by block. The blocks can be predictively encoded with reference to other (already encoded) blocks determined by the coding allocation applied to the corresponding picture of the block. For example, blocks of an I picture can be non-predictively encoded, or blocks of an I picture can be predictively encoded (spatial prediction or intra-frame prediction) with reference to coding blocks of the same picture. Pixel blocks of a P picture can be predictively encoded via spatial prediction or via temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be predictively encoded via spatial prediction or via temporal prediction with reference to one or two previously encoded reference pictures. For other purposes, the source picture or intermediate processed pictures can be subdivided into other types of blocks. The division of coding blocks and other types of blocks may or may not follow the same approach, as will be further described in detail below.

[0092] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In its operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.

[0093] In some example embodiments, the transmitter (640) may send additional data along with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0094] Video can be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits temporal or other correlations between pictures. For example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case of using multiple reference pictures, the motion vector may have a third dimension that identifies the reference picture.

[0095] In some example embodiments, bidirectional prediction techniques may be used for inter-picture prediction. According to such bidirectional prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be jointly predicted by a combination of the first reference block and the second reference block.

[0096] In addition, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0097] According to some example embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, a picture in a video picture sequence is divided into coding tree units (CTUs) for compression, and the CTUs in the picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU may include three parallel coding tree blocks (CTBs): a luminance CTB and two chrominance CTBs. Each CTU may be recursively divided into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU may be divided into a 64×64 pixel CU, or four 32×32 pixel CUs. Each of one or more of the 32×32 blocks may also be divided into four 16×16 pixel CUs. In some example embodiments, each CU is analyzed during encoding to determine the prediction type of the CU among various prediction types, such as an inter-prediction type or an intra-prediction type. Depending on temporal and / or spatial predictability, a CU may be partitioned into one or more prediction units (PUs). Typically, each PU includes a luminance prediction block (PB) and two chrominance PBs. In some embodiments, prediction operations in encoding (encoding / decoding) are performed in units of prediction blocks. Partitioning a CU into PUs (or PBs for different color channels) may be performed in various spatial modes. A luminance or chrominance PB may, for example, include a matrix of values ​​(e.g., luminance values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and the like.

[0098] Figure 7 A diagram of a video encoder (703) according to another example implementation of the present disclosure is shown. The video encoder (703) is configured to receive a processed block (e.g., a prediction block) of sample values ​​within a current video picture in a sequence of video pictures and encode the processed block into an encoded picture that is part of an encoded video sequence. The example video encoder (703) may be used instead of Figure 4 The video encoder (403) in the example.

[0099] For example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) then uses, for example, rate-distortion optimization (RDO) to determine whether to use intra mode, inter mode, or bidirectional prediction mode to best encode the processing block. When it is determined that the processing block is to be encoded in intra mode, the video encoder (703) can encode the processing block into a coded picture using intra prediction techniques; and when it is determined that the processing block is to be encoded in inter mode or bidirectional prediction mode, the video encoder (703) can encode the processing block into a coded picture using inter prediction techniques or bidirectional prediction techniques, respectively. In some example embodiments, a merge mode can be used as an inter-picture prediction sub-mode, wherein a motion vector is derived from one or more motion vector predictors without the aid of an encoded motion vector component external to the prediction value. In some other example embodiments, there can be a motion vector component applicable to the object block. Therefore, the video encoder (703) can include Figure 7 Components not explicitly shown, such as a mode decision module for determining a prediction mode for a processing block.

[0100] exist Figure 7 In the example of FIG. 7 , the video encoder ( 703 ) includes Figure 7 The example arrangement shown in FIG. 1 includes an inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together.

[0101] The inter-frame encoder (730) is configured to: receive samples of a current block (e.g., a processing block); compare the block to one or more reference blocks in a reference picture (e.g., a block in a previous picture and a block in a subsequent picture in display order); generate inter-frame prediction information (e.g., motion vectors, merge mode information, description of redundant information according to an inter-frame coding technique); and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference picture is based on the encoded video information using an embedded Figure 6 The decoding unit 633 in the example encoder of Figure 7 The decoded reference picture is decoded by the residual decoder 728).

[0102] The intra encoder (722) is configured to: receive samples of a current block (e.g., a processing block); compare the block with an already encoded block in the same picture; and generate quantization coefficients after transformation; and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). The intra encoder (722) can calculate an intra prediction result (e.g., a predicted block) based on the intra prediction information and a reference block in the same picture.

[0103] The general controller (721) can be configured to determine general control data and control other components of the video encoder (703) based on the general control data. In an example, the general controller (721) determines a prediction mode of a block and provides a control signal to a switch (726) based on the prediction mode. For example, when the prediction mode is an intra-frame mode, the general controller (721) controls the switch (726) to select an intra-frame mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select intra-frame prediction information and include the intra-frame prediction information in the bitstream; and when the prediction mode of the block is an inter-frame mode, the general controller (721) controls the switch (726) to select an inter-frame prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select inter-frame prediction information and include the inter-frame prediction information in the bitstream.

[0104] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and the prediction result of the block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) may be configured to encode the residual data to generate a transform coefficient. For example, the residual encoder (724) may be configured to convert the residual data from the spatial domain to the frequency domain to generate a transform coefficient. Then, the transform coefficient is quantized to obtain a quantized transform coefficient. In various example embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) may generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) may generate a decoded block based on the decoded residual data and the intra prediction information. The decoded blocks are appropriately processed to generate decoded pictures, and these decoded pictures may be buffered in a memory circuit (not shown) and used as reference pictures.

[0105] The entropy encoder (725) may be configured to format the bitstream to include the coded blocks and perform entropy coding. The entropy encoder (725) may be configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. When the block is encoded in the merge sub-mode of the inter-frame mode or the bidirectional prediction mode, there is no residual information.

[0106] Figure 8 A diagram of a video decoder (810) according to another embodiment of the present disclosure is shown. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an example, the video decoder (810) may be used instead of Figure 4 A video decoder (410) is shown in an example.

[0107] exist Figure 8 In this example, the video decoder (810) includes Figure 8 An entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) coupled together are shown in the example arrangement of.

[0108] The entropy decoder (871) may be configured to reconstruct certain symbols from the coded picture, which represent syntax elements constituting the coded picture. Such symbols may include, for example, a mode for encoding a block (e.g., intra-mode, inter-mode, bidirectional prediction mode, merge sub-mode, or another sub-mode), prediction information (e.g., intra-prediction information or inter-prediction information) that may identify certain samples or metadata used by an intra-frame decoder (872) or an inter-frame decoder (880) for prediction, for example, residual information in the form of quantized transform coefficients, etc. In an example, when the prediction mode is an inter-frame mode or a bidirectional prediction mode, the inter-frame prediction information is provided to the inter-frame decoder (880); and when the prediction type is an intra-frame prediction type, the intra-frame prediction information is provided to the intra-frame decoder (872). The residual information may be inverse quantized and provided to the residual decoder (873).

[0109] The inter-frame decoder (880) may be configured to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.

[0110] The intra decoder (872) may be configured to receive intra prediction information and generate prediction results based on the intra prediction information.

[0111] The residual decoder (873) may be configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize certain control information (including quantizer parameters (Quantizer Parameter, QP)), which may be provided by the entropy decoder (871) (since this may only be a small amount of control information, the data path is not depicted).

[0112] The reconstruction module (874) may be configured to combine the residual output by the residual decoder (873) with the prediction result (output by the inter-frame prediction module or the intra-frame prediction module, as the case may be) in the spatial domain to form a reconstructed block, which forms part of a reconstructed picture, which in turn may be part of a reconstructed video. Note that other suitable operations such as deblocking operations may also be performed to improve visual quality.

[0113] Note that the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using any suitable technology. In some example embodiments, the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using one or more processors executing software instructions.

[0114] Returning to the following intra-frame prediction process: samples in a block (e.g., a luminance prediction block or a chrominance prediction block, or a coding block if it is not further partitioned into prediction blocks) are predicted by samples in adjacent, next adjacent, or other one or more lines, or a combination thereof, to generate a prediction block. The residual between the actual block being encoded and the prediction block can then be processed by quantization after transformation. Various intra-frame prediction modes can be made available, and parameters related to intra-frame mode selection and other parameters can be signaled in the bitstream. Various intra-frame prediction modes can be related to, for example, one or more line positions for prediction samples, the direction of selecting prediction samples from one or more prediction lines along them, and other special intra-frame prediction modes.

[0115] For example, a set of intra-prediction modes (interchangeably referred to as "intra-modes") may include a predetermined number of directional intra-prediction modes. As described above with respect to the example implementation of FIG. 1, these intra-prediction modes may correspond to a predetermined number of directions along which out-of-block samples are selected as predictions for samples predicted in a particular block. In another specific example implementation, eight (8) primary directional modes corresponding to angles from 45 degrees to 207 degrees from the horizontal axis may be supported and predefined.

[0116] In some other implementations of intra prediction, in order to further exploit more variations of spatial redundancy in directional textures, the directional intra mode can also be extended to a set of angles with finer granularity. For example, the 8 angle implementation above can be configured to provide 8 nominal angles, referred to as V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, as Fig. 9 As shown, and for each nominal angle, a predefined number (e.g., 7) of finer angles can be added. With such an extension, a larger total number of directional angles (e.g., 56 in this example) can be used for intra prediction to correspond to the same number of predefined directional intra modes. The predicted angle can be represented by the nominal intra angle plus the angle δ. For the specific example above with 7 finer angle directions for each nominal angle, the angle δ can be a step size of -3 to 3 times 3 degrees.

[0117] The above-mentioned directional intra prediction may also be referred to as unidirectional intra prediction, which is different from the bidirectional intra prediction (also referred to as bidirectional intra prediction) described in the later part of this disclosure.

[0118] In some implementations, as an alternative to or in addition to the above-mentioned directional intra-frame modes, a number of non-directional intra-frame prediction modes can be predefined and made available. For example, 5 non-directional intra-frame modes called smooth intra-frame prediction modes can be specified. These non-directional intra-frame mode prediction modes can be specifically called DC, PAETH, SMOOTH, SMOOTH_V and SMOOTH_H intra-frame modes. Fig.10 The prediction of samples for a particular block in the non-directional mode of these examples is shown in FIG. As an example, Fig.10A 4×4 block 1002 is shown predicted by samples from a top neighboring line and / or a left neighboring line. A particular sample 1010 in the block 1002 may correspond to a sample 1004 directly above the sample 1010 in the top neighboring line of the block 1002, a sample 1006 to the upper left of the sample 1010 as the intersection of the top neighboring line and the left neighboring line, and a sample 1008 directly to the left of the sample 1010 in the left neighboring line of the block 1002. For the example DC intra prediction mode, an average of the left neighboring sample 1008 and the upper neighboring sample 1004 may be used as a prediction value for the sample 1010. For the example PAETH intra prediction mode, top, left, and upper left reference samples 1004, 1008, and 1006 may be obtained, and then the value closest to (top+left-upper left) of the three reference samples may be set as the prediction value for the sample 1010. For the example SMOOTH_V intra prediction mode, sample 1010 may be predicted by quadratic interpolation in the vertical direction of the upper left neighboring sample 1006 and the left neighboring sample 1008. For the example SMOOTH_H intra prediction mode, sample 1010 may be predicted by quadratic interpolation in the horizontal direction of the upper left neighboring sample 1006 and the top neighboring sample 1004. For the example SMOOTH intra prediction mode, sample 1010 may be predicted by the average of the quadratic interpolation in the vertical and horizontal directions. The above non-directional intra mode implementations are illustrated as non-limiting examples only. Other neighboring lines are also contemplated, as well as other non-directional selections of samples, and ways of combining prediction samples to predict specific samples in a prediction block.

[0119] The selection of a specific intra-frame prediction mode from the above-mentioned directional modes or non-directional modes by the encoder at various coding levels (picture, slice, block, unit, etc.) can be signaled in the bitstream. In some exemplary implementations, the exemplary 8 nominal directional modes can be first signaled together with 5 non-angular smoothing modes (a total of 13 choices). Then, if the signaled mode is one of the 8 nominal angular intra-frame modes, an index is further signaled to indicate the selected angle δ to the corresponding nominal angle signaled. In some other exemplary implementations, all intra-frame prediction modes can be indexed together (e.g., 56 directional modes plus 5 non-directional modes to produce 61 intra-frame prediction modes) for signaling.

[0120] In some example implementations, the example 56 or other number of directional intra prediction modes may be implemented with a unified directional predictor that projects each sample of the block to a reference subsample location and interpolates the reference sample via a 2-tap bilinear filter.

[0121] In some implementations, in order to capture the attenuated spatial correlation with the reference on the edges, additional filtering modes called FILTER INTRA modes can be designed. For these modes, in addition to the samples outside the block, the predicted samples within the block can also be used as intra-frame prediction reference samples for certain tiles within the block. For example, these modes can be predefined and can be used for intra-frame prediction of at least luminance blocks (or only luminance blocks). A predefined number (for example, 5) of filter intra-frame modes can be pre-designed, each mode being represented by a set of n-tap filters (for example, 7-tap filters) that reflect the correlation between a sample in, for example, a 4×2 tile and its n adjacent neighboring samples. In other words, the weighting coefficients of the n-tap filter may be position-dependent. Taking 8×8 blocks, 4×2 tiles, and 7-tap filtering as an example, as Fig.11 As shown, the 8×8 block 1102 can be divided into eight 4×2 tiles. Fig.11 In the figure, the blocks are indicated by B0, B1, B1, B3, B4, B5, B6 and B7. Fig.11 The samples in the current block are predicted from the seven neighboring blocks indicated by R0 to R7 in the image. For block B0, all neighboring blocks may have been reconstructed. However, for other blocks, some of the neighboring blocks are in the current block and may not have been reconstructed yet. The predicted values ​​of the adjacent neighboring blocks are used as references. For example, Fig.11 All neighboring tiles of the illustrated tile B7 are not reconstructed, so prediction samples of neighboring tiles are used instead, such as a portion of B4, B5 and / or B6.

[0122] In some implementations of intra prediction, one or more other color components may be used to predict a color component. A color component may be any one of YCrCb, RGB, XYZ color spaces, etc. For example, it may be possible to predict a chroma component (e.g., a chroma block) from a luma component (e.g., a luma reference sample) (referred to as predicting chroma from luma, or CfL (Chroma from Luma)). In some example implementations, cross-color prediction is often only allowed from luma to chroma. For example, chroma samples in a chroma block may be modeled as a linear function of coincident reconstructed luma samples. CfL prediction may be implemented as follows:

[0123] CfL(α)=α×L AC +DC (1)

[0124] Where L ACRepresents the AC contribution of the luma component, α represents the parameters of the linear model, and DC represents the DC contribution of the chroma component. The AC component is obtained, for example, for each sample of the block, while the DC component is obtained for the entire block. Specifically, the reconstructed luma samples can be subsampled to the chroma resolution, and then the average luma value (the DC of luma) is subtracted from each luma value to form the AC contribution of the luma. The AC contribution of the luma is then used in the linear mode of equation (1) to predict the AC value of the chroma component. In order to approximate or predict the chroma AC component from the luma AC contribution, instead of requiring the decoder to calculate the scaling parameters, the example CfL implementation can determine the parameter α based on the original chroma samples and signal it in the bitstream. This reduces the complexity of the decoder and produces more accurate predictions. As for the DC contribution of the chroma component, in some example implementations, it can be calculated using the intra-frame DC mode within the chroma component.

[0125] Back to intra prediction, in some example implementations, sample prediction in a coding block or prediction block can be based on one reference line in a set of reference lines. In other words, instead of always using the nearest line (e.g., the top adjacent line or the left adjacent line of the prediction block shown in FIG. 1 above), it is better to provide multiple reference lines as options for selecting intra prediction. Such intra prediction implementations may be referred to as multiple reference line selection (MRLS). In these implementations, the encoder determines and signals which reference line of multiple reference lines is used to generate intra prediction values. On the decoder side, after parsing the reference line index, the reconstructed reference sample can be identified by looking up the specified reference line according to the intra prediction mode (such as directional, non-directional and other intra prediction modes), thereby generating the intra prediction of the current intra prediction block. In some implementations, the reference line index can be signaled at the coding block level, and only one of the multiple reference lines can be selected and used for intra prediction of a coding block. In some examples, more than one reference line can be selected together for intra prediction. For example, more than one reference line may be combined, averaged, interpolated, or in any other manner, weighted or unweighted, to generate a prediction.In some example implementations, MRLS may be applied only to luma components, and not to chroma components.

[0126] exist Fig.12 In , an example of 4 reference lines MRLS is described. Fig.12As shown in the example, the intra-coded block 1202 can be predicted based on one of four horizontal reference lines 1204, 1206, 1208 and 1210 and four vertical reference lines 1212, 1214, 1216 and 1218. Among these reference lines, 1210 and 1218 are immediately adjacent reference lines. These reference lines can be indexed according to their distance from the coding block. For example, reference lines 1210 and 1218 can be referred to as zero reference lines, while other reference lines can be referred to as non-zero reference lines. Specifically, reference lines 1208 and 1216 can be referred to as first reference lines; reference lines 1206 and 1214 can be referred to as second reference lines; and reference lines 1204 and 1212 can be referred to as third reference lines.

[0127] In addition to the unidirectional intra prediction described in the previous part of this disclosure, reference pixels along two prediction directions may be combined to obtain directional prediction values, which may be referred to as bidirectional intra prediction or intra bi-prediction (IBP).

[0128] In some implementations, for a current block (eg, a coding block or an encoded block), when the directional pattern has an angle less than 90 degrees (eg, but not limited to Fig. 9 In some other implementations, when the direction mode has an angle greater than 180 degrees (for example, but not limited to, Fig. 9 IBP can also be applied to this direction mode when D203_PRED is selected in the direction mode.

[0129] When bidirectional intra prediction is applied, reference pixels from two directions along the directional pattern are selected: one reference pixel is from the top or upper right of the current block, and the other pixel is from the left or lower left of the current block. A weighted average of the two reference pixels can be calculated to achieve the prediction value.

[0130] Fig.13A An example of an IBP for a coding block (1330) is shown, whose directional pattern has a prediction direction (1340) from A (1322) to B (1312). A and B are two reference samples / values, which are also referred to as first prediction values ​​and second prediction values. A is located at the intersection of direction (1340) and the top reference line (1320); and B is located at the intersection of direction (1340) and the left reference line (1310). The prediction of a pixel (x, y) (1332) as in the coding block (1330) is represented as pred(x, y). pred(x, y) can be generated using a weighted combination of two prediction values ​​A and B, as shown in equation (2).

[0131] Pred(x,y)=w*A+(1-w)*B(2)

[0132] A and / or B may be derived using a directional prediction process that includes interpolation of fractional pixel references. Fig. 13B Another example of an IBP for a coding block (1330) is shown with different directional modes for different prediction directions (1350) from A (1314) to B (1324).

[0133] In some implementations of applying IBP, two reference pixels located along the adjacent upper reference line and the adjacent left reference in that direction are weighted to obtain the prediction value. If both IBP and multiple reference line selection (MRLS) are applied to the coding block at the same time, some problems / difficulties may arise. For example, how to apply IBP to the current block when reference samples in one or more non-adjacent reference lines are selected.

[0134] The present disclosure describes various implementations for signaling and / or determining multi-reference line intra prediction in video encoding and / or decoding, thereby addressing at least one of the issues / problems discussed above.

[0135] In some embodiments, reference Fig.14 , a method 1400 for bidirectional intra prediction and multi-reference line intra prediction in video decoding. The method 1400 may include some or all of the following steps: step 1410, receiving a coded video bitstream of a block, the coded video bitstream including mode information of the block; step 1420, determining whether unidirectional intra prediction or bidirectional intra prediction is applicable to the block based on the mode information of the block, the mode information of the block including at least one of a reference line index of the block, an intra prediction mode of the block, and a size of the block; step 1430, if it is determined that unidirectional intra prediction is applicable to the block, performing unidirectional intra prediction on the block; and / or step 1440, if it is determined that bidirectional intra prediction is applicable to the block, performing bidirectional intra prediction on the block.

[0136] In some implementations, when the intra-frame prediction mode does not belong to one of the smooth modes as described above, or when the intra-frame prediction mode generates one or more prediction samples according to a given prediction direction, the intra-frame prediction mode can be classified as one of the directional modes, and / or the intra-frame prediction mode can be referred to as a directional intra-frame prediction mode (or directional mode).

[0137] In some embodiments of the present disclosure, the size of a block (such as but not limited to a coding block, a prediction block, or a transform block) may refer to the width or height of the block. The width or height of the block may be an integer in pixels.

[0138] In some embodiments of the present disclosure, the size of a block (such as but not limited to a coding block, a prediction block, or a transform block) may refer to the area size of the block. The area size of the block may be an integer calculated in pixels by multiplying the width of the block by the height of the block.

[0139] In some various embodiments of the present disclosure, the size of a block (such as but not limited to a coding block, a prediction block, or a transform block) may refer to a maximum value of a width or a height of the block, a minimum value of a width or a height of the block, or an aspect ratio of the block. The aspect ratio of the block may be calculated as the width of the block divided by the height of the block, or may be calculated as the height of the block divided by the width of the block.

[0140] In the present disclosure, a reference line index indicates a reference line among multiple reference lines. In some embodiments, a reference line index of a block of 0 may indicate a reference line adjacent to the block, which is also the reference line closest to the block. Fig.12 , the top reference line (1210) is the top reference line adjacent to the block (1202), which is also the top reference line closest to the block; and the left reference line (1218) is the left reference line adjacent to the block (1202), which is also the left reference line closest to the block. The reference line index of the block is greater than 0, indicating a non-adjacent reference line of the block, which is also the non-closest reference line of the block. For example, reference Fig.12 In block (1202), a reference line index of 1 can indicate a top reference line (1208) and / or a left reference line (1216); a reference line index of 2 can indicate a top reference line (1206) and / or a left reference line (1214); and / or a reference line index of 3 can indicate a top reference line (1204) and / or a left reference line (1212).

[0141] Referring to step 1410, the device may be Figure 5 An electronic device (530) or Figure 8 In some implementations, the device may be a video decoder (810) in Figure 6 In other implementations, the device may be Figure 5 a portion of an electronic device (530) in Figure 8 a portion of a video decoder (810) in, or Figure 6 The encoded video bitstream may be a portion of a decoder (633) in an encoder (620) in FIG. Figure 8 A coded video sequence in Figure 6 or Figure 7 The block can refer to the intermediate coded data in the . This block can refer to the coding block or the coded block.

[0142] In some implementations, when a directional intra prediction mode is selected for a block to generate intra prediction values ​​from samples at non-adjacent reference lines, it may be determined whether to apply unidirectional intra prediction or bidirectional intra prediction to the block, and the determination may depend on mode information of the block. In some implementations, the mode information of the block may include, but is not limited to, a reference line index (e.g., indicating which reference line of a plurality of reference lines), one or more intra prediction angles (e.g., indicating which directional intra prediction mode of a plurality of directional intra prediction modes), and / or a size of the block.

[0143] Referring to step 1420, it may be determined based on the mode information of the block whether unidirectional intra prediction or bidirectional intra prediction is applicable to the block. In some implementations, step 1420 may include, in response to the reference line index of the block indicating a non-adjacent reference line, determining that unidirectional intra prediction is applicable to the block. As an example, unidirectional intra prediction may be applied to non-adjacent reference lines without considering the intra prediction angle of the current block.

[0144] In some embodiments, step 1420 may include, in response to the reference line index of the block indicating a neighboring reference line, determining that bidirectional intra prediction is applicable to the block. As an example, bidirectional intra prediction may only be applicable to neighboring reference lines of the current block.

[0145] In some embodiments, step 1420 may include, in response to a reference line index of the block being greater than a predefined threshold, determining that unidirectional intra prediction is applicable to the block; and / or in response to a reference line index of the block being not greater than a predefined threshold, determining that bidirectional intra prediction is applicable to the block. As an example, determining whether to apply unidirectional intra prediction or bidirectional intra prediction to non-adjacent reference lines, and the determination also depends on whether the value of the reference line index is greater than a predefined value N, where N is a non-negative integer, such as 0, 1, 2, 3, or 4.

[0146] In one example, the predefined value N is 0. When the value of the reference line index of the block is greater than 0, that is, when any non-adjacent reference line (e.g., any line of 1212, 1214, 1216, 1204, 1206 and / or 1208) is used for the block, unidirectional intra prediction is applicable to the block; and / or when the value of the reference line index of the block is not greater than 0, that is, when an adjacent reference line (e.g., Fig.12 When any line in 1218 and / or 1210) is used for the block, bidirectional intra prediction is applicable to the block.

[0147] In another example, the predefined value N is 1. When the value of the reference line index of the block is greater than 1, that is, when any one of the first subset of non-adjacent reference lines (e.g., any one of lines 1212, 1214, 1204, and / or 1206) is used for the block, the unidirectional intra prediction is applicable to the block; and / or when the value of the reference line index of the block is not greater than 1, that is, when the adjacent reference lines (e.g., Fig.12 1218 and / or any one of 1210 in ) and / or any one of the second subset of non-adjacent reference lines (e.g., Fig.12 When any one of lines 1208 and / or 1216 in (a) is used for the block, bidirectional intra prediction is applicable to the block.

[0148] In some embodiments, step 1420 may include, in response to the intra-frame prediction mode of the block belonging to a selected first set of intra-frame prediction modes and the reference line index indicating an adjacent reference line, determining that bidirectional intra-frame prediction is applicable to the block; in response to the intra-frame prediction mode of the block belonging to a selected first set of intra-frame prediction modes and the reference line index indicating a non-adjacent reference line, determining that unidirectional intra-frame prediction is applicable to the block; and / or in response to the intra-frame prediction mode of the block belonging to a selected second set of intra-frame prediction modes, determining that bidirectional intra-frame prediction is applicable to the block. In some implementations, the selected first set of intra-frame prediction modes does not overlap with the selected second set of intra-frame prediction modes. In some other implementations, the selected second set of intra-frame prediction modes may include a diagonal intra-frame prediction mode.

[0149] In some implementations, for a selected set of intra-frame prediction modes, IBP is applied to both adjacent reference lines and non-adjacent reference lines, while for the remaining set of intra-frame prediction modes, IBP is only applied to adjacent reference lines. In some other implementations, the selected set of intra-frame prediction modes may include several intra-frame prediction modes with certain specific direction angles, so that only integer samples are used as reference values ​​for intra-frame prediction. For one example, for some intra-frame prediction modes with diagonal directions (i.e., 45 degrees or 225 degrees), only integer samples are used as reference values. For another example, when intra-frame prediction is applied by using only integer samples for a diagonal intra-frame prediction mode, IBP may also be applied to non-adjacent reference lines, wherein the diagonal intra-frame prediction mode is an intra-frame prediction mode with a diagonal direction (i.e., 45 degrees or 225 degrees).

[0150] In some embodiments, step 1420 may include, in response to the reference line index of the block being an odd integer and / or the reference line index indicating a non-adjacent reference line, determining that unidirectional intra prediction is applicable to the block; and / or in response to the reference line index of the block being an even integer and / or the reference line index indicating an adjacent reference line, determining that bidirectional intra prediction is applicable to the block.

[0151] In some implementations, a determination is made whether to apply unidirectional intra prediction or bidirectional intra prediction to a non-adjacent reference line, and the determination depends on whether the value of the reference line index is an even number or an odd number. For one example, when the reference line index value of the block is an odd number, unidirectional intra prediction is applied to the block; and / or when the reference line index value of the block is an even number, bidirectional intra prediction is applied to the block. For another example, when the reference line index value of the block is an even number, unidirectional intra prediction is applied to the block; and / or when the reference line index value of the block is an odd number, bidirectional intra prediction is applied to the block.

[0152] In some embodiments, there may be only one reference line index for a block. During encoding, the reference line index of the block may be encoded into the encoded bitstream; and / or during decoding, the reference line index of the block may be decoded / extracted from the encoded bitstream. Step 1440 may include determining a first prediction value based on the reference line index of the block; determining a second prediction value based on the reference line index of the block; and / or determining a final prediction value for the block based on a weighted calculation between the first prediction value and the second prediction value.

[0153] In some implementations, when bidirectional intra prediction is applied and the reference line index of the block indicates the use of non-adjacent reference lines for intra prediction, in order to generate an intra prediction value from samples at the non-adjacent reference lines, both a prediction value A and a prediction value B are generated from samples at the non-adjacent reference lines, and then a weighted calculation is performed based on the prediction value A and the prediction value B to generate a final IBP prediction value.

[0154] In some embodiments, there may be more than one reference line index for a block. A first reference line index may indicate a reference line used among multiple top reference lines, and a second reference line index may indicate a reference line used among multiple left reference lines; or vice versa, a first reference line index may indicate a reference line used among multiple left reference lines, and a second reference line index may indicate a reference line used among multiple top reference lines. During encoding, more than one reference line index for a block may be encoded into a coded bitstream; and / or during decoding, more than one reference line index for a block may be decoded / extracted from a coded bitstream.

[0155] In some implementations, step 1440 may include determining a first prediction value based on a first reference line index of the block; determining a second prediction value based on a second reference line index of the block; and / or determining a final prediction value of the block based on a weighted calculation between the first prediction value and the second prediction value.

[0156] In some other implementations, the first reference line index indicates a non-adjacent reference line, and the first prediction value is generated based on a first sample at the non-adjacent reference line; and / or the second reference line index indicates an adjacent reference line, and the second prediction value is generated based on a second sample at the adjacent reference line.

[0157] In reference Fig.15 In one example, when bidirectional intra prediction is applied to a coding block (1530), an intra prediction value (1532) is generated from samples at two reference lines including a non-adjacent reference line and / or one or more adjacent reference lines. Fig.15 As shown, the coding block has a left adjacent reference line (2610), a top adjacent reference line (2620), and at least one top adjacent reference line (1521 and / or 1522). For example, but not limited to, the first reference line index of 2 can indicate the top non-adjacent reference line (1522), so the prediction value A (1528) is generated from the sample at the non-adjacent reference line (2622) selected by the MRLS according to the prediction direction (1550). In some implementations, the second reference line index of 0 can indicate the left adjacent reference line (1510). In some other implementations, the left adjacent reference line (1510) can be indicated / implied without the second reference line index. The prediction value B (1514) is generated from the sample at the adjacent reference line (1510) according to the prediction direction (1550). The prediction value A and the prediction value B are then weighted and combined according to the IBP scheme (for example, according to the formula of equation (2)) to generate the intra-frame prediction value (1532). In some implementations, the weight (w) may depend on the distance between the intra-frame prediction value (1532) and the prediction value A or the prediction value B. In equation (2), the smaller the distance between the intra-frame prediction value (1532) and the prediction value A (1528) (i.e., the closer the prediction value (1532) is to the prediction value A (1528)), the larger the weight (w). Similarly, in equation (2), the larger the distance between the intra-frame prediction value (1532) and the prediction value A (1528) (i.e., the farther the prediction value (1532) is from the prediction value A (1528)), the smaller the weight (w).

[0158] Fig.15 An example with a single left reference line and multiple top reference lines is shown; and similarly, in another example, when bidirectional intra prediction is applied to a coding block, a single top reference line and multiple left reference lines may be applied to the coding block.

[0159] In some embodiments, there may be more than one reference line index for a block. A first reference line index may indicate a reference line used among multiple top reference lines, and a second reference line index may indicate more than one reference line used among multiple left reference lines; or vice versa, a first reference line index may indicate a reference line used among multiple left reference lines, and a second reference line index may indicate more than one reference line used among multiple top reference lines. During encoding, more than one reference line index for a block may be encoded into the encoded bitstream; and / or during decoding, more than one reference line index for a block may be decoded / extracted from the encoded bitstream.

[0160] In some implementations, step 1440 may include, the first reference line index indicating a non-adjacent reference line, and the first prediction value is generated based on a first sample at the non-adjacent reference line; and / or the second reference line index indicating multiple reference lines, and the second prediction value is generated as a linear weighted average of multiple samples, wherein each sample comes from a reference line along a prediction angle among the multiple reference lines.

[0161] In some other implementations, when bidirectional intra prediction is applied to generate intra prediction values ​​from samples at non-adjacent reference lines, samples at multiple (more than 1) reference lines are used to generate one or two prediction values.

[0162] In reference Fig.16 In one example, a linear weighted sum of more than one samples from multiple reference lines along the prediction angle (1650) is used to generate a prediction value: a first sample (1626) from a first reference line (1620), a second sample (1627) from a second reference line (1621), and a third sample (1628) from a third reference line (1622). The weight of each sample can be predefined according to the relative position of the samples (of the multiple reference lines) participating in the derivation of the prediction value.

[0163] In some embodiments, method 1400 may further include: in response to determining that bidirectional intra prediction is applicable to a block and a reference line index of the block indicates a non-adjacent reference line: disabling reference sample filtering processing, in response to reference samples from a non-adjacent reference being available, using the reference samples for bidirectional intra prediction, and / or in response to reference samples from a non-adjacent reference being unavailable, using adjacent available reference samples for bidirectional intra prediction.

[0164] In some implementations, when bidirectional intra prediction is applied to one or more non-adjacent reference lines, reference sample filtering processing may be disabled. Thus, one or more reference samples from one or more non-adjacent (or non-zero) reference lines may be acquired and used directly for bidirectional intra prediction. If one or more reference samples from one or more non-adjacent (or non-zero) reference lines are unavailable, they may be filled from one or more adjacent available reference samples. In some other implementations, filling of unavailable reference samples may include copying from adjacent available reference samples.

[0165] The embodiments of the present disclosure may be used alone or in any order. In addition, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The embodiments of the present disclosure may be applied to luminance blocks or chrominance blocks; and in chrominance blocks, the embodiments may be applied to more than one color component, respectively, or may be applied to more than one color component together.

[0166] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.17 A computer system (2600) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0167] Computer software may be encoded using any suitable machine code or computer language, which may be assembled, compiled, linked, or the like to create code comprising instructions, which may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or the like, or through interpretation, microcode execution, or the like.

[0168] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0169] Fig.17 The components for the computer system (2600) shown in the example are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Neither should the configuration of components be interpreted as having any dependency or requirement relating to any one or combination of components shown in the exemplary embodiment of the computer system (2600).

[0170] The computer system (2600) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, tapping), visual input (e.g., gestures), olfactory input (not shown). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0171] The input human-machine interface devices may include one or more of the following (only one of each described): keyboard (2601), mouse (2602), touchpad (2603), touch screen (2610), data gloves (not shown), joystick (2605), microphone (2606), scanner (2607), camera (2608).

[0172] The computer system (2600) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate one or more senses of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (2610), a data glove (not shown), or a joystick (2605), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (2609), headphones (not shown)), visual output devices (e.g., screens (2610), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may be capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereoscopic image output; virtual reality glasses (not depicted), holographic displays, and smoke generators (not depicted)) and printers (not depicted).

[0173] The computer system (2600) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2620) with CD / DVD etc. media (2621), thumb drives (2622), removable hard drives or solid-state drives (2623), legacy magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD based devices such as security dongles (not depicted), and the like.

[0174] Those skilled in the art will also appreciate that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0175] The computer system (2600) may also include an interface (2654) to one or more communication networks (2655). The network may be, for example, a wireless network, a wired network, an optical network. The network may also be a local area network, a wide area network, a metropolitan area network, an in-vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include: local area networks (e.g., Ethernet, wireless LAN), cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired connections or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CAN buses, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (2649) (e.g., a USB port of the computer system (2600)); other networks are typically integrated into the core of the computer system (2600) by attaching to the system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smart phone computer system). Using any of these networks, the computer system (2600) can communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., a CAN bus to certain CAN bus devices), or bidirectional (e.g., to other computer systems using a local area digital network or a wide area digital network). Certain protocols and protocol stacks may be used on each of these networks and network interfaces as described above.

[0176] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (2640) of the computer system (2600).

[0177] The core (2640) may include one or more central processing units (CPUs) (2641), graphics processing units (GPUs) (2642), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (2643), hardware accelerators (2644) for certain tasks, graphics adapters (2650), etc. These devices, along with read-only memory (ROM) (2645), random access memory (2646), internal mass storage devices (e.g., internal non-user accessible hard drives, SSDs, etc.) (2647) may be connected via a system bus (2648). In some computer systems, the system bus (2648) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus (2648) directly or via a peripheral bus (2649). In an example, a screen (2610) may be connected to a graphics adapter (2650). The architecture of the peripheral bus includes PCI, USB, etc.

[0178] The CPU (2641), GPU (2642), FPGA (2643) and accelerator (2644) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (2645) or RAM (2646). Transition data can also be stored in RAM (2646), while permanent data can be stored in, for example, an internal mass storage device (2647). Fast storage and retrieval of any storage device in the storage device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (2641), GPUs (2642), mass storage devices (2647), ROMs (2645), RAMs (2646), etc.

[0179] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer code may be of a type well known and available to those skilled in the art of computer software.

[0180] As a non-limiting example, a computer system (2600) having an architecture, in particular a core (2640), can provide functionality provided by (one or more) processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software implemented in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as described above, as well as certain storage devices of the core (2640) having non-transitory properties, such as a core internal mass storage device (2647) or ROM (2645). Software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (2640). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the core (2640), in particular the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processing or specific parts of specific processing described herein, including defining data structures stored in RAM (2646) and modifying such data structures according to processing defined by the software. Additionally or alternatively, the computer system may provide functionality provided as a result of hard-wiring logic or otherwise implemented in circuitry (e.g., accelerator (2644)) that may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0181] Although a particular invention has been described with reference to illustrative embodiments, this description is not meant to be limiting. Based on this specification, various modifications of the illustrative embodiments of the present invention and additional embodiments will be apparent to those of ordinary skill in the art. Those skilled in the art will readily recognize that these and various other modifications may be made to the exemplary embodiments shown and described herein without departing from the spirit and scope of the present invention. Therefore, it is contemplated that the appended claims will cover any such modifications and alternative embodiments. Certain proportions in the illustrations may be exaggerated, while other proportions may be minimized. Therefore, the present disclosure and the accompanying drawings should be considered illustrative rather than restrictive.

Claims

1. A method for video decoding, Features The method comprises: receiving a coded video bitstream of a block, wherein the coded video bitstream includes mode information of the block; determining whether unidirectional intra prediction or bidirectional intra prediction is applicable to the block based on mode information of the block, the mode information of the block comprising at least one of a reference line index of the block, an intra prediction mode of the block, and a size of the block; If it is determined that the unidirectional intra prediction is applicable to the block, performing the unidirectional intra prediction on the block; and If it is determined that the bidirectional intra prediction is applicable to the block, the bidirectional intra prediction is performed on the block.

2. The method according to claim 1, in, Determining whether the unidirectional intra prediction or the bidirectional intra prediction is applicable to the block comprises: If the reference line index of the block indicates a non-adjacent reference line, it is determined that the unidirectional intra prediction is applicable to the block.

3. The method according to claim 1, in, Determining whether the unidirectional intra prediction or the bidirectional intra prediction is applicable to the block comprises: If the reference line index of the block indicates a neighboring reference line, it is determined that the bidirectional intra prediction is applicable to the block.

4. The method according to claim 1, in, Determining whether the unidirectional intra prediction or the bidirectional intra prediction is applicable to the block comprises: If the reference line index of the block is greater than a predefined threshold, determining that the unidirectional intra prediction is applicable to the block; and If the reference line index of the block is not greater than the predefined threshold, it is determined that the bidirectional intra prediction is applicable to the block.

5. The method according to claim 4, in, The predefined threshold is a non-negative integer.

6. The method according to claim 1, in, Determining whether the unidirectional intra prediction or the bidirectional intra prediction is applicable to the block comprises: If the intra prediction mode of the block belongs to the selected first set of intra prediction modes and the reference line index indicates an adjacent reference line, determining that the bidirectional intra prediction is applicable to the block; If the intra prediction mode of the block belongs to the selected first set of intra prediction modes and the reference line index indicates a non-adjacent reference line, determining that the unidirectional intra prediction is applicable to the block; and If the intra prediction mode of the block belongs to the selected second set of intra prediction modes, determining that the bidirectional intra prediction is applicable to the block; The selected first group of intra-frame prediction modes and the selected second group of intra-frame prediction modes do not overlap.

7. The method according to claim 6, in, The selected second set of intra prediction modes includes a diagonal intra prediction mode.

8. The method according to claim 1, in, Determining whether the unidirectional intra prediction or the bidirectional intra prediction is applicable to the block comprises: If the reference line index of the block is an odd integer and / or the reference line index indicates a non-adjacent reference line, determining that the unidirectional intra prediction is applicable to the block; and If the reference line index of the block is an even integer and / or the reference line index indicates a neighboring reference line, it is determined that the bidirectional intra prediction is applicable to the block.

9. The method according to claim 1, in, Performing the bidirectional intra prediction on the block includes: determining a first prediction value based on a reference line index of the block; determining a second prediction value based on a reference line index of the block; and A final prediction value of the block is determined based on a weighted calculation between the first prediction value and the second prediction value.

10. The method according to claim 1, in, Performing the bidirectional intra prediction on the block includes: determining a first prediction value based on a first reference line index of the block; determining a second prediction value based on a second reference line index of the block; and A final prediction value of the block is determined based on a weighted calculation between the first prediction value and the second prediction value.

11. The method according to claim 10, in: The first reference line index indicates a non-adjacent reference line, and the first prediction value is generated according to a first sample at the non-adjacent reference line; and The second reference line index indicates an adjacent reference line, and the second prediction value is generated according to a second sample at the adjacent reference line.

12. The method according to claim 10, in, Performing bidirectional intra prediction on the block includes: The first reference line index indicates a non-adjacent reference line, and the first prediction value is generated according to a first sample at the non-adjacent reference line; and The second reference line index indicates a plurality of reference lines, and the second prediction value is generated as a linear weighted average of a plurality of samples, wherein each sample is from a reference line along a prediction angle among the plurality of reference lines.

13. The method according to any one of claims 1 to 12, further comprising: include: If it is determined that the bidirectional intra prediction is applicable to the block and the reference line index of the block indicates a non-adjacent reference line, then: Disable reference sample filtering; If reference samples from the non-adjacent reference line are available, using the reference samples for bidirectional intra prediction; as well as If the reference samples from the non-adjacent reference lines are not available, adjacent available reference samples are used for bidirectional intra prediction.

14. A device for video decoding, Features The device comprises: a receiving module configured to receive a coded video bitstream of a block, the coded video bitstream comprising mode information of the block; a determination module configured to determine whether unidirectional intra prediction or bidirectional intra prediction is applicable to the block based on mode information of the block, the mode information of the block comprising at least one of a reference line index of the block, an intra prediction mode of the block, and a size of the block; a unidirectional intra prediction module configured to perform the unidirectional intra prediction on the block if it is determined that the unidirectional intra prediction is applicable to the block; and A bidirectional intra prediction module is configured to perform the bidirectional intra prediction on the block if it is determined that the bidirectional intra prediction is applicable to the block.

15. A device for video decoding, the device include: a memory for storing instructions; as well as A processor in communication with the memory, wherein when the processor executes the instruction, the processor performs the method according to any one of claims 1 to 13.

16. A non-transitory computer-readable storage medium storing instructions, It is characterized in that When the instructions are executed by a processor, the processor performs the method according to any one of claims 1 to 13.

17. A computer program product comprising a computer program, which, when executed on a computer device, causes the computer device to execute the method according to any one of claims 1 to 9.

18. A method for storing or transmitting a video bitstream, It is characterized in that The video bit stream is decoded according to the decoding method according to any one of claims 1 to 13.