Intra-frame prediction method and related equipment

By employing a multi-reference layer intra-frame prediction method, the problem of low compression efficiency in video data transmission of diverse internet images in the 5G era is solved, achieving high-quality video transmission under limited network resources and improving video coding efficiency.

CN121397239APending Publication Date: 2026-01-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410986928.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

In the 5G era, internet image traffic is growing explosively. Existing intra-frame prediction methods are unable to effectively handle the diversity and differences of new internet images, resulting in low video data transmission compression efficiency. In particular, it is difficult to improve video quality without increasing bit rate under limited network resources.

Method used

The multi-reference layer intra-frame prediction method is adopted. By determining N reference layers of the current block, selecting M reference layers, determining the derivation direction, and fusing along the direction, the N+1th reference layer is generated for intra-frame prediction, thereby improving prediction accuracy and coding efficiency.

Benefits of technology

It improves the accuracy and coding efficiency of intra-frame prediction, and can improve video quality without increasing bit rate under limited network resources. It is suitable for various video encoding and decoding devices and systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397239A_ABST
    Figure CN121397239A_ABST
Patent Text Reader

Abstract

The invention provides an intra-frame prediction method and related equipment, and belongs to the technical field of video coding and decoding. The method comprises the following steps: determining N reference layers of a current block, wherein N is a positive integer greater than 1; selecting M reference layers from the N reference layers, wherein M is a positive integer greater than 1 and less than N; determining derivation directions of the M reference layers; fusing the M reference layers along the derivation direction to obtain the (N + 1) th reference layer of the current block; and performing intra-frame prediction on the current block according to the N reference layers and the (N + 1) th reference layer of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of video coding, and in particular, to an intra prediction method, an intra prediction device, an electronic device, a computer readable storage medium and a computer program product. BACKGROUND

[0002] On one hand, the Internet is about to enter a new era of 5G (5th generation mobile networks or 5th generation wireless systems, 5th-Generation, the fifth generation of mobile communication technology), and the images (videos) in various Internet applications have become the main consumer of Internet bandwidth. In particular, mobile Internet image traffic is increasing day by day, and will have a burst of growth in the 5G era, which will inject a new powerful driving force into the accelerated development of image coding technology. At the same time, it also puts forward many new challenges that have never been encountered in the past for image coding technology. In the 5G era, everything is interconnected, and new types of Internet images generated in various emerging applications have diversity and difference. Therefore, how to research efficient image coding technology according to the characteristics of new types of Internet images with diversity and difference has become an urgent need.

[0003] On the other hand, the amount of video data required to depict even a relatively short film can be substantial, and can create difficulties when the data is to be streamed or otherwise communicated through a communication network with limited bandwidth capacity. Therefore, video data is often compressed before being communicated across modern telecommunication networks. Video compression devices often use software and / or hardware at the source to code the video data, reducing the amount of data needed to represent digital video images, prior to transmission. The compressed data is then received at the destination by a video decompression device that decodes the video data. With limited network resources and ever-increasing demands of higher video quality, improved compression and decompression techniques are needed that improve image quality without increasing the bit rate.

[0004] Therefore, there is a need for a new intra prediction method, an intra prediction device, an electronic device, a computer readable storage medium and a computer program product. SUMMARY

[0005] The embodiment of the present disclosure provides an intra prediction method, comprising: determining N reference layers of a current block, N being a positive integer greater than 1; selecting M reference layers from the N reference layers, M being a positive integer greater than 1 and smaller than N; determining a derivation direction of the M reference layers; fusing the M reference layers along the derivation direction to obtain an (N+1)th reference layer of the current block; and performing intra prediction on the current block according to the N reference layers and the (N+1)th reference layer of the current block.

[0006] The embodiment of the present disclosure provides an intra prediction device, comprising: a processing unit configured to determine N reference layers of a current block, N being a positive integer greater than 1; the processing unit is further configured to select M reference layers from the N reference layers, M being a positive integer greater than 1 and smaller than N; the processing unit is further configured to determine a derivation direction of the M reference layers; the processing unit is further configured to fuse the M reference layers along the derivation direction to obtain an (N+1)th reference layer of the current block; and the processing unit is further configured to perform intra prediction on the current block according to the N reference layers and the (N+1)th reference layer of the current block.

[0007] The embodiment of the present disclosure provides a computer readable storage medium, having a computer program stored thereon, the program being executed by a processor to implement the method described in any embodiment of the present disclosure.

[0008] The embodiment of the present disclosure provides an electronic device, comprising: at least one processor; a storage device configured to store at least one program, when the at least one program is executed by the at least one processor, the at least one processor is caused to implement the method described in any embodiment of the present disclosure.

[0009] The embodiment of the present disclosure provides a computer program product, comprising a computer program, the computer program being executed by a processor to implement the method described in any embodiment of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 A block diagram of an example video coding system is schematically illustrated in accordance with some embodiments of the present disclosure.

[0011] Figure 2 A basic block diagram of an example video encoder is schematically illustrated in accordance with some embodiments of the present disclosure.

[0012] Figure 3 A basic block diagram of an example video decoder is schematically illustrated in accordance with some embodiments of the present disclosure.

[0013] Figure 4 A flowchart of an intra prediction method is schematically illustrated in accordance with some embodiments of the present disclosure.

[0014] Figure 5 A diagram is illustratively shown of a method of multi-reference line based intra prediction according to some embodiments of the present disclosure.

[0015] Figure 6 A diagram is illustratively shown of multi-reference line fusion direction based on intra prediction mode (angle) according to some embodiments of the present disclosure.

[0016] Figure 7 A diagram is illustratively shown of possible candidate derivation directions for the left and top side of a current block according to some embodiments of the present disclosure.

[0017] Figure 8 A diagram is illustratively shown of intra prediction modes according to some embodiments of the present disclosure.

[0018] Figure 9 A diagram is illustratively shown of intra prediction modes according to some other embodiments of the present disclosure.

[0019] Figure 10 A diagram is illustratively shown of intra prediction modes according to some further embodiments of the present disclosure.

[0020] Figure 11 A block diagram of an intra prediction apparatus according to some embodiments of the present disclosure is illustratively shown.

[0021] Figure 12 A structural diagram of an electronic device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0022] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art.

[0023] In embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program having a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented wholly or partially by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the function of the module or unit.

[0024] In the following description, unless otherwise defined, all scientific and technical terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0025] Reference in this disclosure to "one embodiment", "an embodiment", "example embodiment", etc., indicates that a described embodiment can include a particular feature, structure, or characteristic, but every embodiment can not necessarily include the particular feature, structure, or characteristic. Moreover, these phrases are not necessarily referring to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that it is within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0026] It should be understood that although the terms "first" and "second" etc. can be used to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed terms.

[0027] The flow charts shown in the drawings are only illustrative and do not necessarily include all contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be further broken down, and some operations / steps can be combined or partially combined, so the actual execution order can be changed according to actual conditions.

[0028] First, some terms involved in the embodiments of the present disclosure are described.

[0029] Intra(picture)Prediction: Intra prediction.

[0030] Inter(picture)Prediction: Inter prediction.

[0031] SCC: screen content coding, screen content / picture coding. Screen pictures are pictures generated by electronic devices such as computers, mobile phones, televisions, etc., mainly including two types of content: one is computer-generated non-continuous tone content, including a large number of small and sharp line shapes such as text, icons, buttons, and grids; the other is camera-captured content containing a large number of continuous tones, such as movies, television clips, and natural image videos. With the rapid development of cloud computing, mobile communication technology, and wireless display technology, how to display screen pictures on various electronic terminal devices with high quality at low code rate is a problem to be solved by SCC.

[0032] Loop Filtering: Loop filtering.

[0033] QP: Quantization Parameter, quantization parameter.

[0034] LCU: Largest Coding Unit, largest coding unit.

[0035] CTU: Coding Tree Unit, coding tree unit, generally starts from the largest coding unit and is divided downward.

[0036] CU: Coding Unit, coding unit.

[0037] PU: Prediction Unit, prediction unit.

[0038] MV: Motion Vector, motion vector.

[0039] BV: Block Vector, block displacement vector.

[0040] MVP: Motion Vector Prediction, motion vector prediction value.

[0041] AMVP: Advanced Motion Vector Prediction, advanced motion vector prediction.

[0042] MVD: Motion Vector Difference, the difference between the true value of the MVP and the MV.

[0043] ME: Motion Estimation, motion estimation, the process of obtaining the motion vector MV is called motion estimation, which is a technique in motion compensation (MC).

[0044] MC: According to the motion vector and the inter-frame prediction method, the process of obtaining the estimated value of the current image. Motion compensation is a method of describing the difference between adjacent frames (adjacent here means adjacent in the coding relationship, and the two frames in the playing order are not necessarily adjacent). Specifically, it describes how each small block of the previous frame moves to a certain position in the current frame. This method is often used by video compression / video codec to reduce spatial redundancy in video sequences. Adjacent frames are usually similar, that is, they contain a lot of redundancy. The purpose of using motion compensation is to improve the compression ratio by eliminating such redundancy.

[0045] I Slice: Intra Slice, intra slice / piece. The image can be divided into a frame or two fields, and the frame can be divided into one or more pieces (Slice).

[0046] The method provided by the embodiments of the present disclosure can be applied to products using a video codec or video compression, and can be applied to encoding and decoding of lossy data compression and lossless data compression. The data involved in the encoding and decoding process refers to one of the following examples or a combination thereof:

[0047] 1) one-dimensional data;

[0048] 2) two-dimensional data;

[0049] 3) multi-dimensional data;

[0050] 4) graphics;

[0051] 5) images;

[0052] 6) sequences of images;

[0053] 7) videos;

[0054] 8) three-dimensional scenes;

[0055] 9) sequences of continuously changing three-dimensional scenes;

[0056] 10) virtual reality scenes;

[0057] 11) sequences of continuously changing virtual reality scenes;

[0058] 12) images in the form of pixels;

[0059] 13) transform domain data of images;

[0060] 14) two-dimensional or more dimensional byte sets;

[0061] 15) two-dimensional or more dimensional bit sets;

[0062] 16) pixel sets;

[0063] 17) three-component pixel (Y, U, V) sets;

[0064] 18) three-component pixel (Y, Cb, Cr) sets;

[0065] 19) three-component pixel (Y, Cg, Co) sets;

[0066] 20) three-component pixel (R, G, B) sets;

[0067] 21) four-component pixel (C, M, Y, K) sets;

[0068] 22) four-component pixel (R, G, B, A) sets;

[0069] 23) a set of four-component pixels (Y, U, V, A);

[0070] 24) a set of four-component pixels (Y, Cb, Cr, A);

[0071] 25) a set of four-component pixels (Y, Cg, Co, A).

[0072] When the data is an image, or a sequence of images, or a video as listed above, the coding block is a coded region of an image, and should comprise at least one of the following: a group of pictures, a predetermined number of pictures, a picture, a frame of a picture, a field of a picture, a sub-picture of a picture, a slice, a macroblock, a largest coding unit, LCU, a coding tree unit, CTU, a coding unit, CU.

[0073] Figure 1 A block diagram illustrating an example video coding system that can utilize the techniques of this disclosure is shown. As shown, a video coding system can include a source device 110 and a destination device 120. The source device 110 can also be referred to as a video encoding device, and the destination device 120 can also be referred to as a video decoding device. In operation, the source device 110 can be configured to generate encoded video data, and the destination device 120 can be configured to decode the encoded video data generated by the source device 110. The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0074] The video source 112 can include a source such as a video capture device. Examples of video capture devices include, but are not limited to, an interface to receive video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0075] The video data can comprise one or more pictures / images. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that forms a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator and / or a transmitter. The encoded video data can be transmitted directly to the destination device 120 via the I / O interface 116 over the network 130A. The encoded video data can also be stored onto a storage medium / server 130B for access by the destination device 120.

[0076] Destination device 120 can include I / O interface 126, video decoder 124, and display device 122. I / O interface 126 can include a receiver and / or a modem. I / O interface 126 can acquire encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 can decode encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120, which is configured to interface with an external display device.

[0077] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC, also known as H.265) standard, the Versatile Video Coding (VVC, also known as H.266) standard, the H.267 standard, the Audio Video Coding Standard (AVS) standard, and other existing and / or future standards.

[0078] Figure 2 A block diagram of an example of a video encoder according to some embodiments of the disclosure is shown, which can be Figure 1 the example of video encoder 114 in the illustrated system.

[0079] The video encoder can be configured to implement any or all of the techniques of the disclosure. In Figure 2 example, the video encoder includes multiple functional components. The techniques described in this disclosure can be shared among the various functional components of the video encoder. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0080] A video signal can include both camera-captured and computer-generated signals, in terms of the way of acquisition. Due to the difference in statistical characteristics, the corresponding compression encoding manner can also be different.

[0081] Some video coding techniques, such as HEVC, VVC, and AVS, employ a hybrid coding framework. As Figure 2 shown, the pictures in the input video are sequentially encoded, and a series of operations and processes are performed as follows:

[0082] 1) Block partition structure: The input picture is divided into several non-overlapping processing units according to a pre-defined size. Similar compression operation is applied to each processing unit. This processing unit can be referred to as CTU or LCU. CTU or LCU can be further divided into more detailed units, resulting in at least one (one or more) basic coding unit, referred to as CU. Each CU is the most basic element in a coding loop. The following describes various coding modes that can be applied to each CU.

[0083] The video encoder can include a partition unit that can partition an input video picture into one or more video blocks. The video encoder and the video decoder can support various video block sizes.

[0084] 2) Predictive Coding: It includes intra prediction and inter prediction, etc. The original video signal is predicted by the selected reconstructed video signal to obtain the residual video signal. The encoding end needs to determine the most suitable one from the many possible predictive coding modes for the current CU, and inform the decoding end.

[0085] The video encoder can include a mode selection unit / encoding mode decision unit, which can select one of a plurality of coding modes (intra coding or inter coding), for example, based on error results, and use the generated intra coded block or inter coded block to generate residual block data to reconstruct the coding block and can be used as a reference picture / reference image / reference frame. In some examples, the mode selection unit / encoding mode decision unit can select a combination of intra and inter prediction modes, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit / encoding mode decision unit can also select a resolution for the motion vector (e.g., sub-pixel accuracy or integer pixel accuracy) for the block.

[0086] a. Intra prediction: The predicted signal comes from the already coded and reconstructed area within the same picture (part or all of which can be used as a reference area in the embodiments of the present disclosure).

[0087] The basic idea of intra prediction is to remove spatial redundancy by using the correlation of neighboring pixels. In video coding, neighboring pixels refer to the reconstructed pixels of the already coded CUs around the current CU.

[0088] The video encoder can include an intra-prediction unit that can perform intra-prediction for a current video block. When the intra-prediction unit performs intra-prediction for the current video block, the intra-prediction unit can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0089] b. Inter-prediction: the predicted signal comes from other pictures (referred to as reference pictures or reference frames, some or all of which can be referred to as reference regions in embodiments of the disclosure) that have already been coded and are different from the current picture.

[0090] To perform inter-prediction for a current video block, a motion estimation unit can generate motion information for the current video block by comparing one or more reference frames from the decoded picture buffer to the current video block. A motion-compensated prediction unit can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the decoded picture buffer other than the picture associated with the current video block.

[0091] The motion estimation unit and the motion-compensated prediction unit can perform different operations, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" can refer to a portion of a picture composed of macroblocks that are all based on macroblocks within the same picture. Further, as used herein, in some aspects, a "P slice" and a "B slice" can refer to portions of a picture composed of macroblocks that are independent of macroblocks in the same picture.

[0092] In some examples, the motion estimation unit can perform uni-prediction for the current video block, and the motion estimation unit can search a list 0 (L0) or a list 1 (LI) of reference pictures for a reference video block for the current video block. The motion estimation unit can then generate a reference index that indicates a reference picture in the list 0 or the list 1 that contains the reference video block, and a motion vector that indicates a spatial displacement between the current video block and the reference video block. The motion estimation unit can output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. The motion-compensated prediction unit can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.

[0093] Alternatively, in other examples, the motion estimation unit can perform bi-prediction for the current video block. The motion estimation unit can search the reference pictures in list 0 for a reference video block for the current video block and also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit can then generate a plurality of reference indices that indicate a plurality of reference pictures in list 0 and list 1 that contain a plurality of reference video blocks and a plurality of motion vectors that indicate a plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit can output the plurality of reference indices and the plurality of motion vectors for the current video block as motion information for the current video block. The motion-compensated prediction unit can generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information for the current video block.

[0094] In some examples, the motion estimation unit can output a complete set of motion information for use in decoding processing at the decoder. Alternatively, in some embodiments, the motion estimation unit can refer to motion information of another video block to signal the motion information of the current video block. For example, the motion estimation unit can determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0095] In one example, the motion estimation unit can indicate a value in a syntax structure associated with the current video block that indicates that the current video block has the same motion information as another video block. In another example, the motion estimation unit can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0096] As discussed above, a video encoder can signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that can be implemented by a video encoder include advanced motion vector prediction (AMVP) and merge mode signaling.

[0097] Wherein, in different prediction modes and different implementations, the displacement vector can have different names, in the embodiments of the present disclosure, it is described in the following way: 1) the displacement vector in inter prediction is called motion vector (MV); 2) the displacement vector in intra block copy is called block vector (BV) or block displacement vector; 3) the displacement vector in intra string copy is called string vector (SV).

[0098] The video encoder can include a residual generation unit that can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the prediction video block(s) from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample portions of samples in the current video block.

[0099] In other examples, such as in a skip mode, there can be no residual data for the current video block for the current video block, and the residual generation unit can not perform the subtraction operation.

[0100] 3) Transform & Quantization: The residual video signal is transformed into transform domain by DFT (Discrete Fourier Transform), DCT (Discrete Cosine Transform) or other transform operation, and the residual video signal is converted into transform coefficients. The residual video signal in the transform domain is further subjected to a lossy quantization operation, and certain information is lost, so that the quantized signal is conducive to compression expression.

[0101] In some video coding standards, there can be more than one transform mode to choose from, so the encoding end also needs to select one of them for the current CU (current coding CU) to be encoded, and inform the decoding end.

[0102] The finer degree of quantization is usually determined by the quantization parameter (QP). When the QP value is large, the transform coefficients with larger value range will be quantized to the same output, so it usually brings larger distortion and lower code rate. Conversely, when the QP value is small, the transform coefficients with smaller value range will be quantized to the same output, so it usually brings smaller distortion and higher code rate.

[0103] 4) Entropy Coding or Statistical Coding: The quantized transform domain signal will be statistically compressed and coded according to the frequency of each value, and finally output the binary (0 or 1) compressed bitstream or encoded bitstream or encoded video bitstream or encoded video data.

[0104] At the same time, other information generated by encoding, such as selected coding modes, motion vectors, etc., also needs to be entropy coded to reduce the code rate.

[0105] Among them, statistical coding is a kind of lossless coding method, which can effectively reduce the code rate required to express the same signal. Common statistical coding methods include variable length coding (VLC) or context-based binary arithmetic coding (CABAC).

[0106] 5) Loop filtering: the reconstructed decoded picture can be obtained by the inverse operation of the quantization, inverse transformation and prediction compensation operation (steps 2) ~ 4) of the above-mentioned) of the already coded picture. The reconstructed decoded picture is different from the original input picture due to the influence of quantization, and part of the information is different from the original input picture, resulting in distortion. Filtering operation, such as deblocking, SAO (Sample Adaptive Offset) or ALF (Adaptive Loop Filter) filter, can effectively reduce the distortion caused by quantization. Since these filtered reconstructed decoded pictures will be used as references for future coding pictures to predict future signals, the above-mentioned filtering operation is also called loop filtering and filtering operation within the coding loop.

[0107] Figure 2 A basic flowchart of a video encoder is shown. Figure 2 The kth CU (marked as s k [x,y]) is taken as an example. Among them, k is a positive integer greater than or equal to 1 and less than or equal to the number of CUs in the input current picture, s k [x,y] represents a pixel point with coordinates [x,y] in the kth CU, x represents the horizontal coordinate of the pixel point, y represents the vertical coordinate of the pixel point, and x and y are both positive integers greater than or equal to 1, and the upper limit depends on the number of pixels in the kth CU. s k [x,y] is obtained after one of the optimal processing of motion compensation or intra prediction s k [x,y] and Subtracting the residual signal u k [x,y] from the predicted signal s k[x, y] are transformed and quantized, and the quantized output data has two different destinations: one is sent to an entropy encoder for entropy encoding, and the encoded bitstream is output to a buffer for storage and transmission; the other is used for inverse quantization and inverse transformation to obtain the signal u k [x, y]. The signal u k [x, y] is added to to obtain a new prediction signal s * k [x, y]. The signal s * k [x, y] is sent to the buffer of the current image for storage. The signal s * k [x, y] is obtained by intra-image prediction f(s * k [x, y]). The signal s * k [x, y] is obtained by loop filtering s k [x, y], and the signal s k [x, y] is sent to the decoded image buffer for storage to generate the reconstructed video. The signal s k [x, y] is obtained by motion-compensated prediction s r [x + m x , y + m y ], where s r [x + m x , y + m y ] represents a reference block, and m x and m y represent the horizontal and vertical components of the motion vector, respectively.

[0108] According to the above encoding process, it can be seen that at the decoding end, for each CU, the decoder obtains the compressed bitstream, first performs entropy decoding to obtain various mode information (including encoding mode information, such as intra prediction mode or inter prediction mode) and quantized transform coefficients. Each coefficient is subjected to inverse quantization and inverse transformation to obtain a residual signal. On the other hand, according to the known encoding mode information, the prediction signal corresponding to the CU can be obtained, and the two are added to obtain the reconstructed signal. Finally, the reconstructed value of the decoded image needs to be subjected to loop filtering operation to generate the final output signal.

[0109] In other examples, a video encoder can include more, fewer, or different functional components. In one example, an intra block copy (IBC) unit can be included. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0110] Furthermore, although some components, such as the motion estimation unit and the motion-compensated prediction unit, can be integrated, these components are shown separately in the example of FIG. 3 for purposes of explanation. Figure 2

[0111] Figure 3 is a block diagram illustrating an example of a video decoder 300 in accordance with some embodiments of the present disclosure, which can be an example of the video decoder 124 in the system shown in FIG. 1. Figure 1

[0112] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In the example of FIG. 3, the video decoder 300 includes a number of functional components. The techniques described in the present disclosure can be shared among the various functional components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure. Figure 3

[0113] In the example of FIG. 3, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, and a reconstruction unit 306 and buffer 307. In some examples, the video decoder 300 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 200. Figure 3

[0114] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture / list indices, and other motion information. The motion compensation unit 302 can determine this information, for example, by performing advanced motion vector prediction (AMVP) and merge mode. The motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indices, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, “merge mode” can refer to deriving motion information from spatially or temporally neighboring blocks.

[0115] The motion compensation unit 302 can produce a motion-compensated block, which can perform interpolation based on interpolation filtering. An identifier of the interpolation filter used at sub-pixel precision can be included in the syntax elements.

[0116] ​​​​Motion compensation unit 302 can use interpolation filtering used by video encoder 200 during encoding of the video blocks to calculate interpolated values for sub-pixels of the reference blocks.

[0117] Motion compensation unit 302 can determine the interpolation filtering used by video encoder 200 from the received syntax information, and motion compensation unit 302 can use the interpolation filtering to generate the predicted blocks.

[0118] Motion compensation unit 302 can use at least some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and lists of reference frames) for each inter-coded block, and other information to decode the coded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice can be an entire picture, or can also be a region of a picture.

[0119] Intra prediction unit 303 can use, for example, intra prediction modes received in the bitstream to form predicted blocks from spatially neighboring blocks. Dequantization unit 303 dequantizes quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies inverse transforms.

[0120] Reconstruction unit 306 can obtain a decoded block, for example, by adding the residual block to the corresponding predicted block generated by motion compensation unit 202 or intra prediction unit 303. If desired, deblocking filtering can also be applied to filter the decoded block in order to remove blockiness artifacts. The decoded video blocks are then stored in buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and which also produces decoded video for presentation on a display device.

[0121] Some example embodiments of the present disclosure will be described in detail below. Although some embodiments are described with reference to the general video coding or other specific video codecs, the disclosed techniques are applicable to other video coding technologies. Moreover, although some embodiments describe video coding steps in detail, it should be understood that corresponding decoding steps of the coding would be implemented by a decoder. Furthermore, the term video processing includes video coding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bit rates.

[0122] Figure 4A flowchart schematically illustrating an intra prediction method according to some embodiments of the present disclosure is shown. Figure 4 The method provided by the embodiments can be executed by any electronic device with information processing capability, for example, can be executed at the side of a video encoder and / or a video decoder. The method can be applied to existing video coding standards (for example, high efficiency video coding (HEVC)) and future video coding standards (for example, versatile video coding (VVC)) or codecs. As shown in the figure, the method provided by the embodiments of the present disclosure can include the following steps. Figure 4 The method provided by the embodiments of the present disclosure can include the following steps.

[0123] In S410, N reference tiers (also referred to as candidate reference tiers or candidate reference lines) of the current block are determined, N being a positive integer greater than 1.

[0124] The term "block" in the embodiments of the present disclosure can be understood as a prediction block, a prediction unit, a coding block, a decoding block or a coding unit (for example, CU) and the like. In the embodiments of the present disclosure, if the method is executed at the side of a video encoder, the current block refers to a block to be encoded, which can also be referred to as a coding block / block to be encoded / current block to be encoded, for example, a current CU. If the method is executed at the side of a video decoder, the current block refers to a block to be decoded, which can also be referred to as a decoding block / block to be decoded / current block to be decoded, for example, a current CU to be decoded, or a current CU.

[0125] In the embodiments of the present disclosure, the first reference tier in the N reference tiers is referred to as a neighboring reference tier of the current block, that is, the reference pixels in the first reference tier are adjacent to the current block; the second reference tier to the Nth reference tier are referred to as non-neighboring reference tiers of the current block, that is, the reference pixels in the second reference tier to the Nth reference tier are not adjacent to the current block. In some examples, the reference tier is also referred to as a reference line or a reference row. In some examples, the reference row includes a row and / or a column.

[0126] In HEVC and JEM and some other standards such as H.264 / AVC, the reference samples for predicting the current block are limited to the nearest reference line (i.e., the neighboring line or the neighboring reference layer). In standards such as VVC, in addition to the reference pixels of the line and the column (i.e., the first reference layer) adjacent to the current block, the reference pixels of the non-adjacent line (column) (i.e., the second reference layer to the Nth reference layer) are also allowed to be used to predict the current block, so this technology is called the multi-reference line intra prediction method or the multi-reference line based intra prediction method. In the method of multi-reference line intra prediction, the number of candidate reference lines increases from one (i.e., the nearest) to N, where N is an integer greater than 1. That is, the reference samples for predicting the current block are not limited to the nearest line (row and / or column) of the current block. In some examples, in the method of multi-reference line intra prediction, the index number of the candidate reference line increases from zero (i.e., the nearest) to N1. In some examples, the nearest reference line is referred to as the zero reference line or the neighboring reference line or the adjacent reference line, and the other reference lines are referred to as the non-zero reference line or the non-neighboring reference line or the non-adjacent reference line.

[0127] As shown in Figure 5 According to the positional relationship with the current block, the reference pixels of each line and each column are labeled from near to far, which is referred to as the xth line. These reference lines can be marked as the 0th line (i.e., the first reference layer or the first reference layer, which can also be referred to as the neighboring line or the neighboring reference layer), the 1st line (i.e., the second reference layer or the second reference layer, which is a non-neighboring line or a non-neighboring reference layer), and so on.

[0128] In the embodiments of the present disclosure, the reference pixels can also be referred to as "reference samples", which refer to the pixels that have been coded or reconstructed in the current image / current frame to which the current block belongs.

[0129] The multi-reference line intra prediction technology in the related art can only select one of the reference lines (columns) to predict the current block as a whole (i.e., predict all the pixels of the current block), which limits the overall accuracy of the prediction of the current block. To solve this problem, the present disclosure proposes a multi-reference line fusion intra prediction method, which can further improve the accuracy of intra prediction and improve the coding efficiency.

[0130] In S420, M reference layers are selected from N reference layers, where M is a positive integer greater than 1 and less than N.

[0131] In an exemplary embodiment, selecting M reference layers from N reference layers includes: selecting the first reference layer to the Nth reference layer as the M reference layers; or selecting the second reference layer to the Nth reference layer as the M reference layers.

[0132] In this embodiment of the disclosure, M reference layers can be selected from the N reference layers of the current block for fusion to generate a newly added (N+1)th reference layer. In some embodiments, the N reference layers of the current block can be fused to generate the (N+1)th reference layer. This allows the (N+1)th reference layer to simultaneously incorporate information from the N reference layers of the current block. For example, the fused (N+1)th reference layer can be used to replace the N reference layers of the current block, i.e., the fused (N+1)th reference layer can be directly referenced for intra-frame prediction of the current block. Alternatively, both the (N+1)th reference layer and the N reference layers can be used as N+1 candidate reference layers for the current block, and one of them can be arbitrarily selected for intra-frame prediction of the current block.

[0133] In other embodiments, non-nearest reference layers other than adjacent reference layers can be fused to generate the (N+1)th reference layer. In this case, the fused (N+1)th reference layer can replace the second to Nth reference layers of the current block; that is, the first and (N+1)th reference layers are used as candidate reference layers for the current block, and one of them is selected for intra-frame prediction of the current block. Alternatively, the fused (N+1)th reference layer can replace any one of the second to Nth reference layers to generate a new set of N reference layers, and one of these new N reference layers can be arbitrarily selected for intra-frame prediction of the current block.

[0134] In S430, the derivation directions of M reference layers are determined.

[0135] In this embodiment of the disclosure, the derivation direction refers to the direction used to extract M reference pixels from the M reference layers when fusing M reference layers to generate the (N+1)th reference layer, in order to generate the reference pixel at the corresponding position in the (N+1)th reference layer. Therefore, the derivation direction can also be called the fusion direction or the gradient calculation direction.

[0136] by Figure 6 For example, assuming M = N = 3, and the derivation direction is a 45-degree angle to the lower right, then along this derivation direction, we obtain three reference pixels: AL33, AL22, and AL11. These three reference pixels are then fused to obtain the reference pixel corresponding to the diagonal position in the (N+1)th reference layer. As another example, along this derivation direction, we obtain three reference pixels: A31, A21, and A11. These three reference pixels are then fused to obtain the reference pixel located above the current block in the (N+1)th reference layer. As yet another example, along this derivation direction, we obtain three reference pixels: L13, L12, and L11. These three reference pixels are then fused to obtain the reference pixel located to the left of the current block in the (N+1)th reference layer. The acquisition of reference pixels at other positions in the (N+1)th reference layer follows the same principle.

[0137] In an exemplary embodiment, determining the derivation directions of the M reference layers includes: using the default derivation direction as the derivation direction of the M reference layers, where the intra-prediction mode of the current block is not considered, i.e., regardless of whether it adopts an angle prediction mode or a non-angle prediction mode, the pre-selected agreed direction is used as the derivation direction; or, determining the derivation directions of the M reference layers according to the intra-prediction mode of the current block.

[0138] In an exemplary embodiment, determining the inference directions of M reference layers based on the intra-prediction mode of the current block includes: if the intra-prediction mode of the current block is an angle prediction mode, then determining the inference directions of the M reference layers based on the angle corresponding to the angle prediction mode; if the intra-prediction mode of the current block is a non-angle prediction mode, then using the default inference direction as the inference direction of the M reference layers.

[0139] In this embodiment of the disclosure, when determining the derivation direction based on the intra-frame prediction mode of the current block, if an angle prediction mode is used, the derivation direction can be determined according to the angle used in the angle prediction mode. If a non-angle prediction mode is used, the default derivation direction can be used as the derivation direction, that is, regardless of which non-angle prediction mode is used at this time, the pre-selected agreed direction is used as the derivation direction.

[0140] In an exemplary embodiment, the M reference layers include a left reference pixel and an upper reference pixel located in the current block. Determining the derivation direction of the M reference layers based on the angle corresponding to the angle prediction mode includes: for the left reference pixels of the M reference layers, mapping the angle to a candidate horizontal direction as the derivation direction; for the upper reference pixels of the M reference layers, mapping the angle to a candidate vertical direction as the derivation direction.

[0141] For example, such as Figure 6 As shown, assuming the left and top sides of the current block are within the encoded or reconstructed regions of the current image, the M reference layers can include reference pixels located on the left side of the current block (e.g., L13, L12, L11, L23, L22, L21, etc.) and reference pixels on the top (or above) side (e.g., A31, A21, A11, A32, A22, A12, etc.). Correspondingly, the fused N+1th reference layer can also be distinguished as left and top. For example, for reference pixels on the left side, they can be mapped to the corresponding candidate horizontal direction based on the range of angles corresponding to the angle prediction mode adopted by the current block. For reference pixels on the top side, they can be mapped to the corresponding candidate vertical direction based on the range of angles corresponding to the angle prediction mode adopted by the current block.

[0142] In an exemplary embodiment, the candidate horizontal direction includes at least one of a horizontal direction, a horizontal upward angle of 45 degrees, and a horizontal downward angle of 45 degrees; the candidate vertical direction includes at least one of a vertical direction, a vertical left angle of 45 degrees, and a vertical right angle of 45 degrees. However, this disclosure is not limited to this, and the direction can be set according to actual needs. For example, the candidate horizontal direction may also include a horizontal upward angle of 22.5 degrees, a horizontal downward angle of 22.5 degrees, etc., and the candidate vertical direction may also include a vertical left angle of 22.5 degrees, a vertical right angle of 22.5 degrees, etc.

[0143] For example, such as Figure 7 As shown, the left image displays the candidate directions for multi-reference row fusion on the left, including the horizontal direction, 45 degrees upward, and 45 degrees downward. For a reference pixel on the left, if the angle corresponding to the angle prediction mode used by the current block is closest to a candidate horizontal direction, then that candidate horizontal direction is used as the derivation direction for the reference pixel on the left. For example, assuming the angle is 11.25 degrees upward, its derivation direction can be determined as the horizontal direction. Similarly, assuming the angle is horizontal, its derivation direction can be determined as the horizontal direction. The right image displays the candidate directions for multi-reference row fusion on the top, including the vertical direction, 45 degrees to the left, and 45 degrees to the right. For a reference pixel on the top, if the angle corresponding to the angle prediction mode used by the current block is closest to a candidate vertical direction, then that candidate vertical direction is used as the derivation direction for the reference pixel on the top. For example, assuming the angle is 11.25 degrees to the left, its derivation direction can be determined as the vertical direction. Similarly, assuming the angle is vertical, its derivation direction can be determined as the vertical direction.

[0144] In an exemplary embodiment, determining the derivation directions of the M reference layers based on the angles corresponding to the angle prediction mode includes: using the angles as the derivation directions of the M reference layers. That is, in some other embodiments, the angles corresponding to the intra-frame prediction mode adopted by the current block can be directly used as the derivation directions for fusing and generating the (N+1)th reference layer.

[0145] In an exemplary embodiment, the M reference layers include a left reference pixel and an upper reference pixel located in the current block; wherein, the default derivation direction is used as the derivation direction of the M reference layers, including: for the left reference pixel of the M reference layers, the default derivation direction is horizontal; for the upper reference pixel of the M reference layers, the default derivation direction is vertical.

[0146] In S440, along the derivation direction, M reference layers are fused to obtain the (N+1)th reference layer of the current block.

[0147] In an example embodiment, the M reference layers are fused along the derivation direction to obtain the N+1th reference layer of the current block, including: if the angle is not at the integer pixel position of the M reference layers, interpolating the reference pixels of the M reference layers along the angle to generate interpolated reference pixels; and fusing the interpolated reference pixels to obtain the N+1th reference layer of the current block.

[0148] When the angle corresponding to the intra prediction mode adopted by the current block is directly used as the derivation direction for fusing to generate the N+1th reference layer, the angle can be at the integer pixel position or not at the integer pixel position. Still taking the example of Figure 6 If the angle corresponding to the intra prediction mode is 45 degrees to the right and bottom, the angle is at the integer pixel position of the 3 reference layers along the direction of the dashed arrow, such as AL33, AL22, AL21, L13, L22, etc. If the angle corresponding to the intra prediction mode is 22.5 degrees to the right and bottom, the angle is not at the integer pixel position of the 3 reference layers but at the sub-pixel position, and in this case, the two reference pixels adjacent to the angle can be interpolated to obtain interpolated reference pixels, and then the interpolated reference pixels are linearly weighted fused or nonlinearly fused to determine the corresponding reference pixel in the N+1th reference layer.

[0149] In an example embodiment, the M reference layers are fused along the derivation direction to obtain the N+1th reference layer of the current block, including: extracting the reference pixels of the M reference layers along the derivation direction, and linearly weighting the reference pixels of the M reference layers to obtain the N+1th reference layer of the current block. Different weights or the same weights are configured for the reference pixels of different reference layers participating in the weighting.

[0150] In an example embodiment, the M reference layers are fused along the derivation direction to obtain the N+1th reference layer of the current block, including: extracting the reference pixels of the M reference layers along the derivation direction, and nonlinearly fusing the reference pixels of the M reference layers to obtain the N+1th reference layer of the current block.

[0151] In an example embodiment, the nonlinear fusion of the reference pixels of the M reference layers includes: selecting the pixel value with the highest occurrence frequency among the reference pixels of the M reference layers along the derivation direction as the pixel value in the N+1th reference layer; or selecting the median value of the pixel values of the reference pixels of the M reference layers along the derivation direction as the pixel value in the N+1th reference layer; or selecting the average value of the pixel values of the multiple reference pixels closest to each other among the reference pixels of the M reference layers along the derivation direction as the pixel value in the N+1th reference layer.

[0152] The method provided by the embodiments of the present disclosure generates an additional intra prediction reference line (i.e., the N+1th reference line) by fusing information of multiple reference lines (e.g., M reference lines in N reference lines), and performs intra prediction on the current block. That is, on the basis of the original N reference lines (0 to N-1), an additional fused intra prediction reference line is added.

[0153] For example, as shown in FIG. 1, assuming that the size of the current block is m x n, and assuming that N = 3, the reference pixels of the first row (column) adjacent to the current block, the second row (column) and the third row (column) non-adjacent to the current block are represented as follows: Figure 6

[0154] First column to first row (adjacent): L(m+n)1,…,L(m+1)1,Lm1,…,L21,L11,AL11,A11,A12,…,A1n,A1(n+1),…,A1(n+m)…

[0155] Second column to second row (non-adjacent): L(m+n)2,…,L(m+1)2,Lm2,…,L22,L12,AL12,AL22,AL21,A21,A22,…,A2n,A2(n+1),…,A2(n+m)…

[0156] Third column to third row (non-adjacent): L(m+n)3,…,L(m+1)3,Lm3,…,L23,L13,AL13,AL23,AL33,AL32,AL31,A31,A_32,…,A3n,A3(n+1),…,A_3(n+m)…

[0157] It can be understood that the reference pixels of the lower left side and the upper right side of the current block, for example, L(m+n)1,…,L(m+1)1, may not have been encoded, at which time their positions can be filled in some way. In some embodiments, the positions can be filled by copying the reference pixels that have been encoded and are adjacent to the lower left side or the upper right side. For example, L(m+n)1,…,L(m+1)1 can be filled by copying the pixel values of the already encoded Lm1, L(m+n)2,…,L(m+1)2 can be filled by copying the pixel values of the already encoded Lm2, and A1(n+m),…,A11(n+1) can be filled by copying the pixel values of the already encoded A1n. However, the present disclosure is not limited thereto.

[0158] The embodiments of the present disclosure provide various fusion modes of multiple reference lines:

[0159] ​The first fusion manner is a multi-reference line linear weighted average. A derivation direction of multi-reference line fusion is determined first. Along the derivation direction, each line (column) in the M reference layers has a pixel participating in the fusion.

[0160] (1) Vertical / horizontal direction: connecting the left multi-reference column of the current block along the horizontal direction; connecting the above multi-reference line along the vertical direction. That is, in this manner, no matter what the intra prediction mode adopted by the current block is (for example, the angle of the intra prediction mode adopted by the current block at this time can be a 45-degree angle from the upper right), the horizontal direction and the vertical direction are adopted when the N+1 reference line is fused to generate.

[0161] In some embodiments, the different reference lines participating in the weighting are given the same weight.

[0162] For example, taking Figure 6 For example, assuming that the vertical / horizontal direction is adopted, N=M=3, A_21’=(A_31+A_21+A11) / 3, L_12’=(L_13+L_12+_L11) / 3, and so on. That is, the reference pixels of 3 reference lines are fused, and the weights of the 3 reference lines are equal, all being 1 / 3, to generate 1 new fused reference line, A_21’ and L_12’ representing the reference pixels of the corresponding positions in the N+1 reference layer. That is, the weighted average of the reference pixels corresponding to A21 and the reference pixels of the two reference lines before and after A21 can generate a new reference line. In some embodiments, the newly generated N+1 reference layer is different from the original N reference lines, and can be used as a newly added reference line N, that is, any one of the N+1 reference lines can be selected for the intra prediction of the current block.

[0163] In other embodiments, the different reference lines participating in the weighting are given different weights.

[0164] For example, taking Figure 6For example, assuming a vertical / horizontal direction is used, N = M = 3, A_21' = (A_31 + 2*A_21 + A11) / 4, L_12' = (L_13 + 2*L_12 + L_11) / 4, and so on. That is, reference pixels from three reference rows are fused, and the weights of the three reference rows can be unequal. The weights of A_31 and A11 are both 1 / 4, and the weight of A_21 is 2 / 4, to generate a new fused reference row. A_21' and L_12' represent the reference pixels at the corresponding positions in this new fused reference row, i.e., the (N+1)th reference layer. In some embodiments, the newly generated (N+1)th reference layer is different from the original N reference layers and can be used to replace any non-adjacent reference layer among the N reference layers. That is, new N reference rows can be generated, from which one can be randomly selected for intra-frame prediction of the current block. For example, in the example above, A21 has a heavier weight, while the weights of the two adjacent rows are smaller. In this case, the newly generated (N+1)th reference layer can be used to replace the reference layer with a heavier weight. Alternatively, the newly generated (N+1)th reference layer can be used as an additional reference row N, meaning that any one of the N+1 reference rows can be selected for intra-frame prediction of the current block.

[0165] In some other embodiments, different weights are assigned to different reference rows participating in the weighting.

[0166] For example, with Figure 6 For example, assuming a vertical / horizontal direction, N = M = 3, A_21' = (-A_31 + 2*A_21 - A11), L_12' = (-L_13 + 2*L_12 - L11). That is, the weights of A_31 and A11 are both -1, and the weight of A_21 is 2, generating one new fusion reference row. In some embodiments, the newly generated (N+1)th reference layer can be used to replace any non-adjacent reference layer among the N reference layers, i.e., generating N new reference rows from which one can be randomly selected for intra-frame prediction of the current block. For example, in the above example, A21 has a heavier weight, while the weights of the adjacent rows are smaller, so the newly generated (N+1)th reference layer can replace the reference layer with the heavier weight. Alternatively, the newly generated (N+1)th reference layer can be used as an additional reference row N, meaning one of the N+1 reference rows can be randomly selected for intra-frame prediction of the current block.

[0167] (2) Determined by the intra prediction modes (IPM) of the current block: Based on the intra prediction modes of the current block, select a derivation direction for calculating multi-reference row pixel fusion. After determining the derivation direction, when fusing M reference layers along the derivation direction, the weights given to different reference rows participating in the weighting can be any of the methods in (1) above.

[0168] In some embodiments, for the angular prediction mode, the angle is mapped to one of the candidate connection directions (including the candidate horizontal direction and / or the candidate vertical direction) for the left side and the top side respectively. For example, if the derived directions are only horizontal, horizontal 45 degrees up, horizontal 45 degrees down, vertical, vertical 45 degrees left, and vertical 45 degrees right, then for the angles within the corresponding range, the derived direction is mapped to one of the horizontal, horizontal 45 degrees up, horizontal 45 degrees down, vertical, vertical 45 degrees left, and vertical 45 degrees right.

[0169] For example, as shown in FIG. 4A, for the left side reference pixels of the current block, the intra prediction direction of the current block (i.e., the derived direction) is mapped to one of the horizontal, horizontal 45 degrees up, or horizontal 45 degrees down; and for the top side reference pixels of the current block, the intra prediction direction of the current block is mapped to one of the vertical, vertical 45 degrees left, or vertical 45 degrees right. Figure 7

[0170] In some embodiments, for the non-angular prediction mode, a default derived direction can be used, for example, the DC mode (intra DC mode, numbered 1) and the Planar mode (or intra Planar mode, numbered 0) are mapped to: for the left side reference pixels of the current block, the horizontal direction is derived; and for the top side reference pixels of the current block, the vertical direction is derived.

[0171] In other embodiments, for the angular prediction mode, the position of the reference pixel can be derived along the intra prediction angle of the current block (i.e., the angle corresponding to the intra prediction mode). That is, the angle used in the intra prediction mode of the current block is used as the derived direction when generating the reference line N.

[0172] In some embodiments, if the reference pixel of a certain reference line / column is not at the integer pixel position, the sub-pixel reference pixel can be interpolated according to the rules of the angular intra prediction, and then the fusion calculation is performed.

[0173] For example, if the current block uses the upper right 22.5 degrees as the angle in the angular prediction mode, then the derived direction is also the upper right 22.5 degrees. In this case, the reference pixel is not at the integer pixel position, so the reference pixel can be interpolated, for example, for the upper right 22.5 degree derived direction, the adjacent pixels are A1nand A1(n+1), so the reference pixel [A1n+A1(n+1)] / 2, [A2(n+1)+A2(n+2)] / 2, … is interpolated, and then the interpolated reference pixel [A1n+A1(n+1)] / 2, [A2(n+1)+A2(n+2)] / 2, … is used to generate the reference line N (i.e., the N+1threference line). ​

[0174] The second fusion manner is a multi-reference line non-linear fusion manner.

[0175] (1) Connect M reference pixels in the M reference lines in the derivation direction, and select the pixel value of the reference pixel that appears most frequently as the reference pixel at the corresponding position in the N+1 reference line.

[0176] For example, sorting is performed according to the frequency of repeated appearance, and the pixel value with a high frequency is in the front. If the pixel values in the M reference pixels are all different, the pixel in the first reference line can be used as the reference pixel at the corresponding position in the N+1 reference line.

[0177] (2) Connect M reference pixels in the M reference lines in the derivation direction, and select the median value of the pixels as the reference pixel at the corresponding position in the N+1 reference line.

[0178] (3) Connect M reference pixels in the M reference lines in the derivation direction, and select the average value of a plurality of (for example, 2) pixels closest to each other as the reference pixel at the corresponding position in the N+1 reference line. For example, if the pixel values of the M reference pixels are 31, 40, 66, 89, and 255 in turn, the average value of 31 and 40 can be obtained.

[0179] If the M reference pixels are equally close, the average value of the pixels in the front in the sorting is obtained. For example, if the pixel values of the M reference pixels are 2, 3, 4, 5, and 6 in turn, the average value of 2 and 3 can be obtained.

[0180] In S450, the current block is intra-predicted according to the N reference layers and the N+1 reference layer of the current block.

[0181] In an exemplary embodiment, the intra-prediction of the current block according to the N reference layers and the N+1 reference layer of the current block includes: selecting the N+1 reference layer to intra-predict the current block; or selecting one of the first reference layer and the N+1 reference layer to intra-predict the current block; or replacing any one of the second reference layer to the N reference layer with the N+1 reference layer to form a new N reference layer, and selecting one of the new N reference layers to intra-predict the current block.

[0182] In the embodiments of the present disclosure, the use method of the fused reference line can be selected from any one of the following:

[0183] (1) The reference lines closest to the current block (which may include the nearest reference line, i.e., the first reference layer mentioned above) are fused together, and the fused reference line (i.e., the N+1th reference layer) replaces the nearest reference line for single reference line intra-frame prediction. For example, when generating the fused reference line N, reference lines 0 to N-1 are fused at the same time.

[0184] (2) Merge several reference lines closest to the current block (excluding the nearest neighbor reference line). The merged reference line can be used as an additional reference line together with one or more existing reference lines in accordance with the multi-reference line intra-frame prediction method.

[0185] For example, the index value of the fused reference row (i.e., the N+1th reference layer) is set to 1, and the index value of the neighboring reference rows is set to 0, and so on. That is, the reference rows 1 to N-1 are fused to generate the Nth reference row, and then the 0th reference row and the Nth reference row are used as candidates, and one row is selected from them for intra-frame prediction of the current block.

[0186] For example, such as Figure 6 As shown, assuming the current CU uses angle prediction mode with an angle of 45 degrees to the lower right, a prediction value p(x, y) is generated from one of the reference samples AL22' (the fused reference pixel), AL33, AL22, and AL11, where x and y represent the x-coordinate and y-coordinate of the pixel being predicted in the current block, respectively. The encoder can use a signal notification flag (e.g., reference row number, index, or tag) to indicate which reference layer has been selected for intra-frame prediction; that is, the index (mrl_idx) of the selected reference row is signaled and used to generate the intra-frame predictor. The decoder obtains this signal notification flag after decoding to combine it with the intra-frame prediction mode for decoding or reconstruction of the current block.

[0187] The intra-frame prediction method provided in this disclosure selects M reference layers from N reference layers in the current block, determines the derivation direction, and fuses the M reference layers along the derivation direction to obtain the (N+1)th reference layer of the current block. Then, intra-frame prediction is performed on the current block based on the N reference layers and the (N+1)th reference layer. This allows the intra-frame prediction method to fuse multiple reference lines, thereby improving the accuracy and coding efficiency of intra-frame prediction coding. This method can be applied to products equipped with relevant video codecs or video compression.

[0188] Figure 6 This is an example of a multi-reference row fusion direction based on intra-frame prediction mode (angle) (along the lower right 45 degrees). It should be noted that when calculating the fused reference row N, some reference pixels at the edges may not be used, for example, L(m+n)3. When generating the fused reference row N, if the weights of different reference rows are the same, then the upper right 45 degrees and the lower right 45 degrees can be the same.

[0189] In an example, the intra prediction modes can include two non-directional (or non-angular) intra prediction modes and a plurality of directional (or angular) intra prediction modes. The non-directional intra prediction modes can include a planar intra prediction mode numbered 0 and a DC intra prediction mode numbered 1.

[0190] As shown in Figure 8 , the directional intra prediction modes can include 33 intra prediction modes between an intra prediction mode numbered 2 and an intra prediction mode numbered 34. Figure 8 33 angular intra prediction modes are shown. Among them, mode 10 is horizontal mode, mode 26 is vertical mode, and modes 2 / 18 and 34 are diagonal modes.

[0191] To capture arbitrary edge directions present in natural videos, the number of directional intra modes is extended from 33 to 65. And the planar and DC modes remain the same. These denser directional intra prediction modes apply to all block sizes and both luma and chroma intra prediction. As shown in Figure 9 , the directional intra prediction modes can include 65 intra prediction modes between an intra prediction mode numbered 2 and an intra prediction mode numbered 66. The intra prediction modes with horizontal directionality and the intra prediction modes with vertical directionality can be distinguished around the intra prediction mode numbered 34 with left-up diagonal prediction direction. The intra prediction modes numbered 2 to 33 have horizontal directionality, while the intra prediction modes numbered 34 to 66 have vertical directionality. The intra prediction mode numbered 18 and the intra prediction mode numbered 50 can be horizontal intra prediction mode and vertical intra prediction mode, respectively; the intra prediction mode numbered 2 can be referred to as left-down diagonal intra prediction mode; the intra prediction mode numbered 34 can be referred to as left-up diagonal intra prediction mode; and the intra prediction mode numbered 66 can be referred to as right-up diagonal intra prediction mode.

[0192] As shown in Figure 10 , the directional intra prediction modes can include 85 intra prediction modes between an intra prediction mode numbered 2 and an intra prediction mode numbered 76, and numbered -1 to -10. But the present disclosure is not limited thereto, for example, the intra prediction modes can include two non-directional intra prediction modes and 129 directional intra prediction modes. The directional intra prediction modes can include intra prediction modes numbered 2 to 130. Among them, mode 18 is horizontal mode, mode 50 is vertical mode, and modes 2 / 34 and 66 are diagonal modes. Modes -1 to -10 and mode 67 are referred to as wide-angle intra prediction modes.

[0193] Intra-prediction in directional intra-prediction mode can be applied to blocks of all sizes and to both the luma and chroma components. However, this is just an example, and the configuration of intra-prediction mode can vary. Intra-prediction modes can be indexed.

[0194] Intra-prediction types (or additional intra-prediction modes, etc.) may include MRL (Multiple Reference Line intra-prediction). Intra-prediction types can be indicated based on intra-prediction type information, and this information can be implemented in various forms. In one example, intra-prediction type information may include intra-prediction type index information indicating one of the intra-prediction types: reference line information (e.g., intra_luma_ref_idx) relating to whether MRL is applied to the current block and, if so, which reference line is used.

[0195] In this embodiment of the disclosure, to reduce complexity and unnecessary coding overhead, when using non-adjacent reference lines for intra-frame prediction, the encoder and decoder can agree that only intra-frame prediction modes from the MPM (most probable mode, which can be composed of intra-frame prediction modes already used by neighboring blocks of the current block) list can be used. That is, when using MRL technology, since the encoder needs to encode which reference line (which can be an adjacent reference line or a non-adjacent reference line) has been selected, in order to reduce coding costs, it can be stipulated that only intra-frame prediction modes from the MPM list can be selected at this time.

[0196] For example, with Figure 9 For example, suppose there are 67 (or other) intra-prediction modes (DC mode, planar mode, 65 angles) in total, and 6 of them are selected and added to the MPM list. Each intra-prediction mode in the MPM list will have its own MPM index. At this time, the encoder only needs to tell the decoder which MPM index it has selected, without needing to encode all 67 intra-prediction modes, thus reducing the encoding cost.

[0197] The intra-prediction mode in the MPM list can be determined based on the intra-prediction modes used by the neighboring blocks of the current block. For example, if the neighboring blocks A and F of the current block use either DC mode or angle mode, then the MPM list will include that mode. However, this disclosure is not limited to this. The construction of the MPM list can follow agreed-upon rules. Both the encoder and decoder can construct the MPM list. The encoder does not need to encode the MPM list to inform the decoder, and the decoder does not need to perform signal decoding to construct the MPM list itself.

[0198] When an image is divided into blocks, a current block to be coded and neighboring blocks have similar image characteristics. Thus, the current block and the neighboring blocks can have the same or similar intra prediction modes. Accordingly, the encoder can use the intra prediction modes of the neighboring blocks to code the intra prediction mode of the current block.

[0199] In an exemplary embodiment, as the number of intra prediction modes increases, the number of MPM candidates can also increase. Thus, the number of MPM candidates can be different depending on the number of intra prediction modes. However, it is not always the case that the number of MPM candidates increases as the number of intra prediction modes increases. For example, when there are 35 intra prediction modes or when there are 67 intra prediction modes, there can be various numbers of MPM candidates such as 3, 4, 5, or 6 depending on the design. For example, the encoder / decoder can construct an MPM list including 6 MPMs.

[0200] In the present disclosure, a specific term or sentence is used to define a specific information or concept. When the specific term or sentence used to define the specific information or concept is explained throughout the specification, an explanation should not be made limited to the name, but the term should be explained according to the content intended to be expressed by the term, considering various operations, functions, and effects.

[0201] The methods presented herein can be used alone or in any combination. Further, each of the methods (or embodiments), encoders, and decoders can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program that is stored in a non-transitory computer-readable medium.

[0202] The method provided by the embodiments of the present disclosure can be applied to both the encoding end (e.g., a video encoder) and the decoding end (e.g., a video decoder), that is, the same intra prediction method is performed at the encoding end and the decoding end.

[0203] Figure 11 A block diagram of an intra prediction apparatus according to some embodiments of the present disclosure is schematically shown. As Figure 11 shown, the intra prediction apparatus 1100 provided by the embodiments of the present disclosure includes a processing unit 1110.

[0204] The processing unit 1110 is configured to determine N reference layers of a current block, N being a positive integer greater than 1. The processing unit 1110 is further configured to select M reference layers from the N reference layers, M being a positive integer greater than 1 and less than N. The processing unit 1110 is further configured to determine a derivation direction of the M reference layers. The processing unit 1110 is further configured to fuse the M reference layers along the derivation direction to obtain an (N+1)th reference layer of the current block. The processing unit 1110 is further configured to perform intra prediction on the current block according to the N reference layers and the (N+1)th reference layer of the current block.

[0205] In an example embodiment, the processing unit 1110 is further configured to select the first reference layer to the Nth reference layer as the M reference layers, or select the second reference layer to the Nth reference layer as the M reference layers.

[0206] In an example embodiment, the processing unit 1110 is further configured to use a default derivation direction as the derivation direction of the M reference layers, or determine the derivation direction of the M reference layers according to an intra prediction mode of the current block.

[0207] In an example embodiment, the processing unit 1110 is further configured to, if the intra prediction mode of the current block is an angular prediction mode, determine the derivation direction of the M reference layers according to an angle corresponding to the angular prediction mode, or use a default derivation direction as the derivation direction of the M reference layers if the intra prediction mode of the current block is a non-angular prediction mode.

[0208] In an example embodiment, the M reference layers include left-side reference pixels and upper-side reference pixels of the current block. The processing unit 1110 is further configured to, for the left-side reference pixels of the M reference layers, map the angle to a candidate horizontal direction as the derivation direction, and for the upper-side reference pixels of the M reference layers, map the angle to a candidate vertical direction as the derivation direction.

[0209] In an example embodiment, the candidate horizontal direction includes at least one of a horizontal direction, a horizontal upward 45 degrees, and a horizontal downward 45 degrees, and the candidate vertical direction includes at least one of a vertical direction, a vertical leftward 45 degrees, and a vertical rightward 45 degrees.

[0210] In an example embodiment, the processing unit 1110 is further configured to use the angle as the derivation direction of the M reference layers.

[0211] In an example embodiment, the processing unit 1110 is further configured to, if the angle is not at an integer pixel position of the M reference layers, interpolate reference pixels of the M reference layers along the angle to generate interpolated reference pixels, and fuse the interpolated reference pixels to obtain the (N+1)th reference layer of the current block.

[0212] In an example embodiment, the M reference layers include left-side reference pixels and upper-side reference pixels of the current block. The processing unit 1110 is further configured to: for the left-side reference pixels of the M reference layers, the default derivation direction is a horizontal direction; and for the upper-side reference pixels of the M reference layers, the default derivation direction is a vertical direction.

[0213] In an example embodiment, the processing unit 1110 is further configured to: extract reference pixels of the M reference layers along the derivation direction, and perform linear weighting on the reference pixels of the M reference layers to obtain the N+1th reference layer of the current block. Different weights or the same weight are configured for the reference pixels of different reference layers involved in the weighting.

[0214] In an example embodiment, the processing unit 1110 is further configured to: extract reference pixels of the M reference layers along the derivation direction, and perform non-linear fusion on the reference pixels of the M reference layers to obtain the N+1th reference layer of the current block.

[0215] In an example embodiment, the processing unit 1110 is further configured to: select a pixel value that appears most frequently among the reference pixels of the M reference layers along the derivation direction as a pixel value in the N+1th reference layer; or select a median value of the pixel values of the reference pixels of the M reference layers along the derivation direction as the pixel value in the N+1th reference layer; or select an average value of a plurality of pixel values of a plurality of reference pixels closest to each other among the reference pixels of the M reference layers along the derivation direction as the pixel value in the N+1th reference layer.

[0216] In an example embodiment, the processing unit 1110 is further configured to: select the N+1th reference layer to perform intra prediction on the current block; or select one of the first reference layer and the N+1th reference layer to perform intra prediction on the current block; or replace any one of the second reference layer to the Nth reference layer with the N+1th reference layer to form new N reference layers, and select one of the new N reference layers to perform intra prediction on the current block. The first reference layer of the N reference layers is a neighboring reference layer of the current block, and the second reference layer to the Nth reference layer are non-neighboring reference layers of the current block.

[0217] Figure 11 Other contents of the embodiments can refer to the above-described embodiments.

[0218] The embodiments of the present disclosure provide a computer readable storage medium, which stores a computer program. The program is executed by a processor to implement the method described in the above-described embodiments.

[0219] The electronic device provided by the embodiments of the present disclosure includes at least one processor, and a storage device configured to store at least one program, when the at least one program is executed by the at least one processor, the at least one processor implements the method as described in the above embodiments.

[0220] Figure 12 A structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown.

[0221] It should be noted that, Figure 12 The electronic device 120 shown is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.

[0222] As Figure 12 shown, the electronic device 120 includes a central processing unit (CPU) 121, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 122 or programs loaded from a storage portion 128 into a random access memory (RAM) 123. In the RAM 123, various programs and data required for system operation are also stored. The CPU 121, the ROM 122, and the RAM 123 are connected to each other through a bus 124. An input / output (I / O) interface 125 is also connected to the bus 124.

[0223] The following components are connected to the I / O interface 125: an input portion 126 including a keyboard, a mouse, and the like; an output portion 127 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage portion 128 including a hard disk, and the like; and a communication portion 129 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication portion 129 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 125 as necessary. A removable recording medium 1211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 1210 as necessary, so that a computer program read therefrom is installed in the storage portion 128 as necessary.

[0224] In particular, according to embodiments of the present disclosure, the processes described below with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication part 129, and / or installed from the detachable medium 1211. When the computer program is executed by the central processing unit (CPU) 121, various functions defined in the methods and / or apparatuses of the present application are executed.

[0225] It should be noted that the computer-readable storage medium shown in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having at least one wire, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM (Erasable Programmable Read-Only Memory) or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF (Radio Frequency), etc., or any suitable combination of the above.

[0226] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of methods, apparatuses and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment, or a portion of code, which comprises at least one executable instruction for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the drawings. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the reverse order, depending on the functionality involved. It will also be noted that each block in the flowcharts or block diagrams and combinations of blocks in the flowcharts or block diagrams can be implemented by special-purpose hardware-based systems which perform the specified functions or operations, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0227] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware, or by a combination of software and hardware. The names of the units described in the embodiments of the present disclosure do not constitute a limitation on the units per se in some cases.

[0228] As another aspect, the present disclosure also provides a computer-readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device implements the method described in the above embodiments. For example, the electronic device can implement each step in the method provided by any of the embodiments of the present disclosure.

[0229] From the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, or the like) or a network, and includes a number of instructions to make a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) execute the methods according to the embodiments of the present disclosure.

[0230] It should be understood that the present disclosure is not limited to the precise structures described above and illustrated in the drawings and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An intra prediction method, characterized by, The method comprises the following steps: determining N reference layers of a current block, N being a positive integer greater than 1; selecting M reference layers from the N reference layers, M being a positive integer greater than 1 and less than N; determining a derivation direction of the M reference layers; fusing the M reference layers along the derivation direction to obtain an (N+1)th reference layer of the current block; performing intra prediction on the current block according to the N reference layers and the (N+1)th reference layer of the current block.

2. The method of claim 1, wherein, The step of selecting M reference layers from the N reference layers comprises: selecting the first reference layer to the Nth reference layer as the M reference layers; or selecting the second reference layer to the Nth reference layer as the M reference layers.

3. The method of claim 1, wherein, The step of determining a derivation direction of the M reference layers comprises: adopting a default derivation direction as the derivation direction of the M reference layers; or determining the derivation direction of the M reference layers according to an intra prediction mode of the current block.

4. The method of claim 3, wherein, The step of determining the derivation direction of the M reference layers according to the intra prediction mode of the current block comprises: if the intra prediction mode of the current block is an angle prediction mode, determining the derivation direction of the M reference layers according to an angle corresponding to the angle prediction mode; or if the intra prediction mode of the current block is a non-angle prediction mode, adopting a default derivation direction as the derivation direction of the M reference layers.

5. The method of claim 4, wherein, The M reference layers comprise left-side reference pixels and upper-side reference pixels of the current block; and the step of determining the derivation direction of the M reference layers according to the angle corresponding to the angle prediction mode comprises: for the left-side reference pixels of the M reference layers, mapping the angle to a candidate horizontal direction as the derivation direction; and for the upper-side reference pixels of the M reference layers, mapping the angle to a candidate vertical direction as the derivation direction.

6. The method of claim 5, wherein, The candidate horizontal direction comprises at least one of a horizontal direction, a horizontal upward 45-degree direction and a horizontal downward 45-degree direction. The candidate vertical direction comprises at least one of a vertical direction, a vertical leftward 45-degree direction and a vertical rightward 45-degree direction.

7. The method of claim 4, wherein, The step of determining the derivation direction of the M reference layers according to the angle corresponding to the angle prediction mode comprises: adopting the angle as the derivation direction of the M reference layers.

8. The method of claim 7, wherein, The step of fusing the M reference layers along the derivation direction to obtain the (N+1)th reference layer of the current block comprises: if the angle is not located at an integer pixel position of the M reference layers, interpolating reference pixels of the M reference layers along the angle to generate interpolated reference pixels; and fusing the interpolated reference pixels to obtain the (N+1)th reference layer of the current block.

9. The method according to claim 3 or 4, characterized in that, The M reference layers comprise left-side reference pixels and upper-side reference pixels of the current block; and the step of adopting a default derivation direction as the derivation direction of the M reference layers comprises: for the left-side reference pixels of the M reference layers, the default derivation direction is a horizontal direction; and for the upper-side reference pixels of the M reference layers, the default derivation direction is a vertical direction.

10. The method of claim 1, wherein, The step of fusing the M reference layers along the derivation direction to obtain the (N+1)th reference layer of the current block comprises: extracting reference pixels of the M reference layers along the derivation direction, linearly weighting the reference pixels of the M reference layers to obtain the (N+1)th reference layer of the current block. Different weights or same weights are configured to reference pixels of different reference layers participating in the weighting.

11. The method of claim 1, wherein, M reference layers are fused along the derivation direction to obtain the N+1th reference layer of the current block, including: Reference pixels of M reference layers are extracted along the derivation direction, and the reference pixels of M reference layers are nonlinearly fused to obtain the N+1th reference layer of the current block.

12. The method of claim 11, wherein, The nonlinear fusion of the reference pixels of M reference layers includes: The pixel value with the highest frequency of occurrence in the reference pixels of M reference layers along the derivation direction is selected as the pixel value in the N+1th reference layer; or, The median value of the pixel values of the reference pixels of M reference layers along the derivation direction is selected as the pixel value in the N+1th reference layer; or, The average value of the pixel values of the closest multiple reference pixels in the reference pixels of M reference layers along the derivation direction is selected as the pixel value in the N+1th reference layer.

13. The method of claim 1, wherein, The current block is intra-predicted according to the N reference layers and the N+1th reference layer of the current block, including: The N+1th reference layer is selected to intra-predict the current block; or, One of the first reference layer and the N+1th reference layer is selected to intra-predict the current block; or, Any one of the second reference layer to the Nth reference layer is replaced by the N+1th reference layer to form new N reference layers, and one of the new N reference layers is selected to intra-predict the current block; Wherein, the first reference layer in the N reference layers is the neighboring reference layer of the current block, and the second reference layer to the Nth reference layer is the non-neighboring reference layer of the current block.

14. An apparatus for intra prediction, the apparatus comprising: Including: A processing unit is configured to determine N reference layers of a current block, N being a positive integer greater than 1; The processing unit is further configured to select M reference layers from the N reference layers, M being a positive integer greater than 1 and less than N; The processing unit is further configured to determine a derivation direction of the M reference layers; The processing unit is further configured to fuse the M reference layers along the derivation direction to obtain the N+1th reference layer of the current block; The processing unit is further configured to intra-predict the current block according to the N reference layers and the N+1th reference layer of the current block.

15. An electronic device, comprising: Including: At least one processor; A storage device configured to store at least one program, when the at least one program is executed by the at least one processor, the at least one processor implements the method of any one of claims 1 to 13.

16. A computer-readable storage medium, the computer-readable storage medium storing a computer program, characterized in that, When the computer program runs on the computer, the computer is caused to execute the method of any one of claims 1 to 13.

17. A computer program product, characterised in that, Including a computer program, when the computer program is executed by a processor, the method of any one of claims 1 to 13 is implemented.