Method, apparatus and computer program for multi-line intra prediction in video compression
By optimizing intra prediction methods for multiple reference lines and adjusting mode settings based on line indices, the method enhances video coding efficiency and reduces redundancy, addressing inefficiencies in existing technologies like HEVC.
Patent Information
- Application Number
- JP2024083751
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-01-04
- Filing Date
- 2024-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2039-06-27
AI Technical Summary
Existing video coding technologies, such as HEVC, face inefficiencies in multiple-line intra prediction, including reliance on a single reference line, limited application to luma components, inconsistent mode settings for different reference lines, underutilization of adjacent pixel trends, and lack of optimal mode selection and plane or DC mode utilization.
Implement a method and apparatus that utilize a processor to perform intra prediction between multiple reference lines, set modes for zero and non-zero reference lines, signal the most accurate mode flags, and adjust the length of the mode list based on reference line indices, excluding planar and DC modes for non-zero lines.
Enhances video coding efficiency by optimizing mode selection and utilizing multiple reference lines, improving coding gain and reducing redundancy in video compression.
Smart Images

Figure 0007698108000013 
Figure 0007698108000014 
Figure 0007698108000015
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Patent Application No. 16 / 240,388, filed on January 4, 2019, entitled "METHODS AND APPARATUS FOR MULTIPLE LINE INTRA PREDICTION IN VIDEO COMPRESSION", which is a continuation of U.S. Patent Application No. 16 / 234,324, filed on December 27, 2018, which claims the priority of U.S. Provisional Patent Application No. 62 / 694,132, filed on July 5, 2018. Each of the above applications is hereby expressly incorporated by reference in its entirety into this application.
Background Art
[0002] The present disclosure relates to next - generation video coding technology that outperforms HEVC, and more particularly, for example, to improving the intra - prediction method using multiple reference lines.
[0003] The main profile of the video coding standard HEVC (High Efficiency Video Coding) was completed in 2013. Soon after that, the international standardization organizations ITU - T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) began to explore the need to develop a future video coding standard, along with the possibility of significantly enhancing the compression performance compared to the current HEVC standard (including its current extensions). Multiple groups are working together in a joint effort known as the Joint Video Exploration Team (JVET) to evaluate the compression technology designs proposed by experts in this field. To explore video coding technologies that outperform the performance of HEVC, the Joint Exploration Model (JEM) has been developed by JVET, and the current latest version of JEM is JEM - 7.1.
[0004] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published Version 1 of the H. 265 / HEVC (High Efficiency Video Coding) standard in 2013, Version 2 in 2014, Version 3 in 2015, and Version 4 in 2016. Since then, they have continued to research the potential need for standardization of future video coding technologies that have compression performance far exceeding the HEVC standard (including its extensions). In October 2017, a Call for Proposals (CfP) for video compression with performance superior to HEVC was jointly conducted. By February 15, 2018, a total of 22 CfP responses for standard dynamic range (SDR), 12 CfP responses for high dynamic range (HDR), and 12 CfP responses for 360 video classification were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th JVET (Joint Video Exploration Team - Joint Video Experts Team) meeting. After careful evaluation, JVET officially started the standardization of the next-generation video coding superior to HEVC, namely, the so-called Versatile Video Coding (VVC). The current version of the VTM (VVC Test Model) is VTM1.
Summary of the Invention
Problems to be Solved by the Invention
[0005] Even if multiple lines are available, there are various technical problems in this field. For example, there are technical problems in that it has been found that the first reference line is still the most selected line. However, each block having the first reference line always needs to signal one bin to indicate the line index of the current block.
[0006] Also, multiple-line intra prediction is only applied to luma intra prediction. The possible coding gain of multiple-line intra prediction for chroma components is not utilized.
[0007] Also, reference samples with different line indices may have different characteristics, and thus it is not optimal to set the same number of intra prediction modes for different reference lines.
[0008] Also, in multiple-line intra prediction, pixels of multiple adjacent lines are stored and accessed, but the pixels of adjacent lines are not used for smoothing the pixels within the current line.
[0009] Also, in multiple-line intra prediction, the encoder selects one reference line to predict the pixel values of the current block, but the change trend of adjacent pixels is not used for predicting the samples within the current block.
[0010] Also, in multiple-line intra prediction, there is no number of planes or DC mode greater than 1. The search for other versions of the DC or plane mode is not fully utilized.
[0011] Also, multiple line reference pixels are applied to intra prediction, but the coding gain of multiple line reference pixels is not utilized even though there are other places where the reference pixels are used.
[0012] Therefore, a technical solution to such problems is desired.
Means for Solving the Problem
[0013] A method and apparatus are included that comprise a memory configured to store computer program code and a hardware processor or processor configured to access the computer program code and operate as commanded by the computer program code. The computer program is configured to cause the processor to encode or decode a video sequence by performing intra prediction between a plurality of reference lines of the video sequence, an intra prediction code; cause the processor to set an intra prediction mode for a first reference line at a zero reference line closest to a current block of intra prediction between a plurality of non-zero reference lines, an intra prediction mode code; and cause the processor to set one or more most accurate modes for a second reference line at a non-zero reference line, a most accurate mode code.
[0014] According to an exemplary embodiment, the program code further includes signal code configured to cause the processor to signal a reference line index before signaling the most accurate mode flag and the intra mode, and in response to the reference line index being signaled and the signaled index indicating a zero reference line, signal the most accurate mode flag, and in response to the reference line index being signaled and the signaled index indicating at least one of the non-zero reference lines, derive the most accurate mode flag to be true without signaling the most accurate mode flag and signal the most accurate mode index of the current block.
[0015] According to an exemplary embodiment, the most accurate mode code is further configured to cause the processor to include one or more most accurate modes in the most accurate mode list and exclude the planar mode and the DC mode from the most accurate mode list.
[0016] According to an exemplary embodiment, the most accurate mode code causes at least one processor to set the length of the most accurate mode list based on a reference line index value, such that the length of the most accurate mode list further comprises the number of one or more most accurate modes.
[0017] According to an exemplary embodiment, the most accurate mode code further configures at least one processor to set the length of the most accurate mode list to either 1 or 4 in response to detection of a non-zero reference line, and to set the length of the most accurate mode list to 3 or 6 in response to a determination that the current reference line is a zero reference line.
[0018] According to an exemplary embodiment, the most accurate mode code further configures at least one processor to set the length of the most accurate mode list to consist of five most accurate modes in response to detection of a non-zero reference line.
[0019] According to an exemplary embodiment, one of the non-zero reference lines is a line adjacent to the current block and farther away from the current block than the zero reference line.
[0020] According to an exemplary embodiment, the one or more most accurate modes include the most accurate modes at any level from the lowest level to the highest level of the most accurate modes.
[0021] According to an exemplary embodiment, the one or more most accurate modes include only the levels of the most accurate modes permitted for the non-zero reference line.
[0022] Further features, properties, and various advantages of the subject matter of the present disclosure will become more apparent from the following detailed description and the accompanying drawings.
Brief Description of the Drawings
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
[0024] The proposed features described below may be used individually or in combination in any order. Also, the embodiments may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. In the present disclosure, the most probable mode (MPM) can represent the main MPM, the secondary MPM, or both the main MPM and the secondary MPM.
[0025] FIG. 1 shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The communication system 100 includes at least two terminals 102 and 103 interconnected via a network 105. In the case of unidirectional data transmission, the first terminal 103 may encode video data to be transmitted to the other terminal 102 via the network 105 at a local location. The second terminal 102 may receive the encoded video data of the other terminal from the network 105, decode the encoded data, and display the recovered video data. Unidirectional data transmission would be common in applications such as media delivery.
[0026] FIG. 1 shows a second pair of terminals 101 and 104 provided to support two-way transmission of encoded video, which may occur, for example, during a video conference. In the case of two-way data transmission, each of the terminals 101 and 104 may encode video data captured at a local location and transmit it to the other terminal via the network 105. Each of the terminals 101 and 104 may receive the encoded video data transmitted by the other terminal, may decode the encoded data, and may display the recovered video data on a local display device.
[0027] In FIG. 1, the terminals 101, 102, 103, and 104 may be shown as servers, personal computers, and smartphones, but the principles of the present disclosure are not so limited. Embodiments of the present disclosure are also applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network 105 represents any number of networks including wired and / or wireless communication networks that transmit encoded video data between the terminals 101, 102, 103, and 104. The communication network 105 may exchange data over a circuit-switched channel and / or a packet-switched channel. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network 105 may be irrelevant to the operation of the present disclosure unless otherwise described below.
[0028] FIG. 2 shows the arrangement of video encoders and decoders in a streaming environment as an application example of the disclosed subject matter. The disclosed subject matter can be equally applied to other video usage scenarios including, for example, video conferencing, digital television, and storage of compressed video on digital media such as CDs, DVDs, and memory sticks.
[0029] The streaming system may include a capture subsystem 203, which can include a video source 201, such as a digital camera, that creates an uncompressed video sample stream 213, for example. The sample stream 213, shown in thick lines to emphasize the large amount of data compared to the encoded video bitstream, can be processed by an encoder 202 coupled to the camera 201. As will be described in more detail below, the encoder 202 can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter. The encoded video bitstream 204, shown in thin lines to emphasize the small amount of data compared to the sample stream, can be stored in a streaming server 205 for later use. One or more streaming clients 212 and 207 can access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. The client 212 can include a video decoder 211 that decodes the received copy 208 of the encoded video bitstream to generate a transmitted video sample stream 210, which can be displayed on a display device 209 or another display device (not shown). In some streaming systems, the video bitstreams 204, 206, and 208 can be encoded according to several video encoding / compression standards. Examples of such standards are described above and further described herein.
[0030] FIG. 3 is a functional block diagram of a video decoder 300 according to an embodiment of the present invention.
[0031] Receiver 302 may receive one or more codec video sequences decoded by decoder 300, and in the same or another embodiment, may receive one encoded video sequence at the same time. The decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from channel 301, and channel 301 may be hardware / software that is coupled to a storage device that stores the encoded video data. Receiver 302 may receive the encoded video data together with other data such as encoded audio data and / or auxiliary data streams, which may be transferred to entities (not shown) that each uses. Receiver 302 may separate the encoded video sequence from other data. To counter network jitter, buffer memory 303 may be coupled between receiver 302 and entropy decoder / syntax analyzer 304 (hereinafter referred to as "syntax analyzer"). When receiver 302 is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isosynchronous network, buffer 303 may not be necessary or may be made smaller. Buffer 303 may be required for use in a best-effort packet network such as the Internet, is relatively large, and may preferably be of an adaptable size.
[0032] Video decoder 300 may include syntax analyzer 304 to reconstruct symbol 313 from the entropy-encoded video sequence. Such classification of symbols includes information used to manage the operation of decoder 300 and potential information for controlling a display device, such as display device 312, which is not an integral part of the decoder but can be coupled thereto. Control information for the (multiple) display devices is Supplementary Enhancement Information (SEI message), or Video Usability The set of Information, VUI) parameters may be in the form of a fragment (not shown). The syntax analyzer 304 may perform syntax analysis / entropy decoding on the received encoded video sequence. The code of the encoded video sequence may conform to a video encoding technology or standard, and may follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, context-dependent or context-independent arithmetic coding, etc. The syntax analyzer 304 may extract a group of sub-group parameters for at least one of the sub-groups of pixels in the video decoder based on at least one parameter corresponding to a group from the encoded video sequence. The sub-groups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. Also, the entropy decoder / syntax analyzer may extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0033] The syntax analyzer 304 may perform an entropy decoding / syntax analysis operation on the video sequence received from the buffer 303 to generate the symbol 313. The syntax analyzer 304 may receive the encoded data and selectively decode a specific symbol 313. Also, the syntax analyzer 304 may determine whether to provide the specific symbol 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.
[0034] Symbol reconstruction 313 can include a plurality of different units depending on the type of the encoded video or a portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block), as well as other elements. How each unit is included can be controlled by subgroup control information parsed from the video sequence encoded by the parser 304. Such a flow of subgroup control information between the parser 304 and the following plurality of units is not shown for clarity.
[0035] In addition to the function blocks already described, the decoder 200 can be conceptually divided into several functional units as described below. In actual implementations operating under commercial constraints, many of such units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, it is appropriate to conceptually divide it into the following functional units.
[0036] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives quantization transform coefficients and control information, which includes (a plurality of) symbols 313 from the parser 304, such as the transform, block size, quantization factor, quantization scaling matrix to be used, etc. The scaler / inverse transform unit 305 can output a block including sample values, which can be input to the aggregator 310.
[0037] In some cases, the output samples of the scaler / inverse transform unit 305 can be related to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed image can use prediction information from a previously reconstructed part of the current image. Such prediction information can be provided by the intra-image prediction unit 307. In some cases, the intra-image prediction unit 307 uses surrounding already reconstructed information taken from the current (partially reconstructed) image 309 to generate a block of the same size and shape as the block being reconstructed. The aggregator 310 optionally adds the prediction information generated by the intra-prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305 on a sample-by-sample basis.
[0038] In other cases, the output samples of the scaler / inverse transform unit 305 can be related to inter-coded and potentially motion-compensated blocks. In such cases, the motion-compensation prediction unit 306 can access the reference image memory 308 to retrieve the samples to be used for prediction. After motion-compensating the retrieved samples according to the symbols 313 associated with the block, these samples can be added by the aggregator 310 to the output of the scaler / inverse transform unit (in this case called the residual samples or residual signal) to generate the output sample information. The address in the reference image memory from which the motion-compensation unit retrieves the prediction samples can be controlled by the motion vector and is available to the motion-compensation unit in the form of the symbol 313 and can have, for example, X, Y, and reference image components. Motion compensation can also include interpolation of the sample values retrieved from the reference image memory when an exact motion vector of sub-samples is used, a motion vector prediction mechanism, etc.
[0039] The output samples of the aggregation device 310 can be subjected to various loop filtering techniques of the loop filter unit 311. The video compression technique can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and can be made available to the loop filter unit 311 as symbols 313 from the syntax analyzer 304. Furthermore, it can respond to meta information obtained during the decoding of a previous (in decoding order) portion of the encoded image or encoded video sequence and can similarly respond to previously reconstructed and loop-filtered sample values.
[0040] The output of the loop filter unit 311 can be a sample stream that can be output to the display device 312 and stored in the reference image memory 557 for use in subsequent inter-picture prediction.
[0041] Some encoded images can be used as reference images for subsequent prediction once they are fully reconstructed. When an encoded image is fully reconstructed and the encoded image is specified as a reference image (e.g., by the syntax analyzer 304), the current reference image 309 can become part of the reference image buffer 308, and a new current image memory can be reallocated before starting the reconstruction of subsequent encoded images.
[0042] The video decoder 300 may perform a decoding operation in accordance with a predetermined video compression technique that may be described in a standard such as ITU-T Rec. H.265. The encoded video sequence is in accordance with the syntax specified by the video compression technique or standard used, and specifically in the profile described therein, in the sense of complying with the syntax of the video compression technique or standard. What is further required for compliance is that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level may limit the maximum image size, maximum frame rate, maximum reconstructed sample rate (e.g., measured in megasamples per second), maximum reference image size, etc. The limitations set by the level may, in some cases, be further restricted by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0043] In an embodiment, the receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the (multiple) encoded video sequences. The additional data may be used by the video decoder 300 to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.
[0044] FIG. 4 is a functional block diagram of a video encoder 400 according to an embodiment of the present disclosure.
[0045] The encoder 400 may receive video samples from a video source 401 (not part of the encoder) that may capture the (multiple) videos to be encoded by the encoder 400.
[0046] The video source 401 can provide the source video sequence to be encoded by the encoder 400 in the form of a digital video sample stream that can be of any suitable bit depth (such as 8 bits, 10 bits, 12 bits, etc.), any color space (such as BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (such as Y CrCb 4:2:0, Y CrCb 4:4:4, etc.). In a media supply system, the video source 401 may be a storage device that stores previously prepared videos. In a video conferencing system, the video source 401 may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual images that convey motion when viewed in sequence. The image itself may be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. during use. A person skilled in the art will be able to easily understand the relationship between pixels and samples. Hereinafter, the description will focus on samples.
[0047] According to an embodiment, the encoder 400 can encode and compress the images of the source video sequence in real time or under other time constraints required by the application to obtain the encoded video sequence 410. Achieving an appropriate encoding speed is one function of the controller 402. The controller controls other functional units as described later and is functionally coupled to these units. For clarity, the couplings are not shown. The parameters set by the controller can include rate control related parameters (such as picture skip, quantization, lambda value of rate-distortion optimization techniques, etc.), picture size, layout of groups of pictures (GOP), maximum motion vector search range, etc. A person skilled in the art can easily identify other functions of the controller 402 because they may be related to the video encoder 400 optimized for several system designs.
[0048] Some video encoders operate in what is readily recognized by those skilled in the art as an "encoding loop." Although this is an overly simplified explanation, the encoding loop may consist of an encoding section of an encoder 403 (hereinafter referred to as the "source encoder") that is involved in generating symbols based on the input image to be encoded and the (plural) reference images, and a (local) decoder 406 incorporated in an encoder 400 that reconstructs symbols to generate sample data that a (remote) decoder also generates because the compression between the symbols and the encoded video bitstream is reversible in the video compression technology considered in the disclosed subject matter. The reconstructed sample stream is input to a reference image memory 405. When the decoding of the symbol stream results in being bit-exact regardless of the position of the decoder (local or remote), the contents of the reference image buffer are also bit-exact between the local encoder and the remote encoder. In other words, the prediction section of the encoder "considers" the reference image samples to have exactly the same sample values as what the decoder "considers" when using prediction during decoding. This basic principle of reference image synchronicity (and the resulting drift if synchronicity cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.
[0049] The operation of the "local" decoder 406 may be the same as that of the "remote" decoder 300, which has already been described in detail above in connection with FIG. 3. However, referring temporarily to FIG. 4 as well, since the symbols are available and can be reversibly encoded / decoded into an encoded video sequence by an entropy encoder 408 and a syntax analyzer 304, the entropy decoding section of the decoder 300 including a channel 301, a receiver 302, a buffer 303, and a syntax analyzer 304 may not be fully implemented by the local decoder 406.
[0050] What can be considered at present is that all the decoder technologies existing in the decoder, excluding syntax analysis / entropy decoding, must of course exist in the corresponding encoder in almost the same functional form. Since the description of the encoder technology is the reverse of the decoder technology described comprehensively, it can be omitted. More detailed explanations are required and will be described below only in some areas.
[0051] The source encoder 403 may perform motion-compensated predictive coding as part of its operation, and predictively encode the input frame with respect to one or more previously encoded frames from the video sequence designated as the "reference frame". In this method, the encoding engine 407 encodes the difference between the pixel block of the input frame and the pixel blocks of the (plural) reference frames that can be selected as the (plural) prediction references for the input frame.
[0052] The local video decoder 406 may decode the encoded video data of the frame that can be designated as the reference frame based on the symbols generated by the source encoder 403. The operation of the encoding engine 407 may preferably be an irreversible process. When the encoded video data may be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may be a copy of the source video sequence, usually with some errors. The local video decoder 406 may perform the decoding process that may be performed on the reference frame by the video decoder, and may cause the reconstructed reference frame to be stored in the reference image cache 405. In this method, the encoder 400 may locally store a copy of the reconstructed reference frame having the same content as the reconstructed reference frame obtained by the remote video decoder (without transmission errors).
[0053] Predictor 404 may perform a predictive search for the encoding engine 407. That is, predictor 404 may search the reference image memory 405 for sample data (as candidate reference pixel blocks), or some metadata such as reference image motion vectors, block shapes, etc. for the new frame to be encoded, which functions as an appropriate predictive reference for the new image. Predictor 404 may operate on samples on a block-by-pixel block basis to find an appropriate predictive reference. In some cases, the input image may have a predictive reference drawn from a plurality of reference images stored in the reference image memory 405 as determined by the search results obtained by predictor 404.
[0054] Controller 402 may manage the encoding operation of video coder 403, including the setting of parameters and subgroup parameters used for encoding video data, for example.
[0055] The outputs of all the functional units described above may be entropy encoded by entropy coder 408. The entropy coder converts the symbols into an encoded video sequence by reversibly compressing the symbols using techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc. when generated by various functional units.
[0056] Transmitter 409 may buffer the (plural) encoded video sequence when generated by entropy coder 408 in preparation for transmission via communication channel 411, which may be a hardware / software cooperation with a storage device for storing the encoded video data. Transmitter 409 may merge the encoded video data of video coder 403 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0057] The controller 402 may manage the operation of the coder 400. During encoding, the controller 405 may assign several encoded image types to each of the encoded images, which may affect the encoding technique applicable to each image. For example, an image is often assigned to one of the following frame types.
[0058] An intra picture (I picture) can be encoded and decoded without using other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, for example, an Independent Decoder Refresh picture. Those skilled in the art are aware of such variations of I pictures, as well as their respective uses and characteristics.
[0059] A predicted picture (P picture) can be encoded and decoded using intra prediction or inter prediction, using at most one motion vector and a reference index to predict the sample values of each block.
[0060] A bi - directionally predicted picture (B picture) can be encoded and decoded using intra prediction or inter prediction, using at most two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi - predicted picture can use more than two reference pictures and related metadata to reconstruct one block.
[0061] The source image may typically be spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and encoded block by block. The blocks may be encoded predictively with reference to other (already encoded) blocks when determined by the coding assignment applied to each image of the block. For example, blocks of an I picture may be encoded non-predictively, or may be encoded predictively (spatial prediction or intra prediction) with reference to already encoded blocks of the same image. Pixel blocks of a P picture may be encoded non-predictively by spatial prediction or by temporal prediction with reference to one previously encoded reference image. Blocks of a B picture may be encoded non-predictively by spatial prediction or by temporal prediction with reference to one or two previously encoded reference images.
[0062] The video encoder 400 may perform an encoding operation according to a predetermined video encoding technique or standard such as ITU-T Rec. H.265. In that operation, the video encoder 400 may perform various compression operations, which include predictive encoding operations that utilize temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.
[0063] In an embodiment, the transmitter 409 may transmit additional data along with the encoded video. The video encoder 403 may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, and the like.
[0064] FIG. 5 shows the intra prediction modes used in HEVC and JEM. To capture any edge direction appearing in natural video, the number of directional intra modes has been extended from 33 to 65 as used in HEVC. The additional directional modes in JEM, which are on top of HEVC, are indicated by dotted arrows in FIG. 5, and the planar and DC modes remain the same. Such a high density of directional intra prediction modes is applied to all block sizes and both luma and chroma intra predictions. As shown in FIG. 5, the directional intra prediction modes specified by dotted arrows, which are associated with odd intra prediction mode indices, are called odd intra prediction modes. The directional intra prediction modes specified by solid arrows, which are associated with even intra prediction mode indices, are called even intra prediction modes. In this specification, the directional intra prediction modes indicated by solid or dotted arrows in FIG. 5 are also called angular modes.
[0065] In JEM, a total of 67 intra prediction modes are used for luma intra prediction. To encode the intra mode, a size-6 MPM list is constructed based on the intra modes of adjacent blocks. If the intra mode is not from the MPM list, a flag is signaled to indicate whether the intra mode belongs to the selected mode. In JEM-3.0, there are 16 selected modes, which are evenly selected every 4 angular modes. In JVET-D0114 and JVET-G0060, 16 secondary MPMs are derived to replace the evenly selected modes.
[0066] FIG. 6 shows the N reference layers used for intra direction modes. There are block unit 611, segment A601, segment B602, segment C603, segment D604, segment E605, segment F606, first reference layer 610, second reference layer 609, third reference layer 608, and fourth reference layer 607.
[0067] In both HEVC and JEM, as well as in some other standards such as H.264 / AVC, the reference samples used for predicting the current block are restricted to the closest reference line (row or column). In the multiple reference line intra prediction method, the number of candidate reference lines (rows or columns) increases from 1 (i.e., the closest one) to N for the intra direction mode, where N is an integer greater than or equal to 1. Fig. 2 shows a 4x4 prediction unit (PU) as an example to illustrate the concept of the multiple line intra direction prediction method. For the intra direction mode, one of the N reference layers can be arbitrarily selected to generate the predictor. In other words, the predictor p(x,y) is generated from one of the reference samples S1, S2, ~SN. A flag is signaled to indicate which reference layer is selected for the intra direction mode. When N is set to 1, the intra direction prediction method is the same as the conventional method in JEM 2.0. In Fig. 6, the reference lines 610, 609, 608, and 607, together with the top-left reference sample, are composed of six segments 601, 602, 603, 604, 605, and 606. In this specification, the reference layer is also referred to as the reference line. The coordinates of the top-left pixel within the current block unit are (0,0), and the top-left pixel of the first reference line is (-1,-1).
[0068] In JEM, for the luma component, before the generation process, the adjacent samples used for generating the intra prediction samples are filtered. The filtering is controlled by a given intra prediction mode and converts the block size. When the intra prediction mode is DC or the transform block size is 4x4, the adjacent samples are not filtered. When the distance between a given intra prediction mode and the vertical mode (or horizontal mode) is greater than a predefined threshold, the filtering process becomes effective. For filtering the adjacent samples, a [1,2,1] filter and a bilinear filter are used.
[0069] The position dependent intra prediction combination (PDPC) method is an intra prediction method that calls a combination of unfiltered boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. Each prediction sample pred[x][y] located at (x, y) is calculated as follows. pred[x][y]=(wL*R -1,y +wT+R x,-1 +wTL+R -1,-1 ,+(64-wL-wT-wTL)*pred[x][y]+32)>>6 (Equation 2-1) Here, R x,-1 ,R -1,y represent unfiltered reference samples located above and to the left of the current sample (x, y) respectively, and R -1,-1 represents an unfiltered reference sample located at the upper left corner of the current block. The weighting is calculated as follows. wT=32>>((y<<1)>>shift) (Equation 2-2) wL=32>>((x<<1)>>shift) (Equation 2-3) wTL=-(wL>>4)-(wT>>4) (Equation 2-4) shift=(log2(wigth)+log2(height)+2)>>2 (Equation 2-5)
[0070] FIG. 7 shows FIG. 700 in which the weightings (wL, wT, wTL) for the positions (0, 0) and (1, 0) inside one 4x4 block are shown.
[0071] FIG. 8 shows Local Illumination Compensation (LIC) FIG. 800, which is based on a linear model of illumination change and uses a scale factor a and an offset b. Also, it is adaptively enabled or disabled for each of the coding units (CUs) coded in inter mode.
[0072] When LIC is applied to a CU, the least square error method is used to derive parameters a and b using the adjacent samples of the current CU and their corresponding reference samples. More specifically, as shown in FIG. 8, adjacent samples of the subsampled (2:1 subsampling) CU and the corresponding samples in the reference image (identified by the motion information of the current CU or sub-CU) are used. The IC parameters are derived and applied individually to each prediction direction.
[0073] When a CU is encoded in merge mode, the LIC flag is copied from an adjacent block in a manner similar to the motion information copy in merge mode, or a LIC flag is signaled to the CU to indicate whether LIC is applied.
[0074] FIG. 9 shows a flowchart 900 according to an exemplary embodiment.
[0075] In S901, in the case of multiple-line intra prediction, instead of setting the same number of reference layers for all blocks, the number of reference layers for each block may be adaptively selected. Here, the index of the closest reference line is indicated by 1.
[0076] In S902, the block size of the upper / left block can be used to determine the number of reference layers of the current block. For example, when the size of the upper and / or left block is larger than MxN, the number of reference layers for the current block is limited to L. M and N can be 4, 8, 16, 32, 64, 128, 256, and 512. L can be 1 to 8.
[0077] In one embodiment, when M and / or N are 64 or more, L is set to 1.
[0078] In another embodiment, the ratio of the number of upper candidate reference rows to the number of left candidate reference columns is the same as the ratio of the block width to the block height. For example, when the current block size is MxN, the number of candidate upper reference rows is m, the number of candidate left reference columns is n, and M:N = m:n.
[0079] Alternatively, in S903, the positions of the last coefficients of the upper and left blocks can be used to determine the number of reference layers for the current block. For example, if the positions of the last coefficients are within the first MxN region of the upper and / or left blocks, the number of reference layers for the current block is limited to L (e.g., L can be from 1 to 8), and M and N can be from 1 to 1024.
[0080] In one embodiment, when there are no coefficients in the upper and / or left blocks, the number of reference layers for the current block is limited to 1.
[0081] In another embodiment, when the coefficients in the upper and / or left blocks are within the upper left 2x2 region, the number of reference layers for the current block is limited to 1 to 2.
[0082] Alternatively, in S904, the pixel values of the reference samples in the upper and / or left blocks can be used to determine the number of reference layers of the current block. For example, the index L i of the reference line and the index L j of the reference line, and if the difference (L i <L j ) is extremely small, the reference line L j is removed from the reference line list. L i and L j can be from 1 to 8. In some cases, due to the extremely small differences between all reference lines, all reference lines with a number greater than 1 are removed. The method of measuring the difference between two reference lines is not limited to this, but includes gradient, SATD, SAD, MSE, SNR, and PSNR.
[0083] In one embodiment, Li and L j If the average SAD of and L is less than 2, the reference line L j is removed from the reference line list.
[0084] Alternatively, in S905, in order to determine the number of reference layers for the current block, the prediction mode of the upper and / or left mode information can be used.
[0085] In one embodiment, when the prediction mode of the upper and / or left block is the skip mode, the number of reference layers for the current block is limited to L. L can be set to 1 to 8.
[0086] FIG. 10 shows a flowchart 1000 according to an exemplary embodiment.
[0087] In S1001, the chroma reference line index can be derived from luma, whether it is a different tree or the same tree. Here, the index of the closest reference line is indicated by 1.
[0088] In S1002, for the same tree, if the reference line index of the co-located luma block is ≧ 3, the reference line index of the current chroma block is set to 2. Otherwise, the reference line index of the current chroma block is set to 1.
[0089] In S1003, for different trees, when the chroma block covers only one block of the luma component, the reference line index derivation algorithm is the same as 2.a. When the chroma block covers multiple blocks of the luma component, the reference line index derivation algorithm may be any one of the following.
[0090] For blocks juxtaposed with luma components, if the reference line index of the majority of the blocks is less than 3, the reference line index for the current chroma block is derived as 1. Otherwise, the reference line index for the current chroma block is derived as 2. The method of measuring the majority is not limited to this, but can include the area size of the blocks and the number of blocks.
[0091] Alternatively, for blocks juxtaposed with luma components, if the reference line index of one block is 3 or more, the reference line index for the current chroma block is derived as 2. Otherwise, the reference line index for the current chroma block is derived as 1.
[0092] Alternatively, for blocks juxtaposed with luma components, if the reference line index of the majority of the blocks is less than 3, the reference line index for the current chroma block is derived as 1. Otherwise, the reference line index for the current chroma block is derived as 2.
[0093] Alternatively, in S1004, whether to use adaptive selection is considered. If considered, the method of FIG. 9 can also be used to limit the number of reference layers for the current chroma block. After applying the method of FIG. 9, the number of reference layers is L C1 is set. Next, the derivation algorithms illustrated in S1002 and S1003, or S1005 and S1006 of FIG. 10 are further applied to obtain the line index of the current block L C2 of. Thereafter, min(L C1 , L C2 ) becomes the final reference line index for the current chroma block.
[0094] FIG. 11 shows a flowchart 1100 according to an exemplary embodiment.
[0095] In S1101, it is considered that different reference lines have different numbers of intra prediction modes. Here, the index of the closest reference line is indicated by 1.
[0096] For example, the first reference line has 67 modes, the second reference line has 35 modes, the third reference line has 17 modes, and the fourth reference line has 9 modes.
[0097] For example, the first reference line has 67 modes, the second reference line has 33 modes, the third reference line has 17 modes, and the fourth reference line has 9 modes.
[0098] Alternatively, in S1102, reference lines with an index greater than 1 share the same number of intra modes, but are much less than that of the first reference line, such as less than half of the intra prediction modes of the first reference line.
[0099] In S1103, for example, only the directional intra prediction modes with even mode indices are permitted for reference lines with an index greater than 1. As shown in FIG. 5, the directional intra prediction modes with odd mode indices are labeled with dotted arrows, and the directional intra prediction modes with even mode indices are labeled with solid arrows.
[0100] In S1104, in another example, only the directional intra prediction modes with even mode indices, as well as the DC and planar modes, are permitted for reference lines with an index greater than 1.
[0101] In S1105, in another example, only the most probable mode (MPM) is permitted for non-zero reference lines, and the MPM includes both the first-level MPM and the second-level MPM.
[0102] In S1106, in another example, since a reference line index greater than 1 is only valid for the even mode (or odd mode) intra prediction mode, when encoding the intra prediction mode, if a reference line index greater than 1 is signaled, the intra prediction modes such as planar / DC, and the odd (or even) intra prediction mode are excluded from the MPM derivation and the list, excluded from the second level of the MPM derivation and the list, and excluded from the remaining non-MPM mode list.
[0103] In S1107, the reference line index is signaled after signaling the intra prediction mode, and whether to signal the reference line index depends on the signaled intra prediction mode.
[0104] For example, only the directional intra prediction mode with an even mode index is permitted for a reference line with an index greater than 1. If the signaled intra prediction mode is a directional prediction with an even mode index, the selected reference line index is signaled. Otherwise, only one default reference line, such as the closest reference line, is permitted for intra prediction, and the index is not signaled.
[0105] In another example, only the most probable mode (MPM) is permitted for a reference line with an index greater than 1. If the signaled intra prediction is from the MPM, the selected reference line index needs to be signaled. Otherwise, only one default reference line, such as the closest reference line, is permitted for intra prediction, and the index is not signaled.
[0106] In another secondary embodiment, for all directional intra prediction modes, or for all intra prediction modes, a reference line with an index greater than 1 is still valid, and the intra prediction mode index can be used as a context for entropy encoding the reference line index.
[0107] In another embodiment, only the most accurate mode (MPM) is permitted for reference lines with an index greater than 1. In one approach, all MPMs are permitted for reference lines with an index greater than 1. In another approach, a subset of MPMs is permitted for reference lines with an index greater than 1. When MPMs are classified into multiple levels, in one approach, only some levels of MPMs are permitted for reference lines with an index greater than 1. In one example, only the lowest level of MPMs is permitted for reference lines with an index greater than 1. In another example, only the highest level of MPMs is permitted for reference lines with an index greater than 1. In another example, only pre-defined (or signaled / indicated) levels of MPMs are permitted for reference lines with an index greater than 1.
[0108] In another embodiment, only non-MPMs are permitted for reference lines with an index greater than 1. In one approach, all non-MPMs are permitted for reference lines with an index greater than 1. In another approach, a subset of non-MPMs is permitted for reference lines with an index greater than 1. In one example, only non-MPMs associated with even (or odd) indices in descending (or ascending) order of all non-MPM intra-mode indices are permitted for reference lines with an index greater than 1.
[0109] In another embodiment, a plane and a DC mode are assigned to a pre-defined index of the MPM mode list.
[0110] In one example, the pre-defined index further depends on encoded information including, but not limited to, the width and height of the block.
[0111] In another secondary embodiment, for reference lines with an index greater than 1, an MPM with a given index is permitted. The given MPM index can be signaled or specified as a high-level syntax element, for example, in a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), a slice header, or as a common syntax element or parameter for a region of the picture. The reference line index is signaled only when the intra mode of the current block is equal to one of the given MPM indices.
[0112] For example, the length of the MPM list is 6, and the indices of the MPM list are 0, 1, 2, 3, 4, and 5. When the intra mode of the current block is not equal to the modes with MPM indices 0 and 5, reference lines with an index greater than 1 are permitted.
[0113] In S1108, in one embodiment, all intra prediction modes are permitted for the closest reference line of the current block, while only the most accurate mode is permitted (or not permitted) for reference lines with an index greater than 1 (or a specific index value such as 1).
[0114] In S1109, in one embodiment, the most accurate mode includes only the first-level MPMs, such as 3 MPMs in HEVC and 6 MPMs in JEM (or VTM).
[0115] In S1110, in another embodiment, the most accurate mode can be any level of MPM from the lowest level MPM to the highest level MPM.
[0116] In S1111, in another embodiment, only some levels of MPM are permitted for reference lines with an index greater than 1.
[0117] In S1112, in another embodiment, the most accurate mode can be set to only one level of MPM, such as the lowest level of MPM, the highest level of MPM, or a predefined level of MPM.
[0118] In S1113, in another embodiment, the reference line index is signaled before the MPM flag and the intra mode. When the signaled reference line index is 1, the MPM flag is also signaled. When the signaled reference line index is greater than 1, the MPM flag of the current block is not signaled, and the MPM flag of the current block is derived as 1. For reference lines with an index greater than 1, the MPM index of the current block is still signaled.
[0119] In S1114, in one embodiment, the MPM list generation process depends on the reference line index value.
[0120] In one example, the MPM list generation process for reference lines with an index greater than 1 is different from that for reference lines with an index of 1. For reference lines with an index greater than 1, the plane and DC mode are excluded from the MPM list. The length of the MPM list is the same for all reference lines.
[0121] The default MPM used in the MPM list generation process depends on the reference line index. In one example, the default MPM associated with reference lines with an index greater than 1 is different from that associated with reference lines with an index of 1.
[0122] In S1115, in one embodiment, the length of the MPM list, i.e., the number of MPMs, depends on the reference line index value.
[0123] In another embodiment, the length of the MPM list with a reference line index value of 1 is set to be different from that of the MPM list with a reference line index value greater than 1. For example, the length of the MPM list for a reference line with an index greater than 1 is 1 or 2 shorter than the length of the MPM list with a reference line index of 1.
[0124] In another embodiment, the length of the MPM list, i.e., the number of MPMs, is 5 for a reference line index greater than 1. The default MPM for the MPM list generation process is {VER, HOR, 2, 66, 34} when the angular mode of 65 is applied. The order of the default MPMs can be any combination of these 5 listed modes.
[0125] In S1116, for the odd-directional intra prediction mode and / or the angular intra prediction mode (not signaled) in which a reference line index such as planar / DC is derived, multi-line reference samples are used to generate a predictor for the current block.
[0126] For the angular intra prediction mode (not signaled) in which a reference line index is derived, predicted sample values are generated using a weighted sum of a plurality of predictors, and each of the plurality of predictors is a prediction generated using one of the plurality of reference lines.
[0127] In one example, the weighted sum uses a weighting of {3, 1} applied to the predictors respectively generated by the first reference line and the second reference line.
[0128] In another example, the weighting depends on the block size, block width, block height, sample position within the current block to be predicted, and / or the intra prediction mode.
[0129] In one example, for a given angular prediction mode having an odd index, a first reference line is used to generate one prediction block unit Pred1, and a second reference line is used to generate another prediction block unit. As a result, the final prediction value for each pixel of the current block unit is the weighted sum of these two generated prediction block units. This process can be formulated by Equation (4-1), where W i is the same value for all pixels within the same block. For different blocks, W i may be the same, or may depend on the intra prediction mode and block size.
[0130]
Number
[0131] Alternatively, in S1117, the number of intra prediction modes for each reference line is derived from the difference between the reference samples of that line. The methods for measuring the difference include, but are not limited to, gradient, SATD, SAD, MSE, SNR, and PSNR.
[0132] If both the row above and the left column of the reference samples are very similar, the number of modes can be reduced to 4, 9, 17, or 35 modes. The four modes are the planar mode, the DC mode, the vertical mode, and the horizontal mode.
[0133] If only the row above the reference samples is very similar, the modes of the prediction modes close to vertical are downsampled. In a special case, only mode 50 is maintained, and modes 35 to 49 and modes 51 to 66 are excluded. To make the total intra prediction modes 9, 17, or 35, the intra prediction modes in the direction close to horizontal are reduced accordingly.
[0134] If only the left column of the reference samples is very similar, the modes in the prediction mode close to horizontal are downsampled. In a special case, only mode 18 is maintained, and modes 2 to 17 and modes 19 to 33 are excluded. To make the total intra prediction mode 9, 17, or 35, the intra prediction modes in the direction close to vertical are reduced accordingly.
[0135] FIG. 12 shows a flowchart 1200 according to an exemplary embodiment.
[0136] In S1201, smoothing of each sample in the current reference line is performed based on adjacent samples in the current line and the adjacent (plural) reference lines thereto. Here, the index of the closest reference line is indicated by 1.
[0137] In S1202, for each pixel in the current line, all pixels in reference lines 1 to L can be used to smooth the pixels in the current line. L is the maximum number of reference lines allowed for intra prediction, and L may be 1 to 8.
[0138] In S1203, for boundary pixels, they may or may not be filtered. If filtered, each boundary pixel in the same line uses the same filter. Boundary pixels in different lines can use different filters. For example, the boundary pixels in the first reference line can be filtered by a [3, 2, 2, 1] filter, the boundary pixels in the second reference line can be filtered by a [2, 3, 2, 1] filter, the boundary pixels in the third reference line can be filtered by a [1, 2, 3, 2] filter, and the boundary pixels in the fourth reference line can be filtered by a [1, 2, 2, 3] filter.
[0139] In S1204, for other pixels, pixels within each line can use the same filter, and pixels in different lines can use different filters. Alternatively, for other pixels, pixels at different positions can use different filters. However, these filters are predefined, and it is not necessary for the coder to signal the filter index.
[0140] Alternatively, in S1205, the filtering operation for each line may depend on the intra prediction mode and the transform size. The filtering operation becomes effective only when the intra prediction mode and the transform size meet certain conditions. For example, the filtering operation becomes ineffective when the transform size is 4x4 or less.
[0141] Alternatively, in S1206, the filter used for smoothing each pixel may have an irregular filter support shape instead of a rectangular shape. The filter support shape may be predefined and may also depend on any information that can be used by both the coder and the decoder, including but not limited to the reference line index, intra mode, block height, and / or width.
[0142] Alternatively, in S1207, for each pixel in the first reference line, pixels in the first reference line and the second reference line can be used to smooth that pixel. For each pixel in the second reference line, pixels in the first reference line, the second reference line, and the third reference line can be used to smooth that pixel. For each pixel in the third reference line, pixels in the second reference line, the third reference line, and the fourth reference line can be used to smooth that pixel. For each pixel in the fourth reference line, pixels in the third reference line and the fourth reference line can be used to smooth that pixel. In other words, for pixels in the first reference line and the fourth reference line, pixels in two lines are used to filter each pixel, and for pixels in the second reference line and the third reference line, pixels in three lines are used to filter each pixel.
[0143] For example, the filtered pixels in the second reference line and the third reference line can be calculated by Equations 4-2 to 4-5. p’(x,y)=(p(x-1,y)+p(x,y-1)+p(x,y+1)+p(x+1,y)+4*p(x,y))>>3 (Equation 4-2) p’(x,y)=(p(x,y+1)-p(x,y-1)+p(x,y))(Equation 4-3) p’(x,y)=(p(x-1,y)+p(x-1,y-1)+p(x-1,y+1)+p(x,y-1)+p(x,y+1)+p(x+1,y-1)+p(x+1,y)+p(x+1,y+1)+8*p(x,y))>>4(Equation 4-4)
[0144] [Number]
[0145] The filtered pixels in the first reference line can be calculated by Equations 4-6 to 4-10. p'(x, y) = (p(x - 1, y) + p(x, y - 1) + p(x + 1, y) + 5 * p(x, y)) >> 3 (Equation 4-6) p'(x, y) = (p(x - 1, y) + p(x, y - 1) + p(x + 1, y) + 5 * p(x, y)) >> 2 (Equation 4-7) p'(x, y) = (2p(x, y) - p(x, y - 1)) (Equation 4-8) p'(x, y) = (p(x - 1, y) + p(x - 1, y - 1) + p(x, y - 1) + p(x + 1, y - 1) + p(x + 1, y) + 3 * p(x, y)) >> 3 (Equation 4-9)
[0146]
Number
[0147] The filtered pixels in the 4th reference line can be calculated by Equations 4-11 to 4-15. p'(x, y) = (p(x - 1, y) + p(x, y + 1) + p(x + 1, y) + 5 * p(x, y)) >> 3 (Equation 4-11) p'(x, y) = (p(x - 1, y) + p(x, y + 1) + p(x + 1, y) + p(x, y)) >> 2 (Equation 4-12) p'(x, y) = (2p(x, y) - p(x, y + 1)) (Equation 4-13) p'(x, y) = (p(x - 1, y) + p(x - 1, y + 1) + p(x, y + 1) + p(x + 1, y + 1) + p(x + 1, y) + 3 * p(x, y)) >> 3 (Equation 4-14)
[0148]
Number
[0149] Also, rounding such as rounding to zero, positive infinity, or negative infinity may be added to the above mathematical expressions.
[0150] Figure 13 shows a flowchart 1300 according to an exemplary embodiment.
[0151] In S1301, in the current block, samples at different positions may use different combinations of reference samples for different line index predictions. Here, the index of the closest reference line is indicated by 1.
[0152] In S1302, for a given intra prediction mode, each reference line i can generate one prediction block Pred i For each pixel at each position, the mode can use different combinations of these generated prediction blocks Pred i Specifically, for the pixel at position (x, y), Equation 4-16 can be used to calculate the predicted value.
[0153]
Equation
[0154] Here, W i is position-dependent. In other words, the weighting coefficients for the same position are the same, and the weighting coefficients are different for different positions.
[0155] Alternatively, when an intra prediction mode is given for each sample, a series of reference samples from multiple reference lines are selected, and the weighted sum of these selected series of reference samples is calculated as the final predicted value. The selection of the reference samples may depend on the intra mode and the position of the predicted sample, and the weighting may also depend on the intra mode and the position of the predicted sample.
[0156] In S1303, when the reference line x is applied to each sample for intra prediction, the predicted values of line 0 and line x are compared. When line 1 generates significantly different predicted values, the predicted value from line x may be excluded and line 0 may be used instead. The method for measuring the difference between the predicted value of the current position and the predicted values of its adjacent positions is not limited to this, but includes gradient, SATD, SAD, MSE, SNR, and PSNR.
[0157] Alternatively, when two or more predicted values are generated from different reference lines, the intermediate (or average, or the most frequently occurring) value is used as the predicted sample.
[0158] In S1304, when the reference line x is applied to each sample for intra prediction, the predicted values of line 1 and line x are compared. When line 1 generates significantly different predicted values, the predicted value from line x may be excluded and line 1 may be used instead. The method for measuring the difference between the predicted value of the current position and the predicted values of its adjacent positions is not limited to this, but includes gradient, SATD, SAD, MSE, SNR, and PSNR.
[0159] Alternatively, when two or more predicted values are generated from different reference lines, the intermediate (or average, or the most frequently occurring) value is used as the predicted sample.
[0160] FIG. 14 shows a flowchart 1400 according to an exemplary embodiment.
[0161] In S1401, after intra prediction, instead of using only the pixels within the nearest reference line, the pixels within a plurality of lines are used to filter the predicted value of each block. Here, the index of the nearest reference line is indicated by 1.
[0162] For example, in S1402, PDPC may be extended for multi-line intra prediction. Each predicted sample pred[x][y] located at (x, y) is calculated by the following mathematical formula.
[0163] [Number]
[0164] Here, m may be -8 to -2.
[0165] In one example, the reference samples in the two closest lines are used to filter the samples within the current block. For the top-left pixel, only the top-left sample in the first row is used. This can be formulated by Equation 4-18.
[0166] [Number]
[0167] Alternatively, in S1403, the boundary filter may be extended to multiple lines.
[0168] After DC prediction, for the first sequence and the pixels within the first few lines, filtering is performed by adjacent reference pixels. The pixels in the first column can be filtered by the following equation.
[0169] [Number]
[0170] For the pixels in the first row, the filtering operation is as follows.
[0171] [Number]
[0172] In some special cases, the pixels in the first column can be filtered by the following equation. p’(0,y)=p(0,y)+R-1,y -R -2,y (Equation 4-21)
[0173] The pixels in the first row can also be filtered by the following mathematical formula. p’(x,0)=p(x,0)+R x,-1 -R x,-2 (Equation 4-22)
[0174] After vertical prediction, the pixels of the first sequence can be filtered by Equation 4-23.
[0175]
Number
[0176] After horizontal prediction, the pixels of the first few rows can be filtered by Equation 4-24.
[0177]
Number
[0178] In another embodiment, for vertical / horizontal prediction, when a reference line with an index greater than 1 is used to generate prediction samples, the first column / row within the line index greater than 1 and its corresponding pixels are used for boundary filtering. As shown in FIG. 15, in the reference lines 1503, 1502, and the block unit 1501, the second reference line 1503 is used to generate prediction samples for the current block unit, and the pixels having a vertical direction are used for vertical prediction. After vertical prediction, the pixels having a diagonal texture within the reference line 1 and the pixels having a diagonal texture within the reference line 1503 are used for the filtering process of filtering the first column of numbers within the current block unit, which can be formulated by Equation 4-25. m represents the selected line index and can be 2 to 8. n is the number of right shift bits and can be 1 to 8. p’(x,y)=p(x,y)+(p(-1,y)-p(-1,-m))>>n (Equation 4-25)
[0179] For horizontal prediction, the filtering process can be formulated by Equation 4-26. p’(x,y)=p(x,y)+(p(x,-1)-p(-m,-1))>>n (Equation 4-26)
[0180] In another embodiment, when a reference line with an index greater than 1 is used, after diagonal prediction such as Mode 2 and Mode 34 in FIG. 1(a), the pixels along the diagonal direction from the first reference line to the current reference line are used for filtering the pixels within the first column / first few rows of the current block unit. Specifically, after Mode 2 prediction, the pixels within the first few rows can be filtered by Equation 4-27. After Mode 34 prediction, the pixels within the first column can be filtered by Equation 4-28. m represents the reference line index for the current block and can be 2 to 8. n is the number of right shift bits and can be 2 to 8. W i is a weighting coefficient, which is an integer.
[0181]
Number
[0182] Figure 16 shows a flowchart 1600 according to an exemplary embodiment.
[0183] In S1601, in the case of multiple reference line intra prediction, when the reference line index is greater than 1, the modified DC and planar modes are added. Here, the index of the closest reference line is indicated by 1.
[0184] In S1602, in the case of the planar mode, when different reference lines are used, different predefined upper right and lower left reference samples are used to generate prediction samples.
[0185] Alternatively, in S1603, when different reference lines are used, different intra smoothing filters are used.
[0186] In S1604, in the case of the DC mode, for the first reference line, all the pixels in the upper row and the left column are used to calculate the DC value, and when the reference line index is greater than 1, only a part of the pixels is used for the calculation of the DC value.
[0187] For example, the upper pixel in the first reference line is used to calculate the DC value for the second reference line, the left pixel in the first reference line is used to calculate the DC value for the third reference line, and for calculating the DC value of the fourth reference line, half of the left pixel and half of the upper pixel in the first reference line are used.
[0188] In S1605, in the case of the DC mode, all the reference pixels in all available candidate lines (rows and columns) are used to calculate the DC predictor.
[0189] FIG. 17 shows a flowchart 1700 according to an exemplary embodiment.
[0190] S1701 is implemented to extend a plurality of reference lines up to the IC mode. In S1702, a plurality of up / left reference lines are used to calculate IC parameters, and in S1703, a signal is sent indicating which reference lines are used to calculate the IC parameters.
[0191] FIG. 18 shows a flowchart 1800 according to an exemplary embodiment.
[0192] S1801 is implemented to signal a plurality of reference line indexes.
[0193] In one embodiment, in S1802, variable length coding is used to signal the reference line index. The closer the distance to the current block, the shorter the code word. For example, if the reference line indexes are 0, 1, 2, 3, with 0 being closest to the current block and 3 being the farthest, the code words for these are 1, 01, 001, 000, and 0 and 1 can be swapped.
[0194] In another embodiment, in S1806, fixed length coding is used to signal the reference line index. For example, if the reference line indexes are 0, 1, 2, 3, with 0 being closest to the current block and 3 being the farthest, the code words for these are 10, 01, 11, 00, and 0 and 1 can be swapped and the order can be changed.
[0195] In S1803, it is considered whether to use various code word tables. If not, in S1804, in yet another embodiment, variable length coding is used to signal the reference line index, and the order of the indexes in the code word table is (from the shortest code word to the longest code word) 0, 2, 4, … 2k, 1, 3, 5, … 2k+1 (or 2k-1). Index 0 indicates the reference line closest to the current block, and 2k+1 is the farthest.
[0196] In another embodiment, in S1805, a variable - length coding is used to signal the reference line index, and the order of the indexes in the code - word table (from the shortest code - word to the longest) is like the closest, the farthest, the second - closest, the second - farthest, etc. In one specific example, if the reference line indexes are 0, 1, 2, 3, where 0 is the closest to the current block and 3 is the farthest, the code - words for them are 0 for index 0, 10 for index 3, 110 for index 2, and 111 for index 1. The code - words for reference line indexes 1 and 2 may be exchanged. The 0 and 1 within the code - word may be swapped.
[0197] FIG. 19 shows a flowchart 1900 according to an exemplary embodiment.
[0198] In S1901, when the number of upper reference lines (rows) is different from the number of left reference lines (columns), a plurality of reference line indexes are signaled.
[0199] In S1902, in one embodiment, when the number of upper reference lines (rows) is M and the number of left reference lines (columns) is N, the reference line index for max(M, N) may use any of the methods described above, or a combination thereof. The reference line index for min(M, N) takes a subset of the code - words used to indicate the reference line index for max(M, N), usually the shorter one. For example, when M = 4, N = 2, and the code - words used to signal the M(4) reference line indexes {0, 1, 2, 3} are 1, 01, 001, 000, the code - words used to signal the N(2) reference line indexes {0, 1} are 1, 01.
[0200] In another embodiment, in S1903, when the number of upper reference lines (rows) is M, the number of left reference lines (columns) is N, and M and N are different, the reference line indexes for transmitting the upper reference line (row) index and the left reference line (column) index are different, and any of the methods described above, or a combination thereof, may be used independently.
[0201] FIG. 20 shows a flowchart 2000 according to an exemplary embodiment.
[0202] In S2000, it is considered to know the number of reference lines in various coding tools, and in S2001, if possible, in order to save the pixel line buffer, the maximum number of reference lines that can be used for intra prediction, such as deblocking filter or template matching-based intra prediction, may be restricted so as not to exceed the number of reference lines used in other coding tools.
[0203] FIG. 21 shows a flowchart 2100 according to an exemplary embodiment.
[0204] In S2100, the interaction between multi-line intra prediction and other coding tools / modes is performed.
[0205] For example, in S2101, in one embodiment, without limitation, the use and / or signaling of other syntax elements / coding tools / modes, including but not limited to cbf, final position, transform skip, transform type, secondary transform index, primary transform index, PDPC index, may depend on the multi-line reference line index.
[0206] In S2102, in one example, when the multi-line reference index is non-zero, transform skip is not used and the transform skip flag is not signaled.
[0207] In S2103, in another example, the context used for signaling other coding tools, such as transform skip, cbf, primary transform index, secondary transform index, may depend on the value of the multi-line reference index.
[0208] In S2104, in another embodiment, but not limited to this, after other syntax elements including cbf, final position, transform skip, transform type, secondary transform index, primary transform index, PDPC index, the multi-line reference index may be signaled, and the use and / or signaling of the multi-line reference index may depend on other syntax elements.
[0209] FIG. 22 shows a flowchart 2200 according to an exemplary embodiment.
[0210] In S2201, it is contemplated to obtain a reference line index, and in S2202, the reference line index can be used as a context for entropy coding other syntax elements including, but not limited to, intra prediction mode, MPM index, primary transform index, secondary transform index, transform skip flag, coding block flag (CBF), and transform coefficients, and vice versa.
[0211] FIG. 23 shows a flowchart 2300 according to an exemplary embodiment.
[0212] In S2301, it is proposed to include reference line information in the MPM list. That is, when the prediction mode of the current block is the same as one of the candidates in the MPM list, both intra prediction and the selected reference line of the selected candidate are applied to the current block, and the intra prediction mode and the reference line index are not signaled. Also, the number of MPM candidates for different reference line indices is predefined. Here, the closest reference line is indicated by 1.
[0213] In S2302, in one embodiment, the number of MPMs for each reference line index is predefined, which can be signaled as a high-level syntax element such as a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile header, a coding tree unit (CTU) header, etc., or as a common syntax element or parameter of a region of the picture. As a result, the length of the MPM list may vary for different sequences, pictures, slices, tiles, groups of coding blocks, or regions of the picture.
[0214] For example, the number of MPMs for reference line index 1 is 6, and the number of MPMs for each of the other reference line indices is 2. As a result, when the total number of reference lines is 4, the total number of MPM lists is 12.
[0215] In another embodiment, in S2303, all intra prediction modes combined with their reference line indices in the upper, left, upper left, upper right, and lower left blocks are included in the MPM list. As shown in FIG. 2400 of FIG. 24, all adjacent blocks of the current block unit are shown, where A is the lower left block, B, C, D, and E are the left blocks, F is the upper left block, G and H are the upper blocks, and I is the upper right block. This is after adding the modes of the adjacent blocks to the MPM list. If the number of MPM candidates having a given number of reference lines is less than a predefined number, default modes are used to fill the MPM list.
[0216] In another embodiment, in S2304, if the mode of the current block is equal to one of the candidates in the MPM list, the reference line index is not signaled. If the mode of the current block is not equal to any of the candidates in the MPM list, the reference line index is signaled.
[0217] In one example, when line 1 is used for the current block, the second-level MPM mode is still used, but the second-level MPM includes only intra prediction mode information.
[0218] In another example, for other lines, the second-level MPM is not used, and fixed-length coding is used to code the remaining modes.
[0219] FIG. 25 shows a flowchart 2500 according to an exemplary embodiment.
[0220] In S2501, in the current VVC test mode VTM-1.0, the chroma intra coding mode is the same as that of HEVC, including DM (direct copy of the luma mode), and four additional angular intra prediction modes. In the current BMS-1.0, the cross component linear model (CCLM) mode is also applied to chroma intra coding. The CCLM mode includes one LM mode, one multi-model LM (MMLM), and four multi-filter LM (MFLM) modes. This is because only the DM mode is used for chroma blocks when the CCLM mode is not valid, while only the DM and CCLM modes are used for chroma blocks when the CCLM mode is valid.
[0221] In S2502, in one embodiment, only one DM mode is used for chroma blocks, no flag is signaled for chroma blocks, and the chroma mode is derived as the DM mode.
[0222] In another embodiment, in S2503, only one DM and one CCLM mode are used for chroma blocks, and one DM flag is used to signal whether the DM mode or the LM mode is used for the current chroma block.
[0223] In one secondary embodiment, there are three contexts used for signaling the DM flag. When both the left block and the upper block use the DM mode, context 0 is used to signal the DM flag. When only one of the left block and the upper block uses the DM mode, context 1 is used to signal the DM flag. Alternatively, when neither the left block nor the upper block uses the DM mode, context 2 is used to signal the DM flag.
[0224] In another embodiment, in S2504, only the DM and CCLM (when valid) modes are used for small chroma blocks. When the width, or height, or area size (width * height) of the chroma block is less than or equal to Th, the current chroma block is called a small chroma block. Th may be 2, 4, 8, 16, 32, 64, 128, 256, 512, or 1024.
[0225] For example, when the area size of the current chroma block is 8 or less, only the DM and CCLM (when valid) modes are used for the current chroma block.
[0226] In another example, when the area size of the current chroma block is 16 or less, only the DM and CCLM (when valid) modes are used for the current chroma block.
[0227] In another example, for small chroma blocks, only one DM and one CCLM (when valid) mode are used.
[0228] In another embodiment, in S2505, when the intra mode of the luma component is equal to one of the MPM modes, the chroma block can use only the DM mode, no flag is signaled for the chroma mode, or both the DM and CCLM modes are permitted for the chroma block.
[0229] In one example, the MPM mode can be only the first level of MPM.
[0230] In another example, the MPM mode can be only the second level of MPM.
[0231] In another example, the MPM mode can be either the first level of MPM or the second level of MPM.
[0232] In another embodiment, at S2506, when the intra-mode of the luma component is not equal to any of the MPM modes, the chroma block can use the DM mode, no flag is signaled for the chroma mode, or both the DM and CCLM modes are permitted for the chroma block.
[0233] Thus, by the exemplary embodiments described herein, the above-described technical problems can be preferably improved by such technical solutions.
[0234] The foregoing technology can be implemented as computer software using computer-readable instructions, physically stored on one or more computer-readable media, or implemented by one or more hardware processors specifically configured. For example, FIG. 26 shows a computer system 2600 suitable for implementing some embodiments of the disclosed subject matter.
[0235] The computer software can be encoded using any suitable machine code or computer language, which can be used to create code containing instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or through interpretation, execution of microcode, etc., according to mechanisms such as assembly, compilation, and linking.
[0236] The instructions can be executed on various types of computers or their components, such as personal computers, tablet computers, servers, smartphones, game consoles, Internet of Things (IoT) devices, etc.
[0237] The components of the computer system 2600 shown in FIG. 26 are exemplary in nature and are not intended to imply any limitation with respect to the use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having dependencies or requirements with respect to any one of the components shown in the exemplary embodiments of the computer system 2600 or any combination of components.
[0238] The computer system 2600 may include several human interface input devices. Such human interface input devices may respond to input from one or more users, for example, by tactile input (pressing keys, swiping, moving a data glove, etc.), voice input (voice, clapping hands, etc.), visual input (body gestures, etc.), or olfactory input (not shown). The human interface device can also be used to capture several media that are not necessarily directly involved in conscious input by humans, such as voice (speech, music, ambient sound, etc.), images (scanned images, photographic images obtained by a still camera, etc.), video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).
[0239] The input human interface device may include one or more of a keyboard 2602, a mouse 2603, a trackpad 403, a touch screen 2604, a joystick 2605, a microphone 2606, a scanner 2608, and a camera 2607 (only one of each is shown).
[0240] The computer system 2600 may include several human interface output devices. Such human interface output devices can stimulate the senses of one or more users, such as tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen 2610 or a joystick 2605, although there may also be a tactile feedback device that does not function as an input device), sound output devices (speakers 2609, headphones (not shown)), visual output devices (each may or may not have touch screen input capabilities, each may or may not have tactile feedback capabilities, and some of them may be capable of outputting more than three dimensions by means such as two-dimensional video output or three-dimensional output, including screens 2610 such as CRT screens, LCD screens, plasma screens, OLED screens, VR glasses (not shown), holographic displays, smoke generation devices (smoke tank (not shown)), and printers (not shown)).
[0241] The computer system 2600 can further include associated media such as humanly accessible storage devices and optical media including CD / DVD ROM / RW 2612 containing media 2611 such as CD / DVDs, USB memories (thumb-drives) 2613, removable hard drives or solid state drives 2614, legacy magnetic media such as tapes and floppy disks (not shown), and specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0242] It should be further understood by those skilled in the art that the term "computer-readable medium" as used with respect to the subject matter disclosed herein does not include transmission media, carrier waves, or other transient signals.
[0243] The computer system 2600 can further include an interface to one or more communication networks 2615. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, in-vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, mobile communication networks including wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial TV, industrial networks including in-vehicle and CANBus, etc. Some networks generally require an external network interface adapter attached to some general-purpose data ports or peripheral buses 2625 (e.g., the USB ports of the computer system 2600), and other networks are generally integrated into the core of the computer system 2600 by attaching to the system bus as described later (e.g., an Ethernet interface to a PC computer system or a mobile communication network interface to a computer system of a smartphone). Using any such network, the computer system 2600 can communicate with other entities. Such communication can be one-way communication, reception-only communication (e.g., TV broadcast), transmission-only one-way communication (e.g., a device that transmits from CANbus to CANbus), or two-way communication with other computer systems, for example, using a local or wide area digital network. As described above, various protocols and protocol stacks can be used for each such network and each network interface.
[0244] The human interface device, human-accessible storage device, and network interface described above can be attached to the core 2612 of the computer system 2600.
[0245] The core 2612 can include one or more central processing units (CPUs) 2612, a graphics processing unit (GPU) 2622, a specialized programmable processing unit 2624 in the form of an FPGA (Field Programmable Gate Areas), hardware accelerators 2624 for several tasks, etc. Such a device may be connected via a system bus 2626 together with an internal mass storage device 447 such as a read-only memory (ROM) 2619, a random access memory 2618, a hard drive that is not accessible to internal users, an SSD, etc. In some computer systems, the system bus 226 can be accessed in the form of one or more physical plugs so as to be expandable by an additional CPU, GPU, etc. Peripheral devices can be attached directly to the system bus 2626 of the core or via a peripheral bus 2601. The architecture of the peripheral bus includes PCI, USB, etc.
[0246] The CPU 2621, GPU 2622, FPGA 2624, and accelerator 2624 can execute in combination several instructions capable of creating the aforementioned computer code. The computer code can be stored in the ROM 2619 or the RAM 2618. Transient data can also be stored in the RAM 2618, whereas persistent data can be stored, for example, in the internal mass storage device 2620. By using a cache memory, it becomes possible to quickly store and retrieve data in any memory device, and it can be closely associated with one or more CPUs 2621, GPUs 2622, mass storage devices 2620, ROM 2619, RAM 2618, etc.
[0247] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure or can be of the kind well known and available to those skilled in the field of computer software.
[0248] By way of example, and not for purposes of limitation, a computer system having an architecture 2600, specifically a core 2616, can provide functionality as a result of (a plurality of) processors (such as CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be media associated with some storage devices of core 2616 having a non-transitory nature, such as the mass storage devices accessible to a user as introduced above, as well as the core internal mass storage device 2620 or ROM 2619. The software implementing various embodiments of the present disclosure can be stored on such devices and executed by core 2616. The computer-readable media can include one or more memory devices or chips according to specific requirements. The software can cause core 2616 and specifically the processors (including CPUs, GPUs, FPGAs, etc.) therein to execute specific processes described herein, or specific portions of specific processes, including the definition of data structures stored in RAM 2618 and the modification of such data structures according to the processes defined by the software. In addition to or instead of this, the computer system can provide functionality as a result of logic wired in a circuit (such as accelerator 2624) or embodied in a circuit, and can operate instead of or together with software to execute specific processes described herein, or specific portions of specific processes. References to software can include logic as necessary, and vice versa. References to computer-readable media can include circuits (such as integrated circuits (ICs)) that store software to be executed, circuits that embody the logic to be executed, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0249] Although the present disclosure has described several exemplary embodiments, there are changes, substitutions, and various alternative equivalents that are included within the scope of the present disclosure. Accordingly, those skilled in the art will understand that even if not explicitly illustrated or described herein, they can embody the principles of the present disclosure and thus devise many systems and methods that are included within the principles and scope thereof.
[0250] (Appendix 1) A method for decoding a video image executed by at least one processor, comprising: said at least one processor executing the decoding of the video sequence by intra prediction between a plurality of reference lines of the video sequence; setting a plurality of intra prediction modes for the zero reference line closest to the current block of the intra prediction among the plurality of non-zero reference lines; setting at least one most accurate mode for one of the plurality of non-zero reference lines; and a method comprising. (Appendix 2) signaling a reference line index before signaling a most accurate mode flag and an intra prediction mode; signaling the most accurate mode flag in response to the determination that the reference line index is signaled and the signaled index indicates the zero reference line; deriving the most accurate mode flag to be true without signaling the most accurate mode flag and signaling the most accurate mode index of the current block in response to the determination that the reference line index is signaled and the signaled index indicates at least one of the plurality of non-zero reference lines; The method according to Appendix 1, further comprising. (Appendix 3) said at least one most accurate mode for the non-zero reference line is included in a most accurate mode list; The method according to appendix 1, wherein at least one of the planar mode and the DC mode is excluded from the most accurate mode list corresponding to any one of the non-zero reference lines. (Appendix 4) A step of setting the length of the most accurate mode list based on a reference line index value, such that the length of the most accurate mode list includes the number of the at least one most accurate mode. The method according to appendix 3, further comprising. (Appendix 5) The method according to appendix 4, wherein the length of the most accurate list of the reference line with an index value of 1 is set to be different from the length of the most accurate list of the reference line with an index value greater than 1. (Appendix 6) The method according to appendix 4, wherein for a reference line having an index value greater than 1, the length of the most accurate mode list is one shorter than the length of the most accurate mode list of the reference line with an index value of 1. (Appendix 7) A step of setting the length of the most accurate mode list to either 1 or 4 in response to the detection of the non-zero reference line; and A step of setting the length of the most accurate mode list to 3 or 6 in response to a determination that the current reference line is a zero reference line. The method according to appendix 4, further comprising. (Appendix 8) A step of setting the length of the most accurate mode list to consist of five most accurate modes in response to the detection of the non-zero reference line; and A step of setting the length of the corresponding most accurate mode list to 6 in response to the current reference line being a non-zero reference line. The method according to appendix 4, further comprising. (Appendix 9) The method according to any one of appendices 1 to 8, wherein one of the non-zero reference lines is a line adjacent to the current block and is farther away from the current block than the zero reference line. (Appendix 10) The method according to any one of appendices 1 to 9, wherein the at least one most accurate mode consists of the most accurate mode of the first level. (Appendix 11) The method according to any one of appendices 1 to 10, wherein the at least one most accurate mode includes the most accurate modes of any level from the lowest level of the most accurate mode to the highest level of the most accurate mode. (Appendix 12) The method according to any one of appendices 1 to 10, wherein the at least one most accurate mode includes only the levels of the most accurate modes permitted for the non-zero reference line. (Appendix 13) At least one memory configured to store computer program code, At least one hardware processor configured to access the computer program code and operate as instructed by the computer program code to execute the method according to any one of appendices 1 to 12 A device comprising: (Appendix 14) A computer program for causing a computer to execute the method according to any one of appendices 1 to 12.
Explanation of Reference Numerals
[0251] 100 Communication system 101 Terminal 102 Terminal 103 Terminal 104 Terminal 105 Network 200 Decoder 201 Video source 202 Encoder 203 Capture subsystem 204 Video bitstream 205 Streaming server 206 Copy of the encoded video bitstream 207 Streaming client 208 Copy of the encoded video bitstream 209 Display device 210 Video sample stream 211 Video decoder 212 Streaming client 213 Video sample stream 300 Decoder 301 Channel 302 Receiver 303 Buffer memory 304 Syntax analyzer 305 Scaler / inverse transform unit 306 Motion compensation prediction unit 307 Intra-picture prediction unit 308 Reference picture buffer 309 Current reference picture 310 Aggregation device 311 Loop filter unit 312 Display device 313 Symbol 400 Encoder 401 Video source 402 Controller 403 Source encoder 404 Predictor 405 Reference picture memory 406 Decoder 407 Encoding engine 408 Entropy encoder 409 Transmitter 410 Encoded video sequence 411 Communication channel 2600 Computer system 2601 Peripheral bus 2602 Keyboard 2603 Track pad 2604 Mouse 2605 Joystick 2606 Microphone 2607 Camera 2608 Scanner 2609 Audio output device 2610 Touch screen 2611 Media 2612 CD / DVD ROM / RW 2613 USB Memory 2614 Solid State Drive 2615 Communication Network 2616 Core 2617 Graphics Adapter 2618 Random Access Memory (RAM) 2619 Read Only Memory (ROM) 2620 Internal Mass Storage Device 2621 CPU 2622 GPU 2624 FPGA 2624 Accelerator 2625 Peripheral Bus 2626 System Bus
Prior Art Documents
Patent Documents
[0252]
Patent Document 1
Claims
1. A method executed by at least one processor of an encoder, comprising: generating a bitstream by performing an encoding method comprising: performing encoding of a video sequence of a video image by intra prediction between a plurality of reference lines of the video sequence; storing the bitstream in a storage device; The encoding method includes: determining a plurality of intra prediction modes for a zero reference line among the plurality of reference lines that is closest to the current block of the intra prediction; determining a length of a most probable mode list based on a reference line index value; determining at least one most probable mode included in the most probable mode list for one of one or more non-zero reference lines included in the plurality of reference lines; A method comprising:
2. The method of claim 1 , wherein for the non-zero reference line, a most probable mode index of the current block is signaled to a decoder receiving the bitstream.
3. the at least one most probable mode for the non-zero reference line is included in a most probable mode list; The method of claim 1 , wherein at least one of a planar mode and a DC mode is excluded from the most probable mode list corresponding to any one of the non-zero reference lines.
4. The method of claim 1 , wherein a length of the most probable mode list of the zero reference line is different from a length of the most probable mode list of the non-zero reference line.
5. The method of claim 1 , wherein a length of the most probable mode list for the non-zero reference line is one less than a length of a most probable mode list for the zero reference line.
6. 2. The method of claim 1, wherein a most probable mode list for the zero reference line has a length of six, and a most probable mode list for the non-zero reference line has a length of five.
7. at least one memory configured to store computer program code; at least one hardware processor configured to access said computer program code and to act as instructed by said computer program code to perform the method of any one of claims 1 to 6; An apparatus comprising:
8. A computer program product for causing a computer to carry out the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Encoder, decoder and corresponding method for intra prediction
JP2022528050A
JPP7148637B
JPP7495456B
Multi-reference line intra prediction and most probable mode
US20220038684A1
Intra prediction and intra mode coding
WO2016205693A2