Method and apparatus for filling reference sample

By determining the reference sample of the first block based on the encoding mode of the second block in video encoding and performing prediction encoding, the problem of low encoding efficiency in the prior art is solved, and more efficient image or video block encoding is achieved.

CN120077641APending Publication Date: 2025-05-30INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073283.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2023-10-03
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prediction process of images or video blocks, existing video encoding technologies are difficult to effectively utilize reference samples, resulting in low encoding efficiency.

Method used

The reference samples of the first block are determined based on the encoding mode of the second block for reconstructing the image, and used for prediction, thereby encoding or decoding. Specifically, the reference sample belongs to the third block in the reference area of ​​the first block and is different from the second block.

Benefits of technology

The encoding efficiency of image or video blocks is improved, and the encoding error is reduced and the quality of reconstructed images is improved by more efficiently utilizing reference samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077641A_ABST
    Figure CN120077641A_ABST
Patent Text Reader

Abstract

A method and apparatus for encoding or decoding a video in which, for at least one first block to be encoded or decoded, at least one reference sample is determined based on an encoding mode for reconstructing at least one second block. For example, intra filling or motion compensation of a third block to which at least one reference sample belongs is performed using encoded data of at least one second block. A prediction of the at least one first block is then obtained using the at least one reference sample, and the at least one first block is encoded or decoded based on the prediction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of European Application No. 22306598.8 filed on October 21, 2022, which is incorporated herein by reference in its entirety. Technical Field

[0002] This embodiment generally relates to video compression. This embodiment relates to a method and apparatus for encoding or decoding an image or video. More particularly, this embodiment relates to reference sample determination and image or video block prediction. Background Art

[0003] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transformation to exploit the spatial and temporal redundancies in video content. Generally, intra-frame or inter-frame prediction is used to exploit the intra-frame or inter-frame picture correlations, and then the difference between the original block and the predicted block (usually expressed as prediction error or prediction residue) is transformed, quantized, and entropy encoded. To reconstruct the video, the compressed data is decoded through the inverse processes corresponding to entropy encoding, quantization, transformation, and prediction. Summary of the Invention

[0004] According to one aspect, there is provided a method for encoding an image or video. The method includes: determining at least one reference sample of at least one first block of the image based on the coding mode of at least one second block for reconstructing the image; obtaining a prediction of at least one first block using the at least one reference sample; and encoding the at least one first block based on the prediction. The at least one reference sample belongs to a third block in the reference region of the at least one first block and is different from the at least one second block.

[0005] According to another aspect, there is provided an apparatus for encoding an image or video. The apparatus includes one or more processors operable to determine at least one reference sample of at least one first block of the image based on the coding mode of at least one second block for reconstructing the image, obtain a prediction of at least one first block using the at least one reference sample, and encode the at least one first block based on the prediction. The at least one reference sample belongs to a third block in the reference region of the at least one first block and is different from the at least one second block.

[0006] According to another aspect, there is provided a method for decoding an image or video. The method includes: determining at least one reference sample of at least one first block of the image based on the coding mode of at least one second block for reconstructing the image, obtaining a prediction of at least one first block using the at least one reference sample, and decoding the at least one first block based on the prediction. The at least one reference sample belongs to a third block in the reference region of the at least one first block and is different from the at least one second block.

[0007] According to another aspect, there is provided an apparatus for decoding an image or video. The apparatus includes one or more processors operable to determine at least one reference sample of at least one first block of an image based on an encoding mode of at least one second block for reconstructing the image, obtain a prediction of the at least one first block using the at least one reference sample, and decode the at least one first block based on the prediction. The at least one reference sample belongs to a third block in a reference region of the at least one first block and is different from the at least one second block.

[0008] Further embodiments are described herein that may be used alone or in combination.

[0009] In some embodiments, the at least one reference sample belongs to an unreconstructed block or a block having an encoding mode that does not allow using the at least one reference sample to determine a prediction of the at least one first block.

[0010] In some embodiments, intra prediction is used to encode the at least one second block. In other embodiments, inter prediction is used to encode the at least one second block.

[0011] In some embodiments, intra prediction is used to predict the first block. In other embodiments, inter prediction is used to predict the first block.

[0012] In some embodiments, reference samples determined in a right block and / or a bottom block of the first block are used to predict the first block using additional intra prediction directions. In other words, reference samples from a non-causal region of the first block, which is a region not reconstructed at the time of encoding / decoding the first block, are used to predict the first block.

[0013] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform a method for encoding or decoding an image or video according to any of the embodiments described herein. One or more of the present embodiments also provide a non-transitory computer-readable medium and / or a computer-readable storage medium storing instructions for encoding or decoding an image or video according to the method described herein.

[0014] One or more embodiments also provide a computer-readable storage medium storing a bitstream generated according to the method described herein. One or more embodiments also provide a method and an apparatus for transmitting or receiving a bitstream generated according to the above method. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A block diagram of a system in which aspects of the present embodiments may be implemented is illustrated.

[0016] Figure 2 A block diagram of an embodiment of a video encoder in which aspects of the present embodiment can be implemented is illustrated.

[0017] Figure 3 A block diagram of an embodiment of a video decoder in which aspects of the present embodiment can be implemented is illustrated.

[0018] Figure 4 An example of a filled area of a reference picture is illustrated.

[0019] Figure 5 An example of motion-compensated filling of a filled area is illustrated.

[0020] Figure 6 An example of reference samples for intra prediction is illustrated. The pixel value at coordinates (x, y) in the figure is indicated by P(x, y) relative to the current block starting at (0, 0).

[0021] Figure 7 An example of replacement of reference samples for intra prediction is illustrated.

[0022] Figure 8 An example of a method for replacement of reference samples for intra prediction is illustrated.

[0023] Figure 9 Examples of intra prediction directions are shown: Intra prediction directions in HEVC on the left (the numbers indicate the prediction mode indices associated with the corresponding directions. Modes 2 to 17 indicate horizontal directions H-26 to H+32), and modes 18 to 34 indicate vertical directions V-32 to V+32, and intra prediction directions in VVC on the right.

[0024] Figure 10 An example of wide-angle intra prediction is shown.

[0025] Figure 11 An example of intra prediction in planar mode is shown.

[0026] Figure 12 An example of an inter prediction mode using reconstructed reference samples is illustrated.

[0027] Figure 13 An example of a method for encoding a block of an image or video according to an embodiment is illustrated.

[0028] Figure 14 An example of a method for decoding a block of an image or video according to an embodiment is illustrated.

[0029] Figure 15 An example of replacement of reference samples using MC filling according to an embodiment is illustrated.

[0030] Figure 16 Illustrates an example of an intra prediction method for filling missing reference samples using MC filling according to an embodiment.

[0031] Figure 17 Illustrates an example of reference sample substitution using MC filling according to an embodiment, where the upper right reconstruction block is inter-coded and mrlIdx>0.

[0032] Figure 18 Illustrates an example of replacing unavailable reference samples with MC extended by reference blocks according to an embodiment.

[0033] Figure 19 Illustrates an example of reference sample substitution using intra filling according to an embodiment, where the upper reconstruction block is intra-coded.

[0034] Figure 20 Illustrates an example of intra filling applied to picture filling according to an embodiment.

[0035] Figure 21 Illustrates an example of reference sample estimation for intra prediction according to an embodiment.

[0036] Figure 22 Illustrates an example of an intra prediction method according to an embodiment.

[0037] Figure 23 Illustrates an example of intra prediction angles for intra prediction.

[0038] Figure 24 Illustrates an example of an intra prediction method.

[0039] Figure 25A 、 25B and 25C illustrate examples of reference samples for intra prediction substitution according to an embodiment.

[0040] Figure 26 Illustrates an example of a method for encoding blocks of an image or video according to an embodiment.

[0041] Figure 27 Illustrates an example of a method for decoding blocks of an image or video according to an embodiment.

[0042] Figure 28 Illustrates a block diagram of a system in which aspects of this embodiment can be implemented according to another embodiment.

[0043] Figure 29 Illustrates two remote devices communicating via a communication network according to an example of this principle.

[0044] Figure 30Shows the syntax of a signal that is an example in accordance with the present principles. Detailed Description

[0045] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are specifically described and are generally described in a manner that may sound restrictive, at least for purposes of illustrating the respective characteristics. However, this is for purposes of clarity of description and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Additionally, the described aspects can also be combined and interchanged with aspects described in earlier filings.

[0046] The aspects described and contemplated in this application can be implemented in many different forms. The following Figure 1 、 2 and 3 provide some embodiments, but other embodiments are contemplated, and Figure 1 、 2 and the discussion of 3 do not limit the breadth of implementation. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These and other aspects can be implemented as methods, apparatuses, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having a bitstream generated according to any of the methods stored thereon.

[0047] In this application, the terms "reconstruction" and "decoding" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably.

[0048] Various methods are described herein, and each of the methods includes steps or actions for implementing one or more of the steps or actions of the method. Unless a particular order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first", "second", etc. can be used in various embodiments to modify elements, components, steps, operations, etc., such as for example "first decoding" and "second decoding". Unless specifically required, the use of such terms does not imply an ordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or during a time period overlapping with the second decoding.

[0049] The present aspect is not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations (whether pre-existing or future-developed) and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically precluded, the aspects described in the present application can be used alone or in combination.

[0050] Figure 1 A block diagram illustrating an example of a system in which various aspects and embodiments can be implemented. System 100 can be implemented as a device including various components described below and is configured to perform one or more of the aspects described in the present application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 100 can be implemented individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or to other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in the present application.

[0051] System 100 includes at least one processor 110 that is configured to execute instructions loaded therein for implementing, for example, the various aspects described in the present application. Processor 110 can include embedded memory, input / output interfaces, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., volatile memory devices and / or non-volatile memory devices). System 100 includes a storage device 140 that can include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 140 can include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0052] System 100 includes an encoder / decoder module 130 that is configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents the (one or more) modules that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated into the processor 110 as a combination of hardware and software known to those skilled in the art.

[0053] The program code to be loaded onto the processor 110 or the encoder / decoder 130 to execute the aspects described in this application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of the individual entries during the execution of the processes described in this application. Such stored entries may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operation logic.

[0054] In some embodiments, the memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, fast external dynamic volatile memory (such as RAM) is used as the working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by the Joint Video Exploration Team JVET).

[0055] Inputs to the components of system 100 can be provided via various input devices as indicated in block 105. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives RF signals transmitted over the air, for example, by a broadcaster; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high definition multimedia interface (HDMI) input terminals. Figure 1 Other examples not shown include composite video.

[0056] In various embodiments, the input devices of block 105 have corresponding input processing elements associated therewith as known in the art. For example, the RF section can be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal band to one band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band to select a signal band that can be referred to as a channel, for example, in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) de-multiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and de-multiplexer. The RF section can include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near baseband frequency) or down-converting to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and again filtering to a desired frequency band. Various embodiments re-arrange the order of the (and other) elements described above, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0057] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices across the USB and / or HDMI connections. It is to be understood that various aspects of the input processing (e.g., Reed-Solomon error correction) may be implemented, as needed, for example, within a separate input processing IC or within processor 110. Similarly, various aspects of the USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within processor 110. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on the output device.

[0058] The various elements of system 100 may be provided within an integrated housing, within which the various elements may be interconnected using a suitable connection arrangement 115 (e.g., an internal bus known in the art, including an I2C bus, wiring, and printed circuit boards) and data may be transmitted therebetween.

[0059] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented, for example, within a wired and / or wireless medium.

[0060] In various embodiments, a Wi-Fi network (such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream data to system 100. The Wi-Fi signals of these embodiments are received via the communication channel 190 and the communication interface 150 suitable for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or a router that provides access to an external network (including the Internet) to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide the streamed data to system 100, which delivers the data via the HDMI connection of input block 105. Still other embodiments use the RF connection of input block 105 to provide the streamed data to system 100. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0061] System 100 can provide output signals to various output devices, including display 165, speaker 175, and other peripheral devices 185. Displays 165 of various embodiments include, for example, one or more of a touchscreen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. Display 165 can be used for a television, a tablet computer, a laptop computer, a cellular phone (mobile phone), or other devices. Display 165 can also be integrated with other components (e.g., as in a smart phone), or be separate (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 185 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide functions based on the output of system 100. For example, a disk player performs the function of playing the output of system 100.

[0062] In various embodiments, signaling (such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention) is used to convey control signals between system 100 and display 165, speaker 175, or other peripheral devices 185. Output devices can be communicatively coupled to system 100 via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100 using communication channel 190 via communication interface 150. In an electronic device (such as, for example, a television), display 165 and speaker 175 can be integrated with other components of system 100 in a single unit. In various embodiments, display interface 160 includes a display driver, such as, for example, a timing controller (TCon) chip.

[0063] For example, if the RF portion of input 105 is part of a separate set-top box, then display 165 and speaker 175 can alternatively be separate from one or more of the other components. In various embodiments where display 165 and speaker 175 are external components, output signals can be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.

[0064] The described embodiments may be executed by computer software implemented by a processor 110, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. As a non-limiting example, the memory 120 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory). As a non-limiting example, the processor 110 may be of any type suitable for the technical environment and may encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture).

[0065] Figure 2 An encoder 200 is illustrated. Variations of this encoder 200 are contemplated, but for the sake of clarity, the encoder 200 is described below without describing all the expected variations.

[0066] In some embodiments, Figure 2 An encoder that improves the HEVC standard or the VVC standard or an encoder that employs techniques similar to HEVC or VVC is also illustrated, such as the encoder ECM being developed by JVET (Joint Video Exploration Team).

[0067] Before encoding, the video sequence may undergo pre-encoding processing (201). For example, a color transformation (e.g., from RGB 4:4:4 to YCbCr 4:2:0) may be applied to the input color picture, or remapping may be performed on the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of color components), or the picture may be resized (e.g., downscaled). Metadata may be associated with the preprocessing and attached to the bitstream.

[0068] In the encoder 200, pictures are encoded by encoder elements as described below. The picture to be encoded is segmented (202) and processed in units such as CUs (coding units) or blocks. In the present disclosure, different expressions may be used to refer to such units or blocks resulting from the segmentation of a picture. Such expressions may be coding units or CUs, coding blocks or CBs, luminance CBs or blocks... A CTU (coding tree unit) may refer to a group of blocks or a group of units. In some embodiments, the CTU itself may be regarded as a block or a unit.

[0069] For example, each unit is encoded using an intra or inter mode. When a unit is encoded in the intra mode, it performs intra prediction (260). In the inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) which of the intra mode or inter mode to use for encoding the unit and indicates the intra / inter decision via, for example, a prediction mode flag. The encoder may also mix (263) the intra prediction result and the inter prediction result, or mix the results from different intra / inter prediction methods. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.

[0070] The motion refinement module (272) uses the already available reference pictures to refine the motion field of the block without referring to the original block. The motion field of a region can be regarded as a set of motion vectors for all pixels in that region. If the motion vectors are block-based, the motion field can also be represented as a set of motion vectors for all sub-blocks within the region (all pixels within a sub-block have the same motion vector and the motion vectors may vary from sub-block to sub-block). If a single motion vector is used for the region, the motion field of the region can also be represented by a single motion vector (all pixels within the region have the same motion vector).

[0071] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vectors and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder may skip the transformation and directly apply quantization to the untransformed residual signal. The encoder may bypass both the transformation and quantization, i.e., directly encode the residual without applying the transformation or quantization process.

[0072] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce encoding artifacts. The filtered image is stored at the reference picture buffer (280).

[0073] Figure 3 The block diagram of a video decoder 300 is illustrated. In the decoder 300, the bitstream is decoded by decoder elements as described below. The video decoder 300 generally performs a decoding process that is the reverse of the encoding process as described in Figure 2 The encoder 200 generally also performs video decoding as part of encoding the video data.

[0074] In particular, the input to the decoder includes a video bitstream, which may be generated by the video encoder 200. First, entropy decoding (330) is performed on the bitstream to obtain transform coefficients, motion vectors, and other coded information. The picture segmentation information indicates how to segment the picture. Thus, the decoder can partition (335) the picture according to the decoded picture segmentation information. The transform coefficients are dequantized (340) and inverse-transformed (350) to decode the prediction residuals. The decoded prediction residuals and the prediction blocks are combined (355) to reconstruct the image block.

[0075] The prediction block can be obtained (370) from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375). The decoder can blend (373) the intra prediction result and the inter prediction result, or blend the results from multiple intra / inter prediction methods. Before motion compensation, the motion field can be refined (372) by using the already available reference pictures. The loop filter (365) is applied to the reconstructed image. The filtered image is stored at the reference picture buffer (380).

[0076] The decoded picture can further undergo post-decoding processing (385), e.g., inverse color transformation (e.g., from YCbCr 4:2:0 to RGB 4:4:4) or inverse inverse remapping that performs the inverse of the remapping process performed in the pre-coding processing (201), or resizing (e.g., upscaling) the reconstructed picture. The post-decoding processing can use the metadata derived in the pre-coding processing and signaled in the bitstream.

[0077] Some of the embodiments described herein relate to improving the filling of reference samples for blocks of an image or video to be encoded or decoded.

[0078] Reference picture boundary filling

[0079] In a conventional video codec, during the process of reconstructing a picture that is further used as a reference picture, an extended picture region is constructed using the region of the image width / height around the size of the "extSize" column / row, as Figure 4 depicted therein. The samples in the extended region are derived from replicated boundary filling. In inter prediction, when a reference block is partially or fully located outside the picture boundary (OOB), the replicated filled pixels are used for motion compensation (MC). In this way, the reference block can be partially located outside the reconstructed reference picture, as Figure 4 depicted therein, and / or the motion vector is not clipped such that the reference block is fully included in the reference picture. This feature allows for improved coding efficiency.

[0080] Motion compensated picture boundary filling (MC filling)

[0081] Advantageously, as Figure 5As illustrated, it has been proposed to use motion compensation to replace the repetitive filling of the region (530) near the current block (510) encoded inter-frame, and use M columns or M rows to extend (540) the reference block (520). Motion compensation uses the same (one or more) motion vectors MV as those used for reconstructing the current block, as Figure 5 depicted in

[0082] Intra prediction and reference sample substitution

[0083] The intra-frame prediction process in HEVC and VVC consists of three steps:

[0084] · Reference sample generation

[0085] · Intra-frame sample prediction and

[0086] · Post-processing of the predicted samples.

[0087] Figure 6 is illustrated in Reference sample generation the process. The reference samples ref[] are also referred to as L-shaped. For a prediction unit (PU) of size NxN, the top row (2N + 2refIdx) of decoded samples is formed by the previously reconstructed top and top-right pixels of the current PU. Similarly, the left column (2N + 2refIdx) of samples is formed by the reconstructed left pixel and bottom-left pixel. In VVC, the reference rows and columns of samples may be more than one sample away (d = refIdx) from the current block, as Figure 6 depicted in

[0088] Figure 8 illustrates an example of the method (800) for reference sample generation. The corner pixel at the top-left position is also used to fill the gap between the top row and left column references. If some of the samples on the top or left are not available (810), for example because the corresponding CU is not in the same slice, or for example the current CU is at the frame boundary ( Figure 7 at 710 on Figure 7 ), or the current CU is at the bottom-right after quadtree splitting ( Figure 7 at 720 on Figure 8 ), then a method called reference sample substitution is performed, where the missing samples are copied from the available samples in clockwise and counterclockwise directions ( Figure 7 at 700 on Figure 8 ), Figure 7In it, the dashed area corresponds to the area of the picture that has not been reconstructed, and the missing reference samples are represented by dotted lines. At 820, when the top sample or the left sample is available, the reconstructed sample is copied into the reference sample buffer.

[0089] Next, according to the current CU size and prediction mode, a specified filter can be used to filter the reference samples.

[0090] Intra-sample prediction It consists of pixels that predict the target CU based on reference samples. There are different prediction modes: the planar and DC prediction modes are used to predict smooth and gradually changing regions, while the angular (angles defined clockwise from 45 degrees to -135 degrees) prediction mode is used to capture different directional structures. For square blocks, HEVC supports 33 directional prediction modes indexed from 2 to 34. These prediction modes correspond to different prediction directions, as Figure 9 shown in the left side. In VVC, there are 65 angular prediction modes, corresponding to the 33 angular directions defined in HEVC, and another 32 directions, each corresponding to the intermediate direction between an adjacent pair ( Figure 9 right side).

[0091] In VVC, for non-square blocks, the non-allowed regular-direction intra prediction is replaced with additional wide-angle intra prediction modes (see Figure 10 ).

[0092] For a given angular prediction mode, the predictor samples on the reference array are copied along the corresponding direction within the target PU. Some predictor samples may have integer positions, in which case they match the corresponding reference samples; the positions of other predictors will have a fractional part, indicating that their positions will fall between two reference samples. In the latter case, the nearest reference samples are used to interpolate the predictor samples ( Post-processing of predicted samples ). In HEVC, linear interpolation is performed on the two nearest reference samples to calculate the predictor sample value; in VVC, to interpolate the predictor samples, a 4-tap filter fT[] is used, and the filter is selected based on the intra mode direction.

[0093] In addition to the directional mode, the DC mode fills the prediction using the average of the samples in an L-shape (except for rectangular CUs that use the average of the reference samples on the longer side), while the planar mode interpolates the reference samples spatially, as Figure 11 depicted in.

[0094] Other prediction modes using reconstructed reference sample substitution

[0095] There are other coding modes where block prediction is based on reconstructed reference samples located in adjacent templates. For example, Local Illumination Compensation (LIC) derives an illumination compensation model to correct inter-frame predicted samples with a linear model:

[0096] P’(x) = a.P(x) + b

[0097] where P’ is the corrected prediction, P is the inter-frame prediction, x is the sample position, and (a, b) are the illumination compensation parameters (LIC model). The LIC model parameters are derived using some of the reconstructed samples (1210) adjacent to the current block in the current picture and adjacent samples (1235) co-located with the reference block in the reference picture, as Figure 12 depicted. However, some of the reconstructed reference samples in the current picture may be unavailable because additional conditions have restricted (prohibited) access to some of the reconstructed samples (1220) in the current picture for the purpose of reducing implementation complexity (memory access, number of pipeline operations to be reconstructed per block, etc.). For example, such a restriction may be not to access the reconstructed samples of adjacent blocks encoded intra-frame to reconstruct the current block encoded in inter-frame mode. In these cases, reference sample substitution (such as duplicate filling) should be applied appropriately.

[0098] The above MC filling method allows for improved filling of inter-frame encoded pictures (or blocks), however, intra-frame pictures (or blocks) still use duplicate filling because there are no reference blocks that can be extended using inter-frame prediction parameters.

[0099] In some other modes of intra-frame prediction and using reference sample substitution, the reference sample substitution process as described above allows for coping with the unavailability of missing reconstructed samples, but at the cost of reduced coding efficiency due to simple duplication.

[0100] In some embodiments, the reference sample substitution process for intra-frame prediction or inter-frame prediction (also referred to as filling of reference samples) is modified by replacing duplicate filling with motion compensation or intra-sample prediction filling techniques. The embodiments described herein can be applied to any other coding mode where prediction uses adjacent reconstructed reference samples.

[0101] In some embodiments, for the case of samples at the boundary of intra-frame encoded blocks, reference picture boundary filling is improved.

[0102] In other embodiments, reference sample substitution is extended to improve the intra-frame prediction process, such as the intra-frame prediction process known from HEVC or VVC, which has reference samples with an estimate closer to the actual current CU boundary.

[0103] Any one of the embodiments described herein may be implemented, for example, in the intra prediction module 260 or motion estimation 275, motion refinement 272, or motion compensation 270 of the image or video encoder 200, or in the intra prediction module 360 or motion refinement 372 or motion compensation 375 of the image or video decoder 300.

[0104] Figure 13 An example of a method 1300 for encoding a block of an image or video according to an embodiment is illustrated. At 1310, one or more reference samples of a reference region of a block of an image to be encoded are determined. The reference region is, for example, an L-shape on the top and left of the block to be encoded, as Figures 4-5 illustrated in any one of 6-7, 10-12.

[0105] One or more reference samples are determined based on the coding mode of one or more second blocks used to reconstruct the image. The one or more reference samples determined are not part of the one or more second blocks, i.e., they are outside the one or more second blocks. In some embodiments, when the block to be encoded is being processed for encoding, the one or more reference samples belong to blocks that have not yet been reconstructed. In other embodiments, the one or more reference samples belong to blocks that have a coding mode that is not allowed to be used when predicting the block to be encoded using a given coding mode. For example, the one or more reference samples belong to intra-coded blocks, while the block to be encoded will be encoded using an inter prediction coding mode using the LIC tool, in which case, as discussed above, the intra-coded blocks cannot be used to determine the LIC parameters. And depending on the encoder / decoder implementation, the intra-coded blocks in the inter frames have not even been reconstructed when the current block is being processed for inter prediction.

[0106] In some variants, the one or more reference samples belong to blocks adjacent to the one or more second blocks. According to the variants described further below, the one or more second blocks have been reconstructed using an inter prediction mode or using an intra prediction mode.

[0107] When encoding the one or more second blocks using an inter prediction mode, the one or more reference samples are filled with motion compensation data obtained using the motion information of the one or more second blocks.

[0108] When encoding the one or more second blocks using an intra prediction mode, the one or more reference samples are filled with data obtained using the same intra prediction mode as the intra prediction mode used for the one or more second blocks.

[0109] Once one or more reference samples have been determined, at 1320, a prediction of the block to be encoded is obtained using the one or more reference samples, and at 1330, the block is encoded using the prediction.

[0110] Figure 14 illustrates an example of a method 1400 for decoding blocks of an image or video according to an embodiment. At 1410, one or more reference samples of a reference region of a block belonging to an image to be decoded are determined in a similar manner as in Figure 13 1310. Once the one or more reference samples have been determined, at 1420, the one or more reference samples are used to obtain a prediction for the block to be decoded, and at 1430, the block is decoded using the prediction.

[0111] Some variations of the above-mentioned embodiment are further described below.

[0112] Reference sample substitution for intra prediction using adjacent blocks coded in inter mode

[0113] In this variation, the block to be encoded / decoded is intra-predicted, while one or more blocks for determining unavailable reference samples of the block are inter-coded.

[0114] Here, an MC filling process is used to fill in missing reference samples for intra-prediction. Figure 15 and 16 The above figure illustrates this variation, where the dashed region in Figure 15 corresponds to the region of the picture that has not been reconstructed yet, and the missing reference samples are shown as dotted lines.

[0115] Let us consider the current block to be encoded or decoded using intra-prediction, and denote (1510) the rightmost reconstructed block above the current block encoded in an inter-prediction mode. A virtual block (1520) is constructed using data obtained from motion compensation of the right extension (1530) of the reference block used to predict the above-mentioned block (1510). In the case of bi-prediction, a virtual block (1510) is constructed using data obtained from motion compensation and blending of the right extensions of the two reference blocks used to predict the above-mentioned block. The bottom samples of the virtual block (1520) are then used to fill in the missing reference samples to be used for intra-prediction of the current block.

[0116] Similarly, if the left block (1540) reconstructed at the left side of the current block is inter-coded, the same method can be used for the missing reference samples (1550) at the lower left. In this case, the virtual reference block extension (1560) will be located below the reference block used to predict the left block (1540).

[0117] A virtual block (1550) is constructed using data obtained from motion compensation that extends from the bottom of the reference block for predicting the left block (1540). The right samples of the virtual block (1550) are then used to fill in the missing reference samples to be used for intra prediction of the current block. When both the upper-right and lower-left blocks are unavailable and their corresponding right or upper neighboring blocks are inter-coded, two examples of Figure 15 can be combined.

[0118] In a similar manner, when the top block and / or the left block are unavailable, if the corner block (1570) is inter-coded, the same process can be used to determine the missing reference samples of the top block and / or the left block. The virtual reference block extension will be to the right and / or below the reference block of the corner block. In a variant, if both the upper-right block and / or the lower-left block of the current block are unavailable, the virtual reference block extension can be further extended to the right and / or below the reference block of the corner block.

[0119] Figure 8 The reference sample substitution process of Figure 16 is modified according to one of the variants described herein, as illustrated (1600) in

[0120] At 810, it is determined whether the upper-right reference sample and the lower-left reference sample are available, respectively. In other words, it is determined whether the upper-right block and the lower-left block of the current block are still reconstructed, or whether their reconstructed samples can be used to predict the current block. If it is determined that the upper-right reference sample and the lower-left reference sample are available, respectively, at 820, the reconstructed upper-right reference sample and the upper-left reference sample are copied to the reference sample buffer, respectively. Otherwise, at 1610, it is determined whether the left neighboring block of the upper-right block and the top neighboring block of the lower-left block are available and inter-coded reconstructed blocks. If the response is no, at 830, repeated filling is performed to fill the reference samples from the upper-right block and the lower-left block, respectively. Otherwise, at 1620, the reference samples from the upper-right block and the lower-left block are filled into the reference sample buffer, respectively, using motion compensation (MC) filling.

[0121] Then, at 840, using the filled reference sample buffer, intra prediction is used to predict the current block.

[0122] In another variant, to avoid discontinuities between the available (reconstructed) reference samples and the filled-in (missing) reference samples, an offset can be added to the alternative reference samples. The offset is determined as the difference between the last available reference sample (i.e., the reference sample from the reconstructed block coded inter-frame that is closest to the missing reference sample) and the first alternative reference sample. In practice, the available (reconstructed) reference samples consist of the inter-frame prediction plus the residual, while the alternative reference samples consist only of the inter-frame prediction (without the residual).

[0123] Figure 17 Depicts another example where "mrlIdx" is non-zero. In this case, advantageously, the construction of the virtual block (1720) can be performed using motion compensation with an extension (1730) of the reference block used to predict the reconstructed block (1710) spatially close to the missing reference sample. In Figure 17 the example of, the virtual reference block extension (1730) is located below the reference block.

[0124] The selection of the reconstructed block coded inter-frame and the motion parameters used to derive the samples for filling in the missing reference samples can be derived according to different rules. It can be the block closest to the sample to be replaced, or it can follow a given rule, for example: always use the top-leftmost reconstructed block coded inter-frame to fill in the missing sample in the top-right. The given rule can also depend on the priority order of the neighboring blocks used to check the block with the missing sample and the coding mode of the neighboring blocks. For example, if the missing sample is in the top-right block, first check the top-leftmost reconstructed block, and if that block is coded inter-frame, use it to fill in the top-right missing sample, otherwise, check the block above the top-right block, and if that block is coded inter-frame, use it to fill in the missing sample, otherwise use repeated filling to fill in the missing sample.

[0125] Therefore, the rule for selecting the reconstructed block for MC filling can be based on at least one of the spatial distance from the missing reference sample to the reconstructed block, or on the coding mode of the reconstructed block, or on the position of the reconstructed block relative to the current block or the missing reference sample.

[0126] Reference sample substitution for inter prediction using adjacent blocks coded in inter mode

[0127] There are other coding modes where the block prediction mode is based on the reconstructed reference samples located in neighboring templates and which impose some restrictions that make some reference samples unavailable. For these unavailable reference samples, a similar process as described above can be used to replace the conventional reference sample substitution method to replace the missing reference samples. Variants of such embodiments are described in Figure 18 as follows.

[0128] Some reference samples (1820) are not available for the current block for prediction, for example because the reference sample belongs to an intra-coded block and cannot be used for inter prediction of the current block (e.g., for predicting the LIC parameters of the current block), or the block with the unavailable reference sample (1820) has not been reconstructed yet.

[0129] If some reference samples are not available (1820) for the current block to be predicted, then a closer reconstructed sample is searched for from the blocks (1830) adjacent to the block with the unavailable reference sample (1820) and coded in inter mode. Then, the motion parameters (inter prediction 1) used for reconstructing the adjacent block (1830) are used to identify the reference block (1835) of the adjacent block (1830), and it is extended (1840) such that the motion compensation (MC) samples (1850) can be used to replace the unavailable reference samples (1820).

[0130] In a variant, if the current block is coded in inter mode, the motion parameters of the current block can be used (inter prediction 2) to identify the reference block of the current block and extend the reference block to fill the unavailable reference samples. In this case, the extended reference block is the reference block of the current block. In Figure 18 the reference samples (1825) above the reference block of the current block to be predicted are used to fill the unavailable reference samples (1820).

[0131] Reference sample substitution for intra prediction using adjacent blocks coded in intra mode

[0132] In an embodiment, similar to the above embodiment, in the case where the reconstructed adjacent block has been intra-coded, a similar mechanism can be used to fill the missing reference samples for intra prediction. This embodiment is depicted with the example in Figure 19 where the dashed area corresponds to the area of the picture that has not been reconstructed yet, and the missing reference samples are represented by dotted lines. In this example, the top-rightmost reconstructed block (1910) has been intra-coded, and the intra direction is depicted by the gray arrow on the upper left of the figure.

[0133] In this embodiment, a virtual block (1920) is constructed as an intra prediction of the right extension of the upper block (1910). The bottom samples of the virtual block (1920) are used to fill the missing reference samples to be used for intra prediction of the current block. Similarly, if the block reconstructed at the left side of the current block is intra-coded, the same method can be used for the missing reference samples at the lower left of the current block. In this case, the virtual reference block extension (1920) will be located below the reconstructed left block.

[0134] In some variations, the above in - frame filling can be adjusted to some subset of in - frame directions, and otherwise conventional filling is used. For example, it is determined whether the intra - prediction direction of the top - right - most reconstructed block (1910) or the left block is in a given set of intra - prediction modes. If this is the case, the missing reference samples are filled with data obtained using the same intra - prediction direction as the top - right - most reconstructed block (1910) or the left block. Otherwise, conventional filling is used.

[0135] Intra filling for picture filling using adjacent blocks coded in intra mode

[0136] In an embodiment, when the reconstructed boundary block (2010) is intra - coded, the repetitive filling applied at the picture boundary is replaced with in - frame filling, as depicted in the example of Figure 20 The intra - prediction direction used for reconstructing the intra - block (2010) is used to fill the padding samples in the block extension (2020) using the intra - prediction process. Additional reference samples can be used, and the missing reference samples can be filled with conventional methods (such as repetitive filling).

[0137] Estimation of reference samples closer to the current block for intra coding mode

[0138] In another embodiment, the top - right (or bottom - left) reference sample used for intra - prediction is replaced with an estimated reference sample (2150) located near the right edge or bottom edge of the current block, as illustrated above and described with reference to Figure 21 as shown in and with reference to Figure 22 as described above, Figure 22 shows a method 2200 for intra - prediction according to an embodiment. At 2210, it is determined whether the reconstructed sample located at the top - right (or bottom - left) has been coded in an inter - frame mode (2110), and then at 2220, motion information (motion vector and reference index) is used to construct an estimated reference sample (2150). The motion information is used to identify the reference block for reconstructing the sample located at the upper - right (or lower - left), and it is extended downwards (2130) (or upwards on the left, depending on the position of the sample to be estimated relative to the current block). Next, at 2230, the estimated reference sample (2150) is back - projected into the position of the conventional top - right (or bottom - left) reference sample (2155) using interpolation with the intra - prediction direction θ, as Figure 21As depicted on the right side of. For example, at 2230, if the estimated reference sample belongs to the block 2130 located on the right side of the current block, the reference sample on the first column of the block 2130 located on the right side of the current block is back-projected to the last row of the block 2110 located in the upper right of the current block using the intra prediction direction θ. In another example, if the estimated reference sample belongs to the block located at the bottom of the current block, the reference sample on the first row of the block located at the bottom of the current block is back-projected to the last column of the block located in the lower left of the current block using the intra prediction direction θ.

[0139] At 2240, intra prediction is performed using the estimated reference sample with the conventional intra prediction direction.

[0140] In another variant, the estimated reference sample is directly used without backpropagation and interpolation, but the intra prediction process is modified as follows.

[0141] In the conventional intra prediction process (such as in HEVC or VVC), in the case of the horizontal direction, the left reference sample and the upper reference sample are swapped, and intra prediction is applied as in the vertical direction, as Figure 24 depicted in Figure 24 shows an example of the method 2400 for intra prediction. At 2410, it is determined whether to perform intra prediction using the input horizontal intra prediction direction. If this is the case, the left reference sample and the upper reference sample are swapped at 2420. Then, intra prediction is performed at 2430 using the intra prediction direction, which is either the input vertical intra prediction direction or the vertical intra prediction direction corresponding to the input horizontal intra prediction direction. If the input intra prediction direction is horizontal (2410), the intra prediction is flipped at 2440, that is, the prediction obtained from the intra prediction performed at 2430 is mirrored with respect to the diagonal of the current block.

[0142] Figure 23 illustrates an example of the range of intra prediction directions and directions considered to be horizontal intra prediction directions and vertical intra prediction directions.

[0143] In this variant, the estimated (right and / or bottom) reference sample can be swapped with and / or flipped from the conventional left and / or upper reference sample so that the estimated reference sample is located in the upper left of the current block, and conventional intra prediction is performed using the estimated reference sample. Finally, the predicted sample is flipped back to its original position.

[0144] The above variants can replace the conventional intra prediction mode process or can be additional intra prediction modes. In this variant, the modified intra prediction process or the additional intra prediction mode allows intra prediction of the current block based on estimated reference samples located at the right and bottom of the current block.

[0145] For example, if the intra prediction direction angle θ is positive in the vertical direction (e.g., θ = 45°) ( Figure 23 ), then the estimated reference sample (2150) in the right is copied into the left reference sample buffer, and the upper sample (and finally the upper-left sample) is flipped into the top reference sample buffer. Next, the conventional intra prediction process is performed with an intra prediction direction angle equal to θ - 90° (-45° in the example) in the vertical direction. Finally, the predicted samples are horizontally flipped, as Figure 25A depicted in 2501.

[0146] In Figure 25B another variant illustrated in 2502, the estimated reference sample (2150) in the bottom is copied into the upper and upper-right reference sample buffers, and the left sample is flipped into the left reference sample buffer. Finally, the upper-left reference sample is flipped into the lower-left reference sample buffer. Next, the conventional intra prediction process is performed with an intra prediction direction angle equal to θ - 90°. Finally, the predicted samples are vertically flipped. In the variant, if the intra prediction direction angle θ is positive in the horizontal direction (e.g., θ = 45°), then this variant is applied.

[0147] In Figure 25C another variant illustrated in 2503, the estimated reference sample (2150) in the bottom is copied into the upper and upper-right reference sample buffers and flipped, and the right sample is flipped into the left and lower-left reference sample buffers. Next, the conventional intra prediction process is performed with an intra prediction direction angle equal to θ - 90°. Finally, the predicted samples are diagonally flipped. In the variant, if an additional flag is signaled and / or if the intra prediction direction angle θ is negative (e.g., θ = -45°), then this variant is applied.

[0148] These variants can also be extended to the general case of intra prediction where mrlIdx >= 0. In these variants, Example 2501 is then extended to the following example, where the reference samples in the column of the block located above the left of the current block are filled with the estimated reference samples in the corresponding column of the block located above the right of the current block, the reference samples in the row of the block located above the current block are flipped, and the reference samples in the row of the block located above the upper-right of the current block are filled with the flipped reference samples from the corresponding row of the block located above the upper-left of the current block. The corresponding row or column is the row or column with the same row or column distance given by mrlIdx.

[0149] Example 2502 is extended to the following example, where reference samples in the rows of the blocks above the current block are filled with estimated reference samples in the corresponding rows of the block at the bottom of the current block, reference samples in the columns of the blocks on the left side of the current block are flipped, and reference samples in the columns of the blocks in the lower left of the current block are filled with flipped reference samples from the corresponding columns of the block in the upper left of the current block.

[0150] Example 2503 is extended to the following example, where reference samples in the rows of the blocks above the current block and in the corresponding rows of the blocks in the upper right of the current block are filled with flipped estimated reference samples in the corresponding rows of the block at the bottom of the current block and flipped reference samples in the corresponding rows of the block in the lower left of the current block, and reference samples in the columns of the blocks on the left side of the current block and in the columns of the blocks in the lower left of the current block are filled with flipped estimated reference samples in the corresponding columns of the block on the right side of the current block and flipped reference samples in the corresponding columns of the block in the upper right of the current block.

[0151] In another embodiment, an indicator (e.g., a flag) is encoded in the bitstream to indicate whether any of the variants described herein in the embodiments of estimating reference samples is used, or whether conventional intra prediction is used. In the variant, the indicator is encoded only if the variant can be applied. For example, if at least one of the upper right or lower left blocks is encoded in an inter mode, the indicator is encoded, while if the upper right and lower left blocks are encoded in an intra mode, the indicator is not encoded. In another variant, if both the upper right and lower left blocks are encoded in an inter mode, the indicator signals whether conventional intra prediction, swapping of right samples, swapping of bottom samples, or swapping of right and bottom samples should be applied. In the last case, the method allows addressing intra prediction angles up to 135 degrees.

[0152] According to the embodiment depicted as Figure 26 herein, the embodiments described herein can be used in method 2600 for encoding blocks of an image or video. At 2610, one or more reference samples are determined for the blocks on the right side and / or below the block to be encoded. Any of the above variants can be used to determine the reference samples, e.g., based on the coding mode of the blocks adjacent to the right or bottom block. Depending on the coding mode of the adjacent blocks, intra filling or MC filling as described herein can be used.

[0153] Once one or more reference samples have been determined, at 2620, an intra prediction of the block to be coded is obtained using the one or more reference samples determined at 2610, and at 2630, the block is coded using the intra prediction. At 2620, intra prediction may be performed using additional intra prediction directions, such as Figure 23 the intra prediction directions illustrated above are mirrored with respect to the lower left to upper right diagonal.

[0154] In another example, at 2620, intra prediction is performed using conventional intra prediction directions, but where the right buffer and the lower buffer are swapped and flipped, such as Figure 25A , 25B or as described in 25C. Then the intra prediction is flipped according to the intra prediction direction.

[0155] Figure 27 An example of a method 2700 for decoding a block of an image or video according to an embodiment is illustrated. The method 2700 for decoding a block implements the same intra prediction embodiments as the embodiments described with respect to the encoding method 2600. At 2710, one or more reference samples are determined in a similar manner to Figure 26 2610 of

[0156] Once one or more reference samples have been determined, at 2720, an intra prediction of the block to be decoded is obtained using the one or more reference samples (as in 2620), and at 2730, the block is reconstructed using the intra prediction.

[0157] Figure 28 A block diagram of a system in which aspects of the present embodiment may be implemented according to another embodiment is illustrated. Figure 28 An embodiment of an apparatus 2800 for encoding or decoding an image or video according to any of the embodiments described herein is shown. The apparatus includes a processor 2810 and may be interconnected with a memory 2820 through at least one port. Both the processor 2810 and the memory 2820 may also have one or more additional interconnections to the outside.

[0158] The processor 2820 is also configured to determine at least one reference sample of at least one first block of an image based on an encoding mode of at least one second block for reconstructing the image, where the at least one reference sample is outside the at least one second block, obtain a prediction for the at least one first block using the at least one reference sample, and encode the at least one first block based on the prediction using any one of the embodiments described herein. For example, the processor 2820 is configured using a computer program product that includes code instructions implementing any one of the embodiments described herein.

[0159] In another embodiment, the processor 2820 is also configured to determine at least one reference sample of at least one first block of an image based on an encoding mode of at least one second block for reconstructing the image, where the at least one reference sample is outside the at least one second block, obtain a prediction for the at least one first block using the at least one reference sample, and decode the at least one first block based on the prediction using any one of the embodiments described herein. For example, the processor 2820 is configured using a computer program product that includes code instructions implementing any one of the embodiments described herein.

[0160] In Figure 28 the illustrated embodiment, in a transmission context over a communication network NET between two remote devices A and B, device A includes a processor associated with memories RAM and ROM, which is configured to implement the method for encoding an image or video as described with respect to Figures 1-27 and device B includes a processor associated with memories RAM and ROM, which is configured to implement the method for decoding an image or video as described with respect to Figures 1-27 According to an example, the network is a broadcast network, which is adapted to broadcast / transmit the encoded image or video from device A to a decoding device including device B.

[0161] Figure 30 shows an example of the syntax of a signal or bitstream transmitted via a packet-based transport protocol. Each transmitted packet P includes a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may include image or video data according to any one of the above embodiments. In a variant, the signal or bitstream includes data representing any one of the following:

[0162] an indicator indicating whether the missing reference sample for determining the first block is based on the encoding mode for reconstructing the second block,

[0163] an indicator indicating whether repeated filling or MC filling is used for reference sample substitution,

[0164] An indicator indicating whether to use any of the variants for estimating a reference sample described in the embodiments herein or whether to use conventional intra prediction

[0165] An indicator indicating additional intra prediction directions that can be used to encode / decode using the right block and / or bottom block of the first block

[0166] An indicator indicating whether to swap the reference samples of a given block adjacent to the first block

[0167] Various implementations relate to decoding. As used in this application, "decoding" can cover, for example, all or part of the process of performing on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more of the processes typically performed by a decoder, e.g., entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes or alternatively includes processes performed by the decoders of the various implementations described herein, e.g., entropy decoding a sequence of binary symbols to reconstruct image or video data

[0168] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding, and in another embodiment, "decoding" refers to the entire reconstructed picture process including entropy decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally refer to a broader decoding process will be clear based on the context of the specific description and is considered to be well understood by those skilled in the art

[0169] Various implementations relate to encoding. In a manner similar to the discussion above regarding "decoding", as used in this application, "encoding" can cover, for example, all or part of the process of performing on an input video sequence to produce an encoded bitstream. In various embodiments, such a process includes one or more of the processes typically performed by an encoder, e.g., segmentation, differential encoding, transform, quantization, and entropy encoding. In various embodiments, such a process also includes or alternatively includes processes performed by the encoders of the various implementations described herein, e.g., determining resampling filter coefficients, resampling the decoded picture

[0170] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally refer to a broader encoding process will be clear based on the context of the specific description and is considered to be well understood by those skilled in the art.

[0171] Note that the grammatical elements used herein are descriptive terms. Thus, they do not preclude the use of other grammatical element names.

[0172] The present disclosure has described fragments of various information that can be transmitted or stored, such as, for example, syntax. The information can be packaged or arranged in a variety of ways, the variety of ways including, for example, ways common in video standards, such as putting the information into SPS, PPS, NAL units, headers (e.g., NAL unit headers or slice headers), or SEI messages. Other ways are also available, including, for example, ways common in system-level or application-level standards, such as putting the information into one or more of the following:

[0173] a. SDP (Session Description Protocol), for describing the format of a multimedia communication session, for the purposes of session announcement and session invitation, for example as described in RFCs and used in combination with RTP (Real-Time Transport Protocol) transmission.

[0174] b. DASH MPD (Media Presentation Description) descriptor, for example as used in DASH and transmitted via HTTP, the descriptor being associated with a representation or a set of representations to provide additional characteristics to the content representation.

[0175] c. RTP header extensions, for example as used during RTP streaming.

[0176] d. ISO Base Media File Format, for example as used in OMAF and using boxes, which are object-oriented building blocks defined by a unique type identifier and a length (also referred to as "atoms" in some specifications).

[0177] e. HLS (HTTP Live Streaming) manifests transmitted via HTTP. The manifests can be associated, for example, with a version or a set of versions of the content to provide characteristics of the version or the set of versions.

[0178] When the drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding apparatus. Similarly, when the drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding method / process.

[0179] Some embodiments relate to rate - distortion optimization. In particular, during the encoding process, a balance or trade - off between rate and distortion is typically considered, often taking into account constraints on computational complexity. Rate - distortion optimization is typically expressed as minimizing a rate - distortion function, which is a weighted sum of rate and distortion. There are different ways to solve the rate - distortion optimization problem. For example, the method can be based on an exhaustive test of all encoding options (including all considered modes or encoding - parameter values), with a complete evaluation of their encoding cost and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to avoid (save) encoding complexity, especially for the calculation of approximate distortion based on the predicted or prediction - residual signal rather than the reconstructed residual signal. A hybrid of these two methods can also be used, such as by using approximate distortion for only some of the possible encoding options and full distortion for other encoding options. Other methods only evaluate a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform the optimization, but the optimization does not necessarily involve a complete evaluation of both the encoding cost and the associated distortion.

[0180] The implementations and aspects described herein can be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., only as a method), the implementation of the discussed features can be implemented in other forms (e.g., an apparatus or a program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. A method can be implemented in, for example, a processor that generally refers to a processing device, which includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as, for example, a computer, a cellular phone, a portable / personal digital assistant (“PDA”), and other devices that facilitate the transfer of information between end - users.

[0181] References to “one embodiment” or “an embodiment” or “one implementation” or “an implementation” and other variations thereof mean that the specific features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the appearances of the phrases “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation” and any other variations thereof throughout this application do not necessarily all refer to the same embodiment.

[0182] Additionally, this application can relate to “determining” pieces of various information. Determining information can include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from a memory.

[0183] In addition, the present application may relate to "accessing" fragments of various information. Accessing information may include, for example, one or more of the following: receiving information, retrieving information (e.g., retrieving information from a memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information.

[0184] Additionally, the present application may relate to "receiving" fragments of various information. As with "accessing", receiving is intended to be a broad term. Receiving information may include, for example, one or more of the following: accessing information or retrieving information (e.g., retrieving information from a memory). Further, during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information, "receiving" is typically involved in one way or another.

[0185] It is to be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any of the following, " / ", "and / or", and "at least one of...", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first-listed option and the second-listed option (A and B), or only the first-listed option and the third-listed option (A and C), or only the second-listed option and the third-listed option (B and C), or the selection of all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to as many listed items as desired.

[0186] In addition, as used herein, among other things, the word "signal" also refers to indicating something to a corresponding decoder. Thus, in an embodiment, the same parameters are used at both the encoder side and the decoder side. So, for example, an encoder can transmit (explicitly signal) specific parameters to a decoder such that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling can be used without transmission (implicitly signal) to allow only the decoder to know and select the specific parameters. By avoiding transmission of any actual functionality, bit savings are achieved in various embodiments. It is to be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, information is signaled to a corresponding decoder using one or more syntax elements, flags, and the like. Although the foregoing related to the verb form of the word "signal", the word "signal" can also be used as a noun herein.

[0187] As will be apparent to those of ordinary skill in the art, implementations can generate a variety of signals that are formatted to carry information that can, for example, be stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry the bitstream of the described embodiment. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. Signals can be stored on a processor-readable medium.

[0188] Multiple embodiments have been described above. The features of these embodiments can be provided individually or in any combination across various claim categories and types.

Claims

1. A method, comprising: determining at least one reference sample of at least one first block of an image based on an encoding mode of at least one second block for reconstructing the image, the at least one reference sample being outside the at least one second block, obtaining a prediction for the at least one first block using the at least one reference sample, decoding the at least one first block based on the prediction.

2. An apparatus comprising one or more processors, wherein the one or more processors are operable to: determine at least one reference sample of at least one first block of an image based on an encoding mode of at least one second block for reconstructing the image, the at least one reference sample being outside the at least one second block, obtain a prediction for the at least one first block using the at least one reference sample, decode the at least one first block based on the prediction.

3. A method, comprising: determining at least one reference sample of at least one first block of an image based on an encoding mode of at least one second block for reconstructing the image, the at least one reference sample being outside the at least one second block, obtaining a prediction for the at least one first block using the at least one reference sample, encoding the at least one first block based on the prediction.

4. An apparatus comprising one or more processors, wherein the one or more processors are operable to: determine at least one reference sample of at least one first block of an image based on an encoding mode of at least one second block for reconstructing the image, the at least one reference sample being outside the at least one second block, obtain a prediction for the at least one first block using the at least one reference sample, encode the at least one first block based on the prediction.

5. The method according to any one of claims 1 or 3 or the apparatus according to any one of claims 2 or 4, wherein, the at least one reference sample belongs to a non-reconstructed block or belongs to a block having an encoding mode that does not allow using the at least one reference sample to determine a prediction for the at least one first block.

6. The method according to any one of claims 1, 3 or 5 or the apparatus according to any one of claims 2 or 4 - 5, wherein, the at least one reference sample belongs to a block adjacent to the at least one second block.

7. The method according to any one of claims 1, 3 or 5 - 6 or the apparatus according to any one of claims 2 or 4 - 6, wherein, the encoding mode is an intra prediction mode.

8. The method or apparatus according to claim 7, wherein, determining at least one reference sample based on an encoding mode of at least one second block for reconstructing the image includes filling the at least one reference sample with data obtained using the same intra prediction mode as the at least one second block.

9. The method or apparatus according to claim 7, wherein, determining at least one reference sample of at least one first block of an image based on an encoding mode of at least one second block for reconstructing the image includes: In response to determining that the direction of the intra prediction mode is among a first set of intra prediction directions, fill at least one reference sample of at least one first block with data obtained using the same intra prediction mode as at least one second block. Otherwise, use repeated filling to fill at least one reference sample of at least one first block.

10. The method according to any one of claims 1, 3, or 5 - 6 or the apparatus according to any one of claims 2 or 4 - 6, wherein the coding mode is an inter prediction mode.

11. The method or apparatus according to claim 10, wherein, Determining at least one reference sample based on the coding mode of at least one second block for reconstructing an image includes filling at least one reference sample with motion compensation data obtained using the motion information of at least one second block.

12. The method or apparatus according to any one of claims 7, 8, or 11, wherein, Determining at least one reference sample based on the coding mode of at least one second block for reconstructing an image includes adding an offset to the filled at least one reference sample, the offset being determined according to at least one sample of at least one second block and at least one sample of the filled at least one reference sample.

13. The method according to any one of claims 1, 3, or 5 - 12 further includes or the apparatus according to any one of claims 2 or 4 - 12, wherein one or more processors are further configured to: select at least one second block among a plurality of blocks based on at least one selection rule, the at least one selection rule being at least based on one of the following: the spatial distance from at least one reference sample to at least one second block, or the coding mode of at least one second block, or the position of at least one second block.

14. The method according to any one of claims 1, 3, or 5 - 13 or the apparatus according to any one of claims 2 or 4 - 13, wherein, Obtain a prediction for at least one first block using intra prediction.

15. The method according to any one of claims 1, 3, or 5 - 13 or the apparatus according to any one of claims 2 or 4 - 13, wherein, Obtain a prediction for at least one first block using inter prediction.

16. The method or apparatus according to claim 15, wherein, At least one reference sample is used to determine a correction parameter used in obtaining the prediction.

17. The method or apparatus according to any one of claims 7 - 12, wherein, At least one reference sample belongs to a block located on the right side of the first block or a block located at the bottom of the first block.

18. The method or apparatus according to claim 17, wherein, At least one second block is located in the upper right or lower left of the first block.

19. The method or apparatus according to claim 18, wherein, Perform intra prediction on at least one first block using a first prediction direction, and: If at least one reference sample belongs to a block located on the right side of the first block, project the reference samples on the first column of the block located on the right side of the first block to the last row of the block located in the upper right of the first block in the reverse direction using the first prediction direction. If at least one reference sample belongs to a block located at the bottom of the first block, the reference samples on the first row of the block located at the bottom of the first block are back-projected onto the last column of the block located at the lower left of the first block using the first prediction direction.

20. The method or apparatus according to claim 18, performing intra prediction on at least one first block using the first prediction direction, determining at least one reference sample that belongs to a block located on the right side of the first block or belongs to a block located at the bottom of the first block, and in response to the first intra prediction direction, exchanging the reference samples in the blocks located at the right side and at the bottom of the first block with the reference samples in the blocks located at the left side and above the first block, respectively.

21. The method or apparatus according to claim 20, wherein, in response to the first intra prediction direction, flipping at least a portion of the exchanged reference samples.

22. The method or apparatus according to claim 18, performing intra prediction on at least one first block using the first prediction direction, determining at least one reference sample that belongs to a block located on the right side of the first block or belongs to a block located at the bottom of the first block, and in response to the first intra prediction direction, performing at least one of the following operations: filling the reference samples in at least one column of the block located on the left side of the first block with the determined reference samples in at least one corresponding column of the block located on the right side of the first block, flipping the reference samples in at least one row of the block located above the first block, and filling the reference samples in at least one row of the block located at the upper right of the first block with the flipped reference samples from at least one corresponding row of the block located at the upper left of the first block, filling the reference samples in at least one row of the block located above the first block with the determined reference samples in at least one corresponding row of the block located at the bottom of the first block, flipping the reference samples in at least one column of the block located on the left side of the first block, and filling the reference samples in at least one column of the block located at the lower left of the first block with the flipped reference samples from at least one corresponding column of the block located at the upper left of the first block, filling the reference samples in at least one row of the block located above the first block and the reference samples in at least one corresponding row of the block located at the upper right of the first block with the flipped determined reference samples in at least one corresponding row of the block located at the bottom of the first block and the flipped reference samples in at least one corresponding row of the block located at the lower left of the first block, and filling the reference samples in at least one column of the block located on the left side of the first block and the reference samples in at least one column of the block located at the lower left of the first block with the flipped determined reference samples in at least one corresponding column of the block located on the right side of the first block and the flipped reference samples in at least one corresponding column of the block located at the upper right of the first block.

23. The method or apparatus according to any one of claims 21 or 22, wherein, the first prediction direction has a first angle, and a prediction is obtained using an intra prediction direction mode having a second angle corresponding to the first angle minus 90°.

24. The method or apparatus according to claim 23, wherein, Predictions obtained by flipping horizontally or vertically based on a first angle.

25. The method according to any one of claims 1, 3 or 5 - 24 or the apparatus according to any one of claims 2 or 4 - 24, wherein, determining at least one reference sample of at least one first block in response to encoding or decoding of an indicator based on an encoding pattern for reconstructing at least one second block.

26. The method or apparatus according to claim 25, wherein, encoding or decoding of the indicator is based on an encoding pattern of at least one second block.

27. The method or apparatus according to claim 25 and any one of claims 19 - 24, wherein the indicator signals at least one of the following information: whether determining at least one reference sample of at least one first block is based on an encoding pattern for reconstructing at least one second block, or whether to exchange reference samples of a given block adjacent to the first block.

28. A computer program product comprising instructions for causing one or more processors to execute the method according to any one of claims 1, 3, 5 - 27.

29. A non - transitory computer - readable medium storing executable program instructions, the executable program instructions causing a computer executing the instructions to perform the method according to any one of claims 1, 3, 5 - 27.

30. A bitstream comprising data representing an image or video encoded using the method according to any one of claims 1, 3, 5 - 27.

31. A non - transitory computer - readable medium storing the bitstream according to claim 30.

32. An apparatus, comprising: - the apparatus according to any one of claims 2 or 4; and - at least one of the following: (i) an antenna configured to receive a signal, the signal comprising data representing an image or video; (ii) a band limiter configured to limit the signal to a band including data representing an image or video; or (iii) a display configured to display an image or video.

33. The apparatus according to claim 32, wherein, the apparatus comprises at least one of a television, a cellular phone, a tablet computer, a set - top box.