Method, apparatus, and program for video coding

PDPC enhances video coding efficiency by combining intra and inter prediction techniques, reducing bandwidth and storage needs through improved sample value minimization in video encoding/decoding.

JP7704813B2Active Publication Date: 2025-07-08TENCENT AMERICA LLC

Patent Information

Application Number
JP2023134843
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-30
Filing Date
2023-08-22
Publication Date
2025-07-08
Estimated Expiration
2040-01-31

AI Technical Summary

Technical Problem

Conventional video coding techniques, such as those in MPEG-2, H.264, H.265, and JEM, do not effectively utilize intra prediction methods to minimize sample values, leading to inefficiencies in entropy coding and increased bandwidth requirements for video transmission and storage.

Method used

The implementation of position-dependent prediction combination (PDPC) for inter and intra prediction modes in video encoding/decoding, which combines intra and inter prediction techniques to reconstruct video blocks, using weighted neighboring samples based on their positions, and applies filters to enhance prediction accuracy.

Benefits of technology

This approach significantly reduces the bit rate and improves coding efficiency by minimizing sample values, thereby reducing bandwidth and storage requirements while maintaining video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704813000011
    Figure 0007704813000011
  • Figure 0007704813000012
    Figure 0007704813000012
  • Figure 0007704813000013
    Figure 0007704813000013
Patent Text Reader

Abstract

To provide a method and an apparatus for video coding.SOLUTION: There is provided a method and an apparatus for video encoding / decoding according to an embodiment of the present disclosure. In some examples, an apparatus for video decoding includes a processing circuit. For example, the processing circuit decodes prediction information for the current block from an encoded video bitstream. The prediction information indicates that the prediction of the current block is at least partially based on inter prediction. The processing circuit then reconstructs at least the samples of the current block as a combination of the results from the inter prediction and neighboring samples of the block selected on the basis of the position of the samples.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of priority based on U.S. Provisional Application No. 62 / 800,400, filed on February 1, 2019, entitled "ENHANCEMENT FOR POSITION DEPENDENT PREDICTION COMBINATION", and claims the benefit of priority based on U.S. Patent Application No. 16 / 777,339, filed on January 30, 2020, entitled "METHOD AND APPARATUS FOR VIDEO CODING". The entire disclosure of the prior application is incorporated herein by reference in its entirety.

[0002] This disclosure describes embodiments generally related to video coding.

Background Art

[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The research of the inventors, to the extent described in this background art section, and aspects of the description that may not be regarded as prior art at the time of filing, are not admitted as prior art to the present disclosure, either explicitly or implicitly.

[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having spatial dimensions, for example, of luminance samples of 1920×1080 and associated chrominance samples. A series of pictures can have a fixed or variable picture rate, for example, 60 pictures per second or 60 Hz (informally also called the frame rate). Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. To use such video for one hour, a storage area of more than 600 GB is required.

[0005] One purpose of video encoding and decoding can be to reduce the redundancy of the input video signal by compression. Compression can help reduce the aforementioned bandwidth or memory requirements, sometimes by more than two orders of magnitude. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression refers to a technique in which an exact replica of the original signal can be reconstructed from the compressed original signal. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal is useful for the intended application. In the case of video, irreversible compression is widely adopted. The amount of distortion that is acceptable varies depending on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio may reflect that higher compression ratios can be obtained with higher tolerance / acceptable distortion.

[0006] Video encoders and decoders can utilize several broad categories of techniques, such as motion compensation, transformation, quantization, entropy encoding, and the like.

[0007] Video coding techniques can include techniques known as intra coding. In intra coding, sample values are represented without reference to other data from samples or from previously reconstructed reference pictures. In some video coding, pictures are spatially subdivided into blocks of samples. If all blocks of samples are coded in an intra mode, that picture can be an intra picture. Those derivations such as intra pictures and independent decoder refresh pictures can be used to reset the decoder state and thus can be used as the first picture in an encoded video bitstream and video session or as a still image. Samples of an intra block may be subject to transformation and the transform coefficients may be quantized prior to entropy coding. Intra prediction can be a technique that minimizes sample values in a pre-transform region. In some cases, the smaller the post-transform DC value and the smaller the AC coefficients, the fewer bits are required with a given quantization step size to represent the block after entropy coding.

[0008] Conventional intra coding, such as known from MPEG-2 production coding techniques, does not use intra prediction. However, some newer video compression techniques include techniques that attempt, for example, from surrounding sample data and / or metadata obtained during the coding / decoding of blocks of data that are spatially adjacent and precede in decoding order. Such techniques are hereinafter referred to as "intra prediction" techniques. Note that in at least some cases, intra prediction uses only reference data from the currently reconstructed picture and does not use reference data from reference pictures.

[0009] Intra prediction can take many different forms. If two or more of such techniques can be used in a given video coding technology, the technique in use can be coded in an intra prediction mode. In some cases, the mode can have sub - modes and / or parameters, which can be coded individually or can be included in the mode codeword. Which codeword to use for a given combination of mode / sub - mode / parameter can affect the coding efficiency gain through intra prediction and thus can also affect the entropy coding technique used to convert the codeword into the bitstream.

[0010] Certain modes of intra prediction were introduced in H.264, improved in H.265, and further improved in new coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). The predictor block can be formed using neighboring sample values belonging to already available samples. The sample values of the neighboring samples are copied into the predictor block according to a direction. The reference to the direction in use can be coded within the bitstream or can itself be predicted. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0011] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit. For example, the processing circuit decodes prediction information for a current block from an encoded video bitstream. The prediction information indicates that the prediction of the current block is based at least in part on inter prediction. Next, the processing circuit reconstructs at least the samples of the current block as a combination of the result from inter prediction and neighboring samples of a block selected based on the position of the samples.

[0012] In one embodiment, the prediction information indicates an intra-inter prediction mode that uses a combination of intra prediction and inter prediction of the current block, and the processing circuit excludes motion refinement based on bidirectional optical flow from the inter prediction of the current block.

[0013] In some embodiments, the processing circuit determines the use of position-dependent prediction combination (PDPC) for inter prediction based on the prediction information, and reconstructs samples according to the determination of the use of PDPC for inter prediction. In one example, the processing circuit receives a flag indicating the use of PDPC.

[0014] In some examples, the processing circuit decodes a first flag for an intra-inter prediction mode that uses a combination of intra prediction and inter prediction of the current block, and decodes a second flag indicating the use of PDPC if the first flag indicates true for the intra-inter prediction mode. In one example, the processing circuit decodes a second flag indicating the use of PDPC based on the context model of entropy coding.

[0015] In one embodiment, the processing circuit applies a filter to the neighboring samples of the current block and reconstructs the samples of the current block as a combination of the result from inter prediction and the filtered neighboring samples of the current block.

[0016] In other embodiments, the processing circuit reconstructs the samples of the current block using the neighboring samples of the current block without applying a filter to the neighboring samples.

[0017] In some examples, the processing circuit determines whether the current block meets a block size condition that restricts the application of PDPC in the reconstruction of the current block, and excludes the application of PDPC in the reconstruction of at least one sample of the current block if the block size condition is met. In one example, the block size condition depends on the size of the virtual processing data unit.

[0018] In some examples, the processing circuit ignores neighboring samples in the calculation of the combination if the distance of the neighboring samples to the sample is greater than a threshold value.

[0019] In some embodiments, the processing circuit combines the result from inter prediction with neighboring samples selected based on the position of the sample and reconstructed based on inter prediction.

[0020] In one example, when the prediction information indicates the merge mode, the processing circuit decodes a flag indicating the use of PDPC.

[0021] In another example, when the first flag indicates a non-zero residual of the current block, the processing circuit decodes a second flag indicating the use of PDPC.

[0022] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to execute a method for video decoding.

[0023] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Brief Description of the Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 10

Figure 11

Figure 12

Figure 13

Mode for Carrying Out the Invention

[0025] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The communication system (100) includes a plurality of terminal devices that can communicate with each other via, for example, a network (150). For example, the communication system (100) includes a first pair (110) and (120) of terminal devices interconnected via a network (150). In the example of FIG. 1, the first pair (110) and (120) of terminal devices perform unidirectional transmission of data. For example, the terminal device (110) may encode video data (e.g., a stream of video pictures captured by the terminal device (110)) for transmission to another terminal device (120) via the network (150). The encoded video data may be transmitted in the form of one or more encoded video bitstreams. The terminal device (120) may receive the encoded video data from the network (150), decode the encoded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission may be common in media serving applications and the like.

[0026] In other examples, the communication system (100) includes a second pair (130) and (140) of terminal devices that perform bidirectional transmission of encoded video data that may occur, for example, during a video conference. For bidirectional data transmission, in one example, each of the terminal devices (130) and (140) may encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (130) and (140) via the network (150). Each of the terminal devices (130) and (140) may also receive the encoded video data transmitted by the other of the terminal devices (130) and (140), may decode the encoded video data to restore the video pictures, and may display the video pictures on a display device accessible according to the restored video data.

[0027] In the example of FIG. 1, the terminal devices (110), (120), (130), and (140) can be shown as servers, personal computers, and smartphones, but the principles of the present disclosure are not so limited. Embodiments of the present disclosure find use in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (150) represents any number of networks that transfer encoded video data among the terminal devices (110), (120), (130), and (140), including, for example, wired (wired) and / or wireless communication networks. The communication network (150) can exchange data over circuit-switched channels and / or packet-switched channels. Representative networks include communication networks, local area networks, wide area networks, and / or the Internet, among others. For the purposes of this discussion, the architecture and topology of the network (150) may not be important for the operation of the present disclosure, unless otherwise described herein below.

[0028] FIG. 2 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-related applications, including, for example, video conferencing, storage of compressed video on digital media, including digital TV, CD, DVD, memory sticks, and the like.

[0029] A streaming system may include a capture subsystem (213) that can include a video source (201), such as a digital camera, for example, that creates a stream (202) of, for example, uncompressed video pictures. In one example, the stream (202) of video pictures includes samples taken by a digital camera. The stream (202) of video pictures, shown as a thick line to emphasize the high data volume compared to the encoded video data (204) (or encoded video bitstream), can be processed by an electronic device (220) that includes a video encoder (203) coupled to the video source (201). The video encoder (203) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video data (204) (or encoded video bitstream (204)), shown as a thin line to emphasize the lower data volume compared to the stream (202) of video pictures, can be stored in a streaming server (205) for future use. One or more streaming client subsystems, such as the client subsystems (206) and (208) of FIG. 2, can access the streaming server (205) to obtain copies (207) and (209) of the encoded video data (204). The client subsystem (206) can include, for example, a video decoder (210) within an electronic device (230). The video decoder (210) decodes an input copy (207) of the encoded video data and creates an output stream (211) of video pictures that can be rendered on a display (212) (e.g., display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (204), (207), and (209) (e.g., video bitstream) can be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265.In one example, the video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0030] It should be noted that the electronic devices (220) and (230) can include other components (not shown). For example, the electronic device (220) can include a video decoder (not shown), and the electronic device (230) can also include a video encoder (not shown).

[0031] FIG. 3 shows a block diagram of a video decoder (310) according to an embodiment of the present disclosure. The video decoder (310) can be included in an electronic device (330). The electronic device (330) can include a receiver (331) (e.g., a receiving circuit). The video decoder (310) can be used instead of the video decoder (210) in the example of FIG. 2.

[0032] The receiver (331) can receive one or more encoded video sequences decoded by the video decoder (310), and in the same or other embodiments, can receive one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence can be received from a channel (301) that can be a hardware / software link to a storage device storing the encoded video data. The receiver (331) can receive the encoded video data along with other data, such as encoded audio data and / or an auxiliary data stream, that can be transferred to respective using entities (not shown). The receiver (331) can separate the encoded video sequence from other data. To counter network jitter, a buffer memory (315) can be coupled between the receiver (331) and the entropy decoder / parser (320) (hereinafter, "parser (320)"). In certain applications, the buffer memory (315) is part of the video decoder (310). In other cases, it may be external to the video decoder (310) (not shown). In still other cases, for example, there is a buffer memory (not shown) external to the video decoder (310) to counter network jitter, and further, for example, there can be another buffer memory (315) inside the video decoder (310) to handle playback timing. If the receiver (331) is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (315) may not be necessary or may be small. For use in a best-effort packet network such as the Internet, the buffer memory (315) may be required, may be relatively large, advantageously may be of an adaptable size, and may be at least partially implemented in an operating system or similar element (not shown) external to the video decoder (310).

[0033] The video decoder (310) may include a parser (320) for reconstructing symbols (321) from an encoded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (310), and potentially, as shown in FIG. 3, information for controlling a rendering device (312) (e.g., a display screen) that is not an essential part of the electronic device (330) but can be coupled to the electronic device (330). The control information for the rendering device can be in the form of supplementary enhancement information (SEI message) or a video user utility information (VUI) parameter set fragment (not shown). The parser (320) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can follow video coding techniques or video coding standards and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context dependence, etc. The parser (320) can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to a group. The subgroups can include Group of Pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (320) can also extract from the encoded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.

[0034] The parser (320) can perform an entropy decode / parse operation on the video sequence received from the buffer memory (315) to create symbols (321).

[0035] The reconstruction of the symbol (321) may include multiple different units depending on the type of the encoded video picture or a part thereof (such as inter-picture and intra-picture, inter-block and intra-block), and other factors. The units included and the method thereof may be controlled by subgroup control information analyzed from the video sequence encoded by the parser (320). Such a flow of subgroup control information between the parser (320) and the following multiple units is not shown for clarity.

[0036] In addition to the function blocks already described, the video decoder (310) may be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0037] The first unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives, as symbols (321), the quantized transform coefficients and control information including the transform to be used, block size, quantization coefficients, quantization scaling matrix, etc. from the parser (320). The scaler / inverse transform unit (351) may output a block including sample values that can be input to the aggregator (355).

[0038] In some cases, the output samples of the scaler / inverse transform (351) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from previously reconstructed parts of the current picture. Such prediction information may be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses the surrounding already reconstructed information fetched from the current picture buffer (358) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (358) buffers, for example, the partially reconstructed current picture and / or the fully reconstructed current picture. The aggregator (355) may, in some cases, add, sample by sample, the prediction information generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).

[0039] In other cases, the output samples of the scaler / inverse transform unit (351) may relate to inter-coded, potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (353) can access the reference picture memory (357) to fetch the samples used for prediction. After motion compensating the fetched samples according to the symbols (321) related to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) to generate the output sample information (in this case, called residual samples or residual signal). The address in the reference picture memory (357) from which the motion compensation prediction unit (353) fetches the prediction samples can be controlled, for example, by the motion vectors available to the motion compensation prediction unit (353) in the form of symbols (321) having X, Y, and reference picture components. Motion compensation may also include interpolation of the sample values fetched from the reference picture memory (357) when exact sub-sample motion vectors are used, a motion vector prediction mechanism, etc.

[0040] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356). Video compression techniques can be controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream), and can include in-loop filter techniques that are available in the loop filter unit (356) as symbols (321) from the parser (320), but can also be responsive to meta information obtained during the decoding of an encoded picture or a previous (in decoding order) portion of the encoded video sequence, or responsive to previously reconstructed and loop-filtered sample values.

[0041] The output of the loop filter unit (356) can be not only output to the rendering device (312), but can also be a sample stream that can be stored in the reference picture memory (357) for use in future inter-picture prediction.

[0042] Once an encoded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, when the encoded picture corresponding to the current picture is fully reconstructed and identified as a reference picture (e.g., by the parser (320)), the current picture buffer (358) can become part of the reference picture memory (357), and a fresh current picture buffer can be reallocated before starting the reconstruction of the next encoded picture.

[0043] The video decoder (310) can perform a decoding operation according to a predetermined video compression technique in a standard such as ITU-T Rec.H.265. The encoded video sequence may comply with the syntax specified by the video compression technique or standard being used, in the sense that the encoded video sequence complies with both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile can select specific tools from all the tools available in the video compression technique or standard as the only tools available under that profile. Also required for compliance is that the complexity of the encoded video sequence be within the range defined at the level of the video compression technique or standard. In some cases, the level may limit the maximum picture size, maximum frame rate, maximum reconstructed sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limitations set by the level may in some cases be further restricted by the specifications of the Hypothetical Reference Decoder (HRD) and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0044] In one embodiment, the receiver (331) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (310) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) extension layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0045] FIG. 4 shows a block diagram of a video encoder (403) according to an embodiment of the present disclosure. The video encoder (403) is included in an electronic device (420). The electronic device (420) includes a transmitter (440) (e.g., a transmission circuit). The video encoder (403) can be used instead of the video encoder (203) in the example of FIG. 2.

[0046] The video encoder (403) can receive video samples from a video source (401) (not part of the electronic device (420) in the example of FIG. 4) that can capture the video images to be encoded by the video encoder (403). In other examples, the video source (401) is part of the electronic device (420).

[0047] The video source (401) can provide the source video sequence to be encoded by the video encoder (403) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), can be in any color space (e.g., BT.601 Y CrCB, RGB, …), and can have a suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (401) can be a storage device that stores previously prepared video. In a video conferencing system, the video source (401) can be a camera that captures local picture information as a video sequence. The video data can be provided as a plurality of individual pictures that give motion when viewed in sequence. The pictures themselves can be organized as a spatial array of pixels, and each pixel can contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. In the following description, the samples will be mainly described.

[0048] According to one embodiment, the video encoder (403) can encode pictures of a source video sequence in real time or under any other time constraint as required by the application and compress them into an encoded video sequence (443). Enforcing an appropriate encoding speed is one function of the controller (450). In some embodiments, the controller (450) controls other functional units and is functionally coupled to other functional units as described below. For clarity, the couplings are not depicted. Parameters set by the controller (450) may include rate control related parameters (such as picture skip, quantization, lambda value of the rate distortion optimization method), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. The controller (450) can be configured to have other appropriate functions regarding the video encoder (403) optimized for a certain system design.

[0049] In some embodiments, the video coder (403) is configured to operate in an encoding loop. As an overly simplified explanation, in one example, the encoding loop can include a source coder (430) (e.g., responsible for generating symbols such as a symbol stream based on an input picture to be encoded and reference pictures), and a (local) decoder (433) incorporated into the video coder (403). The decoder (433) reconstructs symbols to create sample data in the same way as a (remote) decoder also creates it (in the video compression techniques contemplated by the disclosed subject matter, any compression between the symbols and the encoded video bitstream is reversible). The reconstructed sample stream (sample data) is input into the reference picture memory (434). Since the decoding of the symbol stream yields bit - accurate results regardless of the location of the decoder (local or remote), the content in the reference picture memory (434) is also bit - accurate between the local coder and the remote coder. In other words, the prediction part of the coder "sees" the same sample values as samples of the reference picture that the decoder "refers" to when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related technologies.

[0050] The operation of the "local" decoder (433) can be the same as that of a "remote" decoder (310) such as a video decoder, which has already been described in detail above in relation to FIG. 3. However, referring briefly to FIG. 3 as well, since symbols are available and the encoding / decoding of symbols into the encoded video sequence by the entropy coder (445) and the parser (320) can be lossless, the entropy decoding part of the video decoder (310) including the buffer memory (315) and the parser (320) may not be fully implemented in the local decoder (433).

[0051] The observations that can be made at this point are that decoder technologies other than syntax analysis / entropy decoding existing in the decoder must necessarily exist in a substantially identical functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on the operation of the decoder. The description of encoder technology can be omitted because it is the reverse of the decoder technology described comprehensively. More detailed descriptions are necessary only in specific areas and are provided below.

[0052] In some examples, during operation, the source coder (430) may perform motion-compensated predictive coding that predictively encodes an input picture by referring to one or more previously coded pictures from a video sequence designated as a "reference picture". In this way, the encoding engine (432) encodes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a predictive reference to the input picture.

[0053] The local video decoder (433) may decode the encoded video data of an image that can be designated as a reference picture based on the symbols created by the source coder (430). The operation of the encoding engine (432) may advantageously be an irreversible process. When the encoded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence can typically be a replica of the source video sequence with some errors. The local video decoder (433) can replicate the decoding process that can be performed by the video decoder for the reference picture and store the reconstructed reference picture in the reference picture cache (434). In this way, the video coder (403) can locally store a replica of the reconstructed reference picture having common content as the reconstructed reference picture obtained by the remote video decoder (without transmission errors).

[0054] The predictor (435) can perform a prediction search of the encoding engine (432). That is, for a new picture to be encoded, the predictor (435) can search the reference picture memory (434) for specific metadata that functions as an appropriate prediction reference for the new picture, such as sample data (as candidate reference pixel blocks) or motion vectors and block shapes of reference pictures. The predictor (435) can operate on a sample block - pixel block basis to find an appropriate prediction reference. In some cases, the input picture can have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (434), as determined by the search results obtained by the predictor (435).

[0055] The controller (450) can manage the encoding operation of the source coder (430), including, for example, setting parameters and subgroup parameters used for encoding video data.

[0056] The outputs of all the aforementioned functional units can undergo entropy encoding in the entropy coder (445). The entropy coder (445) converts the symbols generated by various functional units into an encoded video sequence by reversibly compressing the symbols according to techniques such as Huffman coding, variable - length coding, arithmetic coding, etc.

[0057] The transmitter (440) can buffer the encoded video sequence created by the entropy coder (445) and prepare for transmission via a communication channel (460) that can be a hardware / software link to a storage device for storing the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0058] The controller (450) may manage the operation of the video encoder (403). During encoding, the controller (450) may assign a specific encoded picture type to each encoded picture, which may affect the encoding technique that can be applied to each picture. For example, often, a picture may be assigned as one of the following picture types.

[0059] An intra picture (I picture) can be encoded and decoded without using other pictures in the sequence as a source of prediction. In some video encodings, various types of intra pictures can be used, such as, for example, an Independent Decoder Refresh (「IDR」) picture. Those skilled in the art know those variations of I pictures and their respective uses and characteristics.

[0060] A predicted picture (P picture) can be encoded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0061] A bi - directionally predicted picture (B picture) can be encoded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures may use metadata associated with more than two reference pictures for the reconstruction of a single block.

[0062] The source picture is typically spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be encoded block by block. The blocks can be encoded predictively by referring to other (already encoded) blocks as determined by the encoding assignment applied to each picture of the block. For example, blocks of an I picture may be encoded non-predictively, or they may be encoded predictively by referring to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be encoded predictively via spatial prediction or via temporal prediction by referring to one previously encoded reference picture. Blocks of a B picture can be encoded predictively by referring to one or two previously encoded reference pictures via spatial prediction or via temporal prediction.

[0063] The video encoder (403) can perform an encoding operation according to a predetermined video encoding technique or standard such as ITU-T Rec.H.265. In that operation, the video encoder (403) can perform various compression operations including predictive encoding operations that exploit the temporal and spatial redundancy of the input video sequence. Thus, the encoded video data may conform to the syntax specified in the video encoding technique or standard being used.

[0064] In one embodiment, the transmitter (440) can transmit additional data along with the encoded video. The source coder (430) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.

[0065] Videos may be captured in time series as a plurality of source pictures (video pictures). Intra prediction (often abbreviated as intra prediction) utilizes the spatial correlation within a given picture, while inter picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded, referred to as the current picture, is divided into blocks. When a block within the current picture is similar to a reference block within a reference picture that has been previously encoded and buffered in the video, the block within the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and may have a third dimension identifying the reference picture when multiple reference pictures are being used.

[0066] In some embodiments, dual prediction techniques may be used for inter picture prediction. According to the dual prediction technique, two reference pictures such as a first reference picture and a second reference picture are used, both of which are prior to the decoding order of the current picture in the video (however, the display order may be past and future respectively). A block within the current picture can be encoded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0067] Furthermore, to improve the encoding efficiency, merge mode techniques can be used for inter picture prediction.

[0068] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs) which are one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to the temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels.

[0069] FIG. 5 shows a diagram of a video encoder (503) according to another embodiment of the present disclosure. The video encoder (503) receives a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and is configured to encode the processing block into an encoded picture that is part of an encoded video sequence. In one example, the video encoder (503) is used in place of the video encoder (203) of the example of FIG. 2.

[0070] In an example of HEVC, the video encoder (503) receives a matrix of sample values for a processing block, such as a prediction block of 8×8 samples. The video encoder (503) determines whether the processing block is best encoded using an intra mode, an inter mode, or a bi-prediction mode, for example, using rate-distortion optimization. If the processing block is encoded in the intra mode, the video encoder (503) may use intra prediction techniques to encode the processing block into the encoded picture. When the processing block is to be encoded in the inter mode or the bi-prediction mode, the video encoder (503) can use inter prediction techniques or bi-prediction techniques, respectively, to encode the processing block into the encoded picture. In a video encoding technique, the merge mode can be an inter-picture prediction sub-mode in which the motion vector is derived from one or more motion vector predictors without benefiting from the encoded motion vector components outside the predictor. In some other video encoding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (503) includes other components such as a mode decision module (not shown) for determining the mode of the processing block.

[0071] In the example of FIG. 5, the video encoder (503) includes an inter-encoder (530), an intra-encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled to each other as shown in FIG. 5.

[0072] The inter-coder (530) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks within a reference picture (e.g., blocks within a previous picture and a subsequent picture), generate inter-prediction information (e.g., a description of redundant information by inter-coding techniques, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.

[0073] The intra-coder (522) is configured to receive samples of a current block (e.g., a processing block), optionally compare the block with blocks already encoded within the same picture, generate quantized coefficients after transformation, and optionally also generate intra-prediction information (e.g., intra-prediction direction information by one or more intra-coding techniques). In one example, the intra-coder (522) calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks within the same picture.

[0074] The general controller (521) is configured to determine general control data and control other components of the video coder (503) based on the general control data. In one example, the general controller (521) determines the mode of a block and provides a control signal to a switch (526) based on the mode. For example, when the mode is the intra mode, the general controller (521) controls the switch (526) to select the intra-mode result used by the residual calculator (523), and controls the entropy coder (525) to select the intra-prediction information and include it in the bitstream. When the mode is the inter mode, the general controller (521) controls the switch (526) to select the inter-prediction result used by the residual calculator (523), and controls the entropy coder (525) to select the inter-prediction information and include it in the bitstream.

[0075] The residual calculator (523) calculates the difference (residual data) between the received block and the prediction result selected from the intra-coder (522) or the inter-coder (530). The residual coder (524) is configured to operate based on the residual data to encode the residual data in order to generate transform coefficients. In one example, the residual coder (524) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video coder (503) also includes a residual decoder (528). The residual decoder (528) is configured to perform inverse transformation and generate decoded residual data. The decoded residual data can be suitably used in the intra-coder (522) and the inter-coder (530). For example, the inter-coder (530) can generate a decoded block based on the decoded residual data and the inter-prediction information, and the intra-coder (522) can generate a decoded block based on the decoded residual data and the intra-prediction information. In some examples, the decoded block is appropriately processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.

[0076] The entropy coder (525) is configured to format the bitstream to include the encoded block. The entropy coder (525) is configured to include various information according to an appropriate standard such as the HEVC standard. In one example, the entropy coder (525) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Note that according to the disclosed subject matter, there is no residual information when encoding a block in either the merge sub-mode of the inter-mode or the bi-prediction mode.

[0077] FIG. 6 shows a diagram of a video decoder (610) according to another embodiment of the present disclosure. The video decoder (610) is configured to receive an encoded picture that is part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In one example, the video decoder (610) is used in place of the video decoder (210) of the example of FIG. 2.

[0078] In the example of FIG. 6, the video decoder (610) includes an entropy decoder (671), an inter decoder (680), a residual decoder (673), a reconstruction module (674), and an intra decoder (672) coupled to each other as shown in FIG. 6.

[0079] The entropy decoder (671) may be configured to reconstruct from the encoded picture specific symbols representing syntax elements that make up the encoded picture. Such symbols can include, for example, the mode in which a block is encoded (e.g., intra mode, inter mode, bi-prediction mode, the latter two being merge sub-modes or another sub-mode), prediction information (e.g., intra prediction information, inter prediction information, etc.) that can identify specific samples or metadata used for prediction by the intra decoder (672) or the inter decoder (680) respectively, and residual information in the form of, for example, quantized transform coefficients. In one example, when the prediction mode is an inter prediction mode or a bi-prediction mode, the inter prediction information is provided to the inter decoder (680). When the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (672). The residual information can undergo inverse quantization and is provided to the residual decoder (673).

[0080] The inter decoder (680) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0081] The intra decoder (672) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0082] The residual decoder (673) is configured to perform inverse quantization to extract the inverse quantized transform coefficients, and process the inverse quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (673) may also require certain control information (since it includes quantization parameter (QP)), and that information may be provided by the entropy decoder (671) (the data paths not shown as such may be only for low volume control information).

[0083] The reconstruction module (674) is configured to combine, in the spatial domain, the residual as the output by the residual decoder (673) and the prediction result (optionally as the output by the inter or intra prediction module) to form a reconstruction block that can be part of the reconstructed picture, and the reconstruction block can be part of the reconstructed video. Note that other appropriate operations such as a deblocking operation can be performed to improve the visual quality.

[0084] Note that the video encoders (203), (403), and (503), as well as the video decoders (210), (310), and (610) can be implemented using any suitable technology. In one embodiment, the video encoders (203), (403), and (503), as well as the video decoders (210), (310), and (610) can be implemented using one or more integrated circuits. In other embodiments, the video encoders (203), (403), and (403), as well as the video decoders (210), (310), and (610) can be implemented using one or more processors that execute software instructions.

[0085] Aspects of the present disclosure provide enhancements for position-dependent prediction combinations. In various embodiments, the present disclosure provides a set of advanced video coding techniques, particularly a scheme extended for intra prediction mode.

[0086] FIG. 7 shows a diagram of exemplary intra prediction directions and intra prediction modes used in HEVC. There are a total of 35 intra prediction modes (modes 0 to 34) in HEVC. Modes 0 and 1 are non - directional modes, mode 0 is the PLANAR mode, and mode 1 is the DC mode. Modes 2 to 34 are directional modes, mode 10 is the horizontal mode, mode 26 is the vertical mode, and modes 2, 18, and 34 are diagonal modes. In some examples, the intra prediction mode is signaled by three most probable modes (MPM) and 32 remaining modes.

[0087] FIG. 8 shows a diagram of exemplary intra prediction directions and intra prediction modes in some examples (e.g., VVC). There are a total of 87 intra prediction modes (modes - 10 to 76), among which mode 18 is the horizontal mode, mode 50 is the vertical mode, and modes 2, 34, and 66 are diagonal modes. Modes - 1 to - 10 and modes 67 to 76 are called wide - angle intra prediction (WAIP) modes.

[0088] In some examples, HEVC - style intra prediction is based on filtered reference samples. For example, if the intra prediction mode is neither the DC mode nor the PLANAR mode, a filter is applied to the boundary reference samples, and the filtered reference samples are used to predict the values within the current block based on the intra prediction mode.

[0089] In some examples, PDPC combines boundary reference samples with HEVC - style intra prediction. In some embodiments, PDPC is applied to the following intra modes without signaling: PLANAR, DC, WAIP modes, horizontal, vertical, lower - left angular mode (mode 2) and its eight adjacent angular modes (modes 3 to 10), and upper - right angular mode (mode 66) and its eight adjacent angular modes (modes 58 to 65).

[0090] In one example, a predicted sample pred’[x][y] located at position (x, y) is predicted according to Equation 1 using a linear combination of an intra prediction mode (DC, PLANAR, angular) and reference samples. pred’[x][y]=(wL×R(-1,y)+wT×R(x,-1)-wTL×R(-1,-1)+(64 wL-wT+wTL)×pred[x][y]+32)>>6 (Equation 1) In the equation, R(x, -1) and R(-1, y) represent the (unfiltered) reference samples located above and to the left of the current sample (x, y), respectively, R(-1, -1) represents the reference sample located at the upper left corner of the current block, and wT, wL, and wTL represent weights. In the case of the DC mode, the weights are calculated by the following equations. In Equations 2 to 5, width represents the width of the current block, and height represents the height of the current block. wT = 32 >> ((y << 1) >> nScale) (Equation 2) wL = 32 >> ((x << 1) >> nScale) (Equation 3) wTL=(wL >> 4)+(wT >> 4) (Equation 4) nScale=(log 2(width)+log 2(height)-2) >> 2 (Equation 5) In the equation, wT represents the weight coefficient of the reference sample located on the upper reference line having the same horizontal coordinate, wL represents the weight coefficient of the reference sample located on the left reference line having the same vertical coordinate, wTL represents the weight coefficient of the upper left reference sample of the current block, nScale specifies how fast the weight coefficient decreases along the axis (wL decreasing from left to right, or wT decreasing from top to bottom), that is, the decrement rate of the weight coefficient, and in the current design, it is the same along the x-axis (from left to right) and the y-axis (from top to bottom). Also, 32 indicates the initial weight coefficient of the neighboring samples, and the initial weight coefficient is also the weight above (left or upper left) assigned to the upper left sample in the current CB, and the weight coefficient of the neighboring samples in the PDPC process must be less than or equal to this initial weight coefficient.

[0091] In the case of the PLANAR mode, wTL = 0, while in the horizontal mode, wTL = wT, and in the vertical mode, wTL = wL. The PDPC weights can be calculated using addition and shift operations. The value of pred’[x][y] can be calculated in one step using Equation 1.

[0092] Figure 9A shows the weights of the prediction samples at (0,0) in the DC mode. In the example of Figure 9A, since the current block is a 4×4 block with a width of 4 and a height of 4, nScale is 0. Then, wT is 32, wL is 32, and -wTL is -4.

[0093] Figure 9B shows the weights of the prediction samples at (1,0) in the DC mode. In the example of Figure 9B, since the current block is a 4×4 block with a width of 4 and a height of 4, nScale is 0. Then, wT is 32, wL is 8, and -wTL is -2.

[0094] When PDPC is applied to the DC, PLANAR, horizontal, and vertical intra modes, no additional boundary filters such as the HEVC DC mode boundary filter or the horizontal / vertical mode edge filter are required. For example, PDPC combines unfiltered boundary reference samples with HEVC-style intra prediction with filtered boundary reference samples.

[0095] More generally, in some examples, the input to the PDPC process is The intra prediction mode represented by predModeIntra; The width of the current block represented by nTbW; The height of the current block represented by nTbH; The width of the reference samples represented by refW; The height of the reference samples represented by refH; The prediction samples by HEVC-style intra prediction represented by predSamples[x][y], where x = 0..nTbW - 1 and y = 0..nTbH - 1; Unfiltered reference (also called neighborhood) samples p[x][y] where x = -1, y = -1..refH-1, and x = 0..refW-1, y = -1; and the color component of the current block represented by cIdx.

[0096] Furthermore, the output of the PDPC process is the modified prediction samples predSamples’[x][y] where x = 0..nTbW-1, y = 0..nTbH-1.

[0097] And the scaling factor nScale is calculated by Equation 6 which is similar to Equation 5. ((Log 2(nTbW)+Log 2(nTbH)-2)>>2) (Equation 6)

[0098] Furthermore, the reference sample array mainRef[x] with x = 0..refW is defined as the array of unfiltered reference samples above the current block, and another reference sample array sideRef[y] with y = 0..refH is defined as the array of unfiltered reference samples on the left side of the current block according to Equations 7 and 8. mainRef[x]=p[x][-1] (Equation 7) sideRef[y]=p[-1][y] (Equation 8)

[0099] For each position (x, y) within the current block, the PDPC calculation uses the upper reference sample shown as refT[x][y], the left reference sample shown as refL[x][y], and the reference sample at the corner p[-1, -1]. In some examples, the modified prediction samples are calculated according to Equation 9, and the result is appropriately clipped according to the cIdx variable indicating the color component. predSamples’[x][y]=(wL×refL(x,y)+wT×refT(x,y)-wTL×p(-1,-1)+(64 wL-wT+wTL)×predSamples[x][y]+32)>>6 (Equation 9)

[0100] The reference samples refT[x][y], refL[x][y], and the weights wL, wT, and wTL can be determined based on the intra prediction mode predModelIntra.

[0101] In one example, when the intra prediction mode predModeIntra is equal to INTRA_PLANAR (e.g., 0, PLANAR mode, mode 0), INTRA_DC (e.g., 1, DC mode, mode 1), INTRA_ANGULAR 18 (e.g., 18, horizontal mode, mode 18 in the case of 67 intra prediction modes), or INTRA_ANGULAR 50 (e.g., 50, vertical mode, mode 50 in the case of 67 intra prediction modes), the reference samples refT[x][y], refL[x][y], and the weights wL, wT, and wTL can be determined according to Equations 10 to 14. refL[x][y] = p[-1][y] (Equation 10) refT[x][y] = p[x][-1] (Equation 11) wT[y] = 32 >> ((y << 1) >> nScale) (Equation 12) wL[x] = 32 >> ((x << 1) >> nScale) (Equation 13) wTL[x][y] = (predModelntra == INTRA_DC)? ((wL[x] >> 4) + (wT[y] >> 4)) : 0 (Equation 14)

[0102] In another example, when the intra prediction mode predModeIntra is equal to INTRA_ANGULAR 2 (e.g., 2, mode 2 in the case of 67 intra prediction modes) or INTRA_ANGULAR 66 (e.g., 66, mode 66 in the case of 66 intra prediction modes), the reference samples refT[x][y], refL[x][y], and the weights wL, wT, and wTL can be determined according to Equations 15 to 19. refL[x][y] = p[-1][x + y + 1] (Equation 15) refT[x][y] = p[x + y + 1][-1] (Equation 16) wT[y] = 32 >> ((y << 1) >> nScale) (Equation 17) wL[x] = 32 >> ((x << 1) >> nScale) (Equation 18) wTL[x][y] = 0 (Equation 19)

[0103] In another example, when the intra prediction mode predModeIntra is less than or equal to INTRA_ANGULAR 10 (e.g., in the case of the 10, 67 intra prediction mode, mode 10), for location (x, y), variables dXPos[y], dXFrac[y], dXInt[y], and dX[y] are derived based on the variable invAngle, which is a function of the intra prediction mode predModeIntra. In one example, invAngle can be determined based on a look-up table storing invAngle values corresponding to each intra prediction mode. Then, the reference samples refT[x][y], refL[x][y], and the weights wL, wT, and wTL are determined based on the variables dXPos[y], dXFrac[y], dXInt[y], and dX[y].

[0104] For example, the variables dXPos[y], dXFrac[y], dXInt[y], and dX[y] are determined according to Equations 20 - 23. dXPos[y] = ((y + 1) × invAngle + 2) >> 2 (Equation 20) dXFrac[y] = dXPos[y] & 63 (Equation 21) dXInt[y] = dXPos[y] >> 6 (Equation 22) dX[y] = x + dXInt[y] (Equation 23)

[0105] And the reference samples refT[x][y], refL[x][y], and the weights wL, wT, wTL are determined according to Equations 24 - 28. refL[x][y] = 0 (Equation 24) refT[x][y] = (dX[y] < refW - 1)? ((64 - dXFrac[y]) × mainRef[dX[y]] + dXFrac[y] × mainRef[dX[y] + 1] + 32) >> 6 : 0 (Equation 25) wT[y] = (dX[y] < refW - 1)? 32 >> ((y << 1) >> nScale) : 0 (Equation 26) wL[x] = 0 (Equation 27) wTL[x][y] = 0 (Equation 28)

[0106] In another example, when the intra prediction mode predModeIntra is 58 or more in INTRA_ANGULAR (for example, in the case of 58 or 67 intra prediction modes, mode 58), variables dYPos[x], dYFrac[x], dYInt[x], and dY[x] are derived based on the variable invAngle, which is a function of the intra prediction mode predModeIntra. In one example, invAngle can be determined based on a look-up table that stores invAngle values corresponding to each intra prediction mode. Then, the reference samples refT[x][y], refL[x][y], and the weights wL, wT, and wTL are determined based on the variables dYPos[x], dYFrac[x], dYInt[x], and dY[x].

[0107] For example, the variables dYPos[x], dYFrac[x], dYInt[x], and dY[x] are determined according to Equations 29 to 33. dYPos[x] = ((x + 1) × invAngle + 2) >> 2 (Equation 29) dYFrac[x] = dYPos[x] & 63 (Equation 30) dYInt[x] = dYPos[x] >> 6 (Equation 31) dY[x] = x + dYInt[x] (Equation 32)

[0108] And the reference samples refT[x][y], refL[x][y], and the weights wL, wT, wTL are determined according to Equations 33 to 37. refL[x][y] = (dY[x] < refH - 1)? ((64 - dYFrac[x]) × sideRef[dY[x]] + dYFrac[x] × sideRef[dY[x] + 1] + 32) >> 6 : 0 (Equation 33) refT[x][y] = 0 (Equation 34) wT[y] = 0 (Equation 35) wL[x] = (dY[x] < refH - 1)? 32 >> ((x << 1) >> nScale) : 0 (Equation 36) wTL[x][y] = 0 (Equation 37)

[0109] In some examples, when the variable predModeIntra is between 11 and 57 and not equal to either 18 or 50, refL[x][y], refT[x][y], wT[y], wL[y], and wTL[x][y] are all set equal to 0. Next, the values of the filtered samples filtSamples[x][y] where x = 0..nTbW - 1 and y = 0..nTbH - 1 are derived as follows. filtSamples[x][y] = clip1Cmp((refL[x][y] × wL + refT[x][y] × wT - p[-1][-1] × wTL[x][y] + (64wL[x] - wT[y] + wTL[x][y]) × predsample[x][y] + 32) >> 6) (Equation 38)

[0110] Note that some PDPC processes involve non - integer (e.g., floating - point) operations that increase the computational complexity. In some embodiments, the PDPC process involves relatively simple calculations for the PLANAR mode (mode 0), DC mode (mode 1), vertical mode (e.g., mode 50 in the case of 67 intra - prediction modes), horizontal mode (e.g., mode 18 in the case of 67 intra - prediction modes), and diagonal modes (e.g., modes 2, 66, 34 in the case of 67 intra - prediction modes), and the PDPC process involves relatively complex calculations for other modes.

[0111] In some embodiments, for the chroma components of an intra-coded block, the encoder selects the best chroma prediction mode from five modes including the PLANAR mode (mode index 0), the DC mode (mode index 1), the horizontal mode (mode index 18), the vertical mode (mode index 50), the diagonal mode (mode index 66), and a direct copy of the intra prediction mode of the associated luma component, i.e., the DM mode. Table 1 shows the mapping between the intra prediction direction for chroma and the intra prediction mode number.

[0112]

Table 1

[0113] To avoid duplicate modes, in some embodiments, the four modes other than DM are assigned according to the intra prediction mode of the associated luma component. When the intra prediction mode number of the chroma component is 4, the intra prediction direction of the luma component is used for generating the intra prediction samples of the chroma component. When the intra prediction mode number of the chroma component is not 4 and is the same as the intra prediction mode number of the luma component, the intra prediction direction 66 is used for generating the intra prediction samples of the chroma component.

[0114] According to some aspects of the present disclosure, inter-picture prediction (also referred to as inter prediction) includes the merge mode and the skip mode.

[0115] In the merge mode for inter-picture prediction, the motion data (e.g., motion vectors) of a block are inferred instead of being explicitly signaled. In one example, a merge candidate list of candidate motion parameters is first constructed, and then an index identifying the candidate to be used is signaled.

[0116] In some embodiments, the merge candidate list includes a non-sub-CU merge candidate list and a sub-CU merge candidate list. The non-sub-CU merge candidates are constructed based on spatially neighboring motion vectors, collocated temporal motion vectors, and history-based motion vectors. The sub-CU merge candidate list includes affine merge candidates and ATMVP merge candidates. The sub-CU merge candidates are used to derive multiple MVs of the current CU, and different parts of the samples within the current CU can have different motion vectors.

[0117] In skip mode, the motion data of the block is inferred instead of being explicitly signaled, and the prediction residual is 0, i.e., the transform coefficients are not transmitted. At the beginning of each CU in an inter-picture prediction slice, the skip_flag is signaled. The skip_flag indicates that the merge mode is used to derive the motion data and that there is no residual data in the encoded video bitstream.

[0118] According to some aspects of the present disclosure, intra prediction and inter prediction can be appropriately combined, such as in multi-hypothesis intra-inter prediction. Multi-hypothesis intra-inter prediction combines one intra prediction and one prediction with a merge index, and is referred to as the intra-inter prediction mode in the present disclosure. In one example, when the CU is in the merge mode, a specific flag for the intra mode is signaled. If the specific flag is true, the intra mode can be selected from the intra candidate list. For the luma component, the intra candidate list is derived from four intra prediction modes: DC mode, PLANAR mode, horizontal mode, and vertical mode, and the size of the intra candidate list can be 3 or 4 according to the block shape. In one example, if the CU width is greater than twice the CU height, the horizontal mode is removed from the intra mode candidate list, and if the CU height is greater than twice the CU width, the vertical mode is removed from the intra mode candidate list. In some embodiments, the intra prediction is performed based on the intra prediction mode selected by the intra mode index, and the inter prediction is performed based on the merge index. The intra prediction and the inter prediction are combined using a weighted average. In the case of the chroma component, in some examples, DM is always applied without additional signaling.

[0119] In some embodiments, the weights for combining intra prediction and inter prediction can be appropriately determined. In one example, equal weights are applied to the inter prediction and the intra prediction if the DC or PLANAR mode is selected, or if the width or height of the coded block (CB) is less than 4. In another example, for a CB with a CB width and a CB height of 4 or more, if the horizontal / vertical mode is selected, the CB is first divided into four equal-area regions vertically / horizontally. Each region is (w_intra i , w_inter i) having a set of weights shown as, where i ranges from 1 to 4. In one example, the first set of weights (w_intra1, w_inter1) = (6, 2), the second set of weights (w_intra2, w_inter2) = (5, 3), the third set of weights (w_intra3, w_inter3) = (3, 5), and the fourth set of weights (w_intra4, w_inter4) = (2, 6) can be applied to the corresponding regions. For example, the first set of weights (w_intra1, w_inter1) is for the region closest to the reference sample, and the fourth set of weights (w_intra4, w_inter4) is for the region farthest from the reference sample. Then, the two weighted predictions are summed and right-shifted by 3 bits to calculate the composite prediction.

[0120] Also, when intra-coding the neighboring CB, for the following intra-mode coding of the neighboring CB, the intra-prediction mode for the intra-hypothesis of the predictor can be saved.

[0121] According to some aspects of the present disclosure, a motion refinement technique called bidirectional optical flow (BDOF) mode is used in inter-prediction. BDOF is also called BIO in some examples. BDOF is used to improve the dual-prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU when the CU satisfies the following conditions: 1) the height of the CU is not 4 and the size of the CU is not 4×8, 2) the CU is not coded using the affine mode or the ATMVP merge mode. 3) the CU is coded using the "true" dual-prediction mode, that is, one of the two reference pictures is before the current picture in the display order and the other is after the current picture in the display order. In some examples, BDOF is applied only to the luma component.

[0122] Motion refinement in BDOF mode is based on the concept of optical flow that assumes the motion of an object is smooth. For each 4×4 sub-block, motion refinement (v x , v y ) is calculated by minimizing the difference between the L0 prediction sample and the L1 prediction sample. Then, the motion refinement is used to adjust the dual prediction sample values within the 4×4 sub-block. In BDOF processing, the following steps are applied.

[0123] First, the horizontal and vertical gradients of the two prediction signals,

Number

Number

Number

[0124] Next, the autocorrelations and cross-correlations of the gradients S1, S2, S3, S5, and S6 are calculated as follows. S1 = Σ (i,j)∈Ω ψ x (i, j), S3 = Σ (i,j)∈Ω θ(i, j)·ψ x (i, j) S2 = Σ (i,j)∈Ω ψ x (i, j)·ψ y (i, j) S5 = Σ (i,j)∈Ω ψ y (i, j)·ψ y (i, j) S6 = Σ (i,j)∈Ω θ(i, j)·ψ y (i, j) (Equation 40) Here, [Number] Ω is a 6×6 window around a 4×4 sub-block.

[0125] Next, motion refinement (v x , v y ) is derived using cross-correlation terms and auto-correlation terms. [Number] Here, [Number] where [Number] is the floor function.

[0126] Based on motion refinement and gradients, the following adjustments are calculated for each sample within the 4×4 sub-block. [Number]

[0127] Finally, the BDOF samples of the CU are calculated by adjusting the dual-prediction samples as follows. pred BDOF (x, y) = (I (0) (x, y) + I (1) (x, y) + b(x, y) + o offset ) >> shift (Equation 44)

[0128] In the above, n a , n b and [Number] The values are equal to 3, 6, and 12 respectively. These values are selected such that the multiplier in the BDOF process does not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits.

[0129] To derive the gradient values, several predicted samples I within the list k (k = 0, 1) outside the current CU boundary (k) (i, j) can be generated.

[0130] Figure 10 shows an example of an extended CU region within BDOF. In the example of Figure 10, a 4×4 CU 1010 is shown as the shaded area. BDOF uses one extended row / column around the boundary of the CU, and the extended region is shown as a dashed 6×6 block 1020. To control the computational complexity of generating out-of-boundary prediction samples, a bilinear filter is used to generate prediction samples within the extended region (white positions), and a normal 8-tap motion compensation interpolation filter is used to generate prediction samples within the CU (gray positions). These extended sample values are only used in gradient calculations. In the remaining steps of the BDOF process, when samples and gradient values outside the CU boundary are required, they are padded (i.e., repeated) from their nearest neighbors.

[0131] In some examples, triangular partitioning can be used for inter prediction. For example (e.g., VTM 3), a new triangular partitioning mode is introduced for inter prediction. The triangular partitioning mode is applicable only to CUs that are 8×8 or larger and are encoded in skip mode or merge mode. For CUs that meet these conditions, a CU-level flag is signaled to indicate whether the triangular partitioning mode is applied.

[0132] When the triangular partitioning mode is used, the CU is evenly divided into two triangular partitions using either diagonal partitioning or anti-diagonal partitioning.

[0133] FIG. 11 shows the diagonal splitting of a CU and the anti-diagonal splitting of a CU. Each triangular partition within the CU has its own motion information and can perform inter prediction using its own motion. In one example, only single prediction is allowed for each triangular partition. And each partition has one motion vector and one reference index. The single prediction motion constraint is applied to ensure that only two motion compensation predictions per CU are required, just like in conventional dual prediction. The single prediction motion of each partition is derived from a single prediction candidate list constructed using processing.

[0134] In some examples, when a CU level flag indicates that the current CU is encoded using the triangular partition mode, an index within the range of [0, 39] is further signaled. Using this triangular partition index, the direction of the triangular partition (diagonal or anti-diagonal), as well as the motion of each of the partitions, can be obtained via a look-up table. After predicting each of the triangular partitions, the sample values along the diagonal or anti-diagonal edges are adjusted using blend processing with adaptive weights. After the entire CU is predicted, the transform and quantization processing are applied to the entire CU in the same way as in other prediction modes. Finally, the motion field of the CU predicted using the triangular partition mode is stored in 4×4 units.

[0135] In some related examples of the intra prediction mode, PDPC is applied only to intra prediction samples to improve video quality by reducing artifacts. In related examples, there may still be some artifacts for intra prediction samples, which may not result in optimal video results.

[0136] The proposed methods may be used separately or combined in any order.

[0137] According to some aspects of the present disclosure, PDPC is applied to inter-prediction samples (or reconstructed samples of inter-coded CUs), and the use of PDPC filtering techniques in inter-prediction can be referred to as the inter-PDPC mode. In one example, Equation 1 can be appropriately modified for the PDPC between modes. For example, pred[x][y] is modified to indicate the sample value of inter-prediction in the inter-PDPC mode.

[0138] In some examples, a flag (other suitable names can be used for the flag), called interPDPCFlag in the present disclosure, is signaled to indicate whether to apply PDPC to inter-prediction samples. In one example, when the flag interPDPCFlag is true, the inter-prediction samples (or reconstructed samples of inter-coded CUs) are further modified in PDPC processing in a manner similar to PDPC processing for intra-prediction.

[0139] In some embodiments, when the multi-hypothesis intra-inter prediction flag is true, one additional flag, such as interPDPCFlag in the present disclosure, is signaled to indicate whether to apply multi-hypothesis intra-inter prediction or to apply PDPC to inter-prediction samples. In one example, when interPDPCFlag is true, PDPC is directly applied to the inter-prediction samples (or reconstructed samples of inter-coded CUs) to generate the final inter-prediction value (or reconstructed samples of inter-coded CUs). Otherwise, in the example, multi-hypothesis intra-inter prediction is applied to the inter-prediction samples (or reconstructed samples of inter-coded CUs).

[0140] In one embodiment, a fixed context model is used for entropy coding of interPDPCFlag.

[0141] In other embodiments, the selection of the context model used for entropy coding of interPDPCFlag depends on coding information including, but not limited to, whether the neighboring CU is an intra-CU or an inter-CU, or whether it is in an intra mode, the coding block size, etc. The block size can be measured by, for example, block area size, block width, block height, block width + height, block width and height, etc.

[0142] For example, if none of the neighboring modes are intra-coded CUs, the first context model is used. Otherwise, if at least one of the neighboring modes is an intra-coded CU, the second context model is used.

[0143] In another example, if none of the neighboring modes are intra-coded CUs, the first context model is used. If one of the neighboring modes is an intra-coded CU, the second context model is used. If two or more of the neighboring modes are intra-coded CUs, the third context model is used.

[0144] In other embodiments, when PDPC is applied to inter-prediction samples, a smoothing filter is applied to neighboring reconstructed (or predicted) samples.

[0145] In other embodiments, to apply PDPC to inter-prediction samples, neighboring reconstructed (or predicted) samples are directly used for PDPC processing without using a smoothing or interpolation filter.

[0146] In other embodiments, if the block size of the current block is below a threshold, the flag interPDPCFlag can be derived as false. The block size can be measured by, for example, block area size, block width, block height, block width + height, block width and height, etc. In one example, if the block area size is less than 64 samples, the flag interPDPCFlag is derived as false.

[0147] In other embodiments, if the block size of the current block is larger than a threshold such as the size of the VPDU (virtual processing data unit defined as a 64×64 block in the example), the flag interPDPCFlag is not signaled and can be assumed to be false.

[0148] In other embodiments, the PDPC scaling factor (e.g., nScale in the present disclosure) or the weighting factor (e.g., the weight in Equation 1) is restricted such that PDPC (as a filtering technique) does not change the value of the current pixel if the distance from the current pixel to the reconstructed pixel (e.g., Equation 1) is larger than a specific threshold such as the VPDU width / height.

[0149] In some embodiments, to apply PDPC to inter prediction samples (or the reconstructed samples of an inter-coded CU), a default intra prediction mode is assigned to the current block, and as a result, the relevant parameters of the default mode such as the weighting factor and the scaling factor can be used for PDPC processing. In one example, the default intra prediction mode can be the PLANAR mode or the DC mode. In other examples, the default intra prediction mode for the chroma component is the DM prediction mode. In other examples, subsequent blocks can use this default intra prediction mode for intra mode coding and Most Probable Mode (MPM) derivation.

[0150] In other embodiments, to apply PDPC to inter prediction samples, BDOF (or called BIO) is not applied in the process of generating inter prediction samples.

[0151] In other embodiments, if PDPC is applied to inter prediction samples, in one example, only the neighboring reconstructed (or predicted) samples from the inter-coded CU can be used for PDPC processing.

[0152] In other embodiments, interPDPCFlag is signaled for coded blocks coded in merge mode. In other embodiments, PDPC is not applied to skip mode CUs or sub-block merge mode CUs. In other embodiments, PDPC is not applied to triangular partition mode coded CUs.

[0153] In other embodiments, interPDPCFlag is signaled at the TU level and is signaled only if the CBF (Coded Block Flag) of the current TU is not 0. In one example, the th flag interPDPCFlag is not signaled and can be assumed as 0 (false) when the CBF of the current TU is 0. In other examples, the reconstructed samples in the vicinity of the current TU (or current PU / CU) can be used for PDPC.

[0154] In other embodiments, in order to apply PDPC to inter-prediction samples, only the single prediction motion vector can be used to generate intra-prediction samples. If the motion vector of the current block is a bi-prediction motion vector, the motion vector needs to be converted to a single prediction motion vector. In one example, if the motion vector of the current block is a bi-prediction motion vector, the motion vector in List 1 is discarded and the motion vector in List 0 can be used to generate intra-prediction samples.

[0155] In other embodiments, in order to apply PDPC to inter-prediction samples, the transform skip mode (TSM) is neither applied nor signaled.

[0156] In some embodiments, when the intra-inter mode is true, BDOF is not applied to inter-prediction samples.

[0157] In some embodiments, when dual prediction is applied, boundary filtering can be applied to both the forward and backward prediction blocks using spatially neighboring samples in the corresponding reference picture. Next, a dual prediction block is generated by taking the average (or weighted sum) of the forward and backward prediction blocks. In one example, boundary filtering is applied using a PDPC filter. In other examples, the neighboring samples of the forward (or backward) prediction block can include the samples above, to the left, below, and to the right of the spatially neighboring samples in the associated reference picture.

[0158] In some embodiments, when single prediction is applied, filtering is applied to the prediction block using its spatially neighboring samples in the reference picture. In one example, boundary filtering is applied using a PDPC filter. In other examples, the neighboring samples of the forward (or backward) prediction block can include the samples above, to the left, below, and to the right of the spatially neighboring samples in the associated reference picture.

[0159] In some embodiments, PLANAR prediction is applied to an inter prediction block using spatially neighboring samples in the corresponding reference picture. This process is called an inter prediction sample refinement process. In one example, when dual prediction is applied, after the inter prediction sample refinement process, a dual prediction block is generated by averaging (or weighted summing) the prediction blocks with improved accuracy in the forward and backward directions. In other examples, the neighboring samples used in the inter prediction sample refinement process can include the spatially neighboring samples above, to the left, below, and to the right in the associated reference picture.

[0160] FIG. 12 shows a flowchart illustrating an overview of a process (1200) according to an embodiment of the present disclosure. The process (1200) can be used for reconstructing a block to generate a predicted block of the block being reconstructed. In various embodiments, the process (1200) is executed by a processing circuit such as a processing circuit of terminal devices (110), (120), (130), and (140), a processing circuit performing the function of video encoder (203), a processing circuit performing the function of video decoder (210), a processing circuit performing the function of video decoder (310), a processing circuit performing the function of video encoder (403), etc. In some embodiments, the process (1200) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit performs the process (1200). The process starts from (S1201) and proceeds to (S1210).

[0161] (S1210), the prediction information of the current block is decoded from the encoded video bitstream. The prediction information indicates a prediction of the current block based at least in part on inter prediction. In one example, the prediction information indicates an inter prediction mode such as a merge mode, a skip mode, etc. In other examples, the prediction information indicates an intra-inter prediction mode.

[0162] (S1220), based on the prediction information of the current block, a determination is made as to the use of PDPC for inter prediction. In one example, a flag such as interPDPCflag is received and decoded from the encoded bitstream. In other examples, the flag is derived.

[0163] (S1230), the samples of the current block are reconstructed. At least the samples of the current block are reconstructed as a combination of the result from inter prediction and neighboring samples selected based on the positions of the samples. In one example, PDPC is applied to the result of inter prediction in the same manner as PDPC for intra prediction. Then, the process proceeds to (S1299) and ends.

[0164] In one example, the prediction information indicates an intra-inter prediction mode that uses a combination of intra prediction and inter prediction for the current block. Next, it should be noted that motion refinement based on bidirectional optical flow from the inter prediction of the current block can be excluded. In this example, (S1220) can be skipped and the decoding is independent of the PDPC process.

[0165] The above techniques can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 13 shows a computer system (1300) suitable for implementing a particular embodiment of the disclosed subject matter.

[0166] Computer software can be encoded using any suitable machine code or computer language and be the subject of assembly, compilation, linking, or similar mechanisms to create code that includes instructions executable directly, or through interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0167] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0168] The components shown in FIG. 13 for the computer system (1300) are exemplary in nature and are not intended to suggest any limitations regarding the use or functionality of the computer software implementing the embodiments of the present disclosure. Also, the configuration of the components should not be construed as having dependencies or requirements regarding any one or combination of the components shown in the exemplary embodiments of the computer system (1300).

[0169] The computer system (1300) may include a specific human interface input device. Such a human interface input device can respond to input from one or more users, such as, for example, tactile input (keystrokes, swipes, movements of a data glove, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown), etc. Using the human interface device, it is also possible to capture specific media that is not necessarily directly related to conscious input by humans, such as audio (speech, music, ambient sound, etc.), pictures (scanned images, photographic images obtained from a still image camera, etc.), images (2D video, 3D video including stereoscopic video, etc.).

[0170] The input human interface device may include one or more of a keyboard (1301), a mouse (1302), a track pad (1303), a touch screen (1310), a data glove (not shown), a joystick (1305), a microphone (1306), a scanner (1307), a camera (1308) (only one of each shown).

[0171] The computer system (1300) may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., a touch screen (1310), a data glove (not shown), or a joystick (1305) with tactile feedback, although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (1309), headphones (not shown), etc.), visual output devices (whether or not they have a touch screen input function, and whether or not they have a tactile feedback function, including means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown) that can output two-dimensional visual output or three-dimensional or higher-dimensional output, screens (1310) including CRT screens, LCD screens, plasma screens, OLED screens, etc.), and printers (not shown).

[0172] The computer system (1300) may also include a storage device accessible to humans, and related media such as optical media (1321) such as CD / DVD ROM / RW (1320) including CD / DVDs, thumb drives (1322), removable hard drives or solid state drives (1323), legacy magnetic media such as tapes and floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0173] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include a transmission medium, a carrier wave, or other transient signals.

[0174] The computer system (1300) may also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. Further, the network can be local, wide area, metropolitan area, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, vehicle and industrial such as CANBus, etc. For certain networks, generally, an external network interface adapter connected to a specific general data port or peripheral bus (1349) (such as the USB port (1300) of a computer system) is required, and others are generally integrated into the core (1300) of the computer system by connecting to the system bus as described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1300) can communicate with other entities. Such communication can be unidirectional, receive only (such as broadcast TV), transmit only unidirectional (such as from a CANbus to a specific CANbus device), or bidirectional, for example, communication to other computer systems using a local area digital network or a wide area digital network. As described above, specific protocols and protocol stacks can be used for each of these networks and network interfaces.

[0175] The aforementioned human interface device, human-accessible storage device, and network interface can be connected to the core (1340) of the computer system (1300).

[0176] The core (1340) can include special programmable processing devices in the form of one or more central processing units (CPUs) (1341), graphics processing units (GPUs) (1342), field programmable gate arrays (FPGAs) (1343), hardware accelerators for specific tasks (1344), etc. These devices can be connected via a system bus (1348) together with read-only memory (ROM) (1345), random access memory (1346), internal mass storage devices such as internal hard drives and SSDs that are not accessible to the user (1347). In some computer systems, one or more physical plugs can be accessed on the system bus (1348) to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be connected directly to the core's system bus (1348) or via a peripheral bus (1349). Peripheral bus architectures include PCI, USB, etc.

[0177] The CPU (1341), GPU (1342), FPGA (1343), and accelerator (1344) can execute specific instructions that can be combined to form the aforementioned computer code. The computer code can be stored in the ROM (1345) or RAM (1346). Migration data can also be stored in the RAM (1346), while persistent data can be stored, for example, in the internal mass storage device (1347). By using cache memory that can be closely associated with one or more CPUs (1341), GPUs (1342), mass storage devices (1347), ROM (1345), RAM (1346), etc., fast storage and reading from any memory device are made possible.

[0178] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The media and computer code can be specially designed and constructed for the purposes of this disclosure or they can be of the kinds well-known and available to those skilled in the computer software arts.

[0179] By way of example and not limitation, a computer system having an architecture (1300), particularly a core (1340), can provide functionality as a result of software executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.) incorporated in one or more tangible computer-readable media. Such computer-readable media can be associated with the above-introduced mass storage devices accessible to the user, and specific storage devices of the core (1340) having a non-transitory nature, such as the on-core mass storage device (1347) and ROM (1345). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1340). The computer-readable media can include one or more memory devices or chips according to specific needs. The software can cause the core (1340), particularly the processors (including a CPU, GPU, FPGA, etc.) therein, to define data structures stored in the RAM (1346), and change such data structures according to processes defined by the software, so as to execute specific processes or specific parts of specific processes described herein. Additionally, or alternatively, the computer system can provide functionality as a result of logic incorporated in a circuit (e.g., an accelerator (1344)) or otherwise implemented to operate instead of or together with software to execute specific processes or specific parts of specific processes described herein. References to software can include logic, and vice versa as appropriate. References to computer-readable media can, as appropriate, include software for execution, circuits that embody logic for execution, or circuits (such as integrated circuits (ICs)) that store both. The present disclosure encompasses any suitable combination of hardware and software. Appendix A: Acronyms JEM: joint exploration model VVC: versatile video coding BMS: benchmark set Benchmark Set MV: Motion Vector Motion Vector HEVC: High Efficiency Video Coding High Efficiency Video Coding SEI: Supplementary Enhancement Information Supplementary Enhancement Information VUI: Video Usability Information Video Usability Information GOP: Group of Pictures Group of Pictures TU: Transform Units, Transform Units PU: Prediction Units Prediction Units CTU: Coding Tree Units Coding Tree Units CTB: Coding Tree Blocks Coding Tree Blocks PB: Prediction Blocks Prediction Blocks HRD: Hypothetical Reference Decoder Hypothetical Reference Decoder SNR: Signal Noise Ratio Signal Noise Ratio CPU: Central Processing Units Central Processing Units GPU: Graphics Processing Units Graphics Processing Units CRT: Cathode Ray Tube Cathode Ray Tube LCD: Liquid-Crystal Display Liquid-Crystal Display OLED: Organic Light-Emitting Diode Organic Light-Emitting Diode CD: Compact Disc Compact Disc DVD: Digital Video Disc Digital Video Disc ROM: Read-Only Memory Read-Only Memory RAM: Random Access Memory, Random Access Memory ASIC: Application-Specific Integrated Circuit, Application-Specific Integrated Circuit PLD: Programmable Logic Device, Programmable Logic Device LAN: Local Area Network, Local Area Network GSM: Global System for Mobile communications, Global System for Mobile Communications LTE: Long-Term Evolution, Long-Term Evolution CANBus: Controller Area Network Bus, Controller Area Network Bus USB: Universal Serial Bus, Universal Serial Bus PCI: Peripheral Component Interconnect, Peripheral Component Interconnect FPGA: Field Programmable Gate Areas, Field Programmable Gate Areas SSD: solid-state drive, Solid-State Drive IC: Integrated Circuit, Integrated Circuit CU: Coding Unit, Coding Unit

[0180] Although the present disclosure has described several exemplary embodiments, there are changes, substitutions, and various alternative equivalents within the scope of the present disclosure. Therefore, those skilled in the art will understand that although not explicitly shown or described herein, many systems and methods that embody the principles of the present disclosure and are thus within its spirit and scope can be devised.

Explanation of Signs

[0181] 100 Communication System 110 Terminal Device 120 Terminal Device 130 Terminal Device 150 Network 200 Communication System 201 Video Source 202 Video Picture Stream 203 Video Encoder 204 Encoded Video Data 205 Streaming Server 206 Client Subsystem 207 Copy 210 Video Decoder 211 Output Stream of Video Picture 212 Display 213 Capture Subsystem 220 Electronic Device 230 Electronic Device 301 Channel 310 Video Decoder 312 Rendering Device 315 Buffer Memory 320 Entropy Decoder / Parser 321 Symbol 330 Electronic Device 331 Receiver 351 Scaler / Inverse Transformation Unit 352 Intra-Picture Prediction Unit 353 Motion Compensation Prediction Unit 355 Aggregator 356 Loop Filter Unit 357 Reference Picture Memory 358 Picture Buffer 401 Video Source 403 Video Encoder 420 Electronic Device 430 Source Coder 432 Encoding Engine 433 Decoder 434 Reference Picture Memory 435 Predictor 440 Transmitter 443 Encoded video sequence 445 Entropy coder 450 Controller 460 Communication channel 503 Video encoder 521 General controller 522 Intra encoder 523 Residual calculator 524 Residual encoder 525 Entropy encoder 526 Switch 528 Residual decoder 530 Inter encoder 610 Video decoder 671 Entropy decoder 672 Intra decoder 673 Residual decoder 674 Reconstruction module 680 Inter decoder 1200 Processing 1300 Computer system 1301 Keyboard 1302 Mouse 1303 Track pad 1305 Joystick 1306 Microphone 1307 Scanner 1308 Camera 1309 Speaker 1310 Touch screen 1321 Optical media 1322 Thumb drive 1323 Solid state drive 1340 Core 1343 Field programmable gate array (FPGA) 1344 Hardware accelerator for specific tasks 1345 Read only memory (ROM) 1346 Random access memory 1347 Internal large-capacity memory device 1348 System bus 1349 Peripheral bus 1350 Graphics adapter 1354 Network interface

Claims

1. Decoding prediction information of a current block from an encoded video bitstream, wherein the prediction information indicates that the prediction of the current block is at least partially based on an inter prediction; Reconstructing at least samples of the current block as a combination with neighboring samples of the current block selected based on a result from the inter prediction and positions of the samples; Decoding a first flag for an intra-inter prediction mode indicating whether to use a combination of an intra prediction and the inter prediction of the current block; When the first flag is true, directly applying a position-dependent prediction combination (PDPC) to inter prediction samples and excluding motion refinement based on bidirectional optical flow from the inter prediction of the current block; A method for video decoding in a decoder, comprising the above steps.

2. Determining whether the current block meets a block size condition that restricts application of the PDPC to the inter prediction in reconstruction of the current block; When the block size condition is met, excluding the application of the PDPC in reconstruction of at least one sample of the current block; The method according to claim 1, further comprising the above steps.

3. The method according to claim 2, wherein the block size condition depends on a size of a virtual pipeline data unit (VPDU).

4. Reconstructing the samples of the current block as a combination with the filtered neighboring samples of the current block and the result from the inter prediction; The method according to claim 1, further comprising the above step.

5. An apparatus for video decoding configured to perform the method according to any one of claims 1 to 4.

6. A computer program for causing a computer of an apparatus for video decoding to execute the method according to any one of claims 1 to 4.

7. A method for video encoding in an encoder, comprising: Determining whether a current block should be at least partially based on an inter prediction; Encoding a first flag for an intra-inter prediction mode that uses a combination of intra prediction and inter prediction of the current block; Excluding bidirectional optical flow-based motion refinement (BDOF) from the inter prediction of the current block; Encoding at least the samples of the current block as a result of a combination of neighboring samples of the current block selected based on the result of the inter prediction and the positions of the samples; comprising; A method in which a position-dependent prediction combination (PDPC) is directly applied to inter prediction samples when the first flag is true. **Claim 8** The BDOF is applied to the inter prediction when (1) the height of the current block is not 4 and the size is not 4×8, (2) the affine mode or ATMVP merge mode is not used for the current block, and (3) the first reference picture used for the inter prediction is before the current block in the display order and the second reference picture is after the current block in the display order. The method according to claim 7. **Claim 9** The BDOF is applied only to the luma component. The method according to claim 8. **Claim 10** Applying a filter to the neighboring samples of the current block. The method according to claim 7. **Claim 11** When the first flag is not true, multi-hypothesis intra-inter prediction is applied to inter prediction samples. The method according to claim 7. **Claim 12** An apparatus for video encoding configured to perform the method according to any one of claims 7 to 11. **Claim 13** A computer program for causing a computer of an apparatus for video encoding to execute the method according to any one of claims 7 to 11.

Citation Information

Patent Citations

  • Coding apparatus, coding method, program for coding method, and recording medium for recording program for coding method

    JP2007074050A

  • Video coding method, video coding device, computer-readable storage medium and computer program

    JP2022515558A

  • Image encoding device, image decoding device, and programs therefor

    WO2017030198A1

  • System and method for improving combined inter and intra prediction

    WO2020146562A1

Cited By

  • Method, apparatus, and program for video coding

    JP2025138780A