Potential mask blend mode

The latent mask blend mode enhances video coding accuracy by using boundary-recognized composite prediction to weight actual pixel data, addressing the issue of suboptimal predictions from padded reference blocks.

JP2026517542APending Publication Date: 2026-06-02TENCENT AMERICA LLC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2023-09-08
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing video coding methods suffer from reduced accuracy due to padding reference blocks that extend beyond image boundaries, leading to suboptimal predictions and artifacts in decoded video data.

Method used

Implement a latent mask blend mode that uses boundary-recognized composite prediction, weighting actual pixel data more heavily to reconstruct blocks from two reference blocks, even when they partially extend beyond boundaries.

Benefits of technology

Improves coding accuracy by providing a more accurate reconstruction based on actual pixel data, reducing artifacts in decoded video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026517542000001_ABST
    Figure 2026517542000001_ABST
Patent Text Reader

Abstract

The computing system receives a video bitstream containing the current block and a syntax element indicating that the current block will be predicted in blend mode. The current block is encoded using information from a first and a second reference block. When a portion of the current block corresponds to a first region that (i) lies within the corresponding reference boundary in both the first and second reference blocks, or (ii) does not lie within the corresponding reference boundary in both the first and second reference blocks, the system reconstructs the portion by averaging the reference values ​​from the first and second reference blocks. When a portion of the current block corresponds to a second region that lies within the corresponding reference boundary in only one of the first and second reference blocks, the system derives a weighted reference value and reconstructs the portion by combining the weighted reference values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 461,879, filed Apr. 25, 2023, entitled “Implicit Masked Blending Mode,” and is a continuation of, and claims priority to, U.S. Patent Application No. 18 / 241,757, filed Sep. 1, 2023, entitled “Implicit Masked Blending Mode.”

[0002] The disclosed embodiments generally relate to video coding and include, but are not limited to, systems and methods for linear and non - linear blending of block sections in a wedge - based prediction mode.

Background Art

[0003] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Electronic devices transmit and receive or otherwise communicate digital video data over communication networks and / or store digital video data in storage devices. Since the bandwidth capacity of communication networks is limited and the memory resources of storage devices are limited, video coding may be used to compress video data according to one or more video coding standards before the video data is communicated or stored. Video coding can be implemented by hardware and / or software on an electronic / client device or server that provides cloud services.

[0004] Video coding generally utilizes prediction methods (e.g., interpretation, intrapretation) that leverage the inherent redundancy of video data. The goal of video coding is to compress video data into a format that uses a lower bitrate while avoiding or minimizing a decrease in video quality. Several video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended as a successor to HEVC. The ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. A valid version 1.0.0, including Errata 1 of this specification, was released on January 8, 2019. [Overview of the Initiative] [Means for solving the problem]

[0005] This disclosure describes using a potential mask blend mode (e.g., boundary-recognized composite prediction) to blend block sections. For example, in a composite prediction block, the reference block from the reference image may be at least partially outside the boundary. In some systems, the portion of the reference block that is outside the boundary is padded and therefore does not reflect the actual pixel values, which can consequently reduce the accuracy of the prediction. The disclosed method and system provide a more accurate reconstruction based on actual pixel data (e.g., actual pixel data is weighted more heavily), thereby improving the coding accuracy (e.g., by reducing artifacts in the decoded video data).

[0006] According to several embodiments, a method for video decoding is provided. The method includes (a) receiving a video bitstream containing the current image and the current block within the current image; (b) determining, based on syntax elements in the video bitstream, that the current block is encoded using information from a first reference block and a second reference block; (c) determining that at least one of the first reference block and the second reference block is at least partially outside the corresponding reference boundary; and (d) determining that a portion of the current block is (i) within the corresponding reference boundary in both the first reference block and the second reference block, or (ii) within the corresponding reference boundary in both the first reference block and the second reference block. (e) Reconstructing a portion of a block by averaging reference values ​​from a first reference block and a second reference block, in accordance with the determination that the portion of the current block corresponds to a second region that is within the corresponding reference boundary in only one of the first reference block and the second reference block, (i) deriving weighted reference values ​​by applying their respective weights to the reference values ​​of the first reference block and the second reference block, and (ii) reconstructing the portion of the block by combining the weighted reference values ​​of the first reference block and the second reference block.

[0007] According to several embodiments, a method for video encoding is provided. The method includes (a) receiving a current image and a current block within the current image; (b) determining that the current block will be encoded using information from a first reference block and a second reference block; (c) determining that at least one of the first reference block and the second reference block is at least partially outside a corresponding reference boundary; and (d) determining that a portion of the current block is (i) within the corresponding reference boundary in both the first reference block and the second reference block, or (ii) outside the corresponding reference boundary in both the first reference block and the second reference block. Encoding a portion by averaging reference values ​​from a first reference block and a second reference block in accordance with a determination that it corresponds to (e) a second region that lies within the corresponding reference boundary in only one of the first reference block and the second reference block, (i) deriving weighted reference values ​​by applying their respective weights to the reference values ​​of the first reference block and the second reference block, and (ii) encoding the portion by combining the weighted reference values ​​of the first reference block and the second reference block.

[0008] According to some embodiments, computing systems are provided, such as streaming systems, server systems, personal computer systems, or other electronic devices. The computing system includes a control circuit and a memory for storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes encoder components and decoder components (e.g., transcoder components).

[0009] According to some embodiments, a non-temporary computer-readable storage medium is provided. The non-temporary computer-readable storage medium stores one or more instruction sets for execution by a computing system. One or more instruction sets include instructions for performing any of the methods described herein.

[0010] Accordingly, devices and systems are disclosed using methods for encoding and decoding video. Such methods, devices, and systems may complement or replace conventional methods, devices, and systems for video encoding / decoding.

[0011] The features and advantages described herein are not necessarily exhaustive, and in particular, some additional features and advantages will be apparent to those skilled in the art in consideration of the drawings, specification and claims provided herein. Furthermore, it should be noted that the language used herein has been selected primarily for readability and explanatory purposes and has not necessarily been selected to elaborate on or delineate the subject matter described herein.

[0012] To enable a more detailed understanding of this disclosure, a more detailed description can be provided by referring to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the accompanying drawings are merely illustrative of the relevant features of this disclosure and should not be considered limiting, as it is possible to recognize other effective features, as a person skilled in the art will understand by reading this disclosure. [Brief explanation of the drawing]

[0013] [Figure 1] Figure 1 is a block diagram illustrating an exemplary communication system according to several embodiments.

[0014] [Figure 2A] Figure 2A is a block diagram showing exemplary elements of encoder components according to several embodiments.

[0015] [Figure 2B] Figure 2B is a block diagram showing exemplary elements of a decoder component according to some embodiments.

[0016] [Figure 3] Figure 3 is a block diagram showing an exemplary server system according to some embodiments.

[0017] [Figure 4A] Figure 4A is a diagram showing an exemplary coding tree structure according to some embodiments. [Figure 4B] Figure 4B is a diagram showing an exemplary coding tree structure according to some embodiments. [Figure 4C] Figure 4C is a diagram showing an exemplary coding tree structure according to some embodiments. [Figure 4D] Figure 4D is a diagram showing an exemplary coding tree structure according to some embodiments.

[0018] [Figure 5A] Figure 5A is a diagram showing an example of a partition-based prediction mode according to some embodiments.

[0019] [Figure 5B] Figure 5B is a diagram showing an exemplary blend of partitioning modes according to some embodiments. [Figure 5C] Figure 5C is a diagram showing an exemplary blend of partitioning modes according to some embodiments.

[0020] [Figure 5D] Figure 5D is a diagram showing an exemplary wedge-based partitioning according to some embodiments.

[0021] [Figure 5E]Figure 5E shows an exemplary scenario in which the reference block is at least partially outside the corresponding reference image boundary according to several embodiments. [Figure 5F] Figure 5F shows an exemplary scenario in which the reference block, according to several embodiments, is at least partially outside the corresponding reference image boundary.

[0022] [Figure 6A] Figure 6A is a flowchart illustrating exemplary methods for encoding video according to several embodiments.

[0023] [Figure 6B] Figure 6B is a flowchart illustrating exemplary methods for decoding video according to several embodiments. [Modes for carrying out the invention]

[0024] According to common practice, the various features illustrated in the drawings are not necessarily drawn to scale, and it is possible to use the same reference number to represent the same features throughout the specification and drawings.

[0025] This disclosure, in particular, describes the use of various partitioning techniques for partitioning video blocks for better motion prediction and higher quality encoding. This disclosure also describes the use of latent mask blends (e.g., boundary-recognized composite prediction) for reconstructing the encoded current block using information from two reference blocks via a composite mode. As used herein, “latent mask blend” refers to a mask blend that is not explicitly signaled in the video bitstream. As will be described in more detail later, the composite mode can be an effective inter-predictive coding tool. Furthermore, the incorporation of optical flow-based prediction and / or temporal interpolated prediction (TIP) modes can amplify the capabilities of the composite mode. However, in some cases, one or both of the reference blocks extend beyond the corresponding image boundary. In such cases, padding may be applied to the samples outside the image boundary, which is suboptimal for the composite mode because the prediction is no longer based on actual pixel data. The use of latent mask blending can improve the coding accuracy of the composite mode (e.g., reduce artifacts in the decoded video data) by providing a more accurate reconstruction based on actual pixel data (e.g., actual pixel data is weighted more heavily) compared to the current design where portions outside the boundary of the reference block are padded and do not reflect the actual pixel values.

[0026] Exemplary systems and devices Figure 1 is a block diagram showing a communication system 100 according to several embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to 120-m) that are coupled to communicate with one or more networks. In some embodiments, the communication system 100 is a streaming system for use in video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0027] Source device 102 includes a video source 104 (e.g., a camera component or a media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to generate an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from video source 104 may have a larger data volume compared to the encoded video bitstream 108 generated by the encoder component 106. Because the encoded video bitstream 108 has a smaller data volume (less data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared to the video stream from video source 104. In some embodiments, source device 102 does not include the encoder component 106 (e.g., it is configured to transmit uncompressed video data to a network 110).

[0028] One or more networks 110 represent any number of networks that transmit information between the source device 102, the server system 112, and / or the electronic device 120, for example, including wired and / or wireless communication networks. One or more networks 110 may exchange data over circuit switching channels and / or packet switching channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0029] One or more networks 110 include a server system 112 (e.g., a distribution / cloud computing system). In some embodiments, the server system 112 is or includes a streaming server (configured to store and / or distribute video content, such as an encoded video stream from a source device 102). The server system 112 includes a coder component 114 (configured to encode and / or decode video data, for example). In some embodiments, the coder component 114 includes an encoder component and / or a decoder component. In various embodiments, the coder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the coder component 114 is configured to decode an encoded video bitstream 108 and re-encode the video data using different encoding standards and / or methodologies to produce encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video bitstream 108.

[0030] In some embodiments, the server system 112 functions as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to prune an encoded video bitstream 108 to adapt different bitstreams to one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0031] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode encoded video data 116 to generate an output video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., including a media storage device that is communicably coupled to an external display device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access a server system 112 to retrieve encoded video data 116.

[0032] The source device and / or multiple electronic devices 120 may be referred to as “terminal devices” or “user devices.” In some embodiments, one or more of the source device 102 and / or electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smartphone, tablet, or laptop), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0033] In an exemplary operation of the communication system 100, the source device 102 transmits an encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a stream of images captured by the source device. The server system 112 receives the encoded video bitstream 108 and may decode and / or encode the encoded video bitstream 108 using the coder component 114. For example, the server system 112 may apply encoding to the video data that is more optimal for network transmission and / or storage. The server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 to restore the video image and display it as desired.

[0034] In some embodiments, the transmission described above is a one-way data transmission. One-way data transmission may be used in media serving applications, etc. In some embodiments, the transmission described above is a two-way data transmission. Two-way data transmission may be used in video conferencing applications, etc. In some embodiments, the encoded video bitstream 108 and / or encoded video data 116 are encoded and / or decoded according to one of the video encoding / compression standards described herein, such as HEVC, VVC, and / or AV1.

[0035] Figure 2A is a block diagram showing exemplary elements of an encoder component 106 according to several embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source which is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream which can be any preferred bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCB, or RGB), and any preferred sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. Video data may be provided as a series of individual images that give motion when viewed sequentially. The images themselves may be organized as a spatial array of pixels, and each pixel may contain one or more samples depending on the sampling structure, color space, etc., used. Those skilled in the art will readily understand the relationship between pixels and samples. The following description will focus on samples.

[0036] The encoder component 106 is configured to encode and / or compress images from a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. Enforcing an appropriate encoding rate is one function of the controller 204. In some embodiments, the controller 204 controls and is functionally coupled to other functional units, as described below. Parameters set by the controller 204 may include rate control-related parameters (e.g., image skipping, quantizer, and / or lambda values ​​for rate distortion optimization techniques), image size, Group of Pictures (GOP) layout, and maximum motion vector search range. Those skilled in the art will readily be able to identify other functions of the controller 204, as they may relate to encoder components 106 optimized for a particular system design.

[0037] In some embodiments, the encoder component 106 is configured to operate in a coding loop. In a simplified example, the coding loop includes a source coder 202 (responsible for generating symbols, such as a symbol stream, based on, for example, the input image and reference image(s) to be coded) and a (local) decoder 210. The decoder 210 reconstructs the symbols to generate sample data in a manner similar to that of a (remote) decoder (when compression between the symbols and the coded video bitstream is reversible). The reconstructed sample stream (sample data) is input to a reference image memory 208. Since decoding the symbol stream yields bit-accurate results regardless of the decoder position (local or remote), the contents in the reference image memory 208 are also bit-accurate between the local encoder and the remote encoder. In this way, the predictive unit of the encoder interprets the same sample values ​​as reference image samples that the decoder would interpret when using predictions during decoding. This principle of reference image synchronization (and the resulting drift, for example, if synchronization cannot be maintained due to channel errors) is known to those skilled in the art.

[0038] The operation of decoder 210 can be identical to that of a remote decoder, such as decoder component 122, which is described in detail below in relation to Figure 2B. However, referring briefly to Figure 2B, since symbols are available and the encoding / decoding of symbols to the coded video sequence by the entropicorder 214 and parser 254 may be reversible, the entropy decoding portion of decoder component 122, including buffer memory 252, and parser 254 do not need to be fully implemented in local decoder 210.

[0039] The decoder techniques described herein may exist in substantially the same functional form in the corresponding encoders, except for analysis / entropy decoding. For this reason, the subject matter disclosed focuses on decoder operation. The description of encoder techniques may be omitted as it is the reverse of the description of decoder techniques. Only in certain areas is a more detailed explanation required and is provided below.

[0040] As part of its operation, the source coder 202 may perform motion-compensated predictive coding, predictively coding the input frame by referencing one or more previously coded frames from a video sequence designated as reference frames. In this scheme, the coding engine 212 codes the difference between the pixel blocks of the input frame and the pixel blocks of reference frames(s) that may be selected as predictive references for the input frame. The controller 204 may manage the coding operation of the source coder 202, including, for example, setting parameters and subgroup parameters used to encode the video data.

[0041] The decoder 210 decodes the coded video data of a frame that may be designated as a reference frame, based on the symbols generated by the source coder 202. The operation of the coding engine 212 may, advantageously, be a lossy process. When the coded video data is decoded by a video decoder (not shown in Figure 2A), the reconstructed video sequence may be a replica of the source video sequence with some errors. The decoder 210 may reproduce the decoding process that may be performed by a remote video decoder on the reference frame and store the reconstructed reference frame in the reference image memory 208. In this scheme, the encoder component 106 locally stores a copy of the reconstructed reference frame having common content as the reconstructed reference frame that will be acquired by the remote video decoder (without transmission errors).

[0042] The predictor 206 may perform a predictive search of the coding engine 212. That is, with respect to a new frame to be coded, the predictor 206 may search the reference image memory 208 for sample data (as candidate reference pixel blocks) or certain metadata such as motion vectors and block shapes of reference images that may serve as appropriate predictive criteria for the new image. The predictor 206 may operate for each sample block per pixel block to find appropriate predictive criteria. In some cases, the input image may have predictive criteria drawn from multiple reference images stored in the reference image memory 208, as determined by the search results obtained by the predictor 206.

[0043] The outputs of all the aforementioned functional units may be entropicated in the entropicorder 214. The entropicorder 214 converts symbols generated by the various functional units into coded video sequences by reversibly compressing the symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable-length coding, and / or arithmetic coding).

[0044] In some embodiments, the output of the entropicorder 214 is coupled to a transmitter. The transmitter may be configured to buffer the coded video sequence generated by the entropicorder 214 and prepare them for transmission over a communication channel 218, which may be a hardware / software link to a storage device that will store the coded video data. The transmitter may be configured to merge the coded video data from the source coder 202 with other data to be transmitted, such as coded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter may transmit additional data along with the coded video. The source coder 202 may include such data as part of the coded video sequence. The additional data may include time / space / SNR enhancement layers, other forms of redundant data such as redundant images and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0045] The controller 204 may manage the operation of the encoder component 106. During coding, the controller 204 may assign a specific coded image type to each coded image, which may affect the coding technique applied to each image. For example, an image may be assigned as an intra-picture (I-picture), a predictive-picture (P-picture), or a bidirectional predictive-picture (B-picture). Intra-pictures can be coded and decoded without using any other frames in the sequence as a source for prediction. Some video codecs enable different types of intra-pictures, including, for example, Independent Decoder Refresh (IDR) images. Those skilled in the art will recognize their variations of I-pictures and their respective uses and characteristics, and therefore will not repeat them here. Predictive-pictures can be coded and decoded using intra-prediction or inter-prediction, which uses up to one motion vector and reference index to predict the sample value of each block. Bidirectional predictive-pictures can be coded and decoded using intra-prediction or inter-prediction, which uses up to two motion vectors and reference indexes to predict the sample value of each block. Similarly, multiple prediction images can use three or more reference images and associated metadata for the reconstruction of a single block.

[0046] The source image may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively), and each block may be coded. Blocks may be coded predictively by referencing other (already coded) blocks, as determined by the coding assignment applied to each image in the block. For example, a block of image I may be coded non-predictively or predictively by referencing an already coded block of the same image (spatial prediction or intra-prediction). A pixel block of image P may be coded non-predictively via spatial prediction or temporal prediction by referencing one previously coded reference image. A block of image B may be coded non-predictively via spatial prediction or temporal prediction by referencing one or two previously coded reference images.

[0047] The video may be captured in time series as multiple source images (video images). Intra-image prediction (often abbreviated as intra-prediction) utilizes spatial correlations in a given image, while inter-image prediction utilizes (temporal or other) correlations between images. In one example, a particular image being encoded / decoded, called the current image, is partitioned into blocks. When a block in the current image is similar to a reference block in a previously encoded and buffered reference image in the video, the block in the current image can be encoded by a vector called a motion vector. The motion vector points to a reference block in the reference image and may have a third dimension that identifies the reference image if multiple reference images are used.

[0048] The encoder component 106 may perform coding operations in accordance with a predetermined video coding technology or standard, such as any of those described herein. In these operations, the encoder component 106 may perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the coded video data may conform to the syntax specified by the video coding technology or standard being used.

[0049] Figure 2B is a block diagram showing exemplary elements of a decoder component 122 according to several embodiments. The decoder component 122 in Figure 2B is coupled to channel 218 and display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to a loop filter 256 and configured to transmit data to the display 124 (for example, via a wired or wireless connection).

[0050] In some embodiments, the decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more coded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each coded video sequence is independent of other coded video sequences. Each coded video sequence may be received from channel 218, which may be a hardware / software link to a storage device that stores coded video data. The receiver may receive coded video data with other data, each of which may be sent using entities (not shown), e.g., coded audio data and / or auxiliary data streams. The receiver may isolate the coded video sequence from the other data. In some embodiments, the receiver receives additional (redundant) data with coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the decoder component 122 to decode the data and / or to reconstruct the original video data with greater precision. Additional data can take the form of, for example, time, space, or SNR enhancement layers, redundant slices, redundant images, or forward error correction codes.

[0051] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes referred to as an entropy decoder), a scaler / inverse unit 258, an intra-image prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference image memory 266, and a current image memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, the decoder component 122 is implemented at least partially in software.

[0052] Buffer memory 252 is coupled between channel 218 and parser 254 (for example, to counteract network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 within decoder component 122 (for example, configured to handle playout timing), a separate buffer memory is provided outside decoder component 122 (for example, to counteract network jitter). When receiving data from a storage / feed device with sufficient bandwidth and controllability, or from an asynchronous network, buffer memory 252 may not be required or can be small. For use in best-effort packet networks such as the Internet, buffer memory 252 may be required, can be relatively large, and advantageously can be adaptively sized, and may be at least partially implemented in an operating system or similar element (not shown) outside decoder component 122.

[0053] The parser 254 is configured to reconstruct symbols 270 from the coded video sequence. The symbols may include, for example, information used to manage the operation of decoder component 122 and / or information for controlling a rendering device such as a display 124. The control information for the rendering device may be, for example, in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser 254 parses (entropy decodes) the coded video sequence. The coding of the coded video sequence may follow video coding techniques or standards and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, and arithmetic coding with or without context sensitivity. The parser 254 may extract from the coded video sequence a set of at least one subgroup parameters of subgroups of pixels in the video decoder, based on at least one parameter corresponding to a group. Subgroups may include groups of pictures (GOP), images, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), and prediction units (PU). Parser 254 may also extract information from the coded video sequence, such as transform coefficients, quantizer parameter values, and motion vectors.

[0054] The reconstruction of symbol 270 may involve multiple different units, depending on the type of coded video image or part thereof (e.g., inter-image and intra-image, inter-block and intra-block) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information analyzed from the coded video sequence by parser 254. The flow of such subgroup control information between parser 254 and the multiple units described below is not depicted for clarity.

[0055] The decoder component 122 can be conceptually subdivided into several functional units, and in some implementations, these units can interact closely with each other and integrate with each other at least partially. However, for clarity, the conceptual subdivision of the functional units is maintained herein.

[0056] The scaler / inverse unit 258 receives quantized transformation coefficients, as well as control information as symbols 270 (such as which transformation to use, block size, quantization coefficients, and / or quantization scaling matrix), from the parser 254. The scaler / inverse unit 258 can output a block containing sample values ​​that can be input to the aggregator 268.

[0057] In some cases, the output samples of the scaler / inverse transform unit 258 relate to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed images but can use prediction information from previously reconstructed portions of the current image. Such prediction information can be provided by the intra-image prediction unit 262. The intra-image prediction unit 262 may generate blocks of the same size and shape as the block being reconstructed using already reconstructed surrounding information taken from the current (partially reconstructed) image from the current image memory 264. The aggregator 268 may add the prediction information generated by the intra-image prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258, on a sample-by-sample basis.

[0058] In other cases, the output samples of the scaler / inverse unit 258 relate to intercoded, potentially motion-compensated blocks. In such cases, the motion-compensated prediction unit 260 can access the reference image memory 266 to retrieve samples to be used for prediction. After motion-compensating the retrieved samples according to the symbols 270 associated with the blocks, these samples can be added by the aggregator 268 to the output of the scaler / inverse unit 258 (referred to in this case as residual samples or residual signals) to generate output sample information. The address in the reference image memory 266 from which the motion-compensated prediction unit 260 retrieves the prediction samples may be controlled by a motion vector. The motion vector may be available to the motion-compensated prediction unit 260 in the form of a symbol 270, which may have, for example, X, Y, and reference image components. Motion compensation may also include interpolation of sample values ​​retrieved from the reference image memory 266 when the precise motion vectors of the subsamples are used, a motion vector prediction mechanism, and so on.

[0059] The output samples of the aggregator 268 can undergo various loop filtering techniques in the loop filter unit 256. The video compression technique may include in-loop filtering techniques controlled by parameters contained in the coded video bitstream and made available to the loop filter unit 256 as symbols 270 from the parser 254, but may also respond to metadata obtained during decoding of earlier parts (in decoding order) of the coded image or coded video sequence, and may also respond to previously reconstructed and loop-filtered sample values.

[0060] The output of the loop filter unit 256 can be a sample stream that can be output to a rendering device such as the display 124, or it can be stored in the reference image memory 266 for use in future inter-image prediction.

[0061] Once fully reconstructed, a particular coded image can be used as a reference image for future predictions. Once a coded image is fully reconstructed and identified as a reference image (e.g., by parser 254), the current reference image can become part of the reference image memory 266, and new current image memory can be reallocated before starting the reconstruction of subsequent coded images.

[0062] The decoder component 122 may perform decoding operations according to a predetermined video compression technique, which may be documented in a standard such as one of the standards described herein. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that it is faithful to the syntax of the video compression technique or standard, as specified in the video compression technique documentation or standard, particularly the profile documentation therein. In addition, to conform to certain video compression techniques or standards, the complexity of the coded video sequence may be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum image size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference image size, etc. The limits set by the level may be further limited, in some cases, through the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0063] Figure 3 is a block diagram showing a server system 112 according to several embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes one or more field-programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).

[0064] The network interface(s) 304 may be configured to interface with one or more communication networks (e.g., wireless, wired, and / or optical networks). These communication networks may be local, wide-area, urban, vehicle, and industrial, real-time, latency-tolerant, etc. Examples of communication networks include local area networks such as Ethernet® and Wi-Fi; cellular networks such as GSM®, 3G, 4G, 5G, and LTE; wired or wireless wide-area digital television networks such as cable TV, satellite TV, and terrestrial broadcast TV; and vehicle and industrial networks such as CANBus. Such communications may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a specific CANbus device), or bidirectional (e.g., to other computer systems using local or wide-area digital networks). Such communications may include communications to one or more cloud computing networks.

[0065] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device(s) 310 may include one or more of the following: a keyboard, mouse, trackpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. The output device(s) 308 may include one or more of the following: an audio output device (e.g., a speaker), a visual output device (e.g., a display or monitor), etc.

[0066] Memory 314 may include high-speed random-access memory (such as DRAM, SRAM, DDR RAM, and / or other random-access solid-state memory devices) and / or non-volatile memory (such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state memory devices). Memory 314 optionally includes one or more storage devices located remotely from the control circuit 302. Memory 314, or the non-volatile solid-state memory device(s) within Memory 314, includes a non-temporary computer-readable storage medium. In some embodiments, Memory 314, or the non-temporary computer-readable storage medium of Memory 314, stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: ● Operating system 316, which includes procedures for handling various basic system services and for performing hardware-dependent tasks. ● A network communication module 318 used to connect the server system 112 to other computing devices via one or more network interfaces 304 (for example, via wired and / or wireless connections). ● Encoding module 320 for performing various functions with respect to data such as video data Encoding and / or decoding data In some embodiments, the encoding module 320 is an instance of the coder component 114. The encoding module 320 includes, but is not limited to, one or more of the following: o Decoding module 322 for performing various functions related to decoding encoded data as described above, with respect to decoder component 122 o Encoding module 340 for performing various functions related to encoded data as described above with respect to encoder component 106 ●For example, an image memory 352 for storing images and image data for use with the coding module 320. In some embodiments, the image memory 352 includes one or more of the following: a reference image memory 208, a buffer memory 252, a current image memory 264, and a reference image memory 266.

[0067] In some embodiments, the decoding module 322 includes an analysis module 324 (configured to perform the various functions described above with respect to the parser 254, for example), a transformation module 326 (configured to perform the various functions described above with respect to the scalar / inverse transform unit 258, for example), a prediction module 328 (configured to perform the various functions described above with respect to the motion compensation prediction unit 260 and / or intra-image prediction unit 262, for example), and a filter module 330 (configured to perform the various functions described above with respect to the loop filter 256, for example).

[0068] In some embodiments, the coding module 340 includes a coding module 342 (configured to perform the various functions described above with respect to, for example, the source coder 202 and / or the coding engine 212) and a prediction module 344 (configured to perform the various functions described above with respect to, for example, the predictor 206). In some embodiments, the decoding module 322 and / or the coding module 340 includes a subset of the modules shown in Figure 3. For example, a shared prediction module is used by both the decoding module 322 and the coding module 340.

[0069] Each of the identified modules stored in memory 314 corresponds to a set of instructions for performing the functions described herein. The identified modules (e.g., sets of instructions) do not need to be implemented as separate software programs, procedures, or modules, and therefore various subsets of these modules may be combined or rearranged in various embodiments. For example, the coding module 320 may optionally not include separate decoding and coding modules, but rather use the same set of modules to perform both sets of functions. In some embodiments, memory 314 stores a subset of the modules and data structures identified above. In some embodiments, memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0070] In some embodiments, the server system 112 includes web pages and applications implemented using a web or HyperText Transfer Protocol (HTTP) server, a File Transfer Protocol (FTP) server, and a Common Gateway Interface (CGI) script, PHP Hyper-text Preprocessor (PHP), Active Server Pages (ASP), HyperText Markup Language (HTML), Extensible Markup Language (XML), Java®, JavaScript®, Asynchronous JavaScript and XML (AJAX), XHP, Javelin, Wireless Universal Resource File (WURFL), and the like.

[0071] Figure 3 shows a server system 112 according to several embodiments, but Figure 3 is not a schematic diagram of the structure of the embodiments described herein, but rather is primarily intended as a functional description of various features that may be present in one or more server systems. In practice, and as will be recognized by those skilled in the art, items shown separately can be combined, and some items can be separated. For example, some items shown separately in Figure 3 can be implemented on a single server, and a single item can be implemented by one or more servers. The actual number of servers used to implement the server system 112, and how functions are allocated among them, will vary from implementation to implementation and, if desired, will partially depend on the amount of data traffic the server system will handle during peak and average usage periods.

[0072] Examples of coding processes and techniques The coding processes and techniques described below may be implemented in the devices and systems described above (e.g., source device 102, server system 112, and / or electronic device 120). Figures 4A to 4D show exemplary coding tree structures according to several embodiments. As shown in the first coding tree structure (400) in Figure 4A, some coding techniques (e.g., VP9) use a four-way partition tree from 64x64 levels to 4x4 levels, with some additional constraints on the 8x8 block. In Figure 4A, the partition designated as R can be called recursive in that the same partition tree is repeated at lower scales until the lowest 4x4 level is reached.

[0073] As shown in the second coded tree structure (402) in Figure 4B, some coding methods (e.g., AV1) extend the partition tree to a 10-way structure, increasing the maximum size (e.g., referred to as the superblock in VP9 / AV1 parlance) from 128x128. The second coded tree structure includes 4:1 / 1:4 rectangular partitions that are not present in the first coded tree structure. The partition type with three subpartitions in the second row of Figure 4B is referred to as a T-shaped partition. In addition to the coded block size, the coded tree depth can be defined to indicate the partition depth from the root note.

[0074] As an example, a CTU may be partitioned into CUs by using a quad-tree structure represented as a coding tree to adapt to various local characteristics such as HEVC. In some embodiments, the decision of whether to encode an image region using inter-image (time) prediction or intra-image (spatial) prediction is made at the CU level. Each CU can be further partitioned into one, two, or four PUs according to the PU partitioning type. Within a single PU, the same prediction process is applied, and relevant information is transmitted to the decoder on a PU-by-PU basis. After obtaining residual blocks by applying the prediction process based on the PU partitioning type, the CU can be partitioned into TUs according to another quad-tree structure, such as a coding tree for the CU.

[0075] Quad trees with nested multi-type trees using binary and ternary partitioning segmentation structures such as VVC may replace the concept of multiple partition unit types, eliminating the separation of the concepts of CU, PU, ​​and TU, except when necessary for CUs that are too large for the maximum transformation length, for example, and supporting greater flexibility for CU partition shapes. In coded tree structures, CUs can have either a square or rectangular shape. CTUs are initially partitioned by a quadtree (also called a quad tree) structure. Quadtree leaf nodes can be further partitioned by a multi-type tree structure. As shown in the third coded tree structure (404) in Figure 4C, the multi-type tree structure includes four partition types. Multi-type tree leaf nodes are called CUs, and unless a CU is too large for the maximum transformation length, this segmentation is used for prediction and transformation processing without any further partitioning. This means that in most cases, CUs, PUs, and TUs have the same block size in a quad tree with a nested multi-type tree coded block structure. An example of block partitioning for one CTU (406) is shown in Figure 4D, which illustrates an exemplary quad tree with a nested multi-type tree-encoded block structure.

[0076] Motion estimation involves determining the motion vector that describes the transformation from one image to another. The reference image (or block) can be from adjacent frames in a video sequence. The motion vector may be related to the entire image (global motion estimation) or to a specific block. Furthermore, the motion vector can correspond to a translation or deformation model that approximates the motion (e.g., rotation and translation in three dimensions and zoom). Motion estimation can be improved in some situations (e.g., with more complex video objects) by further subdividing the block.

[0077] Geometric Partitioning Mode (GPM) may focus on inter-image prediction blocks (e.g., CUs). When GPM is applied to a block, the block is divided into two parts via a straight partition boundary. The position of the partition boundary may be mathematically defined by an angular parameter φ and an offset parameter ρ. These parameters may be quantized and combined into a GPM partition index lookup table. The GPM partition index of the current block may be encoded into a bitstream. For example, for a CU with size w×h=2k×2l (with respect to luminance samples) where k, l∈{3...6}, 64 partition modes are supported by the GPM of the VVC. For example, since narrow CUs contain few geometrically separated patterns, GPM may be disabled for CUs with aspect ratios greater than 4:1 or less than 1:4.

[0078] After partitioning, the two GPM sections (partitions) contain individual motion information that can be used to predict the corresponding section within the current block. In some embodiments, only unidirectional motion-compensated prediction (MCP) is permitted per GPM section so that the memory bandwidth required for the MCP within the GPM is equal to the memory bandwidth of a normal bidirectional MCP. To simplify the encoding of motion information and reduce the number of possible combinations for the GPM, the motion information can be encoded in merge mode. A GPM merge candidate list can be derived from a merge candidate list to ensure that it contains only unidirectional motion information.

[0079] Figure 5A shows a GPM prediction process according to several embodiments. The current block 510 is divided into a right section and a left section via a partition 516. The right prediction portion of the current block 510 (e.g., CU) of the current image 502 (with dimensions of w × h) is predicted by MV0 from reference block 512 of reference image 504, and the left portion is predicted by MV1 from reference block 514 of reference image 506.

[0080] In some embodiments, when the current block is divided into two partitions via a partition (e.g., partition 516), a blend (corresponding to, for example, “Blend Mode” or “Mask Blend”) is applied to combine the two partitioned blocks. Blending partitions involves applying a weighted sum of predictions to each partition. Figure 5B shows an exemplary blend matrix for a partition (e.g., partition 516) according to some embodiments. In this example, the final GPM prediction (PG) is generated by performing the blend process using integer blend matrices W0 and W1 containing weights in the range of values ​​from 0 to 8, for example. This can be expressed as follows:

[0081]

number

[0082] In Equation 1, J is a matrix of size w × h. The weights of the blend matrix may depend on the displacement between the sample position and the partition boundary. These matrices can be generated on the decoder side to reduce the computational complexity of the blend matrix derivation.

[0083] Next, the generated GPM prediction (PG) can be subtracted from the original signal to produce the residual. The residual may be converted to a bitstream, quantized, and encoded using, for example, a standard VVC conversion, quantization, and entropycoding engine. On the decoder side, the signal is reconstructed by adding the residual to the GPM prediction PG. For example, when the residual is negligible, the GPM can support a skip mode. For example, the residual is dropped by the encoder, and the GPM prediction PG is used directly by the decoder as the reconstructed signal.

[0084] Wedge-based prediction is an example of a blend mode. Wedge-based prediction is a composite prediction mode similar to GPM (in AV1, for example). Wedge-based prediction can be used for both inter-internal and inter-intranal combinations. The boundaries of moving objects are often difficult to approximate with block partitions on a grid. A solution is to predefine a codebook of possible wedge partitions (e.g., 16) and signal the wedge index in the bitstream when the coded units will be further partitioned in such a way. In the current wedge design in AV1, 16 modes are supported because, using the multi-symbol adaptive context coding used in AV1, up to 16 symbols can be signaled with a single syntax element. A shape codebook of 16 variables, including any partition direction of horizontal, vertical, or diagonal (e.g., tilt ±2 or ±0.5), is designed for both square blocks 540 and rectangular blocks 542, as shown in Figure 5D. In many cases, to mitigate the spurious high-frequency components generated by directly juxtaposing two predictors, a soft-cliff shaped 2-D wedge mask may be employed to smooth the edges around the intended partition (e.g., m(i,j) is close to 0.5 around the edge and gradually transforms to binary weights at both ends).

[0085] The wedge modes in AV1 may be extended to allow the use of wedge modes for 64x64, 32x64, 64x32, 16x64, and 64x16 blocks. Furthermore, the wedge modes may be defined in Hessenorm form, as shown in Figure 5C, where the angle φ indicates the direction of the division boundary and the distance ρ indicates the offset of the division boundary from the center of the block. The angle may be quantized to a value (e.g., 20 values) using tangent values. The distance may be quantized based on the block size. For example, three distances may be used for angles greater than 180 degrees, and for angles of 0 or 90 degrees. For other angles, four distances may be used. In this way, 8 × 4 + 12 × 3 = 68 modes may be supported. Since more than 16 modes are supported, the wedge index may be signaled with three syntactic elements, for example, angular direction, angle, and distance. The angular direction indicates whether the angle is less than 180 degrees. Depending on the angular direction, the actual angle may be signaled. Depending on the signaled angle, the distance may be signaled.

[0086] The wedge blend mask may be quantized directly from the distance between the sample position and the partition boundary. Using a partition boundary definition in Hessenorm form, the distance can be defined as follows:

[0087]

number

[0088] During the ceremony,

[0089]

number

[0090] This is the distance from the center,

[0091]

number

[0092] is the partition angle. The angle and distance may be quantized using the tangent value and block size. Therefore, only a lookup table and transition operation may be used to calculate the quantized d(m, n), as shown in Equation 3 below.

[0093]

number

[0094] The blend weights at the corresponding positions may be derived using Equation 4.

[0095]

number

[0096] Blend weights may be calculated on the fly (for example, due to their low computational complexity) or they may be stored in advance (for example, similar to AV1 wedge mode design).

[0097] In some embodiments, a composite prediction block is predicted by two or more reference blocks, each having a corresponding reference image. Some current systems generate a composite prediction block by combining values ​​from two reference blocks using a simple average or a weighted average. In some cases, at least one of the reference blocks is at least partially outside the boundary. Figures 5E and 5F illustrate exemplary scenarios in some embodiments where a composite prediction block is predicted using a partially out-of-bounds reference block. In Figure 5E, a composite prediction block is predicted using values ​​from reference block 556 of reference image 552 and values ​​from reference block 558 of reference image 554. The shaded region of reference block 556 is within the corresponding reference boundary 566 (e.g., the reference image boundary), but the unshaded region of reference block 556 is outside the corresponding reference boundary 566. The shaded region of reference block 558 is within the corresponding reference boundary 568, but the unshaded region of reference block 556 is outside the corresponding reference boundary 568. In Figure 5F, a composite predicted block is predicted using values ​​from reference block 576 in reference image 572 and values ​​from reference block 578 in reference image 574. The shaded areas of reference block 576 are within the corresponding reference boundary 586, while the unshaded areas of reference block 576 are outside the corresponding reference boundary 586. The shaded areas of reference block 578 are within the corresponding reference boundary 588, while the unshaded areas of reference block 578 are outside the corresponding reference boundary 588. In some current systems, the portion(s) outside the boundary of a reference block are padded and therefore do not reflect the actual pixel values, which can consequently reduce the accuracy of the prediction.

[0098] In some embodiments, a potential mask is generated based on a determination that at least one predictor of the composite mode is partially or completely outside the corresponding image boundary. In some embodiments, the mask is generated using the following exemplary code snippet, and the masked composite mode is applied.

[0099]

number

[0100] During the ceremony,

[0101]

number

[0102] This indicates the sample position within the image coordinates.

[0103]

number

[0104] and

[0105]

number

[0106] teeth,

[0107]

number

[0108] and

[0109]

number

[0110] This variable indicates whether the sample in the corresponding image boundary is located outside the corresponding image boundary.

[0111]

number

[0112] The generated mask

[0113]

number

[0114] This represents the weighting coefficients in [the specified context].

[0115] In some embodiments, the predictor for the mask synthesis mode is as shown in Equation 5.

[0116]

number

[0117] This is generated.

[0118]

number

[0119] In some embodiments, the potential mask generation process described above is applied to the synthesis mode, the optical flow-based synthesis mode, and / or the time interpolation prediction (TIP) mode. In some embodiments, in the case of the optical flow mode (including optical flow for TIP), the potential mask generation process is applied only to the final predicted signal after the refinement process and is not involved during refinement.

[0120] The methods and processes described below detail potential mask blend modes. According to some embodiments, potential mask blend modes are implemented when at least a portion of the reference block used to predict the current block lies outside the corresponding boundary (e.g., the reference image boundary). Potential mask blend modes may be used in conjunction with the GPM and / or wedge-based prediction processes described above. Several exemplary embodiments are referenced below.

[0121] Figure 6A is a flowchart illustrating a method 600 for encoding video according to several embodiments. The method 600 may be implemented in a computing system (e.g., a server system 112, a source device 102, or an electronic device 120) having a control circuit and a memory for storing instructions for execution by the control circuit. In some embodiments, the method 600 is implemented by executing instructions stored in the memory of the computing system (e.g., memory 314). The system receives the current image and the current block within the current image (602). The system determines that the current block will be encoded using information from a first reference block and a second reference block (604). The system determines that at least one of the first reference block and the second reference block is at least partially outside the corresponding reference boundary (606). If a portion of the current block corresponds to a first region that (i) lies within the corresponding reference boundary in both the first and second reference blocks (e.g., region 560 in Figure 5E), or (ii) does not lie within the corresponding reference boundary in both the first and second reference blocks (e.g., region 564 in Figure 5E), the system encodes the portion by averaging the reference values ​​from the first and second reference blocks (608). If a portion of the current block corresponds to a second region that lies within the corresponding reference boundary in only one of the first and second reference blocks (e.g., region 562 in Figure 5E), the system encodes the portion by (i) deriving a weighted reference value by applying the respective weights to the reference values ​​of the first and second reference blocks, and (ii) combining the weighted reference values ​​from the first and second reference blocks (610).

[0122] Figure 6B is a flowchart illustrating a method 650 for decoding video according to several embodiments. The method 650 may be implemented in a computing system (e.g., a server system 112, a source device 102, or an electronic device 120) having a control circuit and a memory for storing instructions for execution by the control circuit. In some embodiments, the method 650 is implemented by executing instructions stored in the memory of the computing system (e.g., memory 314). The system receives a video bitstream containing the current image and the current block within the current image (652). In some embodiments, the system receives syntax elements indicating that the current block will be predicted in blend mode. In some embodiments, the system determines, based on the syntax elements in the video bitstream, that the current block is encoded using information from a first reference block and a second reference block (654). The system determines that at least one of the first reference block and the second reference block is at least partially outside the corresponding reference boundary (656). In accordance with the determination that a portion of the current block corresponds to a first region that (i) lies within the corresponding reference boundary in both the first and second reference blocks (e.g., region 560), or (ii) does not lie within the corresponding reference boundary in both the first and second reference blocks (e.g., region 564), the system reconstructs the portion by averaging the reference values ​​from the first and second reference blocks (658). In accordance with the determination that a portion of the current block corresponds to a second region that lies within the corresponding reference boundary in only one of the first and second reference blocks (e.g., region 562), the system (i) derives a weighted reference value by applying the respective weights to the reference values ​​of the first and second reference blocks, and (ii) reconstructs the portion by combining the weighted reference values ​​from the first and second reference blocks (660).

[0123] Figures 6A and 6B illustrate several logical stages in a specific order, but the stages that are not order-dependent may be rearranged, and the other stages may be combined or separated. Since several rearrangements or other groupings not specifically mentioned are obvious to those skilled in the art, the rearrangements and groupings presented herein are not exhaustive. Furthermore, it should be recognized that the stages may be implemented in hardware, firmware, software, or any combination thereof.

[0124] In some embodiments, some shaded regions are inside the reference image boundary, while other transparent regions are outside the reference image boundary. The common region outside the dashed line (i.e., where both reference predictor p0 and reference predictor p1 are inside the reference image boundary, or both are outside the reference image boundary) is simply averaged using average, and the region between the dashed lines (i.e., where only one reference predictor is inside the reference image and the other reference block is outside the reference) is combined using latent mask weighting. For example, as illustrated in reference block 556 and reference block 558 in Figure 5E (or reference block 576 and reference block 578 in Figure 5F, where reference block 578 is at the corner of reference image 574), the shaded regions of the reference blocks are inside the respective boundaries of the corresponding reference images (e.g., reference images 552 and 554), and the unshaded regions are outside the respective boundaries of the corresponding reference images. In some embodiments, regions common to both reference blocks 556 and 558 located inside their respective reference boundaries (e.g., region 560), or regions common to both reference blocks 556 and 558 located outside their respective reference boundaries (e.g., region 564), are averaged using a simple average (e.g., an unweighted average). In some embodiments, regions containing (a) one reference block inside each reference boundary and (b) the other reference block outside each reference boundary (e.g., region 562) are combined using latent mask weighting, meaning that the respective weights are not signaled in the video bitstream (e.g., derived in the decoder components). See, for example, B1 below.

[0125] In some embodiments, the weighting coefficients are fixed, and the sum of the weighting coefficients is equal to 2 to the power of n. For example, the weighting coefficient P0 of a first predictor (e.g., corresponding to reference block 556) may be expressed as w1. The weighting coefficient P1 of a second predictor (e.g., corresponding to reference block 568) may be expressed as w2. In some cases (e.g., when P1 is within the boundary and P0 is outside the boundary), w1 is equal to 0 and w2 is equal to 64 (or 32), meaning that only P1 (e.g., reference block 556) is used as the composite predictor. In some cases (e.g., when P1 is within the boundary and P0 is outside the boundary), w1 is equal to 1 and w2 is equal to 63 (or 31), meaning that P1 (e.g., reference block 556) is more important, but the padded value provided by P0 (e.g., reference block 568) is still used.

[0126] In some embodiments, a number of fixed weighting coefficient values ​​are predefined, and the index of the best coefficient value (e.g., the coefficient value with the lowest cost and / or the coefficient value that results in the most accurate reconstruction) is signaled at a high level of syntax (e.g., in the image or tile header).

[0127] In some embodiments, the weighting coefficients for P0 and P1 depend on the distance between two reference frames and the current frame. In some embodiments, a higher weighting coefficient is assigned to P0 if the distance between the reference images for P0 is smaller than the distance between the reference frames for P1. In some embodiments, a plurality of predetermined weighting coefficient pairs are stored in a lookup table. In some embodiments, the decision of which weighting coefficient pair to use is potentially derived based on the distance between the two reference frames and the current frame, and / or whether the current region is outside the corresponding reference images for P0 or P1.

[0128] In some embodiments, the weighting coefficients gradually increase / decrease, starting from the reference boundary (or reference image boundary) of P0 (e.g., reference boundary 566) or P1 (e.g., reference boundary 568). For example, in the case of P0, the first row from the boundary uses the weighting coefficient w0, the second row uses the weighting coefficient w0+d, and the third row uses w0+2d, while in P1, the corresponding first row uses 64 w0, the second row uses 64-w0-d, and the third row uses 64-w0-2d.

[0129] In some embodiments, weighting coefficients are generated using a wedge mask generation equation (e.g., Equation 2) to calculate the distance between the sample position (e.g., within region 562) and the reference boundary at P0 (or P1). Depending on the blend function, the corresponding weighting coefficients may be calculated using a lookup table and / or clamp function (e.g., Equation 4).

[0130] In some embodiments, nonlinear blend functions such as sigmoid functions, hyperbolic tangent functions, trigonometric functions, exponential functions, or polynomial functions are used.

[0131] In some embodiments, regions inside the boundary use the same weight for P0 / P1 (for example, to reduce complexity). For example, P0 may have a weight of 1 and P1 may have a weight of 63.

[0132] (A1) In one embodiment, several embodiments include a method for video encoding (e.g., method 600). In some embodiments, the method is implemented in a computing system having memory and control circuits (e.g., server system 112). In some embodiments, the method is implemented in a coding module (e.g., coding module 320). In some embodiments, the method is implemented in a source coding component (e.g., source coder 202), a coding engine (e.g., coding engine 212), and / or an entropicorder (e.g., entropicorder 214). This method involves (i) receiving the current image and the current block within the current image, wherein the current block will be encoded using information from the first and second reference blocks (e.g., a composite prediction mode); (ii) determining that at least one of the first and second reference blocks is at least partially outside a corresponding reference boundary (e.g., reference boundary 566 or reference boundary 558); and (iii) determining whether a portion of the current block is (a) within the corresponding reference boundary in both the first and second reference blocks (e.g., region 560), or (b) within the corresponding reference boundary in both the first and second reference blocks. (iv) Encoding a portion by averaging reference values ​​from a first reference block and a second reference block, in accordance with the determination that it corresponds to a first region that is not within the boundary (e.g., region 564); (a) Deriving a weighted reference value by applying the respective weights to the reference values ​​of the first reference block and the second reference block, in accordance with the determination that the portion of the current block corresponds to a second region that is within the corresponding reference boundary in only one of the first reference block and the second reference block (e.g., region 562); and (b) Encoding the portion by combining the weighted reference values ​​of the first reference block and the second reference block. For example, the corresponding reference boundary is an image boundary, a slice boundary, a sub-image boundary, or a tile boundary.

[0133] (A2) In some embodiments of A1, the method further includes transmitting the encoded portion over a video bitstream.

[0134] (A3) In some embodiments of A1 or A2, each weight is a fixed weight of a first set of fixed weights among a plurality of sets of fixed weights, the fixed weight of the first set is determined based on the measurement or estimation accuracy of the fixed weights of each set.

[0135] (A4) In some embodiments of A1 to A3, encoding a portion by averaging reference values ​​from a first reference block and a second reference block includes (i) deriving a second weighted reference value by applying a second weight to each of the reference values ​​of the first reference block and the second reference block, the second weight being based on the distance between the first reference block and the second reference block and the current block, and (ii) encoding the portion by combining the second weighted reference values.

[0136] (A5) In some embodiments of A4, the first reference block is closer to the current block than the second reference block, the reference value of the first reference block is assigned a higher relative weight, and the reference value of the second reference block is assigned a lower relative weight.

[0137] (A6) In some embodiments of A4 or A5, each second weight is a fixed weight of a first set of fixed weights from a plurality of sets of fixed weights, the fixed weights of the first set are determined based on the measurement or estimation accuracy of the fixed weights of each set.

[0138] (A7) In some embodiments of A1 to A6, each weight is determined based on the distance between the second region and the corresponding reference boundary.

[0139] (A8) In some embodiments of A1 to A7, each weight is generated using one or more wedge mask generation equations.

[0140] (A9) In some embodiments of A1 to A8, the first and second regions are blended using a nonlinear blend function to determine weighting coefficients based on the distance from the boundary between the first and second regions.

[0141] (A10) In some embodiments of A1 to A9, the respective weights are fixed. For example, each weight of the first reference block is assigned a value of 0, and each weight of the second reference block is assigned a value of n, in which case n is the largest available weight. As another example, each weight of the first reference block is assigned a value of 1, and each weight of the second reference block is assigned a value of n-1, in which case n is the largest available weight.

[0142] (A11) In some embodiments of A1-A3 and A7-A10, encoding a portion by averaging reference values ​​from a first reference block and a second reference block includes averaging unweighted reference values ​​from a first reference block and a second reference block.

[0143] (B1) In other embodiments, some embodiments include a method for video decoding (e.g., method 650). In some embodiments, the method is implemented in a computing system having memory and control circuits (e.g., server system 112). In some embodiments, the method is implemented in a coding module (e.g., coding module 320). In some embodiments, the method is implemented in a parser (e.g., parser 254), a motion prediction component (e.g., motion compensation prediction unit 260), and / or an intra-prediction component (e.g., intra-image prediction unit 262). The method involves (i) receiving a video bitstream (e.g., coded video sequence) containing the current image (e.g., current image 502) and the current block within the current image, wherein the current block is encoded using information from a first reference block (e.g., reference block 514) and a second reference block (e.g., reference block 512); (ii) determining that at least one of the first reference block and the second reference block is at least partially outside the corresponding reference boundary; and (iii) determining that a portion of the current block is (a) within the corresponding reference boundary in both the first reference block and the second reference block, or (b) within the first reference block and (iv) Reconstructing a portion by averaging reference values ​​from the first and second reference blocks, in accordance with the determination that the portion of the current block corresponds to a first region that is not within the corresponding reference boundary in both of the second reference blocks; (a) deriving weighted reference values ​​by applying the respective weights to the reference values ​​of the first and second reference blocks; and (b) reconstructing the portion by combining the weighted reference values ​​of the first and second reference blocks, in accordance with the determination that the portion of the current block corresponds to a second region that is within the corresponding reference boundary in only one of the first and second reference blocks. In some embodiments, receiving a video bitstream includes receiving syntax elements indicating that the current block will be predicted in blend mode.For example, the first region is reconstructed using a simple average, and the second region is reconstructed using latent mask weighting. In some embodiments, the respective weights are not signaled within the video bitstream (e.g., derived in the decoder components). In some embodiments, determining that the current block is encoded using information from the first and second reference blocks includes determining that the current block will be predicted in synthetic prediction mode (e.g., synthetic interpretation mode).

[0144] (B2) In some embodiments of B1, each weight is a fixed weight of a first set of fixed weights from a set of multiple sets. The fixed weights of the first set are identified from a second syntax element in the video bitstream. For example, a set of fixed weighting coefficient values ​​are predefined, and the index of the best coefficient value (such as determined by an encoder component) is signaled in high-level syntax (e.g., image or tile header).

[0145] (B3) In some embodiments of B1 or B2, reconstructing a portion by averaging reference values ​​from a first reference block and a second reference block includes (i) deriving a second weighted reference value by applying a second weight to the reference values ​​of the first reference block and the second reference block, the second weight being based on the distance between the first reference block and the second reference block and the current block, and (ii) reconstructing the portion by combining the second weighted reference values. For example, the weights assigned to P0 (e.g., reference block 556) and P1 (e.g., reference block 558) in the first region depend on the distance between the two reference frames and the current frame. In some embodiments, the second weights are fixed (e.g., identical in all regions within the corresponding reference boundary).

[0146] (B4) In some embodiments of B3, the first reference block is closer to the current block than the second reference block, and the reference value of the first reference block is assigned a higher relative weight, while the reference value of the second reference block is assigned a lower relative weight. For example, if the distance between the reference image of P0 and the current image is smaller than the distance between the reference image of P1 and the current image, then P0 is assigned a higher weighting coefficient.

[0147] (B5) In some embodiments of B3 or B4, each second weight is a fixed weight of a first set of fixed weights from a plurality of sets of fixed weights, the fixed weights of the first set are identified from syntax elements in the video bitstream. For example, a plurality of predetermined weighting coefficient pairs are stored in a lookup table, and the decision of which weighting coefficient pair to use is potentially derived based on the distance between two reference frames and the current frame, and / or whether the portion corresponds to a region outside the reference boundary of P0 or P1.

[0148] (B6) In some embodiments of B1 to B5, each weight is determined based on the distance between the second region and the corresponding reference boundary. In some embodiments, the weighting coefficients gradually increase / decrease, starting from the reference boundary (or boundary in the reference image) in P0 (or P1). For example, in P0, the first row outside the reference boundary uses the first weighting coefficient w0, the second row outside the reference boundary uses the second weighting coefficient w0+d, the third row outside the reference boundary uses the third weighting coefficient w0+2d, and so on. In this example, the corresponding first row in P1 may use 64-w0, the second row may use 64-w0-d, and the third row may use 64-w0-2d.

[0149] (B7) In some embodiments of B1 to B6, each weight is generated using one or more wedge mask generation equations. For example, the weighting coefficients are generated using wedge mask generation equations such that the distance between the sample position in a first region at P0 (or P1) and the reference boundary is calculated. In this example, depending on the blend function, the corresponding weighting coefficients are calculated using a lookup table or a clamp function.

[0150] (B8) In some embodiments of B1 to B7, the first and second regions are blended using a nonlinear blend function to determine weighting coefficients based on the distance from the boundary between the first and second regions. For example, the blend function is a nonlinear blend function such as a sigmoid function, a hyperbolic tangent function, a trigonometric function, an exponential function, or a polynomial function.

[0151] (B9) In some embodiments of B1 to B8, the respective weights are fixed. In some embodiments, the sum of the respective weights is equal to 2 to the power of n. For example, the weighting coefficient of the first reference block is set to be equal to w1 for the region outside the corresponding reference boundary, and the weighting coefficient of the second reference block is set to be equal to w2 for the region outside the corresponding reference boundary. In some embodiments, all regions inside the corresponding reference boundary use the same weight for only one of the first and second reference blocks. For example, the prediction block from P0 uses a weight of 1, and the prediction block from P1 uses a weight of 63.

[0152] (B10) In some embodiments of B1 to B9, each weight of the first reference block is assigned a value of 0, and each weight of the second reference block is assigned a value of n, in which case n is the largest available weight. For example, with respect to a given portion, w1 for P0 is equal to 0, and w2 for P1 is equal to 64 (or 32), which means that only P1 contributes to the composite predictor.

[0153] (B11) In some embodiments of B1 to B9, each weight of the first reference block is assigned a value of 1, and each weight of the second reference block is assigned a value of n-1, where n is the largest available weight. For example, with respect to a given portion, w1 for P0 is equal to 1, and w2 for P1 is equal to 63 (or 31), meaning that P1 is more important, but P0 is still used (e.g., P0 corresponds to a padded value).

[0154] (B12) In some embodiments of B1-B2 and B6-B11, reconstructing the portion by averaging the reference values ​​from a first reference block and a second reference block includes averaging the unweighted reference values ​​from a first reference block and a second reference block.

[0155] (B13) In some embodiments of B1 to B12, the corresponding reference boundary is an image boundary, a slice boundary, a sub-image boundary, or a tile boundary.

[0156] In other embodiments, some embodiments include a computing system (e.g., server system 112) which includes a control circuit (e.g., control circuit 302) and a memory coupled to the control circuit (e.g., memory 314), the memory storing one or more sets of instructions configured to be executed by the control circuit, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A11 and B1 to B13 above).

[0157] In yet another embodiment, some embodiments include a non-temporary computer-readable storage medium that stores one or more sets of instructions for execution by a control circuit of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1-A11 and B1-B13 above).

[0158] In this specification, terms such as “first,” “second,” etc., may be used to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. The terms used herein are for the purpose of describing only specific embodiments and are not intended to limit the scope of the claims. Where used in the description of embodiments and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural form unless the context clearly indicates otherwise. Where used herein, the term “and / or” should also be understood to refer to and encompass all possible combinations of one or more of the associated enumerated items. Where used herein, the terms “comprises” and / or “comprising,” when used herein, specify the presence of the described features, integers, steps, actions, elements, and / or components, but should be understood not to exclude the presence or addition of one or more other features, integers, steps, actions, elements, components, and / or groups thereof.

[0159] As used herein, the term "if" can be interpreted, depending on the context, as meaning "when," "upon," "in response to determining," "in accordance with a determination," or "in response to detecting" that the stated premise is true. Similarly, the phrase "[if it is determined that the stated premise is true]," "[if the stated premise is true]," or "[when the stated premise is true]" can be interpreted, depending on the context, as meaning "in response to determining," "in response to determining," "in accordance with a determination," "in response to detecting," or "in response to detecting."

[0160] The foregoing description is provided with reference to specific embodiments for illustrative purposes. However, the above exemplary discussion is not intended to be exhaustive or to limit the claims to the exact form disclosed. In light of the above teachings, many modifications and variations are possible. The embodiments have been selected and described so as to be available to those skilled in the art in the best way to illustrate the principles of operation and practical applications.

Claims

1. A video decoding method implemented in a computing system having memory and one or more processors, Receiving a video bitstream containing the current image and the current block within the current image, and a syntax element indicating that the current block will be predicted in blend mode, Based on the syntax elements in the video bitstream, it is determined that the current block is encoded using information from the first reference block and the second reference block. Determining that at least one of the first reference block and the second reference block is at least partially outside the corresponding reference boundary, Reconstructing the portion of the current block by averaging reference values ​​from the first and second reference blocks, in accordance with the determination that (i) the portion of the current block corresponds to a first region that is within the corresponding reference boundary in both the first and second reference blocks, or (ii) the portion of the current block that is not within the corresponding reference boundary in both the first and second reference blocks. In accordance with the determination that the portion of the current block corresponds to a second region that is within the corresponding reference boundary in only one of the first reference block and the second reference block, A weighted reference value is derived by applying the respective weights to the reference values ​​of the first reference block and the second reference block. Reconstructing the portion by combining the weighted reference values ​​of the first reference block and the second reference block, Methods that include...

2. The method according to claim 1, wherein each of the aforementioned weights is a fixed weight of a first set of fixed weights of a plurality of sets, the fixed weights of the first set are identified from a second syntax element in the video bitstream.

3. Reconstructing the portion by averaging the reference values ​​from the first reference block and the second reference block is: A second weighted reference value is derived by applying a second weight, based on the distance between the first reference block and the second reference block and the current block, to the reference value of the first reference block and the second reference block. Reconstructing the portion by combining the second weighted reference values, The method according to claim 1, including the method described in claim 1.

4. The method according to claim 3, wherein the first reference block is closer to the current block than the second reference block, the reference value of the first reference block is assigned a higher relative weight, and the reference value of the second reference block is assigned a lower relative weight.

5. The method according to claim 3, wherein each of the second weights is a fixed weight of a first set of fixed weights of a plurality of sets, the fixed weights of the first set are identified from a second syntax element in the video bitstream.

6. The method according to claim 1, wherein each of the aforementioned weights is determined based on the distance between the second region and the corresponding reference boundary.

7. The method according to claim 1, wherein each of the aforementioned weights is generated using one or more wedge mask generation equations.

8. The method according to claim 1, wherein the first region and the second region are blended using a nonlinear blend function for determining weighting coefficients based on the distance from the boundary between the first region and the second region.

9. The method according to claim 1, wherein each of the aforementioned weights is fixed.

10. The method according to claim 1, wherein each of the weights of the first reference block is assigned a value of 0, and each of the weights of the second reference block is assigned a value of n, where n is the largest available weight.

11. The method according to claim 1, wherein each of the weights of the first reference block is assigned a value of 1, and each of the weights of the second reference block is assigned a value of n-1, where n is the largest available weight.

12. The method according to claim 1, wherein reconstructing the portion by averaging the reference values ​​from the first reference block and the second reference block includes averaging the unweighted reference values ​​from the first reference block and the second reference block.

13. The method according to claim 1, wherein the corresponding reference boundary is an image boundary, a slice boundary, a sub-image boundary, or a tile boundary.

14. Control circuit and Memory and One or more sets of instructions, which are stored in the memory and configured to be executed by the control circuit, Receiving a video bitstream containing the current image and the current block within the current image, and a syntax element indicating that the current block will be predicted in blend mode, Based on the syntax elements in the video bitstream, it is determined that the current block is encoded using information from the first reference block and the second reference block. Determining that at least one of the first reference block and the second reference block is at least partially outside the corresponding reference boundary, Reconstructing the portion of the current block by averaging reference values ​​from the first and second reference blocks, in accordance with the determination that (i) the portion of the current block corresponds to a first region that is within the corresponding reference boundary in both the first and second reference blocks, or (ii) the portion of the current block that is not within the corresponding reference boundary in both the first and second reference blocks. In accordance with the determination that the portion of the current block corresponds to a second region that is within the corresponding reference boundary in only one of the first reference block and the second reference block, A weighted reference value is derived by applying the respective weights to the reference values ​​of the first reference block and the second reference block. Reconstructing the portion by combining the weighted reference values ​​of the first reference block and the second reference block, One or more sets of instructions, including instructions for performing the following: A computing system equipped with [the following features].

15. The computing system according to claim 14, wherein each of the aforementioned weights is a fixed weight of a first set of fixed weights of a plurality of sets, the fixed weights of the first set are identified from a second syntax element in the video bitstream.

16. Reconstructing the portion by averaging the reference values ​​from the first reference block and the second reference block is: A second weighted reference value is derived by applying a second weight, based on the distance between the first reference block and the second reference block and the current block, to the reference value of the first reference block and the second reference block, Reconstructing the portion by combining the second weighted reference values, The computing system according to claim 14, including the following:

17. The computing system according to claim 14, wherein each of the aforementioned weights is determined based on the distance between the second region and the corresponding reference boundary.

18. A non-temporary computer-readable storage medium for storing one or more sets of instructions configured to be executed by a computing device having a control circuit and memory, wherein the one or more sets of instructions are: Receiving a video bitstream containing the current image and the current block within the current image, and a syntax element indicating that the current block will be predicted in blend mode, Based on the syntax elements in the video bitstream, it is determined that the current block is encoded using information from the first reference block and the second reference block. Determining that at least one of the first reference block and the second reference block is at least partially outside the corresponding reference boundary, Reconstructing the portion of the current block by averaging reference values ​​from the first and second reference blocks, in accordance with the determination that (i) the portion of the current block corresponds to a first region that is within the corresponding reference boundary in both the first and second reference blocks, or (ii) the portion of the current block that is not within the corresponding reference boundary in both the first and second reference blocks. In accordance with the determination that the portion of the current block corresponds to a second region that is within the corresponding reference boundary in only one of the first reference block and the second reference block, A weighted reference value is derived by applying the respective weights to the reference values ​​of the first reference block and the second reference block. Reconstructing the portion by combining the weighted reference values ​​of the first reference block and the second reference block, A non-temporary computer-readable storage medium containing instructions for performing a certain action.

19. The non-temporary computer-readable storage medium according to claim 18, wherein each of the aforementioned weights is a fixed weight of a first set of fixed weights of a plurality of sets, the fixed weights of the first set being identified from a second syntax element in the video bitstream.

20. Reconstructing the portion by averaging the reference values ​​from the first reference block and the second reference block is: A second weighted reference value is derived by applying a second weight, based on the distance between the first reference block and the second reference block and the current block, to the reference value of the first reference block and the second reference block. Reconstructing the portion by combining the second weighted reference values, A non-temporary computer-readable storage medium according to claim 18, including the following: