Method and apparatus for double prediction using sample adaptation weights
Patent Information
- Application Number
- JP2024521873
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-03
- Filing Date
- 2022-11-08
- Publication Date
- 2025-11-17
AI Technical Summary
Current coding formats like AV1 do not adequately account for statistical variations within prediction blocks, leading to suboptimal video compression and decoding efficiency.
Implementing a method and apparatus for dual prediction using sample-adaptive weights, where each sub-block in reference pictures is assigned weights based on predefined patterns or calculated minimization of cost functions, allowing for position-dependent weighting adjustments.
Improves video compression efficiency by optimizing weight application to individual sub-blocks, enhancing decoding accuracy and reducing data redundancy.
Smart Images

Figure 00000031_0000 
Figure 00000031_0001 
Figure 00000031_0002
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 359,764, filed on July 8, 2022, and U.S. Patent Application No. 17 / 980,294, filed on November 3, 2022, the disclosures of which are hereby incorporated by reference in their entireties. Technical Field
[0002] The present disclosure generally relates to communication systems, and more particularly, to methods and apparatuses for double prediction using sample - adaptive weights.
Background Art
[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. This coding format was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium established in 2015 that includes semiconductor companies, video - on - demand providers, video content production companies, software development companies, and web browser vendors. Many of the components of the AV1 project were provided from previous research efforts by the Alliance's members. Individual contributors started experimental technical platforms several years ago; for example, Xiph's / Mozilla's Daala had already made its code public in 2010, Google's experimental VP9 evolution project VP10 was announced on September 12, 2014, and Cisco's Thor was made public on August 11, 2015. Based on the construction of the VP9 codebase, AV1 incorporates additional technologies, some of which were developed in these experimental formats. The first version 0.1.0 of the AV1 reference codec was made public on April 7, 2016. The alliance announced the release of the AV1 bitstream specification on March 28, 2018, along with reference software - based encoders and decoders. On June 25, 2018, an effective version 1.0.0 of the specification was released. On January 8, 2019, an effective version 1.0.0 with Errata 1 of this specification was released. The AV1 bitstream specification includes a reference video codec. Current coding for biprediction by sharing the same weighting for all samples within a prediction block does not properly account for statistical variations at different positions within the prediction block. Summary of the Invention Means for Solving the Problems
[0004] The following presents a simplified summary of one or more embodiments of the present disclosure in order to provide a basic understanding of such embodiments. This summary is not an extensive overview of all contemplated embodiments, nor is it intended to identify key or critical elements of all embodiments or to delineate the scope of any or all embodiments. Its sole purpose is to present some concepts of one or more embodiments of the present disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[0005] A method, an apparatus, and a non-transitory computer-readable medium for dual prediction using sample adaptive weights are disclosed by the present disclosure.
[0006] According to an exemplary embodiment, a method executed by at least one processor of a video decoder includes receiving a coded video bitstream including a current picture, a first reference picture, and a second reference picture, where the current picture includes a current block divided into a plurality of sub-blocks. The method includes determining that the current picture is predicted using a dual prediction mode or a composite prediction mode based on the first reference picture and the second reference picture. The method includes obtaining a plurality of predetermined weighting patterns, each weighting pattern being signaled as an index value. The method includes selecting a weighting pattern based on a predetermined condition. The method includes deriving a first weight to be applied to a first sub-block in the first reference picture and a second weight to be applied to a second sub-block in the second reference picture based on the index value corresponding to the selected weighting pattern. The method includes assigning the first weight to the first sub-block and the second weight to the second sub-block based on the selected weighting pattern. The method further includes decoding the current block by weighted dual prediction based at least on the first sub-block weighted by the first weight and the second sub-block weighted by the second weight.
[0007] According to an exemplary embodiment, a video decoder includes at least one memory configured to store computer program code, and at least one processor configured to access the computer program code and operate as commanded by the computer program code. The computer program code includes reception code configured to cause the at least one processor to receive a coded video bitstream including a current picture, a first reference picture, and a second reference picture, where the current picture includes a current block divided into a plurality of sub-blocks. The computer program code includes determination code configured to cause the at least one processor to determine that the current picture is predicted using a bi-prediction mode or a composite prediction mode based on the first reference picture and the second reference picture. The computer program code includes acquisition code configured to cause the at least one processor to acquire a plurality of predetermined weighting patterns, where each weighting pattern is signaled as an index value. The computer program code includes selection code configured to cause the at least one processor to select a weighting pattern based on a predetermined condition. The computer program code includes derivation code configured to cause the at least one processor to derive a first weight to be applied to a first sub-block in the first reference picture and a second weight to be applied to a second sub-block in the second reference picture based on the index value corresponding to the selected weighting pattern. The computer program code includes assignment code configured to cause the at least one processor to assign the first weight to the first sub-block and the second weight to the second sub-block based on the selected weighting pattern. The computer program code includes decoding code configured to cause the at least one processor to decode the current block by weighted bi-prediction based at least on the first sub-block weighted by the first weight and the second sub-block weighted by the second weight.
[0008] According to an exemplary embodiment, a non-transitory computer-readable medium storing instructions that, when executed by a processor in a video decoder, cause the processor to execute a method, the method including receiving a coded video bitstream including a current picture, a first reference picture, and a second reference picture, the current picture including a current block divided into a plurality of sub-blocks. The method includes determining that the current picture is to be predicted using a bi-prediction mode or a composite prediction mode based on the first reference picture and the second reference picture. The method includes obtaining a plurality of predetermined weighting patterns, each weighting pattern being signaled as an index value. The method includes selecting a weighting pattern based on a predetermined condition. The method includes deriving a first weight to be applied to a first sub-block in the first reference picture and a second weight to be applied to a second sub-block in the second reference picture based on the index value corresponding to the selected weighting pattern. The method includes assigning the first weight to the first sub-block and the second weight to the second sub-block based on the selected weighting pattern. The method further includes decoding the current block by weighted bi-prediction based at least on the first sub-block weighted by the first weight and the second sub-block weighted by the second weight.
[0009] Additional embodiments are described in the following description, and will be apparent in part from the description, and may be learned by practice of the presented embodiments of the disclosure.
[0010] The above and other aspects, features, and aspects of the embodiments of the present disclosure will become apparent from the following description when taken in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6(A)
Figure 6(B)
Figure 7(A)
Figure 7(B)
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
DETAILED DESCRIPTION OF THE INVENTION
[0012] The following detailed description of the exemplary embodiments refers to the accompanying drawings. Like reference numerals in different drawings may identify the same or similar elements.
[0013] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementation forms. Further, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, in the flowcharts and descriptions of operations provided below, one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least in part), and the order of one or more operations may be rearranged, which will be understood.
[0014] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specific control hardware or software code used to implement these systems and / or methods does not limit the embodiments. Thus, the operations and behaviors of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0015] Even if a particular combination of features is recited in the claims and / or disclosed herein, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features may not be specifically recited in the claims and / or may be combined in ways not disclosed herein. Each dependent claim listed below can depend directly on only one claim, but combinations of each dependent claim with any other claim in the set of claims are included in the disclosure of possible implementations.
[0016] Elements, acts, or instructions used in this application are not to be construed as being critical or essential, unless so specified. Also, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” Where only one item is intended, the term “one” or similar language is used. Also, the terms “has,” “have,” “having,” “include,” “including,” and the like as used herein are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “at least partially based on” unless otherwise specified. Additionally, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.
[0017] Throughout this specification, references to “an embodiment,” “embodiment,” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the solution. Thus, the phrases “in an embodiment,” “in embodiments,” and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.
[0018] Furthermore, the described features, advantages, and characteristics of the present disclosure may be combined in any suitable manner in one or more embodiments. Those skilled in the art will recognize, in light of the description herein, that the present disclosure may be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments of the present disclosure.
[0019] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) may include at least two terminals (110, 120) interconnected via a network (150). In the case of unidirectional data transmission, the first terminal (110) may code video data at a local location for transmission to the other terminal (120) via the network (150). The second terminal (120) may receive the coded video data of the other terminal from the network (150), decode the coded data, and display the restored video data. Unidirectional data transmission may be common, for example, in media providing applications.
[0020] FIG. 1 shows a second pair of terminals (130, 140) provided to support bidirectional transmission of coded video that may occur, for example, during a video conference. In the case of bidirectional data transmission, each terminal (130, 140) may code video data captured at a local location for transmission to the other terminal via the network (150). Each terminal (130, 140) may also receive the coded video data transmitted by the other terminal, may decode the coded data, and may display the restored video data on a local display device.
[0021] In FIG. 1, the terminals (110-140) may be shown as servers, personal computers, and smartphones, and / or any other type of terminal. For example, the terminals (110-140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (150) represents any number of networks that transmit coded video data among the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data over circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (150) may not be important for the operation of the present disclosure, unless otherwise described below.
[0022] FIG. 2 shows the arrangement of video encoders and decoders in a streaming environment as an example of the use of the disclosed subject matter. The subject matter of the present disclosure can be equally applied to other video-related uses, such as video conferencing, digital television, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0023] As shown in FIG. 2, the streaming system (200) may include a capture subsystem (213) that can include a video source (201) and an encoder (203). The video source (201) may be, for example, a digital camera and may be configured to create an uncompressed video sample stream (202). The uncompressed video sample stream (202) may provide a high data volume compared to an encoded video bit stream and can be processed by an encoder (203) coupled to the camera (201). The encoder (203) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video bit stream (204) may include a lower data volume compared to the sample stream and can be stored at a streaming server (205) for future use. One or more streaming clients (206) can access the streaming server (205) to obtain a video bit stream (209) that can be a copy of the encoded video bit stream (204).
[0024] In an embodiment, the streaming server (205) may also function as a Media-Aware Network Element (MANE). For example, the streaming server (205) may be configured to prune the encoded video bit stream (204) to match one or more of the streaming clients (206) with potentially different bit streams. In an embodiment, the MANE may be provided separately from the streaming server (205) in the streaming system (200).
[0025] The streaming client (206) can include a video decoder (210) and a display (212). The video decoder (210) can decode, for example, a video bitstream (209) that is an input copy of the encoded video bitstream (204) and generate an output video sample stream (211) that can be rendered on the display (212) or another rendering device (not shown). In some streaming systems, the video bitstreams (204, 209) can be encoded according to a particular video encoding / compression standard. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. A video coding standard, informally known as Versatile Video Coding (VVC), is under development. Embodiments of the present disclosure can be used in the context of VVC.
[0026] Figure 3 shows an exemplary functional block diagram of a video decoder (210) attached to a display (212) according to an embodiment of the present disclosure. The video decoder (210) can include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra prediction unit (352), a motion compensation prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (). In at least one embodiment, the video decoder (210) can include an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The video decoder (210) may also be partially or fully embodied in software operating on one or more CPUs having associated memory.
[0027] In this and other embodiments, the receiver (310) may receive one or more coded video sequences to be decoded by the decoder (210) one coded video sequence at a time, and the decoding of each coded video sequence is independent from other coded video sequences. The coded video sequences may be received from a channel (312) that may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data along with other data that may be transferred to respective using entities (not shown), such as coded audio data and / or auxiliary data streams. The receiver (310) may separate the coded video sequences from the other data. To counter network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter "parser"). When the receiver (310) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous network, the buffer (315) may be unused or may be made small. When used in a best-effort packet network such as the Internet, the buffer (315) may be required, may be relatively large, and may be of an adaptable size.
[0028] Video decoder (210) may include a parser (320) for reconstructing symbols (321) from an entropy-coded video sequence. The categories of these symbols include, for example, information used to manage the operation of decoder (210) and potentially information for controlling a rendering device such as a display (212) that may be coupled to the decoder as shown in FIG. 2. The control information for the (one or more) rendering devices may be in the form of a Supplemental Enhancement Information (SEI) message or a Video User Utility Information (VUI) parameter set fragment (not shown). Parser (320) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence can be in accordance with a video coding technology or video coding standard and can follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context dependence, etc. Parser (320) may extract a set of subgroup parameters of at least one subgroup of pixels within the video decoder based on at least one parameter corresponding to a group from the coded video sequence. Subgroups can include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. Parser (320) may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.
[0029] The parser (320) can perform an entropy decoding / parsing operation on the video sequence received from the buffer (315) to create symbols (321). The reconstruction of the symbols (321) can involve multiple different units depending on the type of the coded video picture or a part thereof (such as inter and intra pictures, inter and intra blocks, etc.), as well as other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the video sequence coded by the parser (320). Such a flow of subgroup control information between the parser (320) and a plurality of the following units is not shown for clarity.
[0030] Beyond the function blocks already described, the decoder (210) can be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate.
[0031] One unit can be the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) can receive quantization transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. from the parser (320) as one or more symbols (321). The scaler / inverse transform unit (351) can output a block including sample values that can be input to the aggregator (355).
[0032] In some cases, the output samples of the scaler / inverse transform (351) may be related to blocks that are intra-coded, i.e., do not use prediction information from previously reconstructed pictures, but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) may use surrounding already reconstructed information fetched from the current (partially reconstructed) picture in the current picture memory (358) to generate a block of the same size and shape as the block being reconstructed. The aggregator (355) may, in some cases, add, sample by sample, the prediction information generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).
[0033] In other cases, the output samples of the scaler / inverse transform unit (351) may be related to inter-coded, potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (353) can access the reference picture memory (357) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbols (321) associated with the block, these samples can be added to the output of the scaler / inverse transform unit (351) by the aggregator (355) to generate the output sample information (in this case, called residual samples or a residual signal). The address in the reference picture memory (357) from which the motion compensation prediction unit (353) fetches the prediction samples can be controlled by the motion vector. The motion vector can be in the form of, for example, a symbol (321) having X, Y, and reference picture components and may be available to the motion compensation prediction unit (353). Motion compensation can also include interpolation of the sample values fetched from the reference picture memory (357) when an exact motion vector with sub-samples is used, a motion vector prediction mechanism, etc.
[0034] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356). The video compression technique is controlled by parameters included in the coded video bitstream and can include in-loop filter techniques made available to the loop filter unit (356) as symbols (321) from the parser (320), but can also respond to meta information obtained during decoding of a previous part (in decoding order) of the coded picture or coded video sequence, and can also respond to previously reconstructed and loop filtered sample values.
[0035] The output of the loop filter unit (356) can be an output sample stream that can be output to a rendering device such as a display (212) and can also be stored in the reference picture memory (357) for use in future inter-picture prediction.
[0036] Once fully reconstructed, a particular coded picture can be used as a reference picture for future prediction. When a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (320)), the current reference picture can become part of the reference picture memory (357), and a new current picture memory can be reallocated before starting the reconstruction of the next coded picture.
[0037] The video decoder (210) can perform a decoding operation according to a predetermined video compression technique that can be documented in a standard such as ITU-T Rec.H.265. The coded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that it is faithful to the syntax of the video compression technique or standard as specified in the video compression technique document or standard, specifically the profile document therein. Also, in order to conform to some video compression techniques or standards, the complexity of the coded video sequence can also be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstructed sample rate (measured, for example, in megasamples per second), maximum reference picture size, etc. The limits set by the level may in some cases be further restricted by the virtual reference decoder (HRD) specification and the metadata of the HRD buffer management signaled in the coded video sequence.
[0038] In one embodiment, the receiver (310) can receive additional (redundant) data along with the encoded video. The additional data can be included as part of the (one or more) coded video sequences. The additional data can be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, a temporal layer, a spatial layer, or an SNR enhancement layer, a redundant slice, a redundant picture, a forward error correction code, etc.
[0039] Figure 4 shows an exemplary functional block diagram of a video encoder (203) associated with a video source (201) according to an embodiment of the present disclosure. The video encoder (203) may include, for example, an encoder that is a source coder (430), a coding engine (432), a (local) decoder (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy coder (445), a controller (450), and a channel (460).
[0040] The encoder (203) may receive video samples from a video source (201) (which is not part of the encoder) that may capture (one or more) video images to be coded by the encoder (203). The video source (201) may provide the source video sequence to be coded by the encoder (203) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCB, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media supply system, the video source (201) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that give motion when viewed in sequence. Each picture itself may be organized as a spatial array of pixels, and each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0041] According to one embodiment, the encoder (203) can code and compress pictures of the source video sequence into a coded video sequence (443) in real time or under any other arbitrary time constraints required by the application. Enforcing an appropriate coding speed is one function of the controller (450). The controller (450) may also control other functional units as described hereinafter and may be functionally coupled to these units. The coupling is not shown for clarity. The parameters set by the controller (450) can include rate control related parameters (such as picture skip, quantizer, lambda value of rate distortion optimization techniques), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller (450) as they may relate to the video encoder (203) optimized for a specific system design.
[0042] Some video encoders operate in a manner that is readily recognizable by those skilled in the art as a "coded loop." As an overly simplified explanation, the coded loop can consist of an encoding portion of a source coder (430) that is responsible for creating symbols based on an input picture to be coded and (one or more) reference pictures, and a (local) decoder (433) incorporated in an encoder (203) that reconstructs symbols to create sample data that a (remote) decoder would also create when the compression between the symbols and the coded video bitstream is reversible in a particular video compression technique. The reconstructed sample stream can be input into a reference picture memory (434). Since bit-exact results are obtained regardless of the location of the decoder (local or remote) by decoding the symbol stream, the contents of the reference picture memory are also bit-exact between the local encoder and the remote encoder. In other words, the prediction portion of the encoder "sees" the same sample values as reference picture samples that the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronicity (and the resulting drift if synchronicity cannot be maintained, for example due to channel errors) is known to those skilled in the art.
[0043] The operation of the "local" decoder (433) may be the same as that of the "remote" decoder (210), which has already been described in detail above in connection with FIG. 3. However, since the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder (445) and the parser (320) can be reversible, the entropy decoding portion of the decoder (210) including the channel (312), the receiver (310), the buffer (315), and the parser (320) may not be fully implemented in the local decoder (433).
[0044] What can be said at this point is that any decoder technology other than the parse / entropy decoding existing in the decoder may need to exist in substantially the same functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technologies can be omitted since they can be the reverse of the decoder technologies described comprehensively. More detailed explanations are needed and will be described below only in some areas.
[0045] As part of the operation, the source coder (430) may perform motion compensation predictive coding to predictively code an input frame by referring to one or more previously coded frames from a video sequence designated as a "reference frame". In this way, the coding engine (432) codes the difference between a pixel block of the input frame and a pixel block of the (one or more) reference frames that can be selected as the (one or more) prediction references to the input frame.
[0046] The local video decoder (433) may decode the coded video data of a frame that can be designated as a reference frame based on the symbols created by the source coder (430). The operation of the coding engine (432) may advantageously be an irreversible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence can usually be a reproduction of the source video sequence with some errors. The local video decoder (433) may reproduce the decoding process that can be performed by the video decoder for the reference frame and store the reconstructed reference frame in the reference picture memory (434). In this way, the encoder (203) may locally store a copy of the reconstructed reference frame having the same content as the reconstructed reference frame obtained by the remote video decoder (without transmission errors).
[0047] Predictor (435) can perform predictive search of the coding engine (432). That is, for a new frame to be coded, the predictor (435) can obtain sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors and block shapes that can serve as appropriate predictive references for the new picture, and search the reference picture memory (434). The predictor (435) can operate on one sample block for each pixel block to find an appropriate predictive reference. In some cases, the input picture can have predictive references drawn from a plurality of reference pictures stored in the reference picture memory (434), as determined by the search results obtained by the predictor (435).
[0048] Controller (450) can manage the coding operations of the video coder (430), including, for example, setting parameters and subgroup parameters used for encoding video data. The outputs of all the aforementioned functional units can be entropy-coded by the entropy coder (445). The entropy coder converts the symbols generated by various functional units into a coded video sequence by reversibly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable-length coding, arithmetic coding, etc.
[0049] The transmitter (440) may buffer for transmission via a communication channel (460), which may be a hardware / software link to a storage device that will store the encoded video data, the (one or more) coded video sequences created by the entropy coder (445). The transmitter (440) may merge the coded video data from the video coder (430) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown). The controller (450) may manage the operation of the encoder (203). During coding, the controller (450) may assign a particular coded picture type to each coded picture, which may affect the coding technique applicable to each picture. For example, a picture may often be assigned as an intra picture (I picture), a predicted picture (P picture), or a bidirectional predicted picture (B picture).
[0050] An intra picture (I picture) may be a picture that can be coded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs allow for various types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those skilled in the art will recognize these variations of I pictures and their respective uses and characteristics.
[0051] A predicted picture (P picture) may be a picture that can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0052] A bi-directional predicted picture (B picture) can be coded and decoded using intra prediction or inter prediction that uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0053] A source picture can generally be spatially subdivided into a plurality of sample blocks (e.g., blocks of samples of 4×4, 8×8, 4×8, or 16×16 each) and coded block by block. The blocks can be predictively coded by referring to other (already coded) blocks as determined by the coding assignment applied to each block of the picture. For example, blocks of an I picture may be coded non-predictively or predictively by referring to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be coded non-predictively via spatial prediction or via temporal prediction referring to one previously coded reference picture. Pixel blocks of a B picture can be coded non-predictively via spatial prediction or via temporal prediction referring to one or two previously coded reference pictures.
[0054] The video coder (203) can perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec.H.265. In its operation, the video coder (203) can perform various compression operations including predictive coding operations that utilize temporal and spatial redundancies in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technology or standard being used.
[0055] In one embodiment, the transmitter (440) may transmit additional data along with the encoded video. The video coder (430) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual user utility information (VUI) parameter set fragments, and the like.
[0056] Before explaining specific aspects of embodiments of the present disclosure in more detail, some terms that are referred to in the remainder of this specification are introduced below.
[0057] Hereinafter, a "sub-picture" refers to a rectangular arrangement of samples, blocks, macroblocks, coding units, or similar entities that may be meaningfully grouped and independently coded at a modified resolution in some cases. One or more sub-pictures may form a picture. One or more coded sub-pictures may form a coded picture. One or more sub-pictures may be assembled into a picture, or one or more sub-pictures may be extracted from a picture. In certain environments, one or more coded sub-pictures may be assembled in the compressed domain without transcoding to the coded picture at the sample level, and in the same or certain other cases, one or more coded sub-pictures may be extracted from the coded picture in the compressed domain.
[0058] Hereinafter, "adaptive resolution change" (ARC) refers to a mechanism that enables changing the resolution of a picture or sub-picture within an encoded video sequence, for example, by resampling a reference picture. Hereinafter, an "ARC parameter" refers to the control information required to perform an adaptive resolution change and may include, for example, filter parameters, scaling factors, the resolution of the output and / or reference pictures, various control flags, and the like.
[0059] VP9 uses a four-way partitioning tree from the 64×64 level down to the 4×4 level with some additional restrictions for blocks 8×8 and below, as shown in the upper half of FIG. 5 which illustrates the partitioning of a 64×64 block (500). The partition designated as R refers to a recursive partitioning where the same partitioning tree is repeated at a lower scale until the minimum 4×4 level is reached.
[0060] AV1 not only expands the partitioning tree to a ten-way structure as shown in FIG. 5, but also increases the maximum size (referred to as a superblock in VP9 / AV1 terminology) to start from a 128×128 block (502). This partitioning includes 4:1 / 1:4 rectangular partitions that did not exist in VP9. In one example, none of the rectangular partitions can be further subdivided. Additionally, AV1 adds further flexibility in the use of partitions below the 8×8 level in that 2×2 chroma inter prediction is possible in certain cases.
[0061] In HEVC, a coding tree unit (CTU) can be divided into coding units (CUs) by using a quadtree structure represented as a coding tree so as to adapt to various local characteristics. The decision of whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction can be made at the CU level. Each CU can be further divided into one, two, or four prediction units (PUs) according to the PU partition type. Within one PU, the same prediction process may be applied, and related information can be sent to the decoder on a PU basis. After obtaining a residual block by applying a prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another quadtree structure such as the coding tree of the CU. One of the important features of the HEVC structure is that this structure has multiple partition concepts including CUs, PUs, and TUs. In HEVC, a CU or TU can only be of a square shape, while a PU can be of a square or rectangular shape for an inter-predicted block. In HEVC, one coding block may be further divided into four square sub-blocks, and a transform process may be executed for each sub-block (e.g., TU). Each TU can be recursively further divided into smaller TUs called residual quadtree (RQT) (e.g., using quadtree partitioning). At the picture boundary, HEVC may adopt implicit quadtree partitioning, so the block maintains quadtree partitioning until its size fits within the picture boundary.
[0062] In HEVC, a CTU can be divided into CUs by using a quadtree structure represented as a coding tree so as to conform to various local characteristics. The determination of whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction can be made at the CU level. Each CU can be further divided into one, two, or four PUs depending on the PU partition type. Within one PU, the same prediction process may be applied, and related information may be sent to the decoder on a PU basis. After obtaining a residual block by applying a prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another quadtree structure such as the coding tree of the CU. One of the important features of the HEVC structure is that this structure has multiple partition concepts including CUs, PUs, and TUs.
[0063] The QTBT structure removes the concepts of multiple partition types (e.g., the QTBT structure removes the separation of the concepts of CU, PU, and TU), enhancing the flexibility of the CU partition shape. In the QTBT block structure, the CU can have either a square or rectangular shape. As shown in FIGS. 6(A) and 6(B), a coding tree unit (CTU) can first be partitioned by a quadtree structure. A quadtree leaf node can be further partitioned by a binary tree structure. There can be two types of binary tree partitions: symmetric horizontal partition and symmetric vertical partition. The binary tree leaf node may be called a coding unit (CU), and its segmentation can be used for prediction and transformation processing without further partitioning. This means that the CU, PU, and TU have the same block size in the QTBT coding block structure. In JEM, a CU may be composed of coding blocks (CBs) of different color components. For example, one CU may include one luma CB and two chroma CBs in the case of P slices and B slices in a 4:2:0 chroma format, and may sometimes consist of a single-component CB. For example, in the case of an I slice, one CU includes only one luma CB or only two chroma CBs.
[0064] The following parameters are defined for the QTBT partitioning method. - CTU size: The size of the quadtree root node, the same concept as in HEVC - MinQTSize: The minimum allowable quadtree leaf node size - MaxBTSize: The maximum allowable binary tree root node size - MaxBTDepth: The maximum allowable binary tree depth - MinBTSize: The minimum allowable binary tree leaf node size
[0065] In an example of the QTBT partitioning structure, the CTU size may be set as 128×128 luma samples, there are two corresponding chroma samples of 64×64 blocks, MinQTSize may be set as 16×16, MaxBTSize may be set as 64×64, MinBTSize (both width and height) may be set to 4×4, and MaxBTDepth may be set to 4. The quadtree partitioning may be first applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes may have sizes ranging from 16×16 (e.g., MinQTSize) to 128×128 (e.g., CTU size). When the leaf quadtree node is 128×128, since the size exceeds MaxBTSize (e.g., 64×64), the node may not be further divided by the binary tree. Otherwise, the leaf quadtree node may be further partitioned by the binary tree. Therefore, the quadtree leaf node may also be the root node of a binary tree with a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (e.g., 4), no further division is considered. When the width of the binary tree node is equal to MinBTSize (e.g., 4), no further horizontal division is considered. Similarly, when the height of the binary tree node is equal to MinBTSize, no further vertical division is considered. The leaf nodes of the binary tree are further processed by prediction and transformation processing without further partitioning. In JEM, the maximum CTU size is 256×256 luma samples.
[0066] FIG. 6(A) shows an example of block partitioning using QTBT, and FIG. 6(B) shows the corresponding tree representation. The solid lines indicate quadtree divisions, and the dotted lines indicate binary tree divisions. At each division (e.g., non-leaf) node of the binary tree, one flag indicating the division type (e.g., horizontal or vertical) used is signaled, where 0 indicates a horizontal division and 1 indicates a vertical division. In the case of quadtree division, since the quadtree division always divides the block in both the horizontal and vertical directions to generate four sub-blocks of the same size, there is no need to specify the division type.
[0067] In addition, the QTBT scheme can support the flexibility for the luma and chroma to have separate QTBT structures. Currently, for P slices and B slices, the luma CTB and chroma CTB within one CTU share the same QTBT structure. However, for I slices, the luma CTB is partitioned into CUs by the QTBT structure, and the chroma CTB is partitioned into chroma CUs by a different QTBT structure. This means that the CU of an I slice may be composed of coding blocks of a luma component or coding blocks of two chroma components, and the CU of a P slice or B slice may be composed of coding blocks of all three color components.
[0068] In HEVC, the inter prediction of small blocks may be restricted such that dual prediction is not supported for 4×8 blocks and 8×4 blocks in order to reduce the memory access for motion compensation, and the inter prediction is not supported for 4×4 blocks. In the QTBT implemented in JEM-7.0, these restrictions can be removed.
[0069] In VVC, a multi-type tree (MTT) structure may be included, which further adds horizontal and vertical center-side ternary trees on top of the QTBT as shown in FIGS. 7(A) and 7(B).
[0070] The main advantages of the ternary tree partitioning include (i) including the complement to the quadtree and binary tree partitioning, and the ternary tree partitioning can capture the object located at the block center while the quadtree and binary tree always divide along the block center, and (ii) the width and height of the proposed ternary tree partitions are always powers of 2 so that no additional transformation is required. The design of the two-level tree is mainly motivated by the reduction of complexity. The complexity of the tree traversal is T D where T represents the number of split types and D represents the depth of the tree.
[0071] In HEVC, a bi-predicted signal can be generated by averaging two predicted signals obtained from two different reference pictures and / or by using two different motion vectors. In VVC, the bi-prediction mode can be extended beyond simple averaging to enable weighted averaging of two predicted signals. Equation (1) P bi-pred =((8 - w)*P0 + w*P1 + 4) >> 3
[0072] In weighted average bi-prediction, five weights w ∈ {-2, 3, 4, 5, 10} can be allowed. For each bi-predicted CU, the weight w can be determined in one of the following two ways: (1) For non-merge CUs, the weight index is signaled after the motion vector difference, and (2) For merge CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. BCW can be applied only to CUs having 256 or more luma samples (e.g., CU width × CU height is 256 or more). In the case of low-delay pictures, all five weights can be used. In the case of non-low-delay pictures, only three weights (w ∈ {3, 4, 5}) can be used.
[0073] In the encoder, a fast search algorithm can be applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, unequal weights can be checked conditionally only for 1-pel and 4-pel motion vector accuracies when the current picture is a low-delay picture. When combined with Affine, Affine ME can be performed for unequal weights only if the Affine mode is selected as the current best mode. When the two reference pictures in bi-prediction are the same, unequal weights can be checked conditionally only. Depending on the POC distance, coding QP, and temporal level between the current picture and its reference picture, unequal weights may not need to be searched when certain conditions are met.
[0074] The BCW weight index can be coded using a bypass-coded bin following a context-coded bin. The first context-coded bin may indicate whether equal weights are used. If unequal weights are used, additional bins may be signaled using bypass coding to indicate which unequal weights are used. Weighted prediction (WP) is a coding tool supported by the H.264 / AVC and HEVC standards for efficiently coding video content using fading. Support for WP has also been added to the VVC standard. WP may allow weight parameters (weights and offsets) to be signaled for each reference picture in each of the reference picture lists L0 and L1. In that case, during motion compensation, the weights and offsets of the corresponding reference picture may be applied. WP and BCW can be designed for different types of video content. To avoid the interaction between WP and BCW that complicates the VVC decoder design, when a CU uses WP, the BCW weight index is not signaled and w is assumed to be 4 (i.e., equal weights are applied).
[0075] For a merge CU, the weight index can be inferred from neighboring blocks based on the merge candidate index. This feature can be applied to both the normal merge mode and the inherited affine merge mode. For the constructed affine merge mode, the affine motion information can be constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine merge mode can be set equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be applied to a CU together. When a CU is coded in the CIIP mode, the BCW index of the current CU can be set to 2 (e.g., equal weights).
[0076] In the current design of BCW, the weighting applied to the two prediction blocks is either explicitly signaled or inherited from neighboring blocks. However, all samples within a prediction block share the same weighting. This sharing is sub-optimal because there can be statistical variations at different positions within the prediction block. Therefore, when BCW is applied to coded blocks, a sample-adaptive weighting (or position-dependent weighting) can be used to derive the final predictor. This concept can also be extended to situations where the weighting can be estimated using reconstructed (or predicted) samples in the vicinity of the coded blocks in order to save signaling overhead.
[0077] Embodiments of the present disclosure relate to a set of advanced image and video coding techniques. More specifically, embodiments of the present disclosure relate to a dual prediction method that uses sample-adaptive weights for inter-coding. Embodiments of the present disclosure can be applied to dual prediction motion compensation on top of VVC or composite prediction mode on top of AV1 because both dual prediction motion compensation and composite prediction mode use multiple reference frames.
[0078] Embodiments of the present disclosure may be used separately or combined in any order. Further, each of the embodiments that utilize an encoder or a decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. Hereinafter, the term block may be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU. The term block here may also be used to refer to a transform block.
[0079] Reconstructed samples in the vicinity of the current block, forward prediction block, and / or backward prediction block may also be referred to as the template of the current block, the template of the forward prediction block, and / or the template of the backward prediction block. An example of a template is shown in FIG. 8, which shows the current block (800), P0 block (802), and P1 block (804) together with their corresponding templates. The template can represent the reconstructed samples in the vicinity shown as the texture portion.
[0080] In some embodiments, the weighting applied to the prediction blocks in list 0 and / or list 1 in dual-prediction motion compensation may depend on the position of the samples within the prediction block. In some embodiments (first weighting mode), a group of weighting patterns may be predefined, and an index value may be associated with each weighting pattern within the group. The index value may be signaled in the bitstream. The decoder may apply the weighting pattern associated with the index for dual-prediction motion compensation.
[0081] In some embodiments (second weighting mode), a group of weighting patterns is predefined, and the weighting pattern that minimizes a predefined cost measure calculated using the templates of the current block and the prediction blocks may be selected without signaling. In both the encoder and the decoder, the weighting may be calculated directly using the reconstructed samples in the vicinity of the current block as well as the reconstructed samples in the forward and / or backward vicinity.
[0082] In one example, the weighting may be derived using the minimum mean square error based on the reconstructed samples in the vicinity. The samples within the templates of P0, P1, and the current block are vectors
Number
Number
[0083] In another example, the weighting may be derived using the minimum mean square error based on the neighboring reconstruction samples. The samples in the templates of P0, P1, and the current block are vectors [Number] , [Number] and [Number] can be represented as. To find the best weights a and (1 - a) applied to P0 and P1 to generate the prediction block, the following cost can be minimized: [Number] Here, N is the total number of samples in the template. The solution can be given as follows: [Number] [Number] is.
[0084] In some embodiments (the third weighting mode), the weighting pattern of the current block may be inherited from neighboring blocks according to the weighting pattern. For example, if the current block is coded in NEARMV or merge mode, the weighting pattern may be inherited from one of the neighboring blocks. The rules for selecting neighboring blocks may be predefined.
[0085] In some embodiments, the final prediction value of the current block may be generated using one of the aforementioned first weighting mode, second weighting mode, and third weighting mode. For example, the bitstream may include a flag or indicator indicating one of the first weighting mode, second weighting mode, and third weighting mode of the block.
[0086] In some embodiments, one flag may be signaled to indicate whether equal weights are applied to combine the prediction samples of list 0 and list 1. If unequal weights are selected / indicated, one of the first weighting mode and the second weighting mode may be further applied to indicate the weighting for each sample.
[0087] In some embodiments, one flag may be signaled to indicate whether equal weights are applied to combine the prediction samples of list 0 and list 1. If unequal weighting is selected / indicated, a second flag may be signaled (or derived) to indicate whether a block-level weighting factor (e.g., all samples within one block share the same weighting factor) or a sample-position-dependent weighting factor is used. If a sample-position-dependent weighting factor is used / indicated, one of the first weighting mode and the second weighting mode may be further applied.
[0088] In some embodiments, the weight values applied to the samples in the prediction blocks of list 0 and list 1 are for a given sample (pij shown as) and the central sample (p c shown as) and depends on the distance between them. As an example, the distance may be p ij and p c and may be measured by the maximum absolute value of the difference between the abscissa and ordinate of. FIG. 9 shows a block (900) divided into a plurality of sub-blocks. Each sub-block may be referred to as a sample. As shown in FIG. 9, samples labeled with the same index value may be applied with the same weight value.
[0089] In another example, the distance may be p ij and p c and may be measured by the quantization distance value between the abscissa and ordinate of. As shown in FIG. 10, samples of the block (1000) labeled with the same index value may be applied with the same weight value.
[0090] In some embodiments, the position of the sample closer to the center position of the block may be associated with a weighting that is further (or closer) from the equal weighting (0.5). In some embodiments, the position of the sample closer to the template position (the neighbor above or to the left of the current block) may be associated with a weighting that is further (or closer) from the equal weighting (0.5). In some embodiments, there may be multiple patterns of weighting for each sample, and the selection may be signaled or derived implicitly.
[0091] In some embodiments, the weight value may depend on the abscissa. In some embodiments, the weight value may depend on the ordinate. In some embodiments, the weighting may depend on the sum or difference between the abscissa and the ordinate.
[0092] In some embodiments, the above-described weighting method may be excluded from frame or higher-level weighting prediction, BDOF, wedge-based prediction (or GPM), or DMVR (but not limited thereto). That is, if those modes are enabled for the current coding block, a block-level simple average may be used instead of the introduced adaptive sample weighting.
[0093] FIG. 11 shows a flowchart of an embodiment of a process (1100) for performing dual prediction using adaptive weighting. The process (1100) may be performed by a decoder such as decoder (210). The process may start with an operation (1102) in which a coded video bitstream is received. The bitstream may include a current picture, a first reference picture, and a second reference picture. The current picture may include a current block divided into a plurality of sub-blocks.
[0094] The process proceeds to operation (1104), where a first motion vector is determined that points to a first sub-block of a first block in the first reference picture from at least one sub-block within the current block. The process proceeds to operation (1106), where a second motion vector is determined that points to a second sub-block of a second block in the second reference picture from at least one sub-block within the current block.
[0095] Then, the process proceeds to operation (1108), where a weighting pattern is selected based on a predetermined condition. For example, one of the first weighting mode, the second weighting mode, and the third weighting mode described above may be selected. The process proceeds to operation (1110), where the first sub-block and the second sub-block are weighted based on the selected weighting pattern. The process proceeds to operation (1112), where at least one sub-block is decoded based on the weighted first sub-block and the weighted second sub-block.
[0096] The techniques of the embodiments of the present disclosure described above can be implemented as computer software physically stored on one or more computer-readable media using computer-readable instructions. For example, FIG. 12 shows a computer system (1200) suitable for implementing embodiments of the disclosed subject matter.
[0097] The computer software can be coded using any suitable machine code or computer language that can be subject to mechanisms such as assembly, compilation, and linking to create code including instructions that can be executed directly, or via interpretation, execution of microcode, etc., by a computer central processing unit (CPU), a graphics processing unit (GPU), etc.
[0098] The instructions can be executed on various types of computers or computer components including, for example, personal computers, tablet computers, servers, smartphones, game consoles, Internet of Things devices, etc.
[0099] The components shown in FIG. 12 with respect to the computer system (1200) are exemplary in nature and are not intended to suggest any limitation regarding the use or functionality scope of the computer software implementing the embodiments of the present disclosure. Also, regarding the configuration of the components, no reliance or requirement should be construed with respect to any one or combination of the components shown in the exemplary embodiments of the computer system (1200).
[0100] The computer system (1200) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users, for example, through tactile input (such as keystrokes, swipes, movements of a data glove, etc.), voice input (such as voice, clapping, etc.), visual input (such as gestures, etc.), and olfactory input (not shown). The human interface device can also be used to capture specific media that is not necessarily directly related to conscious input by humans, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (including 2D video, 3D video such as stereoscopic video, etc.).
[0101] The input human interface device may include one or more of a keyboard (1201), a mouse (1202), a trackpad (1203), a touch screen (1210), a data glove, a joystick (1205), a microphone (1206), a scanner (1207), and a camera (1208) (only one of each is shown).
[0102] The computer system (1200) may also include some kind of human interface output device. Such a human interface output device can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such a human interface output device may include a tactile output device (for example, tactile feedback by a touch screen (1210), a data glove, or a joystick (1205), although there may also be a tactile feedback device that does not function as an input device). For example, such a device may include an audio output device (such as a speaker (1209), headphones (not shown), etc.), a visual output device (such as a screen (1210) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, regardless of whether it has a touch screen input function, and regardless of whether it has a tactile feedback function, some of which can output two-dimensional visual output or output beyond three dimensions through means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown).
[0103] The computer system (1200) can also include an optical medium including a CD / DVD ROM / RW (1220) with a CD / DVD or similar medium (1221), a thumb drive (1222), a removable hard drive or solid state drive (1223), legacy magnetic media such as tapes and floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc., and storage devices and related media that can be handled by humans.
[0104] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.
[0105] The computer system (1200) can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital networks for television including cable television, satellite television and terrestrial television, vehicle and industrial including CANBus, etc. Certain networks generally require an external network interface adapter connected to a specific general-purpose data port or peripheral bus (1249) (such as a USB port of the computer system (1200)), while others are generally integrated into the core of the computer system (1200) by connection to the system bus as described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1200) can communicate with a counterpart. Such communication can be, for example, unidirectional, receive only (such as broadcast television), transmit only unidirectional (such as CANbus to a specific CANbus device), or bidirectional using a local or wide area digital network. Such communication can include communication to a cloud computing environment (1255). Specific protocols and protocol stacks can be used for each of those networks and network interfaces as described above.
[0106] The aforementioned human interface device, human-accessible storage device, and network interface (1254) can be connected to the core (1240) of the computer system (1200).
[0107] The core (1240) can include one or more central processing units (CPUs) (1241), a graphics processing unit (GPU) (1242), a special programmable processing device in the form of a field programmable gate array (FPGA) (1243), a hardware accelerator (1244) for specific tasks, etc. These devices may be connected through a system bus (1248) together with a read-only memory (ROM) (1245), a random access memory (1246), an internal hard drive that is not accessible to the user, an internal mass storage such as an SSD (1247). In some computer systems, the system bus (1248) can be accessible in the form of one or more physical plugs, enabling expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (1248) of the core or through a peripheral bus (1249). Architectures for peripheral buses include PCI, USB, etc. A graphics adapter (1250) may be included in the core (1240).
[0108] The CPU (1241), GPU (1242), FPGA (1243) and accelerator (1244) can execute some instructions that can be combined to form the above computer code. The computer code can be stored in the ROM (1245) or the RAM (1246). Transient data can also be stored in the RAM (1246), while persistent data can be stored, for example, in the internal mass storage (1247). By using a cache memory that can be closely associated with one or more CPUs (1241), GPUs (1242), mass storage (1247), ROM (1245), RAM (1246), etc., it can be made possible to perform high-speed storage and reading on any of the memory devices.
[0109] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure or can be of the kinds available to and well known to those of ordinary skill in the computer software arts.
[0110] By way of example and not limitation, a computer system (1200) having an architecture, and in particular a core (1240), can function as a result of software executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.) in one or more tangible computer-readable media. Such computer-readable media can be not only media associated with the above-introduced user-accessible mass storage, but also specific storage of the core (1240) with non-transitory characteristics, such as internal core mass storage (1247) and ROM (1245). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1240). The computer-readable media can include one or more memory devices or chips in response to individual requests. The software can cause the core (1240), and in particular a processor (including a CPU, GPU, FPGA, etc.) within the core (1240), to execute a specific process or a specific part of a specific process described herein, which includes determining a data structure stored in the RAM (1246) and modifying such a data structure according to a process defined by the software. In addition to or instead of this, the computer system can function as a result of logic by connection or be implemented in other ways in a circuit (such as an accelerator (1244)) to function, which can operate instead of or in cooperation with software to execute a specific process or a specific part of a specific process described in this application. If necessary, when described as "software", it can include logic and vice versa. References to computer-readable media can, if necessary, include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0111] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementation forms.
[0112] It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed herein are examples of exemplary approaches. It is understood that, based on design preferences, the specific order or hierarchy of blocks in a process / flowchart can be rearranged. Also, some blocks may be combined, omitted, etc. The appended method claims present the elements of the various blocks in a sample order and are not meant to be limited to the specific order or hierarchy presented.
[0113] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of integration of technical details. Further, one or more of the components described above may be stored on a computer-readable medium and implemented as instructions executable by at least one processor (and / or may include at least one processor). The computer-readable medium may include a computer-readable non-transitory storage medium having computer-readable program instructions for causing a processor to perform operations.
[0114] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes the following, namely, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or raised structures in grooves in which instructions are recorded, and any suitable combination of the foregoing. A computer-readable storage medium as used herein should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0115] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical transmission fiber, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each respective computing / processing device.
[0116] The computer-readable program code / instructions for performing the operations may be in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in an object-oriented programming language such as Smalltalk or C++, and a procedural programming language such as the "C" programming language or a similar programming language. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for performing the aspects or operations.
[0117] These computer-readable program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, causing the instructions executed by the processor of the computer or other programmable data processing apparatus to create means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram, thereby generating a machine. These computer-readable program instructions can also be stored in a computer-readable storage medium having stored instructions, the computer-readable storage medium including a product that includes instructions for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram, such that the computer, programmable data processing apparatus, and / or other devices are instructed to function in a particular manner.
[0118] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operation steps are executed on the computer, other programmable apparatus, or other device to generate a computer-implemented process, causing the instructions executed on the computer, other programmable apparatus, or other device to implement the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0119] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of one or more executable instructions for implementing the specified (one or more) logical functions. The methods, computer systems, and computer-readable media may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to those shown in the figures. In some alternative implementations, the functions described in the blocks may be executed in an order different from that described in the figures. For example, two blocks shown in succession may actually be executed simultaneously or substantially simultaneously, or the blocks may be executed in the reverse order depending on the related functions, in some cases. It should also be noted that each block of the block diagrams and / or flowchart diagrams, as well as combinations of blocks of the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified functions or operations or that realizes a combination of dedicated hardware and computer instructions.
[0120] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. It is understood that the actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting of the embodiments. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.
[0121] The above disclosure also encompasses the embodiments listed below.
[0122] (1) A method executed by at least one processor of a video decoder, the method comprising: receiving a coded video bitstream including a current picture, a first reference picture, and a second reference picture, wherein the current picture includes a current block divided into a plurality of sub-blocks; determining that the current picture is predicted using a bi-prediction mode or a composite prediction mode based on the first reference picture and the second reference picture; obtaining a plurality of predetermined weighting patterns, each weighting pattern being signaled as an index value; selecting a weighting pattern based on a predetermined condition; deriving a first weight to be applied to a first sub-block in the first reference picture and a second weight to be applied to a second sub-block in the second reference picture based on the index value corresponding to the selected weighting pattern; assigning the first weight to the first sub-block and the second weight to the second sub-block based on the selected weighting pattern; and decoding the current block by weighted bi-prediction based at least on the first sub-block weighted by the first weight and the second sub-block weighted by the second weight.
[0123] (2) The method according to feature (1), wherein the predetermined condition specifies a selection index included in the coded video bitstream, and the weighting pattern is selected from a plurality of weighting patterns based on the selection index.
[0124] (3) The method according to feature (1), wherein the predetermined condition specifies a minimum cost measurement value for selecting a weighting pattern from a plurality of weighting patterns, and the minimum cost measurement value is calculated based on a template associated with at least one sub-block, a template associated with the first sub-block, and a template associated with the second sub-block.
[0125] (4) The predetermined condition is the method according to feature (1), which indicates the weighting pattern of the neighboring sub-blocks in the vicinity of at least one sub-block.
[0126] (5) The predetermined condition includes: (i) a first weighting mode in which the weighting pattern is selected from a plurality of weighting patterns based on a selection index included in the bit stream; (ii) a second weighting mode in which the weighting pattern that minimizes a cost measurement value calculated based on a template associated with at least one sub-block, a template associated with a first sub-block, and a template associated with a second sub-block is selected from a plurality of weighting patterns; and (iii) a third weighting mode in which the weighting pattern is selected based on the weighting pattern of the neighboring sub-blocks in the vicinity of at least one sub-block. The method is according to feature (1), which indicates a plurality of weighting modes.
[0127] (6) The selection of one of the plurality of weighting modes is based on an indicator included in the bit stream. The method is according to feature (5).
[0128] (7) The bit stream includes a first flag indicating whether the selected weighting pattern applies equal weighting to the first sub-block and the second sub-block. Based on the determination that the first flag indicates unequal weighting, a second flag is included in or derived from the bit stream. The second flag indicates the selection of one of the first weighting mode and the second weighting mode. The method is according to feature (5).
[0129] (8) The bitstream includes a first flag indicating whether the selected weighting pattern applies equal weighting to the first sub-block and the second sub-block. Based on the determination that the first flag indicates unequal weighting, a second flag is included in or derived from the bitstream. The second flag indicates whether all of the at least one sub-blocks share the same weighting pattern. Based on the determination that the second flag indicates that not all of the sub-blocks share the same weighting pattern, one of a first weighting mode and a second weighting mode is selected, the method according to feature (7).
[0130] (9) The predetermined condition indicates the distance of at least one sub-block to the center of the current block, and the weighting pattern is selected based on the distance, the method according to feature (7).
[0131] (10) The distance is measured by the maximum absolute value of the difference in the horizontal and vertical coordinates between at least one sub-block and the center of the current block, the method according to feature (9).
[0132] (11) The distance is measured by the quantized distance value of the horizontal and vertical coordinates between at least one sub-block and the center of the current block, the method according to feature (9).
[0133] (12) A first distance between at least one sub-block and the center of the current block that is closer to the center of the current block than a second distance has a first weighting pattern that is closer to the equal weighting of the first sub-block and the second sub-block than a second weighting pattern associated with the second distance, the method according to feature (9).
[0134] The method according to any one of features (1) to (12), further comprising: determining a first motion vector pointing to a first sub-block of a first block in a first reference picture from at least one sub-block in a current block; and determining a second motion vector pointing to a second sub-block of a second block in a second reference picture from at least one sub-block in the current block.
[0135] At least one memory configured to store computer program code, and at least one processor configured to access the computer program code and operate as commanded by the computer program code, wherein the computer program code comprises: reception code configured to cause the at least one processor to receive a coded video bitstream including a current picture, a first reference picture, and a second reference picture, wherein the current picture includes a current block divided into a plurality of sub-blocks; determination code configured to cause the at least one processor to determine that the current picture is predicted using a dual prediction mode or a composite prediction mode based on the first reference picture and the second reference picture; acquisition code configured to cause the at least one processor to acquire a plurality of predetermined weighting patterns, each weighting pattern being signaled as an index value; selection code configured to cause the at least one processor to select a weighting pattern based on a predetermined condition; derivation code configured to cause the at least one processor to derive a first weight to be applied to a first sub-block in the first reference picture and a second weight to be applied to a second sub-block in the second reference picture based on the index value corresponding to the selected weighting pattern; assignment code configured to cause the at least one processor to assign the first weight to the first sub-block and the second weight to the second sub-block based on the selected weighting pattern; and decoding code configured to cause the at least one processor to decode the current block by weighted dual prediction based at least on the first sub-block weighted by the first weight and the second sub-block weighted by the second weight. A video decoder comprising at least one processor.
[0136] (15) The predetermined condition specifies a selection index included in the coded video bitstream, and the weighting pattern is selected from a plurality of weighting patterns based on the selection index, for the video decoder according to feature (14).
[0137] (16) The predetermined condition specifies a minimum cost measurement value for selecting a weighting pattern from a plurality of weighting patterns, and the minimum cost measurement value is calculated based on a template associated with at least one sub-block, a template associated with a first sub-block, and a template associated with a second sub-block, for the video decoder according to feature (14).
[0138] (17) The predetermined condition indicates the weighting pattern of a neighboring sub-block in the vicinity of at least one sub-block, for the video decoder according to feature (14).
[0139] (18) The predetermined condition includes: (i) a first weighting mode in which a weighting pattern is selected from a plurality of weighting patterns based on a selection index included in the bitstream; (ii) a second weighting mode in which a weighting pattern that minimizes a cost measurement value calculated based on a template associated with at least one sub-block, a template associated with a first sub-block, and a template associated with a second sub-block is selected from a plurality of weighting patterns; and (iii) a third weighting mode in which a weighting pattern is selected based on the weighting pattern of a neighboring sub-block in the vicinity of at least one sub-block, for the video decoder according to feature (14).
[0140] (19) The selection of one of the plurality of weighting modes is based on an indicator included in the bitstream, for the video decoder according to feature (18).
[0141] When executed by a processor in a video decoder, receiving a coded video bitstream including a current picture, a first reference picture, and a second reference picture, wherein the current picture includes a current block divided into a plurality of sub-blocks; determining that the current picture is predicted using a bi-prediction mode or a composite prediction mode based on the first reference picture and the second reference picture; obtaining a plurality of predetermined weighting patterns, each weighting pattern being signaled as an index value; selecting a weighting pattern based on a predetermined condition; deriving a first weight to be applied to a first sub-block in the first reference picture and a second weight to be applied to a second sub-block in the second reference picture based on the index value corresponding to the selected weighting pattern; assigning the first weight to the first sub-block and the second weight to the second sub-block based on the selected weighting pattern; and decoding the current block by weighted bi-prediction based at least on the first sub-block weighted by the first weight and the second sub-block weighted by the second weight. A non-transitory computer-readable medium storing instructions for causing a processor to execute a method including the steps above.
Description of the Signs
[0142] 100 Communication system, 110 First terminal, 120 Second terminal, 130 Terminal, 140 Terminal, 150 Communication network, 200 Streaming system, 201 Video source, camera, 202 Uncompressed video sample stream, 203 Video coder, video source, video encoder, 204 Video bitstream, 205 Streaming server, 206 Streaming client, 209 Video bitstream, 210 Video decoder, 211 Video sample stream, 212 Display, 213 Capture subsystem, 310 Receiver, 312 Channel, 315 Buffer memory, 320 Entropy decoder / parser, 321 Symbol, 351 Scaler / inverse transform unit, 352 Intra-picture prediction unit, 353 Motion compensation prediction unit, 355 Aggregator, 356 Loop filter unit, 357 Reference picture memory, 358 Current picture memory, 430 Source coder, video coder, 432 Coding engine, 433 Local video decoder, 434 Reference picture memory, 435 Predictor, 440 Transmitter, 443 Coded video sequence, 445 Entropy coder, 450 Controller, 460 Communication channel, 500 64×64 block, 502 128×128 block, 800 Current block, 802 P0 block, 804 P1 block, 900 Block, 1000 Block, 1100 Process, 1200 Computer system, 1201 Keyboard, 1202 Mouse, 1203 Trackpad, 1205 Joystick, 1206 Microphone, 1207 Scanner, 1208 Camera, 1209 Audio output device speaker, 1210 Touch screen, 1220 CD / DVD ROM / RW, 1221 Medium, 1222 Thumb drive, 1223 Solid state drive, 1240 Core, 1241 Central processing unit (CPU), 1242 Graphics processing unit (GPU), 1243 Field programmable gate array (FPGA), 1244 Hardware accelerator, 1245 Read only memory (ROM), 1246 Random access memory, 1247 On-core large capacity storage, 1248System bus, 1249 Peripheral bus, 1250 Graphics adapter, 1254 Network interface, 1255 Cloud computing environment
Claims
1. 1. A method executed by at least one processor of a video decoder, the method comprising: receiving a coded video bitstream including a current picture, a first reference picture, and a second reference picture, the current picture including a current block divided into a plurality of sub-blocks; determining that the current picture is predicted using a bi-prediction mode or a combined prediction mode based on the first reference picture and the second reference picture; obtaining a plurality of predetermined weighting patterns, each weighting pattern being signaled as an index value; selecting a weighting pattern based on a predetermined condition; deriving a first weight to be applied to a first sub-block in the first reference picture and a second weight to be applied to a second sub-block in the second reference picture based on the index value corresponding to the selected weighting pattern; assigning the first weight to the first sub-block and the second weight to the second sub-block based on the selected weighting pattern; and decoding the current block by weighted bi-prediction based at least on the first sub-block weighted by the first weight and the second sub-block weighted by the second weight.
2. The method of claim 1 , wherein the predetermined condition specifies a selection index to be included in the coded video bitstream, and the weighting pattern is selected from a plurality of the weighting patterns based on the selection index.
3. 2. The method of claim 1, wherein the predetermined condition specifies a minimum cost measure for selecting the weighting pattern from a plurality of weighting patterns, the minimum cost measure being calculated based on a template associated with the at least one sub-block, a template associated with the first sub-block, and a template associated with the second sub-block.
4. The method of claim 1 , wherein the predetermined condition dictates a weighting pattern for neighboring sub-blocks in a neighborhood of the at least one sub-block.
5. The predetermined condition is: (i) a first weighting mode, in which the weighting pattern is selected from a plurality of weighting patterns based on a selection index included in the coded video bitstream; (ii) a second weighting mode, in which the weighting pattern that minimizes a cost measure calculated based on the template associated with the at least one sub-block, the template associated with the first sub-block, and the template associated with the second sub-block is selected from a plurality of weighting patterns; and and (iii) a third weighting mode in which the weighting pattern is selected based on weighting patterns of neighboring sub-blocks in a neighborhood of the at least one sub-block.
6. The method of claim 5 , wherein the selection of the one of the plurality of weighting modes is based on an indicator included in the coded video bitstream.
7. 6. The method of claim 5, wherein the coded video bitstream includes a first flag indicating whether the selected weighting pattern applies equal weighting to the first sub-block and the second sub-block, and wherein a second flag is included in or derived from the coded video bitstream based on a determination that the first flag indicates unequal weighting, the second flag indicating a selection of one of the first weighting mode and the second weighting mode.
8. the coded video bitstream includes a first flag indicating whether the selected weighting pattern applies equal weighting to the first sub-block and the second sub-block; a second flag is included in or derived from the coded video bitstream based on a determination that the first flag indicates unequal weighting, the second flag indicating whether all sub-blocks of the at least one sub-block share the same weighting pattern; 8. The method of claim 7, wherein one of the first weighting mode and the second weighting mode is selected based on a determination that the second flag indicates that not all of the sub-blocks share the same weighting pattern.
9. The method of claim 7 , wherein the predetermined condition indicates a distance of the at least one sub-block to a center of the current block, and the weighting pattern is selected based on the distance.
10. The method of claim 9 , wherein the distance is measured by the maximum absolute value of the difference between the horizontal and vertical coordinates of the at least one sub-block and the center of the current block.
11. The method of claim 9 , wherein the distance is measured by quantized distance values of abscissas and ordinates of the at least one sub-block and the center of the current block.
12. 10. The method of claim 9, wherein a first distance between the at least one sub-block and the center of the current block that is closer to the center of the current block than a second distance has a first weighting pattern that is closer to equal weighting of the first sub-block and the second sub-block than a second weighting pattern associated with the second distance.
13. determining a first motion vector from at least one sub-block in the current block to a first sub-block of a first block in the first reference picture; determining a second motion vector from the at least one sub-block in the current block to a second sub-block of a second block in the second reference picture; The method of claim 1 further comprising:
14. A video decoder configured to perform a method according to any one of claims 1 to 13.
15. A computer program product for causing at least one processor to carry out the method of any one of claims 1 to 13.