Method, apparatus and computer program for adaptive multiple hypothesis probability model for arithmetic coding
Patent Information
- Application Number
- JP2023564595
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-08
- Filing Date
- 2022-11-09
- Publication Date
- 2025-11-17
AI Technical Summary
Current coding standards for multiple hypothetical probability models in communication systems, such as those used in the AV1 video coding format, do not adequately consider frame type, block size, and prediction mode, leading to inefficiencies in arithmetic coding.
The implementation of adaptive multiple hypothetical probability modeling for arithmetic coding, which involves receiving a coded video bitstream, decoding syntax elements, selecting a probability update rate based on conditions like frame type or block size, updating probability models, and decoding blocks within pictures.
This approach enhances the accuracy of probability estimation and improves coding performance by adaptively updating probability models based on specific conditions within the video bitstream, leading to more efficient arithmetic coding.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims priority to U.S. Provisional Patent Application No. 63 / 318,588, filed March 10, 2022, and U.S. Patent Application No. 18 / 053,494, filed November 8, 2022, the disclosures of which are incorporated by reference in their entireties herein.
[0002] The present disclosure relates generally to communication systems, and more particularly, to a method and apparatus for adaptive multiple hypothesis probability models for arithmetic coding. [Background technology]
[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. This coding format was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium founded in 2015 that includes semiconductor companies, video-on-demand providers, video content producers, software developers, and web browser vendors. Many of the components of the AV1 project were contributed by previous research efforts by members of the Alliance. Individual contributors started experimental technology platforms several years ago, namely Xiph's / Mozilla's Daala, which already released its code in 2010, Google's experimental VP9 evolution project VP10, announced on September 12, 2014, and Cisco's Thor, published on August 11, 2015. Building on the VP9 code base, AV1 incorporates additional technologies, some of which were developed in these experimental formats. The first version 0.1.0 of the AV1 reference codec was published on April 7, 2016. The Alliance announced the release of the AV1 bitstream specification on March 28, 2018, along with a reference software-based encoder and decoder. The specification's effective version 1.0.0 was released on June 25, 2018. The specification's effective version 1.0.0 with Errata 1 was released on January 8, 2019. The AV1 bitstream specification includes a reference video codec. Current coding standards for the multiple hypothesis probability model do not adequately take into account other information such as frame type, block size, and prediction mode. Summary of the Invention [Means for solving the problem]
[0004] The following presents a simplified summary of one or more embodiments of the present disclosure in order to provide a basic understanding of such embodiments. This summary is not an extensive overview of all contemplated embodiments, and is not intended to identify key or critical elements of all embodiments or to delineate the scope of any or all embodiments. Its sole purpose is to present some concepts of one or more embodiments of the present disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[0005] SUMMARY OF THE DISCLOSURE A method, apparatus, and non-transitory computer-readable medium for adaptive multi-hypothesis probability modeling for arithmetic coding is disclosed by the present disclosure.
[0006] According to an exemplary embodiment, a method performed by at least one processor of a video decoder includes receiving a coded video bitstream including at least one picture and one or more syntax elements encoded according to a multiple hypothesis arithmetic coding. The method further includes decoding each syntax element from the one or more syntax elements based on the multiple hypothesis arithmetic coding. The method further includes selecting a probability update rate from a plurality of probability update rates based on a predetermined condition, the plurality of probability update rates including a first probability update rate higher than a second probability update rate. The method further includes updating at least one probability model utilized in the multiple hypothesis arithmetic coding based on the selected probability update rate. The method further includes decoding at least one block in the at least one picture based on the decoded one or more syntax elements.
[0007] According to an exemplary embodiment, a video decoder includes at least one memory configured to store computer program code and at least one processor configured to access the computer program code and operate as instructed by the computer program code. The computer program code includes a receiving code configured to cause the at least one processor to receive a coded video bitstream including at least one picture and one or more syntax elements encoded according to a multiple hypothesis arithmetic coding. The computer program code further includes a first decoding code configured to cause the at least one processor to decode each syntax element from the one or more syntax elements based on the multiple hypothesis arithmetic coding. The computer program code further includes a selection code configured to cause the at least one processor to select a probability update rate from a plurality of probability update rates based on a predetermined condition, the plurality of probability update rates including a first probability update rate higher than a second probability update rate. The computer program code further includes an update code configured to cause the at least one processor to update at least one probability model utilized in the multiple hypothesis arithmetic coding based on the selected probability update rate. The computer program code further includes second decoding code configured to cause the at least one processor to decode at least one block in the at least one picture based on the decoded one or more syntax elements.
[0008] According to an exemplary embodiment, a non-transitory computer-readable medium is stored with instructions that, when executed by a processor in a video decoder, cause the processor to perform a method including receiving a coded video bitstream including at least one picture encoded according to multiple hypothesis arithmetic coding and one or more syntax elements. The method further includes decoding each syntax element from the one or more syntax elements based on the multiple hypothesis arithmetic coding. The method further includes selecting a probability update rate from a plurality of probability update rates based on a predetermined condition, the plurality of probability update rates including a first probability update rate higher than a second probability update rate. The method further includes updating at least one probability model utilized in the multiple hypothesis arithmetic coding based on the selected probability update rate. The method further includes decoding at least one block in the at least one picture based on the decoded one or more syntax elements.
[0009] Additional embodiments will be set forth in the description that follows, and in part will be obvious from the description, and / or may be learned by practice of presented embodiments of the present disclosure.
[0010] The above and other aspects, features, and aspects of embodiments of the present disclosure will become apparent from the following description taken in conjunction with the accompanying drawings. [Brief description of the drawings]
[0011] [Figure 1] 1 is a schematic block diagram of a communication system in accordance with various embodiments of the present disclosure. [Diagram 2] 1 is a schematic block diagram of a communication system in accordance with various embodiments of the present disclosure. [Diagram 3] FIG. 2 is a schematic block diagram of a decoder, according to various embodiments of the present disclosure. [Figure 4] 2 is a block diagram of an encoder according to various embodiments of the present disclosure. [Diagram 5] 1 is an example flowchart of a process for decoding bins according to various embodiments of the present disclosure. [Figure 6] 1 is an example flow chart of a process for performing adaptive multi-hypothesis probability modeling for arithmetic coding. [Figure 7] FIG. 1 is a diagram of a computer system in accordance with various embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] The following detailed description of the exemplary embodiments refers to the accompanying drawings, in which the same reference numbers in different drawings may identify the same or similar elements.
[0013] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations. Moreover, one or more features or components of one embodiment may be incorporated or combined with another embodiment (or one or more features of another embodiment). Furthermore, in the flow charts and descriptions of the operations provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may occur (at least partially) simultaneously, and one or more operations may be reordered.
[0014] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not intended to limit the embodiments. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it will be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0015] Although particular feature combinations are recited in the claims and / or disclosed herein, these combinations are not intended to limit the disclosure of possible embodiments. Indeed, many of these features may be combined in ways not specifically recited in the claims and / or disclosed herein. Although each dependent claim listed below may depend directly on only one claim, the disclosure of possible embodiments includes each dependent claim in combination with all other claims in the claim set.
[0016] No element, act, or instruction used herein should be construed as critical or required unless expressly described as such. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." When only one item is intended, the term "one" or similar words are used. Also, as used herein, terms such as "has," "have," "having," "include," "including," and the like are intended to be open-ended terms. Furthermore, the phrase "based on" is intended to mean "based at least in part on," unless otherwise specified. Furthermore, phrases such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood as including only A, only B, or both A and B.
[0017] Throughout this specification, references to "one embodiment," "an embodiment," or similar language mean that a particular feature, structure, or characteristic described in connection with the illustrated embodiment is included in at least one embodiment of the solution. Thus, the phrases "in one embodiment," "in an embodiment," and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.
[0018] Furthermore, the described features, advantages, and characteristics of the present disclosure may be combined in any suitable manner in one or more embodiments. In light of the description herein, those skilled in the art will recognize that the present disclosure may be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments of the present disclosure.
[0019] FIG. 1 shows a simplified block diagram of a communication system (100) according to one embodiment of the present disclosure. The system (100) may include at least two terminals (110, 120) interconnected via a network (150). In the case of unidirectional data transmission, a first terminal (110) may code video data at a local location for transmission to a counterpart terminal (120) via the network (150). The second terminal (120) may receive the coded video data of the counterpart terminal from the network (150), decode the coded data, and display the reconstructed video data. Unidirectional data transmission may be common in media serving applications, etc.
[0020] 1 illustrates a second pair of terminals (130, 140) arranged to support bidirectional transmission of coded video, such as may occur during a video conference. For bidirectional transmission of data, each terminal (130, 140) may code video data captured at a local location for transmission to the other terminal over a network (150). Each terminal (130, 140) may also receive coded video data transmitted by the other terminal, decode the coded data, and display the recovered video data on a local display device.
[0021] In FIG. 1, the terminals (110-140) may be depicted as servers, personal computers, and smartphones, and / or any other type of terminal. For example, the terminals (110-140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (150) represents any number of networks that convey coded video data between the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of the network (150) may not be important to the operation of the present disclosure, unless described below.
[0022] 2 shows the arrangement of video encoders and decoders in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, such as, for example, video conferencing, digital television, storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.
[0023] As shown in FIG. 2, the streaming system (200) may include a capture subsystem (213) that may include a video source (201) and an encoder (203). The video source (201) may be, for example, a digital camera and may be configured to create an uncompressed video sample stream (202). The uncompressed video sample stream (202) may provide a high amount of data compared to an encoded video bitstream and may be processed by an encoder (203) coupled to the camera (201). The encoder (203) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream (204) may include a low amount of data compared to the sample stream and may be stored in a streaming server (205) for future use. One or more streaming clients (206) may access the streaming server (205) to obtain a video bitstream (209), which may be a copy of the encoded video bitstream (204).
[0024] In an embodiment, the streaming server (205) may also function as a Media-Aware Network Element (MANE). For example, the streaming server (205) may be configured to prune the encoded video bitstream (204) to tailor a potentially different bitstream to one or more of the streaming clients (206). In an embodiment, a MANE may be provided separately from the streaming server (205) in the streaming system (200).
[0025] The streaming client (206) may include a video decoder (210) and a display (212). The video decoder (210) may, for example, decode a video bitstream (209), which is an incoming copy of the encoded video bitstream (204), and create an outgoing video sample stream (211) that may be rendered on a display (212) or other rendering device (not shown). In some streaming systems, the video bitstreams (204, 209) may be encoded according to some video coding / compression standard. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. A video coding standard informally known as Versatile Video Coding (VVC) is under development. Embodiments of the present disclosure may be used in the context of VVC.
[0026] 3 illustrates an exemplary functional block diagram of a video decoder (210) attached to a display (212) according to one embodiment of the disclosure. The video decoder (210) may include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra prediction unit (352), a motion compensation prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (). In at least one embodiment, the video decoder (210) may include an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The video decoder (210) may also be embodied partially or entirely in software running on one or more CPUs with associated memory.
[0027] In this and other embodiments, the receiver (310) may receive one or more coded video sequences to be decoded by the decoder (210), one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences may be received from a channel (312), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (310) may receive the encoded video data along with other data, e.g., coded audio data and / or auxiliary data streams, that may be forwarded to a respective using entity (not shown). The receiver (310) may separate the coded video sequences from the other data. To combat network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereafter "parser"). When the receiver (310) is receiving data from a store-and-forward device of sufficient bandwidth and controllability, or from an isosynchronous network, the buffer (315) may not be used or may be small. For use with best-effort packet networks such as the Internet, the buffer (315) may be required and may be relatively large and of adaptive size.
[0028] The video decoder (210) may include a parser (320) for reconstructing symbols (321) from the entropy coded video sequence. These symbol categories include, for example, information used to manage the operation of the decoder (210) and potentially information for controlling a rendering device such as a display (212) that may be coupled to the decoder as shown in FIG. 2. The control information for the rendering device(s) may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (320) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser (320) may extract from the coded video sequence a set of subgroup parameters for at least one of a subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (320) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.
[0029] The parser (320) may perform entropy decoding / parsing operations on the video sequence received from the buffer (315) to create symbols (321). The reconstruction of the symbols (321) may involve a number of different units, depending on the type of video picture or portion thereof coded (inter-picture and intra-picture, inter-block and intra-block, etc.), and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by the parser (320). The flow of such subgroup control information between the parser (320) and the following units is not shown for clarity.
[0030] In addition to the functional blocks already mentioned, the decoder (210) may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:
[0031] One unit may be a scalar / inverse transform unit (351), which may receive quantized transform coefficients as well as control information including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. as symbol(s) (321) from the parser (320). The scalar / inverse transform unit (351) may output a block containing sample values that may be input to an aggregator (355).
[0032] In some cases, the output samples of the scaler / inverse transform (351) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) generates blocks of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the current (partially reconstructed) picture from a current picture memory (358). The aggregator (355) adds, possibly on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).
[0033] In other cases, the output samples of the scalar / inverse transform unit (351) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensated prediction unit (353) may access the reference picture memory (357) to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols (321) associated with the block, these samples may be added to the output of the scalar / inverse transform unit (351) by the aggregator (355) to generate output sample information (in this case referred to as residual samples or residual signals). The addresses in the reference picture memory (357) from which the motion compensated prediction unit (353) fetches the prediction samples may be controlled by a motion vector. The motion vector may be available to the motion compensated prediction unit (353) in the form of a symbol (321), which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory (357) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.
[0034] The output samples of the aggregator (355) may be subjected to various loop filtering techniques in a loop filter unit (356). Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video bitstream and made available to the loop filter unit (356) as symbols (321) from the parser (320), but may also be responsive to meta-information obtained during decoding of a coded picture or previous (in decoding order) portion of a coded video sequence, or to previously reconstructed and loop filtered sample values.
[0035] The output of the loop filter unit (356) may be a sample stream that may be output to a rendering device, such as a display (212), and may also be stored in a reference picture memory (357) for use in future inter-picture prediction.
[0036] Once fully reconstructed, a particular coded picture may be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture has been identified as a reference picture (e.g., by the parser (320)), the current reference picture may become part of the reference picture memory (357), and the new current picture memory may be reallocated before beginning reconstruction of the next coded picture.
[0037] The video decoder (210) may perform decoding operations according to a given video compression technique, which may be documented in a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense of adhering to the syntax of the video compression technique or standard, as specified in the video compression technique document or standard, specifically in a profile document therein. Also, to conform to some video compression techniques or standards, the complexity of the coded video sequence may be within a range prescribed by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further limited in some cases by a Hypothetical Reference Decoder (HRD) specification and HRD buffer management metadata signaled in the coded video sequence.
[0038] In one embodiment, the receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal, spatial, or SNR enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.
[0039] 4 shows an example functional block diagram of a video encoder (203) associated with a video source (201) according to one embodiment of this disclosure. The video encoder (203) may include an encoder, e.g., a source coder (430), a coding engine (432), a (local) decoder (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy coder (445), a controller (450), and a channel (460).
[0040] The encoder (203) may receive video samples from a video source (201) (not part of the encoder) that may capture the video image(s) to be coded by the encoder (203). The video source (201) may provide a source video sequence to be coded by the encoder (203) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...) and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media delivery system, the video source (201) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The picture itself may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc. used. Those skilled in the art can readily understand the relationship between pixels and samples. The following discussion focuses on samples.
[0041] According to one embodiment, the encoder (203) may code and compress pictures of a source video sequence into a coded video sequence (443) in real time or under any other time constraint required by the application. Enforcing an appropriate coding rate is one function of the controller (450). The controller (450) may also control and be operatively coupled to other functional units as described below. Coupling is not shown for clarity. Parameters set by the controller (450) may include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, ...), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art may readily identify other functions of the controller (450) as they may relate to a video encoder (203) optimized for a particular system design.
[0042] Some video encoders operate in what those skilled in the art will readily recognize as a "coding loop." As an oversimplified explanation, the coding loop may consist of an encoding part of a source coder (430) (responsible for creating symbols based on the input picture to be coded and the reference picture(s)) and a (local) decoder (433) built into the encoder (203) that reconstructs the symbols to create sample data that a (remote) decoder would also create when the compression between the symbols and the coded video bitstream is lossless in a particular video compression technique. The reconstructed sample stream may be input to a reference picture memory (434). The decoding of the symbol stream results in bit-exact results regardless of the location of the decoder (local or remote), so the contents of the reference picture memory are also bit-exact between the local and remote encoders. In other words, the predictive part of the encoder "sees" exactly the same sample values as the decoder would "see" when using prediction during decoding as reference picture samples. This basic principle of reference picture synchrony (and the resulting drift if synchrony cannot be maintained, for example due to channel errors) is known to those skilled in the art.
[0043] The operation of the "local" decoder (433) may be the same as that of the "remote" decoder (210), which has already been described in detail above in connection with Figure 3. However, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (445) and parser (320) may be lossless, the entropy decoding portion of the decoder (210), including the channel (312), receiver (310), buffer (315), and parser (320), may not be fully implemented in the local decoder (433).
[0044] An observation that can be made at this point is that any decoder technique, except for parsing / entropy decoding, present in the decoder may need to be present in substantially the same functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on the operation of the decoder. Descriptions of encoder techniques may be omitted, as they may be the inverse of the decoder techniques described generically. Only in certain areas are more detailed descriptions necessary and are provided below.
[0045] As part of its operation, the source coder (430) may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence designated as “reference frames.” In this manner, the coding engine (432) codes differences between pixel blocks of the input frame and pixel blocks of reference frame(s) that may be selected as predictive reference(s) to the input frame.
[0046] The local video decoder (433) may decode the coded video data of a frame that may be designated as a reference frame based on the symbols created by the source coder (430). The operation of the coding engine (432) may advantageously be a lossy process. When the coded video data may be decoded in a video decoder (not shown in FIG. 4), the reconstructed video sequence may be a copy of the source video sequence, usually with some errors. The local video decoder (433) may reproduce the decoding process that may be performed by the video decoder on the reference frame and store the reconstructed reference frame in the reference picture memory (434). In this way, the encoder (203) may locally store (without transmission errors) a copy of the reconstructed reference frame that has a common content with the reconstructed reference frame obtained by the far-end video decoder.
[0047] The predictor (435) may perform the prediction search of the coding engine (432). That is, for a new frame to be coded, the predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that may serve as suitable prediction references for the new picture. The predictor (435) may operate on one sample block per pixel block to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (435), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (434).
[0048] The controller (450) may manage the coding operations of the video coder (430), including, for example, setting parameters and subgroup parameters used to encode the video data. The output of all the aforementioned functional units may be entropy coded in the entropy coder (445). The entropy coder converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as, for example, Huffman coding, variable length coding, arithmetic coding, etc.
[0049] The transmitter (440) may buffer the coded video sequence(s) created by the entropy coder (445) in preparation for transmission over a communication channel (460), which may be a hardware / software link to a storage device that will store the coded video data. The transmitter (440) may merge the coded video data from the video coder (430) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown). The controller (450) may manage the operation of the encoder (203). During coding, the controller (450) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as intra-pictures (I-pictures), predicted pictures (P-pictures), or bidirectionally predicted pictures (B-pictures).
[0050] An intra picture (I-picture) may be a picture that can be coded and decoded without using any other frame in a sequence as a source of prediction. Some video codecs allow various types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.
[0051] A predictive picture (P picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict sample values for each block.
[0052] A bidirectionally predicted picture (B-picture) may be one that can be coded and decoded using intra- or inter-prediction that uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predicted picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0053] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). Pixel blocks of a P-picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Pixel blocks of a B-picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0054] The video coder (203) may perform coding operations in accordance with a given video coding technique or standard, such as ITU-T Rec. H.265. In its operations, the video coder (203) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.
[0055] In one embodiment, the transmitter (440) may transmit additional data along with the encoded video. The video coder (430) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.
[0056] Before describing certain aspects of the embodiments of the present disclosure in more detail, certain terms that will be referenced in the remainder of the specification are introduced below.
[0057] A "subpicture" hereafter refers to a rectangular arrangement of samples, blocks, macroblocks, coding units, or similar entities that may be semantically grouped and coded independently at a modified resolution, possibly. One or more subpictures may form a picture. One or more coded subpictures may form a coded picture. One or more subpictures may be assembled into a picture, and one or more subpictures may be extracted from a picture. In certain circumstances, one or more coded subpictures may be assembled in the compressed domain without transcoding to the coded picture to the sample level, and in the same or certain other cases, one or more coded subpictures may be extracted from a coded picture in the compressed domain.
[0058] "Adaptive Resolution Change" (ARC) hereinafter refers to a mechanism that enables changing the resolution of pictures or sub-pictures in a coded video sequence, e.g., by resampling of reference pictures. "ARC parameters" hereinafter refers to the control information needed to perform adaptive resolution change, and may include, e.g., filter parameters, scaling factors, output and / or reference picture resolutions, various control flags, etc.
[0059] The context-adaptive arithmetic coding (CABAC) engine in HEVC and VVC may use a table-based probability transition process between 64 different representative probability states. In HEVC, the range ivlCurrRange representing the state of the coding engine may be quantized to a set of four values before the calculation of a new interval range. HEVC state transitions may be implemented using a table containing all 64×4 8-bit pre-calculated values to approximate the value of ivlCurrRange*pLPS(pStateIdx), where pLPS is the probability of the least probable symbol (LPS) and pStateIdx is the index of the current state. The decoding decision may be implemented using a pre-calculated LUT. A first ivlLpsRange may be obtained using the LUT as follows: ivlLpsRange may then be used to update ivlCurrRange and calculate the output binVal. Formula (1) ivlLpsRange=rangeTabLps[pStateIdx][qRangeIdx]
[0060] In VVC, the probability may be linearly represented by the probability index pStateIdx. Thus, all calculations may be done in formulas without LUT operations. To improve the accuracy of the probability estimation, a multi-hypothesis probability update model may be applied. The pStateIdx used in interval subdivision in the binary arithmetic coder may be a combination of two probabilities pStateIdx0 and pStateIdx1. The two probabilities may be associated with each context model and may be updated independently with different adaptation rates. The adaptation rates of pStateIdx0 and pStateIdx1 for each context model may be pre-trained based on the statistics of the associated bin. The probability estimate pStateIdx may be the average of the estimates from the two hypotheses.
[0061] 5 illustrates one embodiment of a process (500) for decoding a single binary decision. The process (500) may begin with an operation (502) to determine a value of the variable ivlCurrRange. In operation (504), if the variable ivlCurrRange is less than or equal to the variable ivlOffset, the process proceeds to operation (506) to update the values of the variables binVal, ivlOffset, and ivlCurrRange. If the variable ivlCurrRange is greater than or equal to the value of the variable ivlOffset, the process proceeds to operation (508) to update the value of the variable binVal. The process proceeds from either operation (506) or operation (508) to operation (510) to update the variables pStateIdx0 and pStateIdx1. The process proceeds to operation (512) to perform the RenormD process.
[0062] As done in HEVC, VVC CABAC may also have a QP-dependent initialization process that is invoked at the beginning of each slice. Given an initial value of the luma QP for a slice, the initial probability state of the context model, denoted as preCtxState, may be derived as follows: Formula (2) m=slopeIdx×5-45 Formula (3) n=(offsetIdx<<3)+7 Equation (4) preCtxState=Clip3(1,127,((m×(QP-32))>>4)+n), Here, slopeIdx and offsetIdx are limited to 3 bits, and the total initialization value can be represented with 6-bit precision. The probability state preCtxState directly represents the probability in the linear domain. Therefore, preCtxState only needs an appropriate shift operation before being input to the arithmetic coding engine, and the logarithmic mapping to the linear domain and a 256-byte table are stored. Equation (5) pStateIdx0=preCtxState<<3 Equation (6) pStateIdx1=preCtxState<<7
[0063] In AV1, an M-ary arithmetic coding engine may be used to entropy code the syntax elements. Each syntax element may be associated with an alphabet of M elements, where M may be any integer value between 2 and 16. The input to the encoding may be the M-ary symbols and a coding context, which may include a set of M probabilities represented by a cumulative distribution function (CDF). The probabilities may be updated after coding / parsing each syntax element. The probability update rate refers to how often the probabilities are updated after coding or parsing each syntax element. The cumulative distribution function may be an array of M 15-bit integers as follows: Equation (7) C=[c0,c1,...,c (M-2) ,2 15 ], where c n / 32768 is the probability that the symbol is less than or equal to n.
[0064] The probability update may be performed using the following formula: Formula (8)
number
number
[0065] From the above equation, the probability update rate has a larger value initially and then saturates after 32 occurrences.
[0066] The multiple hypothesis probability model for encoding M-ary symbols may include modifications to the AV1 arithmetic coding engine for multiple hypothesis estimation and regularization.
[0067] In multi-hypothesis estimation, AV1 may use a data-adaptive model for probability updates, with the update rate being higher for fewer occurrences of the syntax element and lower for more observations. However, only a single probability model is used by the engine. Several studies have shown that multi-hypothesis estimation, where each syntax element maintains two or more probability tables with different update rates, can result in additional compression efficiencies. Thus, the multi-hypothesis probability model may be implemented with two update rates: Formula (10)
number
number
[0068] In regularization, a faster update rate α2 may result in a strongly biased distribution where certain symbol probabilities are reduced to near zero. Probabilities close to zero may result in a BDRATE loss. To counter this effect, a regularization technique is used that at the end of each probability update, m ) is the threshold (P min ), if p m P min A regularization term may be applied to all probabilities such that the probability is moved to: The regularization term may be taken from a uniform distribution or may depend on the sample space of the syntax element.
[0069] The proposed design of the multi-hypothesis probability model for the arithmetic coding engine uses the average of two hypotheses to update the probability of all contexts. However, conventional multi-hypothesis probability modeling is independent of frame type (key frame, inter frame, intra-only frame, etc.), coded block size, coded block prediction mode information, etc. For different syntaxes, the optimal design of the probability update model can utilize this prior information to rapidly adapt the arithmetic coding engine to the current symbol statistics, thereby improving the coding performance.
[0070] The embodiments of the present disclosure relate to a set of advanced video coding techniques including an adaptive multiple hypothesis probability model for arithmetic coding. The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the method (or embodiment), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium. In the following, the term block may be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU. The term block here may also be used to refer to a transform block. In the following items, when referring to a block size, it may refer to either the width or height of the block, or the maximum value of the width and height, or the minimum value of the width and height, or the size of the area (width*height), or the aspect ratio of the block (width:height, or height:width).
[0071] The embodiments of the present disclosure may also apply on a slice or tile basis. For example, if the following methods apply to frame types, the same methods may also apply to slice and / or tile types. In some embodiments, references to a faster update rate may refer to a larger value of the update rate, or a smaller probability update window size.
[0072] In some embodiments, the probability update rate used to update the probability model of a syntax element may depend on the frame type of the current picture being reconstructed. In some embodiments, key frames and / or intra-only frames use a different update rate compared to the update rate used for other frame types. For example, key frames and / or intra-only frames may use a faster update rate (e.g., α2) for all or selected syntax elements. In some embodiments, inter frames may use a slower update rate (e.g., α1) for all or selected syntax elements.
[0073] In some embodiments, the final probabilistic model is updated with update rates α1, α2, α3, ..α n When the weights are computed as a linear combination of two or more hypotheses with weights w1, w2, w3, ... w n For example, the probability hypotheses may be used for update rates α1, α2, α3, ..α n Using p1, p2, p3, ..p n where the final probability update is the weighted sum
number
[0074] In some embodiments, the probability update rate used to update the probability model of a syntax element may depend on the coded block size. As an example, for smaller block sizes, a faster update rate (e.g., α2) may be used for all syntax elements. For example, blocks of 4x4, 4x8, 8x4, 8x8, etc. may use a faster update rate for their syntax elements. In another example, for larger block sizes, a smaller update rate (e.g., α2) may be used for all syntax elements. For example, blocks of sizes 16x16 to 256x256, etc. may use a faster update rate for syntax elements for these blocks.
[0075] In some embodiments, the final probabilistic model is updated with update rates α1, α2, α3, ..α nWhen the weights are computed as a linear combination of two or more hypotheses with weights w1, w2, w3, ... w n etc. may be used. For example, the weights used depend on the block size. In some embodiments, n If corresponds to a faster update rate, then smaller blocks have larger weights w n Similarly, for larger blocks, larger weights may be utilized due to the slower update rate.
[0076] In some embodiments, the probability update rate used to update the probability model of a syntax element may depend on the prediction mode information of the coded block. As an example, for intra-predicted blocks, a faster update rate (e.g., α2) may be used for all or selected syntax elements. In another example, for inter-predicted blocks, a smaller update rate (e.g., α2) may be used for all or selected syntax elements.
[0077] In some embodiments, the final probabilistic model is updated with update rates α1, α2, α3, ..α n When the weights are computed as a linear combination of two or more hypotheses with weights w1, w2, w3, ... w n For example, the probability hypotheses may be used for update rates α1, α2, α3, ..α n Using p1, p2, p3, ..p n where the final probability update is the weighted sum
number
[0078] In some embodiments, the probability update rate used to update the probability model of a syntax element may depend on other coding information, including but not limited to the quantization step size, QP, temporal layer, frame resolution, content type (e.g., whether a screen content coding tool is used), etc. In some embodiments, the probability update rate used to update the probability model of a syntax element is specified in a high level syntax, including but not limited to the VPS, PPS, APS, SPS, picture header, slice header, tile header, CTU (or superblock) header.
[0079] 6 is an example flow chart of a process (600) for performing adaptive multi-hypothesis probability modeling for arithmetic coding. The process (600) may be performed by a decoder, such as the decoder (210). The process may begin at an operation (602) in which a coded video bitstream is received. The bitstream may include at least one picture and one or more syntax elements encoded according to multi-hypothesis arithmetic coding. The process proceeds to an operation (604) in which each syntax element is decoded based on the multi-hypothesis arithmetic coding.
[0080] The process continues with operation (606) where a probability update rate is selected from a plurality of probability update rates based on a predetermined condition. For example, the predetermined condition may specify a frame type of at least one picture in the bitstream, and the probability update rate is selected based on whether the frame type is a key frame, an intra frame, or an inter frame. As another example, the predetermined condition may specify a block size of at least one block in the picture. As another example, the predetermined condition may specify a prediction mode of the at least one block.
[0081] The process proceeds to operation (608) where at least one probability model utilized in the multiple hypothesis arithmetic coding is updated based on the selected probability update rate. The process proceeds to operation (610) where at least one block in at least one picture is decoded based on the one or more decoded syntax elements.
[0082] The techniques of the embodiments of the present disclosure described above can be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. For example, Figure 7 shows a computer system (700) suitable for implementing embodiments of the disclosed subject matter.
[0083] The computer software may be coded using any suitable machine code or computer language that may be subjected to mechanisms such as assembly, compilation, linking, etc. to produce code containing instructions that may be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly, or via interpretation, microcode execution, etc.
[0084] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0085] 7 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having a dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system (700).
[0086] The computer system (700) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (such as keystrokes, swipes, data glove movements), audio input (such as voice, clapping), visual input (such as gestures), or olfactory input (not depicted). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as sound (speech, music, ambient sounds, etc.), images (scanned images, photographic images obtained from still image cameras, etc.), and video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).
[0087] The input human interface devices may include one or more (only one of each shown) of a keyboard (701), a mouse (702), a trackpad (703), a touch screen (710), a data glove, a joystick (705), a microphone (706), a scanner (707), and a camera (708).
[0088] The computer system (700) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (710), data gloves, or joystick (705), although there may also be haptic feedback devices that do not function as input devices). For example, such devices may be audio output devices (such as speakers (709), headphones (not shown)), visual output devices (such as screens (710), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capability, each with or without haptic feedback capability, some of which may output two-dimensional visual output or output in more than three dimensions, such as through means of stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0089] The computer system (700) may also include human-accessible storage devices and their associated media, such as optical media, including CD / DVD ROM / RW (720) with CD / DVD or similar media (721), thumb drives (722), removable hard drives or solid state drives (723), legacy magnetic media such as tapes and floppy disks (not depicted), and specialized ROM / ASIC / PLD based devices (not depicted) such as security dongles.
[0090] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0091] The computer system (700) may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable, satellite and terrestrial television, vehicular and industrial including CANBus, etc. Certain networks generally require an external network interface adapter connected to a specific general-purpose data port or peripheral bus (749) (e.g., a USB port on the computer system (700)), while other networks are generally integrated into the core of the computer system (700) by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (700) may communicate with other entities. Such communications may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a particular CANbus device), or two-way, such as to other computer systems using local or wide area digital networks. Such communications may include communications to cloud computing environments (755). Specific protocols and protocol stacks may be used in each of those networks and network interfaces, as described above.
[0092] The aforementioned human interface devices, human accessible storage devices, and network interfaces (754) may be attached to the core (740) of the computer system (700).
[0093] The core (740) may include one or more central processing units (CPUs) (741), graphics processing units (GPUs) (742), specialized programmable processing units in the form of field programmable gate areas (FPGAs) (743), hardware accelerators for specific tasks (744), etc. These devices may be connected via a system bus (748), along with read only memory (ROM) (745), random access memory (746), internal mass storage (747) such as an internal hard drive or SSD that is not user accessible. In some computer systems, the system bus (748) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (748) or via a peripheral bus (749). Architectures for peripheral buses include PCI, USB, etc. A graphics adapter (750) may be included in the core (740).
[0094] The CPU (741), GPU (742), FPGA (743), and accelerator (744) may execute certain instructions that may combine to constitute the above-mentioned computer code. The computer code may be stored in ROM (745) or RAM (746). Temporary data may also be stored in RAM (746), while persistent data may be stored, for example, in internal mass storage (747). Rapid storage and retrieval from any memory device may be enabled by the use of cache memory, which may be closely associated with one or more of the CPU (741), GPU (742), mass storage (747), ROM (745), RAM (746), etc.
[0095] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the type well known and available to those skilled in the computer software arts.
[0096] By way of example and not limitation, a computer system having the architecture (700), and in particular the core (740), may provide functionality as a result of a processor(s) (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be user-accessible mass storage as described above, as well as media associated with specific storage of the core (740) of a non-transitory nature, such as the core internal mass storage (747) or ROM (745). Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core (740). The computer-readable media may include one or more memory devices or chips, depending on particular needs. The software may cause the core (740), and in particular the processors therein (including CPU, GPU, FPGA, etc.) to perform certain processes, or certain portions of certain processes, described herein, including defining data structures stored in RAM (746) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (744)), which may operate in place of or in conjunction with software to perform particular processes, or particular portions of particular processes, described herein. References to software may encompass logic, and vice versa, as appropriate. References to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that store software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0097] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations.
[0098] It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed herein is illustrative of example approaches. It is understood that the specific order or hierarchy of blocks within the processes / flowcharts may be rearranged based on design preferences. Also, some blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in a sample order, and are not meant to be limited to the specific order or hierarchy presented.
[0099] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of integration of technical details. Furthermore, one or more of the above-mentioned components may be implemented as instructions stored on a computer-readable medium and executable by at least one processor (and / or may include at least one processor). The computer-readable medium may include a computer-readable non-transitory storage medium having computer-readable program instructions for causing a processor to perform operations.
[0100] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves with instructions recorded on them, and any suitable combination of the above. As used herein, a computer-readable storage medium should not be construed as a transitory signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over electrical wires.
[0101] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium into each computing / processing device, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fiber, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.
[0102] The computer readable program code / instructions for performing the operations may be either assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk or C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or a connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuit to perform an aspect or operation.
[0103] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to generate a machine such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram blocks. These computer-readable program instructions may also be stored on a computer-readable storage medium that may direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium having stored instructions includes an article of manufacture including instructions that implement aspects of the functions / acts specified in the flowchart and / or block diagram blocks.
[0104] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device such that a series of operational steps are executed on the computer, other programmable apparatus, or other device to generate a computer-implemented process such that the instructions, executed on the computer, other programmable apparatus, or other device, implement the functions / acts specified in the flowchart and / or block diagram blocks.
[0105] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or part of instructions that includes one or more executable instructions for implementing a specified logical function(s). The methods, computer systems, and computer-readable media may include additional, fewer, different, or differently arranged blocks compared to those shown in the figures. In some alternative implementations, the functions described in the blocks may be performed in a different order than described in the figures. For example, two blocks shown in succession may in fact be executed simultaneously or substantially simultaneously, or the blocks may be executed in reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, as well as combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations or realizes a combination of dedicated hardware and computer instructions.
[0106] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not intended to limit the embodiments. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it will be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0107] The above disclosure also encompasses the embodiments listed below.
[0108] (1) A method performed by at least one processor of a video decoder, the method comprising: receiving a coded video bitstream including at least one picture and one or more syntax elements encoded according to multiple hypothesis arithmetic coding; decoding each syntax element from the one or more syntax elements based on the multiple hypothesis arithmetic coding; selecting a probability update rate from a plurality of probability update rates based on a predetermined condition, the plurality of probability update rates including a first probability update rate higher than a second probability update rate; updating at least one probability model utilized in the multiple hypothesis arithmetic coding based on the selected probability update rate; and decoding at least one block in the at least one picture based on the decoded one or more syntax elements.
[0109] (2) The method according to feature (1), wherein the predetermined condition specifies a frame type of the at least one picture.
[0110] (3) The method of feature (2), wherein the selecting step includes selecting the first probability update rate in response to determining that a frame type of the at least one picture is one of a key frame and an intra frame.
[0111] (4) The method of any one of features (2) to (3), wherein the selecting step includes selecting the second probability update rate in response to determining that a frame type of the at least one picture is interframe.
[0112] (5) The method according to any one of features (2) to (4), wherein the multiple hypothesis arithmetic coding includes a linear combination of a plurality of probability models, each probability model being derived from the plurality of probability update rates using a corresponding probability update rate, and the linear combination is a weighted sum using a plurality of weights, each weight depending on the frame type of the at least one picture.
[0113] (6) The method of feature (5), wherein in response to a determination that the frame type is one of a key frame and an intra frame, a greater weight is used compared to a determination that the frame type is an inter frame.
[0114] (7) The method according to any one of features (1) to (6), wherein the predetermined condition specifies a block size of the at least one block.
[0115] (8) The method of feature (7), wherein in response to determining that the at least one block is an N×M block, the first probability update rate is selected, N being one of 4 and 8, and M being one of 4 and 8.
[0116] (9) The method of any one of features (7) to (8), wherein the second probability update rate is selected in response to determining that the at least one block is an N×M block, where N is 16 or greater and M is 16 or greater.
[0117] (10) The method according to any one of features (7) to (9), wherein the multiple hypothesis arithmetic coding includes a linear combination of a plurality of probability models, each probability model being derived from the plurality of probability update rates using a corresponding probability update rate, and the linear combination is a weighted sum using a plurality of weights, each weight depending on the block size of the at least one block.
[0118] (11) The method of feature (10), wherein a first weight from the plurality of weights for a first block is greater than a second weight from the plurality of weights for a second block, and the first block is smaller than the second block.
[0119] (12) The method according to any one of features (2) to (11), wherein the predetermined condition specifies a prediction mode of the at least one block.
[0120] (13) The method of feature (12), wherein the first probability update rate is selected in response to determining that the prediction mode of the at least one block is an intra prediction mode.
[0121] (14) The method of any one of features (12) to (13), wherein the second probability update rate is selected in response to determining that the prediction mode of the at least one block is an inter prediction mode.
[0122] (15) The method according to any one of features (12) to (14), wherein the multiple hypothesis arithmetic coding includes a linear combination of a plurality of probability models, each probability model being derived from the plurality of probability update rates using a corresponding probability update rate, and the linear combination is a weighted sum using a plurality of weights, each weight depending on the prediction mode of the at least one block.
[0123] (16) The method of feature (15), wherein in response to a determination that the prediction mode is an intra prediction mode, the weighting is greater compared to a determination that the prediction mode is an inter prediction mode.
[0124] (17) A video decoder, comprising: at least one memory configured to store computer program code; and at least one processor configured to access the computer program code and operate as instructed by the computer program code, the computer program code comprising: receiving code configured to cause the at least one processor to receive a coded video bitstream including at least one picture and one or more syntax elements encoded according to multiple hypothesis arithmetic coding; and configuring the at least one processor to decode each syntax element from the one or more syntax elements based on the multiple hypothesis arithmetic coding. 13. A video decoder comprising: a first decoding code generated by decoding a plurality of probability update rates based on a predetermined condition; a selection code configured to cause the at least one processor to select a probability update rate from a plurality of probability update rates based on a predetermined condition, the plurality of probability update rates including a first probability update rate that is higher than a second probability update rate; an update code configured to cause the at least one processor to update at least one probability model utilized in the multiple hypothesis arithmetic coding based on the selected probability update rate; and a second decoding code configured to cause the at least one processor to decode at least one block in the at least one picture based on the decoded one or more syntax elements.
[0125] (18) The video decoder of feature (17), wherein the predetermined condition specifies a frame type of the at least one picture.
[0126] (19) The video decoder of feature (18), wherein the selection code is further configured to cause the at least one processor to select the first probability update rate in response to determining that a frame type of the at least one picture is one of a key frame and an intra frame.
[0127] (20) A non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor in a video decoder, cause the processor to perform a method comprising: receiving a coded video bitstream including at least one picture and one or more syntax elements encoded according to multiple hypothesis arithmetic coding; decoding each syntax element from the one or more syntax elements based on the multiple hypothesis arithmetic coding; selecting a probability update rate from a plurality of probability update rates based on a predetermined condition, the plurality of probability update rates including a first probability update rate that is higher than a second probability update rate; updating at least one probability model utilized in the multiple hypothesis arithmetic coding based on the selected probability update rate; and decoding at least one block in the at least one picture based on the decoded one or more syntax elements. [Explanation of symbols]
[0128] 100 Communication Systems 110 First Terminal 120 Second Terminal 130 terminals 140 terminals 150 Network 200 Streaming System 201 Video Sources 202 Uncompressed Video Sample Stream 203 Encoder 204 encoded video bitstream 205 Streaming Server 206 Streaming Client 209 Video Bitstream 210 Video Decoder 211 outgoing video sample stream 212 Display 213 Capture Subsystem 310 Receiver 312 Channels 315 Buffer Memory 320 Entropy Decoder / Parser 321 Symbols 351 Scaler / Descaler Unit 352 Intra Prediction Units 353 Motion Compensation Prediction Unit 355 Aggregator 356 Loop Filter Unit 357 Reference Picture Memory 358 Current Picture Memory 430 Source Coder 432 Coding Engine 433 (local) decoder 434 Reference Picture Memory 435 Predictor 440 Transmitter 443 coded video sequence 445 Entropy Coder 450 Controller 460 Channels 700 Computer Systems 701 Keyboard 702 Mouse 703 Trackpad 705 Joystick 706 Microphone 707 Scanner 708 Camera 710 Touchscreen 720 CD / DVD ROM / RW 721 Medium 722 Thumb Drive 723 Solid State Drive 740 cores 741 Central Processing Unit (CPU) 742 Graphics Processing Unit (GPU) 743 Field Programmable Gate Area (FPGA) 744 Hardware Accelerator 745 Read-Only Memory (ROM) 746 Random Access Memory (RAM) 747 Large internal storage capacity 748 System Bus 749 Surrounding Bus 750 Graphics Adapter 754 Network Interface 755 Cloud Computing Environment
Claims
1. 1. A method performed by at least one processor of a video decoder, the method comprising: receiving a coded video bitstream including at least one picture and one or more syntax elements encoded according to multiple hypothesis arithmetic coding; decoding each syntax element from the one or more syntax elements based on the multiple hypothesis arithmetic coding; selecting a probability update rate from a plurality of probability update rates based on a predetermined condition, the plurality of probability update rates including a first probability update rate that is higher than a second probability update rate; updating at least one probability model utilized in the multiple hypothesis arithmetic coding based on the selected probability update rate; and decoding at least one block in the at least one picture based on the decoded one or more syntax elements.
2. The method of claim 1 , wherein the predetermined condition specifies a frame type of the at least one picture.
3. 3. The method of claim 2, wherein the selecting step comprises selecting the first probability update rate in response to determining that a frame type of the at least one picture is one of a key frame and an intra frame.
4. The method of claim 2 , wherein the selecting step comprises selecting the second probability update rate in response to determining that a frame type of the at least one picture is interframe.
5. the multiple-hypothesis arithmetic coding includes a linear combination of a plurality of probability models, each probability model being derived using a corresponding probability update rate from the plurality of probability update rates; The method of claim 2 , wherein the linear combination is a weighted sum using a plurality of weights, each weight depending on the frame type of the at least one picture.
6. The method of claim 5 , wherein in response to a determination that the frame type is one of a key frame and an intra frame, a greater weight is used compared to a determination that the frame type is an inter frame.
7. The method of claim 1 , wherein the predetermined condition specifies a block size of the at least one block.
8. 8. The method of claim 7, wherein the first probability update rate is selected in response to determining that the at least one block is an N×M block, wherein N is one of 4 and 8, and M is one of 4 and 8.
9. 8. The method of claim 7, wherein the second probability update rate is selected in response to determining that the at least one block is an N×M block, wherein N is greater than or equal to 16 and M is greater than or equal to 16.
10. the multiple-hypothesis arithmetic coding includes a linear combination of a plurality of probability models, each probability model being derived using a corresponding probability update rate from the plurality of probability update rates; The method of claim 7 , wherein the linear combination is a weighted sum using a plurality of weights, each weight depending on the block size of the at least one block.
11. 11. The method of claim 10, wherein a first weight from the plurality of weights for a first block is greater than a second weight from the plurality of weights for a second block, and the first block is smaller than the second block.
12. The method of claim 2 , wherein the predetermined condition specifies a prediction mode for the at least one block.
13. The method of claim 12 , wherein the first probability update rate is selected in response to determining that the prediction mode of the at least one block is an intra-prediction mode.
14. The method of claim 12 , wherein the second probability update rate is selected in response to determining that the prediction mode of the at least one block is an inter prediction mode.
15. the multiple-hypothesis arithmetic coding includes a linear combination of a plurality of probability models, each probability model being derived using a corresponding probability update rate from the plurality of probability update rates; The method of claim 12 , wherein the linear combination is a weighted sum using a plurality of weights, each weight depending on the prediction mode of the at least one block.
16. The method of claim 15 , wherein the weight is greater in response to a determination that the prediction mode is an intra-prediction mode compared to a determination that the prediction mode is an inter-prediction mode.
17. A video decoder configured to perform a method according to any one of claims 1 to 16.
18. A computer program product for causing a processor to carry out a method according to any one of claims 1 to 16.