Video decoding method, video decoder, and non-transitory computer readable medium

CN117044201BActive Publication Date: 2026-09-11TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280017065.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-11-08
Filing Date
2022-11-09
Publication Date
2026-09-11
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

当前用于多假设概率模型的编码标准没有充分考虑诸如帧类型、块大小、预测模式等的其他信息

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117044201B_ABST
    Figure CN117044201B_ABST
Patent Text Reader

Abstract

A method performed by at least one processor of a video decoder includes receiving a coded video bitstream including at least one picture and one or more syntax elements encoded according to multi-hypothesis arithmetic coding. The method further includes decoding each of the one or more syntax elements based on the multi-hypothesis arithmetic coding. The method further includes selecting a probability update rate from a plurality of probability update rates based on a predetermined condition, the plurality of probability update rates including a first probability update rate higher than a second probability update rate. The method further includes updating at least one probability model used in the multi-hypothesis arithmetic coding based on the selected probability update rate. The method further includes decoding at least one block in the at least one picture based on the one or more syntax elements decoded.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the priority benefit of U.S. Provisional Patent Application No. 63 / 318,588, filed March 10, 2022, and U.S. Patent Application No. 18 / 053,494, filed November 8, 2022, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to communication systems, and more specifically to methods and apparatus for adaptive multi-hypothesis probabilistic models for arithmetic coding. Background Technology

[0004] The Alliance for Open Media (AOMedia) Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. Developed by AOMedia to succeed VP9, ​​this coding format is a video coding format that the alliance, founded in 2015, includes semiconductor companies, video-on-demand providers, video content producers, software development companies, and web browser vendors. Many components of the AV1 project draw from previous research by alliance members. Individual contributors began experimenting with technology platforms several years ago: the Xiph / Mozilla Foundation's Daala project released code in 2010, Google's experimental VP9 evolution project VP10 was released on September 12, 2014, and Cisco's Thor project was released on August 11, 2015. Building on the VP9 codebase, AV1 incorporates other technologies, some of which were developed using these experimental formats. The first version (0.1.0) of the AV1 reference codec was released on April 7, 2016. The Alliance announced the release of the AV1 stream specification, along with software-based reference encoders and decoders, on March 28, 2018. A validation version (1.0.0) of the specification was released on June 25, 2018. A validation version (1.0.0) including Errata Table 1 was released on January 8, 2019. The AV1 stream specification includes a reference video codec. Current coding standards used for multi-hypothesis probabilistic models do not adequately consider other information such as frame type, block size, and prediction mode. Summary of the Invention

[0005] The following presents a simplified overview of one or more embodiments of this disclosure to provide a basic understanding of these embodiments. This overview is not a comprehensive summary of all contemplated embodiments, nor is it intended to identify key or critical elements of all embodiments, nor to depict any or all scopes of embodiments. Its sole purpose is to present some concepts of one or more embodiments of this disclosure in a simplified form as a prelude to the more detailed description that follows.

[0006] This disclosure discloses methods, apparatus, and non-transitory computer-readable media for performing adaptive multi-hypothesis probabilistic modeling for arithmetic coding.

[0007] According to an exemplary embodiment, a method executed by at least one processor of a video decoder includes: receiving an encoded video stream, the encoded video stream including at least one picture encoded according to multi-hypothesis arithmetic coding and one or more syntax elements. The method further includes: decoding each of the one or more syntax elements based on the multi-hypothesis arithmetic coding. The method further includes: selecting a probability update rate from a plurality of probability update rates based on predetermined conditions, the plurality of probability update rates including a first probability update rate higher than a second probability update rate. The method further includes: updating at least one probabilistic model used in the multi-hypothesis arithmetic coding based on the selected probability update rate. The method further includes: decoding at least one block in the at least one picture based on the decoded one or more syntax elements.

[0008] According to an exemplary embodiment, a video decoder includes at least one memory configured to store computer program code, and at least one processor configured to access and operate according to the computer program code instructions. The computer program code includes: receiving code configured to cause the at least one processor to receive an encoded video stream, the encoded video stream including at least one picture encoded according to multi-hypothesis arithmetic coding and one or more syntax elements. The computer program code further includes: first decoding code configured to cause the at least one processor to decode each of the one or more syntax elements based on the multi-hypothesis arithmetic coding. The computer program code further includes: selection code configured to cause the at least one processor to select a probability update rate from a plurality of probability update rates based on predetermined conditions, the plurality of probability update rates including a first probability update rate higher than a second probability update rate. The computer program code further includes: update code configured to cause the at least one processor to update at least one probabilistic model used in the multi-hypothesis arithmetic coding based on the selected probability update rate. The computer program code further includes: second decoding code configured to cause the at least one processor to decode at least one block in the at least one picture based on the decoded one or more syntax elements.

[0009] According to an exemplary embodiment, a non-transitory computer-readable medium storing instructions, which, when executed by a processor in a video decoder, cause the processor to perform a method, the method comprising: receiving an encoded video stream, the encoded video stream comprising at least one picture and one or more syntax elements encoded according to multi-hypothesis arithmetic coding. The method further comprises: decoding each of the one or more syntax elements based on the multi-hypothesis arithmetic coding. The method further comprises: selecting a probability update rate from a plurality of probability update rates based on predetermined conditions, the plurality of probability update rates including a first probability update rate higher than a second probability update rate. The method further comprises: updating at least one probabilistic model used in the multi-hypothesis arithmetic coding based on the selected probability update rate. The method further comprises: decoding at least one block in the at least one picture based on the decoded one or more syntax elements.

[0010] Additional embodiments will be set forth in the description below (and in part will become more apparent from the description), and / or can be learned by practicing the embodiments presented in this disclosure. Attached Figure Description

[0011] The above and other aspects and features of the embodiments of this disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which: Figure 1 This is a schematic diagram of a communication system according to various embodiments of the present disclosure; Figure 2 This is a schematic diagram of a communication system according to various embodiments of the present disclosure; Figure 3 This is a schematic diagram of a block diagram of a decoder according to various embodiments of the present disclosure; Figure 4 This is a block diagram of an encoder according to various embodiments of the present disclosure; Figure 5 Example flowcharts of the process for decoding binary items (bin) according to various embodiments of the present disclosure are shown; Figure 6 An example flowchart is shown for performing adaptive multihypothesis probabilistic modeling for arithmetic coding; Figure 7 This is a computer system diagram according to various embodiments of the present disclosure. Detailed Implementation

[0012] The following detailed description of exemplary embodiments is with reference to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0013] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit implementation to the precise forms disclosed. Modifications and variations are possible based on the foregoing disclosure, or may be obtained from practice of the implementation. Furthermore, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, as will be understood from the flowcharts and operational descriptions provided below, one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be switched.

[0014] It is evident that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit these implementations. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to any specific software code. It should be understood that software and hardware can be designed to implement the systems and / or methods based on the descriptions herein.

[0015] Even if a particular combination of features is stated in the claims and / or disclosed in the specification, these combinations are not intended to limit the possible disclosure. In fact, many of these features can be combined in ways not specifically listed in the claims and / or not disclosed in the specification. Although each dependent claim listed below may depend directly on only one claim, the possible disclosure includes combinations of each dependent claim with every other claim in the claim set.

[0016] Unless explicitly stated otherwise, no element, action, or instruction used herein should be construed as critical or necessary. Furthermore, as used herein, the article “a” is intended to include one or more items and may be used interchangeably with “a or more.” The term “a” or similar language is used if intended to refer to a single item. Additionally, as used herein, the terms “having,” “including,” etc., are intended to indicate open-ended terms. Furthermore, unless explicitly stated otherwise, the word “based on” is intended to mean “at least partially based on.” Moreover, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.

[0017] Throughout this specification, references to "an embodiment," "embodiment," or similar language mean that a particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of this solution. Therefore, the phrases "in one embodiment," "in an embodiment," and similar language throughout this specification may, but not necessarily all, refer to the same embodiment.

[0018] Furthermore, the features, advantages, and characteristics described herein can be combined in any suitable manner in one or more embodiments. Based on the description herein, those skilled in the art will recognize that this disclosure can be practiced without one or more specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may be recognized in some embodiments may not be present in all embodiments of this disclosure.

[0019] Figure 1 A simplified block diagram of a communication system (100) according to an embodiment of the present disclosure is shown. The system (100) may include at least two terminals (110, 120) interconnected via a network (150). For unidirectional data transmission, the first terminal (110) may encode video data at a local location for transmission to the other terminal (120) via the network (150). The second terminal (120) may receive the encoded video data from the other terminal from the network (150), decode the encoded data, and display the recovered video data. Unidirectional data transmission is common in applications such as media services.

[0020] Figure 1 A second pair of terminals (130, 140) is shown provided to support bidirectional transmission of encoded video, which may occur, for example, during a video conference. For bidirectional data transmission, each terminal (130, 140) can encode video data captured at a local location for transmission to the other terminal via a network (150). Each terminal (130, 140) can also receive encoded video data sent by the other terminal, can decode the encoded data, and can display the recovered video data on a local display device.

[0021] exist Figure 1 In this context, terminals (110-140) can be represented as servers, personal computers and smartphones, and / or any other type of terminal. For example, terminals (110-140) can be laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. A network (150) refers to any number of networks that transmit encoded video data between terminals (110-140), including, for example, wired and / or wireless communication networks.

[0022] The communication network (150) can exchange data in circuit-switched and / or packet-switched channels. Representative networks may include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this disclosure, unless explained below, the architecture and topology of the network (150) may be irrelevant to the operation of this disclosure.

[0023] As an example of the application of the disclosed topic. Figure 2The placement of video encoders and decoders in a streaming environment is illustrated. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0024] like Figure 2 As shown, the streaming media system (200) may include a capture subsystem (213), which may include a video source (201) and an encoder (203). The video source (201) may be, for example, a digital camera and may be configured to create an uncompressed sample video stream (202). The uncompressed sample video stream (202) may provide a high amount of data compared to an encoded video stream and may be processed by an encoder (203) coupled to the camera (201). The encoder (203) may include hardware, software, or a combination of hardware and software to implement or carry out aspects of the disclosed subject matter as described in more detail below. The encoded video stream (204) may include a lower amount of data compared to the sample stream and may be stored on a streaming media server (205) for future use. One or more streaming media clients (206) may access the streaming media server (205) to retrieve a video stream (209), which may be a copy of the encoded video stream (204).

[0025] In this embodiment, the streaming server (205) can also be used as a Media-Aware Network Element (MANE). For example, the streaming server (205) can be configured to trim the encoded video stream (204) to customize potentially different streams for one or more streaming clients (206). In this embodiment, the MANE can be provided separately from the streaming server (205) in the streaming system (200).

[0026] A streaming client (206) may include a video decoder (210) and a display (212). For example, the video decoder (210) may decode a video stream (209) that is an incoming copy of an encoded video stream (204) and create an outgoing sample stream (211) that can be displayed on the display (212) or another presentation device (not shown). In some streaming systems, the video streams (204, 209) may be encoded according to certain video encoding / compression standards. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. The video encoding standard under development is informally referred to as Universal Video Coding (VVC). Embodiments of this disclosure may be used in the context of VVC.

[0027] Figure 3An example functional block diagram of a video decoder (210) attached to a display (212) according to an embodiment of the present disclosure is shown. The video decoder (210) may include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra-frame prediction unit (352), a motion compensation prediction unit (353), an aggregator (355), a loop filter unit (356), a reference image memory (357), and a current image memory (358). In at least one embodiment, the video decoder (210) may include an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. The video decoder (210) may also be partially or wholly implemented in software running on one or more CPUs having associated memory.

[0028] In this and other embodiments, the receiver (310) may receive one or more encoded video sequences, which will be decoded one encoded video column at a time by the decoder (210), wherein the decoding of each encoded video column is independent of the other encoded video columns. The encoded video sequences may be received from a channel (312), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive encoded video data, as well as other data such as encoded audio data and / or auxiliary data streams, that can be forwarded to their respective user entities (not indicated). The receiver (310) may separate the encoded video sequences from other data. To prevent network jitter, a buffer (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter referred to as the "parser"). The buffer (315) may not be used, or it may be made smaller, when the receiver (310) receives data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isochronous synchronization network. For use on best-effort packet networks such as the Internet, the buffer (315) may be necessary, may be relatively large, and may have an adaptive size.

[0029] The video decoder (210) may include a parser (320) to reconstruct symbols (321) from the entropy-coded video sequence. These symbols include, for example, information for managing the operation of the decoder (210), and potentially information for controlling rendering devices, such as a display (212) that may be coupled to the decoder, such as... Figure 2As shown. The control information for the display device may be a parameter set fragment (not shown) of Supplementary Enhancement Information (SEI) messages or Video Usability Information (VUI). The parser (320) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (320) may extract a subgroup parameter set of at least one subgroup of pixels in the subgroups of pixels in the encoded video sequence for use in the video decoder based on at least one parameter corresponding to a group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (320) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc. The parser (320) performs entropy decoding / parsing operations on the video sequence received from the buffer (315) to create symbols (321). The reconstruction of symbols (321) may involve multiple different units, depending on the type of encoded video pictures or a subset of encoded video pictures (e.g., inter-frame and intra-frame pictures, inter-frame and intra-frame blocks) and other factors. Which units are involved and how they are involved can be controlled by the subgroup control information parsed by the parser (320) from the encoded video sequence. For brevity, the flow of such subgroup control information between the parser (320) and the various units described below is not described.

[0030] In addition to the functional blocks already mentioned, the decoder (210) can be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0031] A unit can be a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) can receive quantization transform coefficients as symbols (321) from the parser (320) as well as control information, including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (351) can output a block containing sample values, which can be input into the aggregator (355).

[0032] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed images, but can use predictive information from previously reconstructed portions of the current image. Such predictive information may be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses surrounding reconstructed information extracted from the current (partially reconstructed) image in the current image memory (358) to generate blocks of the same size and shape as the blocks being reconstructed. In some cases, the aggregator (355) adds the predictive information generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) based on each sample. In other cases, the output samples of the scaler / inverse transform unit (351) may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit (353) can access the reference image memory (357) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (321) belonging to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (referred to in this case as residual samples or residual signals) to generate output sample information. The address in the reference image memory (357) from which the motion compensation prediction unit (353) extracts the predicted samples can be controlled by motion vectors. The motion vectors can be used by the motion compensation prediction unit (353) in the form of symbols (321), which may have, for example, X, Y and reference image components. Motion compensation may also include interpolation of sample values ​​extracted from the reference image memory (357) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.

[0033] The output samples of the aggregator (355) can be subjected to various loop filtering techniques in the loop filter unit (356). The video compression technique may include an in-loop filter technique controlled by parameters included in the encoded video bitstream and available to the loop filter unit (356) as symbols (321) from the parser (320). However, the video compression technique may also respond to metadata obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, as well as to previously reconstructed and loop-filtered sample values.

[0034] The output of the loop filter unit (356) can be a sample stream that can be output to a rendering device such as a display (212) and stored in a reference image memory (357) for subsequent inter-frame image prediction.

[0035] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. Once the encoded images have been fully reconstructed and the encoded images (by, for example, the parser (320)) are identified as reference images, the current reference image can become part of the reference image memory (357), and a new current image memory can be reallocated before the reconstruction of subsequent encoded images begins.

[0036] The video decoder (210) can perform decoding operations according to a predetermined video compression technique, which may be documented in a standard such as ITU-T Rec.H.265. The encoded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard (as specified in the video compression technique document or standard, and specifically in its configuration file). Furthermore, the complexity of the encoded video sequence may be within the range defined by the hierarchy of the video compression technique or standard in order to conform to some video compression techniques or standards. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.

[0037] In this embodiment, the receiver (310) may receive additional (redundant) data along with the encoded video. This additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0038] Figure 4An example functional block diagram of a video encoder (203) associated with a video source (201) according to an embodiment of the present disclosure is shown. The video encoder (203) may include, for example, an encoder (432) as a source encoder (430), an encoding engine (432), a (local) decoder (433), a reference image memory (434), a predictor (435), a transmitter (440), an entropy encoder (445), a controller (450), and a channel (460).

[0039] The encoder (203) can receive video samples from a video source (201) (not part of the encoder), which can capture video images to be encoded by the encoder (203). The video source (201) can provide a source video sequence in the form of a digital video sample stream to be encoded by the encoder (203), which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (201) can be a storage device storing previously prepared video. In a video conferencing system, the video source (203) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following text focuses on describing samples.

[0040] According to an embodiment, the encoder (203) can encode and compress images of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate is a function of the controller (450). The controller (450) can also control other functional units as described below and can be functionally coupled to these units. For simplicity, coupling is not shown in the figures. Parameters set by the controller (450) may include rate control related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. Other functions of the controller (450) can be readily identified by those skilled in the art, as they may involve a video encoder (203) optimized for a particular system design.

[0041] Some video encoders operate within an “encoding loop” readily recognizable to those skilled in the art. As an oversimplification, the encoding loop may include the encoding portion of a source encoder (430) responsible for creating symbols based on the input image to be encoded and a reference image, and a (local) decoder (433) embedded in an encoder (203), which reconstructs the symbols to create sampled data that the (remote) decoder will also create when compression between the symbols and the encoded video stream is lossless in some video compression techniques. The reconstructed sample stream can be fed into a reference image memory (434). Since decoding of the symbol stream produces bit-accurate results independent of the decoder’s location (local or remote), the contents of the reference image memory are also bit-accurately corresponding between the local and remote encoders. In other words, the reference image samples “seen” by the encoder’s prediction portion are exactly the same sample values ​​that the decoder will “see” during the prediction process. This fundamental principle of reference image synchronicity (and the drift that occurs when synchronicity cannot be maintained, for example, due to channel errors) is known to those skilled in the art.

[0042] The operation of the "local" decoder (433) can be combined with the above. Figure 3 The “remote” decoder is the same as the video decoder (210) described in detail. However, since symbols are available and the entropy encoder (445) and parser (320) are able to encode / decode symbols into an encoded video sequence without loss, the entropy decoding portion of the decoder (210), which includes the channel (312), receiver (310), buffer (315), and parser (320), may not be fully implemented in the local decoder (433).

[0043] It can be observed that any decoder technique other than parsing / entropy decoding, which exists in the decoder, may need to exist in the corresponding encoder in essentially the same functional form. For this reason, the disclosed subject focuses on decoder operation. The description of encoder techniques can be simplified, as encoder techniques can be inverses of the fully described decoder techniques. More detailed descriptions are only required in certain areas, and are provided below.

[0044] As part of its operation, the source encoder (430) performs motion-compensated predictive coding. This motion-compensated predictive coding predictively encodes the input frame, referencing one or more previously encoded frames in the video sequence designated as "reference frames." In this manner, the encoding engine (432) encodes the differences between pixel blocks of the input frame and pixel blocks of the reference frame, which can be selected as a prediction reference for the input frame.

[0045] The local video decoder (433) can decode encoded video data of a frame that can be designated as a reference frame, based on symbols created by the source encoder (430). The operation of the encoding engine (432) can advantageously be a lossy process. When the encoded video data can be decoded by the video decoder (433), Figure 4 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (433) replicates the decoding process, which can be performed by the video decoder on the reference frame, and allows the reconstructed reference frame to be stored in the reference image memory (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame that shares the same content (no transmission errors) as the reconstructed reference frame that will be obtained by the remote video decoder.

[0046] The predictor (435) can perform a prediction search against the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) can search in the reference image memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. The predictor (435) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (435), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (434).

[0047] The controller (450) manages the encoding operations of the video encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data. Entropy encoding is performed on the outputs of all the aforementioned functional units in the entropy encoder (445). The entropy encoder performs lossless compression on the symbols generated by the various functional units according to techniques known to those skilled in the art, such as Huffman coding, variable-length coding, and arithmetic coding, thereby converting the symbols into an encoded video sequence.

[0048] The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) in preparation for its transmission via a communication channel (460), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) can combine the encoded video data from the video encoder (430) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown). The controller (450) can manage the operation of the encoder (203). During encoding, the controller (450) can assign a particular encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures may typically be assigned as intra-frame pictures (I-pictures), predictive pictures (P-pictures), or bidirectional predictive pictures (B-pictures). An intra-frame picture (I-picture) is a picture that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are familiar with variations of I-pictures and their corresponding applications and characteristics. A predictive picture (P-picture) can be a picture that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block.

[0049] Bidirectional predictive images (B-images) can be images that can be encoded and decoded using intra-frame or inter-frame prediction, which uses at most two motion vectors and a reference index to predict sample values ​​for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata to reconstruct a single block.

[0050] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, which are determined by the coding assignment of the corresponding images applied to the block. For example, a block of an I image can be nonpredictively coded, or the block can be predictively coded (spatial prediction or intra-frame prediction) with reference to already coded blocks of the same image. A pixel block of a P image can be nonpredictively coded with reference to a previously coded reference image via spatial prediction or temporal prediction. A block of a B image can be nonpredictively coded with reference to one or two previously coded reference images via spatial prediction or temporal prediction.

[0051] The video encoder (203) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. In operation, the video encoder (203) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0052] In this embodiment, the transmitter (440) may transmit additional data while transmitting encoded video. The video encoder (430) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0053] Before describing certain aspects of the embodiments of this disclosure in more detail, several terms that are referenced in the remainder of this specification are introduced below.

[0054] In some cases, "sub-picture" in the following text refers to a rectangular arrangement of samples, blocks, macroblocks, coding units, or similar entities that are semantically grouped and can be independently encoded at varying resolutions. One or more sub-pictures can form an image. One or more encoded sub-pictures can form an encoded image. One or more sub-pictures can be assembled into an image, and one or more sub-pictures can be extracted from an image. In some environments, one or more encoded sub-pictures can be assembled in a compressed domain without transcoding them to the sample level as encoded images, and in the same or some other cases, one or more encoded sub-pictures can be extracted from an encoded image in a compressed domain.

[0055] In the following text, "Adaptive Resolution Change (ARC)" refers to a mechanism that enables the resolution of images or sub-images in an encoded video sequence to be changed, for example, by resampling a reference image. "ARC parameters" refer to the control information required to perform adaptive resolution change, which may include, for example, filter parameters, scaling factors, the resolution of the output and / or reference images, various control flags, etc.

[0056] The Context-Adaptive Arithmetic Coding (CABAC) engine in High Efficiency Video Coding (HEVC) and VVC can utilize a table-based probability transformation process between 64 distinct representative probabilistic states. In HEVC, the range representing the states of the coding engine, ivlCurrRange, can be quantized into a set of four values ​​before computing the new range. The ivlCurrRange can be approximated using a table containing all 64 × 4 8-bit pre-computed values. The value of pLPS(pStateIdx) is used to implement HEVC state transitions, where pLPS is the probability of the Least Probable Symbol (LPS), and pStateIdx is the index of the current state. Decoding decisions can be implemented using a pre-computed LUT. The first ivlLpsRange can be obtained using the LUT, as shown below. Then, the ivlLpsRange can be used to update the ivlCurrRange and compute the output binVal (bit value).

[0057] Formula (1) ivlLpsRange = rangeTabLps[pStateIdx][qRangeIdx] In VVC, probabilities can be linearly represented by the probability index pStateIdx. Therefore, all computations can be performed using the formula without LUT operations. To improve the accuracy of probability estimation, multiple hypothesis probabilities can be applied to update the model. The pStateIdx used in interval subdivision in the binary arithmetic encoder can be a combination of two probabilities: pStateIdx0 and pStateIdx1. These two probabilities can be associated with each context model and can be updated independently with different adaptation rates. The adaptation rates of pStateIdx0 and pStateIdx1 for each context model can be pre-trained based on statistics of the associated binary terms. The probability estimate pStateIdx can be the average of the estimates from the two hypotheses.

[0058] Figure 5An embodiment of a process (500) for decoding a single binary decision is shown. The process (500) may begin at step (502) to determine the value of the variable ivlCurrRange. At step (504), if the variable ivlCurrRange is less than or equal to the variable ivlOffset, the process proceeds to step (506) to update the values ​​of variables binVal, ivlOffset, and ivlCurrRange. If the variable ivlCurrRange is greater than the value of the variable ivlOffset, the process proceeds to step (508) to update the value of the variable binVal. The process proceeds from either step (506) or step (508) to step (510) to update variables pStateIdx0 and pStateIdx1. The process proceeds to step (512) to execute the RenormD process.

[0059] As executed in HEVC, VVC CABAC can also have a QP-dependent initialization process called at the beginning of each slice. Given the initial value of the brightness QP for a slice, the initial probability state of the context model (denoted as preCtxState) can be derived as follows: Formula (2) m = slopeIdx × 5 – 45 Formula (3) n = (offsetIdx<<3) + 7 Formula (4) preCtxState = Clip3(1, 127, ((m × (QP) 32))>>4) + n), The slopeIdx and offsetIdx can be restricted to 3 bits, and the total initialization value can be represented with 6 bits of precision. The probability state preCtxState directly represents the probability in the linear domain. Therefore, before being input into the arithmetic coding engine, preCtxState may only require appropriate shift operations, and the mapping from the logarithmic domain to the linear domain and a 256-byte table are stored.

[0060] Formula (5) pStateIdx0 = preCtxState<<3 Formula (6) pStateIdx1 = preCtxState<<7 In AV1, an M-ary arithmetic coding engine can be used for entropy coding of syntax elements. Each syntax element can be associated with an alphabet of M elements, where M can be any integer value between 2 and 16. The input to the encoding can be an M-ary symbol and an encoding context, which can include a set of M probabilities represented by a Cumulative Distribution Function (CDF). The probabilities can be updated after each syntax element is encoded / parsed. The probability update rate refers to the frequency at which the probabilities are updated after each corresponding syntax element is encoded or parsed. The Cumulative Distribution Function can be an array of M 15-bit integers, as shown below: Formula (7) , in, c n / 32768 is the probability that the sign is less than or equal to n.

[0061] Probability updates can be performed using the following formula: Formula (8) , Here, α is an adaptive probability update rate based on the number of times a symbol has been decoded (up to 32 times), and m is the index of an element in the CDF. This adaptability of α enables faster probability updates when encoding / parsing syntax elements. The M-ary arithmetic coding process can follow the design of a traditional arithmetic coding engine. However, only the most valid 9 bits of the 15-bit probability value can be input to the arithmetic encoder / decoder. The probability update rate α associated with a symbol can be calculated based on the number of times the associated symbol appears during bitstream parsing, and the value of α can be reset at the beginning of a frame or tile using the following formula: Formula (9) , According to the equation above, the probability update rate has a large value at the beginning, and then saturates after 32 occurrences.

[0062] The multi-hypothesis probabilistic model used to encode M-ary symbols can include modifications to the AV1 arithmetic coding engine regarding multi-hypothesis estimation and regularization.

[0063] In multi-hypothesis estimation, AV1 can use a data-adaptive model for probability updates, where the fewer occurrences of a syntax element, the higher the update rate; conversely, the more observations, the lower the update rate. However, the engine uses only a single probability model. Multiple studies have shown that multi-hypothesis estimation can provide additional compression efficiency, where each syntax element maintains two or more probability tables with different update rates. Therefore, a multi-hypothesis probabilistic model can be implemented with two update rates: Formula (10) , Formula (11) , in, Modeling for faster updates, while Modeling slower updates. The final probabilistic model can be computed as a linear combination of hypotheses. Currently, the average of two hypotheses is used.

[0064] In regularization, a faster update rate This can lead to a strongly skewed distribution where the probabilities of some symbols decrease to near zero. Probabilities close to zero can cause a loss in the BD-rate (Bjontegaard-Delta rate). To counteract this effect, a regularization method is used, where at the end of each probability update, if the probability ( p m ) less than the threshold ( P min Then the regularization term can be applied to all probabilities, making p m Move to P min The regularization term can be taken from a uniform distribution and can depend on the sample space of the syntax elements.

[0065] The proposed multi-hypothesis probabilistic model for arithmetic coding engines uses the average of two hypotheses to update the probabilities of all contexts. However, traditional multi-hypothesis probabilistic modeling is agnostic to frame type (keyframe, inter-frame, intra-frame only, etc.), block size, and block prediction mode information. For different syntaxes, the optimal design of the probability update model can utilize this prior information to enable the arithmetic coding engine to adapt more quickly to the current symbol statistics, thereby improving coding performance.

[0066] This disclosure relates to a set of advanced video coding techniques, including an adaptive multiple hypothesis probabilistic model for arithmetic coding. Embodiments of this disclosure can be used individually or in any combination. Furthermore, each method (or embodiment), encoder, and decoder can be implemented using processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. In the following, an item block can be interpreted as a prediction block, coding block, or coding unit (i.e., CU). An item block can also be used to refer to a transform block. In the following items, when referring to block size, it can refer to the width or height of the block, or the maximum of the width and height, or the minimum of the width and height, or the area size (width). (height), or the aspect ratio of the block (width:height, or height:width).

[0067] This disclosure can also be applied based on slices or tiles. For example, when the following method is applied to a frame type, the same method can also be applied to slice types and / or tile types. In some embodiments, a faster update rate may refer to a larger update rate value, or a lower probability of updating the window size.

[0068] In some embodiments, the probability update rate of the probabilistic model used to update syntax elements may depend on the frame type of the current image in the reconstruction. In some embodiments, keyframes and / or intra-frames only use different update rates compared to update rates used for other frame types. For example, keyframes and / or intra-frames only may use a faster update rate for all or selected syntax elements (e.g., In some embodiments, inter-frame frames may use a slower update rate for all or selected syntax elements (e.g., ).

[0069] In some embodiments, when the final probability model is computed with an update rate When two or more hypotheses are linearly combined, the weights This can be used for the linear combination. For example, the probabilistic assumption can use the update rate. Export as The final probability update is derived as a weighted sum: In some embodiments, the weights used may depend on the frame type. For example, if Corresponding to a faster update rate, keyframes and / or only intra-frames can use larger weights. Similarly, for inter-frame frames, larger weights can be used for slower update rates.

[0070] In some embodiments, the probability update rate of the probabilistic model used to update syntax elements can depend on the coded block size. For example, for smaller block sizes, a faster update rate can be used for all syntax elements (e.g., ...). For example, blocks of 4×4, 4×8, 8×4, and 8×8 can use a faster update rate for their syntax elements. In another example, for larger block sizes, a smaller update rate can be used for all syntax elements (e.g., ...). For example, for blocks of sizes ranging from 16×16 to 256×256, a faster update rate can be used for the syntax elements of these blocks.

[0071] In some embodiments, when the final probability model is computed with an update rate When the weights are linear combinations of two or more assumptions, the weights This can be used for the linear combination. For example, the weights used depend on the block size. In some embodiments, if A faster update rate allows smaller blocks to use larger weights. Similarly, for larger blocks, larger weights can be used for slower update rates.

[0072] In some embodiments, the probability update rate of the probabilistic model used to update syntax elements may depend on the prediction mode information of the coded block. For example, for intra-prediction blocks, a faster update rate may be used for all or selected syntax elements (e.g., In another example, for inter-frame prediction blocks, a smaller update rate can be used for all or selected syntax elements (e.g., ).

[0073] In some embodiments, when the final probability model is computed with an update rate When the weights are linear combinations of two or more assumptions, the weights This can be used for the linear combination. For example, the probabilistic assumption can use the update rate. Export as The final probability update can be derived as a weighted sum: In some embodiments, the weights used may depend on the prediction model. For example, if Corresponding to a faster update rate, intra-frame prediction blocks can use larger weights. Similarly, for inter-frame prediction blocks, larger weights can be used for slower update rates.

[0074] In some embodiments, the probability update rate of the probabilistic model used to update syntax elements may depend on other encoding information, including but not limited to quantization step size, QP, time layer, frame resolution, and content type (e.g., whether a screen content encoding tool is used). In some embodiments, the probability update rate of the probabilistic model used to update syntax elements is specified in the high-level syntax, including but not limited to VPS, PPS, APS, SPS, image header, slice header, tile header, and CTU (or superblock) header.

[0075] Figure 6 An example flowchart is shown for performing an adaptive multi-hypothesis probabilistic modeling process (600) for arithmetic coding. Process (600) can be performed by a decoder such as a decoder (210). The process can begin at step (602) where an encoded video stream is received. The stream may include at least one image and one or more syntax elements encoded according to multi-hypothesis arithmetic coding. The process proceeds to step (604), where each syntax element is decoded based on multi-hypothesis arithmetic coding.

[0076] The process proceeds to step (606), where a probability update rate is selected from multiple probability update rates based on predetermined conditions. For example, the predetermined conditions may specify the frame type of at least one image in the bitstream, wherein the probability update rate is selected based on whether the frame type is a keyframe, an intra-frame frame, or an inter-frame frame. As another example, the predetermined conditions may specify the block size of at least one block in the image. As yet another example, the predetermined conditions may specify the prediction mode for at least one block.

[0077] The process proceeds to step (608), where at least one probabilistic model used in multi-hypothesis arithmetic coding is updated based on a selected probability update rate. The process proceeds to step (610), where at least one block in at least one image is decoded based on one or more decoded syntax elements.

[0078] The techniques described in the embodiments of this disclosure can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 7 A computer system (700) suitable for implementing embodiments of the disclosed subject matter is shown.

[0079] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly processed to create code containing instructions that can be executed directly by a computer's central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode, or similar means.

[0080] The instructions can be executed on various types of computers or their components, including personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things devices.

[0081] Figure 7 The components of the computer system (700) shown are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any one or a combination of components shown in the exemplary embodiments of the computer system (700). The computer system (700) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to one or more human users, for example, through input such as: tactile input (e.g., keystrokes, swipes, movement of a data glove), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images acquired from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0082] The input human-machine interface device may include one or more of the following (only one of each is shown): keyboard (701), mouse (702), touchpad (703), touch screen (710), data glove, joystick (705), microphone (706), scanner (707), and camera (708).

[0083] The computer system (700) may also include certain human-machine interface output devices. Such human-machine interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of a touchscreen (710), data gloves, or joysticks (705), but may also be tactile feedback devices that are not input devices). For example, such devices may be audio output devices (e.g., speakers (709), headphones (not shown)), visual output devices (e.g., screens (710) including CRT screens, LCD screens, plasma screens, OLED screens, each screen may or may not have touchscreen input functionality, each screen may or may not have tactile feedback functionality - some of these screens are capable of outputting two-dimensional or more three-dimensional visual outputs through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).

[0084] The computer system (700) may also include human-accessible storage devices and their associated media: such as optical media including CD / DVD ROM / RW (720) having media such as CD / DVD (721), finger drives (722), removable hard disk drives or solid-state drives (723), conventional magnetic media such as magnetic tapes and floppy disks (not shown), devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown), etc. Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0085] The computer system (700) may also include interfaces to one or more communication networks. These networks may be, for example, wireless networks, wired networks, or optical networks. Further, networks may be local area networks, wide area networks, metropolitan area networks, vehicle and industrial networks, real-time networks, latency-tolerant networks, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital television networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including CANBus, etc. Some networks typically require external network interface adapters (e.g., USB ports of the computer system (700)) to connect to certain general-purpose data ports or peripheral buses (749); other network interfaces are typically integrated into the core of the computer system (700) by connecting to the system bus (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). The computer system (700) can use any of these networks to communicate with other entities. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., CANbus connected to certain CANbus devices), or bidirectional, such as connecting to other computer systems using a local area network or wide area network (WAN) digital network. This communication can include communication to a cloud computing environment (755). As mentioned above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.

[0086] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface (754) can be attached to the kernel (740) of the computer system (700).

[0087] The core (740) may include one or more central processing units (CPU) (741), graphics processing units (GPUs) (742), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (743), hardware accelerators (744) for certain tasks, etc. These devices, as well as read-only memory (ROM) (745), random access memory (746), and internal mass storage (747) such as internal non-user-accessible hard disk drives, SSDs, etc., may be connected via a system bus (748). In some computer systems, the system bus (748) may be accessed in the form of one or more physical plugs to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the core's system bus (748) or connected to the core's system bus via a peripheral bus (749). Peripheral bus architectures include PCI, USB, etc. A graphics adapter (750) may be included in the core (740).

[0088] The CPU (741), GPU (742), FPGA (743), and accelerator (744) can execute certain instructions that can be combined to form the aforementioned computer code. This computer code can be stored in ROM (745) or RAM (746). Transient data can also be stored in RAM (746), while permanent data can be stored, for example, in internal mass storage (747). Fast storage and retrieval to any storage device can be achieved using a cache, which can be closely associated with one or more CPUs (741), GPUs (742), mass storage (747), ROM (745), RAM (746), etc.

[0089] Computer-readable media may have computer code thereon for performing various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software. As a non-limiting example, a computer system having an architecture (700), particularly a kernel (740), can be made functional by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, and some non-transitory memory of the kernel (740), such as internal kernel mass storage (747) or ROM (745). Software implementing various embodiments of this disclosure can be stored in such means and executed by the kernel (740). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the kernel (740), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures (746) stored in RAM and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality through hard-wired or otherwise embodied logic in circuitry (e.g., the accelerator (744)), which may replace or operate with the software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0090] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit implementation to the precise forms disclosed. Modifications and variations are possible based on the foregoing disclosure, or may be obtained from practice of the embodiments.

[0091] It should be understood that the specific order or hierarchy of blocks in the process / flowcharts disclosed herein is illustrative of the exemplary method. Based on design preferences, it is understood that the specific order or hierarchy of blocks in the process / flowcharts may be rearranged. Furthermore, some blocks may be combined or omitted. The appended method claims present elements of various blocks in a sample order and are not intended to limit one to the specific order or hierarchy presented.

[0092] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of integration technical detail. Furthermore, one or more of the above components may be implemented as instructions stored on a computer-readable medium and executable by at least one processor (and / or may include at least one processor). The computer-readable medium may include a computer-readable non-transitory storage medium having computer-readable program instructions thereon for causing the processor to perform operations.

[0093] Computer-readable storage media can be tangible devices that can retain and store instructions for use by an instruction execution device. For example, a computer-readable storage medium can be, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanical encoding devices (e.g., punched cards or raised structures in recesses on which instructions are recorded), and any suitable combination of the foregoing. The computer-readable storage medium used in this document should not be construed as a transient signal, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., optical pulses through fiber optic cables), or electrical signals transmitted through wires.

[0094] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or downloaded to an external computer or external storage device. This network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the corresponding computing / processing device.

[0095] Computer-readable program code / instructions used to perform operations can be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and procedural programming languages ​​such as the "C" programming language or similar programming languages. Computer-readable program instructions can execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the internet through an internet service provider). In some embodiments, electronic circuits including programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) can execute computer-readable program instructions by utilizing state information of computer-readable program instructions to personalize the electronic circuits, thereby performing various aspects or operations.

[0096] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in the flowchart and / or one or more block diagram blocks. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other apparatus to operate in a particular manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture of instructions implementing aspects of the functions / actions specified in the flowchart and / or one or more block diagram blocks.

[0097] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device, thereby producing a computer-implemented process that enables the function / action specified in the flowchart and / or one or more block diagram blocks to be implemented on the computer, other programmable apparatus or other device.

[0098] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. The method, computer system, and computer-readable medium may include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in the figures. In some alternative implementations, the functions indicated in the blocks may appear in the order indicated in the figures. For example, in fact, two blocks shown consecutively may be executed simultaneously or substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware that performs the specified function or action, or a combination of dedicated hardware and computer instructions.

[0099] It is evident that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit these implementations. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to any specific software code. It should be understood that software and hardware can be designed to implement the systems and / or methods based on the descriptions herein.

[0100] The above disclosure also includes the following embodiments: (1) A method executed by at least one processor of a video decoder, the method comprising: receiving an encoded video stream, the encoded video stream comprising at least one picture encoded according to a multi-hypothesis arithmetic coding and one or more syntax elements; decoding each of the one or more syntax elements based on the multi-hypothesis arithmetic coding; selecting a probability update rate from a plurality of probability update rates based on predetermined conditions, the plurality of probability update rates including a first probability update rate higher than a second probability update rate; updating at least one probability model used in the multi-hypothesis arithmetic coding based on the selected probability update rate; and decoding at least one block in the at least one picture based on the decoded one or more syntax elements.

[0101] (2) The method according to feature (1), wherein the predetermined condition specifies the frame type of the at least one image.

[0102] (3) The method according to feature (2), wherein the selection includes: selecting the first probability update rate in response to determining that the frame type of the at least one image is either a keyframe or an intraframe.

[0103] (4) The method according to feature (2) or (3), wherein the selection includes: selecting the second probability update rate in response to determining that the frame type of the at least one image is an inter-frame frame.

[0104] (5) The method according to any one of features (2) to (4), wherein the multi-hypothesis arithmetic coding comprises a linear combination of a plurality of probability models, wherein each probability model is derived using a probability update rate corresponding to the plurality of probability update rates, and wherein the linear combination is a weighted sum using a plurality of weights, wherein each weight depends on the frame type of the at least one image.

[0105] (6) The method according to feature (5), wherein when determining that the frame type is either a key frame or an intra frame, a greater weight is used than when determining that the frame type is an inter-frame frame.

[0106] (7) The method according to any one of features (1) to (6), wherein the predetermined condition specifies the block size of the at least one block.

[0107] (8) The method according to feature (7), wherein, in response to determining that the at least one block is an N×M block, the first probability update rate is selected, wherein N is one of 4 and 8, and M is one of 4 and 8.

[0108] (9) The method according to feature (7) or (8), wherein, in response to determining that the at least one block is an N×M block, the second probability update rate is selected, wherein N is greater than or equal to 16 and M is greater than or equal to 16.

[0109] (10) The method according to any one of features (7) to (9), wherein the multi-hypothesis arithmetic coding comprises a linear combination of a plurality of probability models, wherein each probability model is derived using a probability update rate corresponding to the plurality of probability update rates, and wherein the linear combination is a weighted sum using a plurality of weights, wherein each weight depends on the block size of the at least one block. (11) The method according to feature (10), wherein a first weight among the plurality of weights for the first block is greater than a second weight among the plurality of weights for the second block, and the first block is smaller than the second block.

[0110] (12) The method according to any one of features (2) to (11), wherein the predetermined condition specifies the prediction mode of the at least one block.

[0111] (13) The method according to feature (12), wherein, in response to determining that the prediction mode of the at least one block is an intra-frame prediction mode, the first probability update rate is selected.

[0112] (13) The method according to feature (12) or (13), wherein, in response to determining that the prediction mode of the at least one block is an inter-frame prediction mode, the second probability update rate is selected.

[0113] (15) The method according to any one of features (12) to (14), wherein the multi-hypothesis arithmetic coding comprises a linear combination of a plurality of probability models, wherein each probability model is derived using a probability update rate corresponding to the plurality of probability update rates, and wherein the linear combination is a weighted sum using a plurality of weights, wherein each weight depends on the prediction mode of the at least one block.

[0114] (16) The method according to feature (15), wherein the weight is greater when the prediction mode is determined to be an intra-frame prediction mode than when the prediction mode is determined to be an inter-frame prediction mode.

[0115] (17) A video decoder, comprising: at least one memory configured to store computer program code; and at least one processor configured to access and operate according to instructions of the computer program code, the computer program code comprising: receiving code configured to cause the at least one processor to receive an encoded video stream, the encoded video stream comprising at least one picture encoded according to a multi-hypothesis arithmetic coding and one or more syntax elements; a first decoding code configured to cause the at least one processor to decode each of the one or more syntax elements based on the multi-hypothesis arithmetic coding; a selection code configured to cause the at least one processor to select a probability update rate from a plurality of probability update rates based on predetermined conditions, the plurality of probability update rates including a first probability update rate higher than a second probability update rate; an update code configured to cause the at least one processor to update at least one probability model used in the multi-hypothesis arithmetic coding based on the selected probability update rate; and a second decoding code configured to cause the at least one processor to decode at least one block in the at least one picture based on the decoded one or more syntax elements.

[0116] (18) The video decoder according to feature (17), wherein the predetermined condition specifies the frame type of the at least one picture.

[0117] (19) The video decoder according to feature (18), wherein the selection code is further configured such that the at least one processor: in response to determining that the frame type of the at least one image is either a keyframe or an intraframe, selects the first probability update rate.

[0118] (20) A non-transitory computer-readable medium storing instructions, which, when executed by a processor in a video decoder, cause the processor to perform a method comprising: receiving an encoded video stream, the encoded video stream comprising at least one picture encoded according to a multi-hypothesis arithmetic coding and one or more syntax elements; decoding each of the one or more syntax elements based on the multi-hypothesis arithmetic coding; selecting a probability update rate from a plurality of probability update rates based on predetermined conditions, the plurality of probability update rates including a first probability update rate higher than a second probability update rate; updating at least one probability model used in the multi-hypothesis arithmetic coding based on the selected probability update rate; and decoding at least one block in the at least one picture based on the decoded one or more syntax elements.

Claims

1. A method of video decoding, the method comprising: The method includes: Receive an encoded video stream, the encoded video stream comprising at least one image and one or more syntax elements encoded according to multiple hypothesis arithmetic coding; Decode each of the one or more syntax elements based on the multi-hypothesis arithmetic encoding; A probability update rate is selected from multiple probability update rates based on predetermined conditions, wherein the multiple probability update rates include a first probability update rate that is higher than a second probability update rate; the predetermined conditions include one or more of the following: frame type of at least one image, block size of at least one block, prediction mode of at least one block, quantization step size, quantization parameter QP, temporal layer, frame resolution, and content type; each of the multiple probability update rates refers to the frequency at which the probability is updated after each corresponding syntax element is encoded or parsed. Update at least one probabilistic model used in the multi-hypothesis arithmetic coding based on the selected probability update rate; and At least one block in the at least one image is decoded based on one or more syntax elements of the decoded image; the multi-hypothesis arithmetic coding includes a linear combination of a plurality of the probability models, wherein each of the probability models is derived using a probability update rate corresponding to one of the plurality of probability update rates, and the linear combination is a weighted sum using a plurality of weights, each of the weights depending at least on the frame type of the at least one image.

2. The method of claim 1, wherein, The selection includes: selecting the first probability update rate in response to determining that the frame type of the at least one image is either a keyframe or an intraframe.

3. The method according to claim 1, characterized in that, The selection includes: in response to determining that the frame type of the at least one image is an inter-frame frame, selecting the second probability update rate.

4. The method according to claim 1, characterized in that, When determining that the frame type is either a keyframe or an intraframe, a greater weight is used than when determining that the frame type is an interframe frame.

5. The method according to claim 1, characterized in that, In response to determining that the at least one block is an N×M block, the first probability update rate is selected, wherein N is one of 4 and 8, and M is one of 4 and 8.

6. The method according to claim 1, characterized in that, In response to determining that the at least one block is an N×M block, a second probability update rate is selected, wherein N is greater than or equal to 16 and M is greater than or equal to 16.

7. The method according to claim 1, characterized in that, Each of the weights also depends on the block size of the at least one block.

8. The method according to claim 7, characterized in that, The first weight among the plurality of weights for the first block is greater than the second weight among the plurality of weights for the second block, and the first block is less than the second block.

9. The method according to claim 1, characterized in that, In response to determining that the prediction mode of the at least one block is an intra-frame prediction mode, the first probability update rate is selected.

10. The method according to claim 1, characterized in that, In response to determining that the prediction mode of the at least one block is an inter-frame prediction mode, the second probability update rate is selected.

11. The method according to claim 1, characterized in that, Each of the weights also depends on the prediction mode of the at least one block.

12. The method according to claim 11, characterized in that, The weight is greater when the prediction mode is determined to be an intra-frame prediction mode than when the prediction mode is determined to be an inter-frame prediction mode.

13. A video decoder, characterized in that, include: At least one memory is configured to store computer program code; as well as At least one processor is configured to access the computer program code and operate in accordance with the instructions of the computer program code to implement the method of any one of claims 1 to 12.

14. A non-transitory computer-readable medium, characterized in that, The device stores instructions that, when executed by a processor in a video decoder, cause the processor to perform the method of any one of claims 1 to 12.

Citation Information

Patent Citations

  • Binary arithmetic coding with progressive modification of adaptation parameters

    US20190110080A1