Method, device, and storage medium for multi-symbol arithmetic coding
Multiple-symbol arithmetic coding with cumulative distribution function updates addresses inefficiencies in video coding, enhancing compression and quality through dynamic probability adaptation.
Patent Information
- Application Number
- JP2025061869
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-21
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-03
AI Technical Summary
Existing video coding technologies face inefficiencies in reducing redundancy and achieving high compression ratios while maintaining acceptable video quality, particularly in the context of intra and inter-picture prediction methods.
Implement multiple-symbol arithmetic coding in video decoding, involving the use of an array of cumulative distribution functions to update probability rates for each extracted symbol, enhancing the efficiency of entropy coding.
Improves compression efficiency and flexibility in video coding by dynamically adapting probability estimates, leading to more effective reduction of redundancy and improved video quality.
Smart Images

Figure 2025100580000001_ABST
Abstract
Description
Technical Field
[0001] Related Applications This application claims the benefit of priority based on U.S. Provisional Application No. 63 / 306,386, filed on February 3, 2022, and U.S. Non-Provisional Application No. 17 / 869,964, filed on July 21, 2022. The entire disclosures of both applications are incorporated herein by reference.
[0002] The present disclosure relates to video coding and / or decoding techniques, and more particularly, to improved designs and / or implementations of multiple-symbol arithmetic coding and / or decoding.
Background Art
[0003] The description of the background art herein is for the purpose of generally indicating the background of the disclosure. In the context of the description of the background art, the achievements of the inventors whose names are currently listed, and other aspects of the description that cannot obtain the status of prior art as of the filing date of the present application, are not to be recognized as prior art to the present disclosure, either explicitly or implicitly.
[0004] Video coding and decoding can be performed using inter-picture prediction together with motion compensation. Uncompressed digital video may include a sequence of pictures, each picture having a given spatial size, for example a spatial size of 1920×1080 luminance samples, and associated full chrominance samples or chrominance samples by subsampling. The picture sequence has a picture rate (also known as frame rate) that may be constant or variable, for example having a picture rate of 60 pictures per second, i.e., 60 frames per second. Uncompressed video has specific bitrate requirements for streaming and data processing. For example, video with a pixel resolution of 1920×1080, a frame rate of 60 per second, and chroma subsampling of 4:2:0 with 8 bits per pixel per color channel requires a bandwidth of about 1.5 Gbit / s. Such video requires more than 600 GBytes of storage space for one hour.
[0005] One of the purposes of video coding and decoding can be said to be to reduce the redundancy of the input video signal compressed by compression. Compression can be said to be useful for relaxing the above-mentioned bandwidth requirements and / or storage area requirements, and in some cases, it can be said to be useful for relaxing by two digits or more. Both reversible compression and irreversible compression, and combinations thereof, can be used. Reversible compression refers to a method in which an exact replica of the original signal can be reconstructed from the compressed original signal through a decoding process. Irreversible compression refers to a coding / decoding process in which the original video information is not completely retained during coding and cannot be completely restored during decoding. When using irreversible compression, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough that, despite some loss of information, the reconstructed signal can be used for the intended purpose. In the case of video, irreversible compression is widely used for many applications. The amount of acceptable distortion depends on the application. For example, users of certain consumer streaming applications may tolerate greater distortion than users of movie or television broadcast applications. The compression ratio achievable by a specific coding algorithm can be selected or adjusted to reflect various distortion tolerances. That is, a large acceptable distortion usually enables a coding algorithm that results in large-scale losses and a high compression ratio.
[0006] For example, techniques obtained from several broad categories and steps, including motion compensation, Fourier transform, quantization, and entropy coding, can be utilized in video encoders and decoders.
[0007] The video codec technology can include techniques known as intracoding. In intracoding, sample values are expressed without referring to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in an intra mode, the picture can be referred to as an intra picture. Since the decoder state can be reset using an intra picture and its derivatives such as an independent decoder refresh picture, this can be used as the first picture of a coded video bitstream or video session, or as a still picture. Thereafter, the samples of the blocks after intra prediction can be subjected to a transform to the frequency domain, and the transform coefficients thus generated can be quantized prior to entropy coding. Intra prediction is representative of techniques for minimizing sample values in the domain before transformation. In some cases, the smaller the DC value and the AC coefficients after transformation, the smaller the size of the bits required for a given quantization step representing the block after entropy coding.
[0008] For example, conventional intracoding such as coding known in MPEG-2 generation coding techniques does not use intra prediction. However, some more recent video compression techniques include techniques that attempt to code / decoder blocks based on surrounding sample data and / or metadata obtained during encoding and / or decoding of spatially neighboring ones, and / or surrounding sample data and / or metadata preceding in decoding order in a block of intracoded or decoded data. Hereinafter, such techniques will be referred to as "intra prediction" techniques. Note that in at least some examples, intra prediction uses only reference data of the current picture being reconstructed and does not use reference data of other reference pictures.
[0009] It is possible that there are many different forms of intra prediction. If one or more of such techniques are available for a given video coding technique, the technique(s) used can be referred to as an intra prediction mode. One or more intra prediction modes may be performed in a specific codec. In some examples, a mode can have sub-modes and / or be associated with various parameters, and the mode / sub-mode information and the intra coding parameters of the video block can be coded separately or included together in a mode codeword. Depending on which codeword is used for a given combination of mode, sub-mode, and / or parameters, it can affect the improvement of coding efficiency by intra prediction, and thus can affect the entropy coding technique used to convert the codeword into a bitstream.
[0010] A specific mode of intra prediction was introduced from H.264, improved in H.265, and further improved in new coding techniques such as joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). Usually, in intra prediction, a predicted value block can be formed using the values of neighboring samples that are available. For example, the available values of a specific set of neighboring samples arranged along a predetermined direction and / or line can be replicated and put into the predicted value block. The reference value in the direction being used can be coded in the bitstream, or the reference value itself can be predicted.
[0011] Referring to FIG. 1A, a subset of nine prediction directions specified in 33 possible intra prediction value directions of H.265 (corresponding to 33 angular modes out of the 35 intra modes described in H.265) is shown in the lower right. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample at 101 is predicted using neighboring samples. For example, arrow (102) indicates that the sample (101) is predicted from one or more neighboring samples diagonally up to the right at an angle of 45 degrees from the horizontal direction. Similarly, arrow (103) indicates that the sample (101) is predicted from one or more neighboring samples diagonally down to the left of the sample (101) at an angle of 22.5 degrees from the horizontal direction.
[0012] Continuing to refer to FIG. 1A, a square block (104) of 4×4 samples is shown in the upper left (indicated by the thick dashed line). The square block (104) contains 16 samples, each of which is labeled with an index including "S", the Y - dimensional position (e.g., row index) of the sample, and the X - dimensional position (e.g., column index) of the sample. For example, sample S21 is the second sample in the Y - dimension (from the top) and the first sample in the X - dimension (from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y - dimension and the X - dimension. If the block size is 4×4 samples, S44 is in the lower right. Further examples of reference samples following the same numbering scheme are shown. The reference samples are labeled with an index including "R" and the Y - position (e.g., row index) and X - position (column index) of the reference sample relative to the block (104). In both H.264 and H.265, neighboring predicted samples adjacent to the block during reconstruction are used.
[0013] Picture prediction in block 104 can be started by replicating the reference sample value from the neighboring samples according to the signaled prediction direction. For example, for block 104, assume that the coding video bitstream contains signaling indicating the prediction direction of arrow (102). That is, assume that the coding video bitstream contains signaling indicating that samples are predicted from one or more prediction samples diagonally up to the right at an angle of 45 degrees from the horizontal direction. In such a case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.
[0014] In some examples, especially when the directions cannot be evenly divided at 45 degrees, the values of multiple reference samples may be combined, for example, through interpolation, to calculate the reference sample.
[0015] As video coding technology continues to evolve, the number of possible directions has increased. In H.264 (2003), for example, nine different directions are available for intra prediction. This increased to 33 in H.265 (2013), but JEM / VVC / BMS can support up to 65 directions at the time of this disclosure. Although experimental considerations have been made to assist in identifying the optimal intra prediction direction, some techniques of entropy coding can be used while tolerating some disadvantage of bits for the direction to code such an optimal direction in the case of a small number of bits. Further, in some cases, the direction itself can be predicted from the neighboring directions used in the intra prediction of the decoded neighboring blocks.
[0016] FIG. 1B shows a schematic diagram (180) showing 65 intra prediction directions according to JEM, and many prediction directions of various coding technologies are shown.
[0017] The way of mapping the bits representing the intra prediction direction in the coded video bitstream to the prediction direction may vary depending on the video coding technology, and this mapping method can range from, for example, a simple direct mapping of the prediction direction to techniques such as a complex adaptation method that results in an intra prediction mode, a codeword, and the most likely mode. However, in all cases, there may be some directions of intro prediction that are statistically less likely to occur in the video content than some other directions. Since the purpose of video compression is to reduce redundancy, such less likely directions may be represented by a larger number of bits than the more likely directions if the video coding technology is properly designed.
[0018] Inter-picture prediction (i.e., inter prediction) may be based on motion compensation. In motion compensation, sample data obtained from a previously reconstructed picture or a part thereof (reference picture) is spatially shifted in the direction indicated by a motion vector (hereinafter, MV), and then used for predicting a newly reconstructed picture or a part of the picture (e.g., a block). In some cases, the reference picture may be the same as the picture at the time when the reconstruction is being performed. The MV may have two dimensions X and Y or three dimensions, and the third dimension is an index related to the reference picture being used (similar to the temporal dimension).
[0019] In some video compression techniques, the current motion vector (MV) applicable to a region of sample data can be predicted from other MVs. For example, it can be predicted from another MV that is related to another region of sample data that is spatially adjacent to the region being reconstructed and that precedes the current MV in decoding order. By doing so, the total amount of data required to code the MVs can be substantially reduced as the redundancy of the correlated MVs is removed, thereby increasing the compression efficiency. For example, when coding an input video signal obtained from a camera (referred to as raw video), since a region larger than the region to which one MV is applicable moves in the same direction in the video sequence, in some cases, there is a statistical likelihood that the region can be predicted using a similar motion vector derived from the MVs of neighboring regions, so MV prediction can function effectively. As a result, the actual MV of a given region may become similar to or the same as the predicted MV from surrounding MVs. After entropy coding, the above MV can be represented with fewer bits than the number of bits that would be used if the MV were coded directly without being predicted from one or more neighboring MVs. In some examples, MV prediction can be an example where lossless compression of a signal (i.e., an MV) is derived from an original signal (i.e., a sample stream). In other examples, for instance, MV prediction itself may become irreversible due to rounding errors when calculating a predicted value from several surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, “High Efficiency Video Coding”, December 2016). Among the many MV prediction mechanisms described in H.265, the technique hereinafter referred to as “spatial merge” will be described below.
[0021] Specifically, referring to FIG. 2, the current block (201) comprises samples that are predictable from a previous block of the same size that has been spatially shifted and detected by an encoder during a motion search process. Instead of directly coding the MV, the MV can be derived from metadata associated with one or more reference pictures. For example, the MV can be derived from the latest (in decoding order) reference picture using the MV associated with any one of five surrounding samples shown at A0, A1 and B0, B1, B2 (202-206 respectively). In H.265, in MV prediction, a predicted value obtained from the same reference picture used by neighboring blocks can be used.
Summary of the Invention
Means for Solving the Problems
[0022] In the present disclosure, various embodiments of methods, apparatuses, and computer-readable storage media for video encoding and / or decoding are described.
[0023] According to one aspect, an embodiment of the present disclosure provides a method for multiple symbol arithmetic coding in video decoding. The method includes steps in which a device receives a coded video bitstream. The device includes a memory for storing instructions and a processor in communication with the memory. The method also includes steps in which the device obtains an array of cumulative distribution functions and corresponding M symbols, where M is an integer greater than 1, and the device performs arithmetic decoding on the coded video bitstream to extract at least one symbol based on the array of cumulative distribution functions, and for each extracted symbol, the device updates the array of cumulative distribution functions according to the extracted symbol based on at least one probability update rate, and the device continues to perform arithmetic decoding on the coded video bitstream to extract the next symbol based on the updated array of cumulative distribution functions.
[0024] In another aspect, in an embodiment of the present disclosure, an apparatus for multiple symbol arithmetic coding in video decoding is provided. The apparatus includes a memory for storing instructions and a processor communicating with the memory. When the processor executes the instructions, the processor is configured to cause the apparatus to perform the above method for video decoding and / or encoding.
[0025] In another aspect, in an embodiment of the present disclosure, a non-transitory computer-readable medium storing instructions for causing a computer to perform the above method for video decoding and / or encoding when executed by the computer is provided.
[0026] The above aspects and other aspects, as well as examples of realizing these aspects, are described in more detail in the drawings, the description, and the claims.
[0027] From the following detailed description and the accompanying drawings, further features, properties, and various effects of the disclosed subject matter will become more apparent.
Brief Description of the Drawings
[0028]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
DETAILED DESCRIPTION OF THE INVENTION
[0029] Hereinafter, a part of the present invention will be described in detail with reference to the accompanying drawings showing specific examples of embodiments as illustrative examples. On the other hand, the present invention may be implemented in various different forms, and thus, it is intended that the protected subject matter covered or claimed is not limited to any of the embodiments described below. Also, note that the present invention may be implemented as a method, a device, a component, or a system. Therefore, embodiments of the present invention may take the form of, for example, hardware, software, firmware, or any combination thereof.
[0030] Throughout the specification and claims, terms may have special meanings suggested or implied by the context beyond the explicitly stated meaning. The phrase "in one embodiment" or the phrase "in some embodiments" used in this disclosure does not necessarily refer to the same embodiment, and the phrase "in another embodiment" or "in other embodiments" used in this disclosure does not necessarily refer to different embodiments. Similarly, the phrase "in one implementation" or the phrase "in some implementations" used in this disclosure does not necessarily refer to the same implementation, and the phrase "in another implementation" or "in other implementations" used in this disclosure does not necessarily refer to different implementations. The intention is, for example, that the claimed protected subject matter includes, in whole or in part, a combination of typical embodiments / implementations.
[0031] Overall, the grammar of a term may be at least partially determined by its usage in context. For example, terms such as "and", "or", or "and / or" as used in this disclosure may include various meanings that may depend at least in part on the context in which the term is used. Typically, when "or" is used to relate a list such as A, B, or C, "or" is intended to mean A, B, or C, and in this case it is used in an inclusive sense. Additionally, "or" may be intended to mean either A, B, or C, and in this case it is used in an exclusive sense. In addition, the term "one or more" or the term "at least one" as used in this disclosure may, at least in part, depending on the context, be used to describe any singular feature, structure, or characteristic, or may be used to describe a combination of plural features, structures, or characteristics. Similarly, terms such as "a", "an", or "the" may also be understood to recognize singular or plural usage, at least in part, depending on the context. In addition to the above, the term "based on" or the term "determined by" may also, in this case, at least in part, depending on the context, not necessarily be understood to necessarily intend to recognize an exclusive set of factors, and furthermore, there may be room for the existence of other factors that are not necessarily explicitly described.
[0032] FIG. 3 shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) may perform unidirectional data transmission. For example, the terminal device (310) may code video data (for example, video data of a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The coded video data can be transmitted as one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to restore video pictures, and display video pictures corresponding to the restored video data. Unidirectional data transmission may be performed in a media providing application or the like.
[0033] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of coded video data, which may be implemented, for example, when applied to a video conference. In the case of bidirectional data transmission, as an example, each of the terminal devices (330) and (340) may code video data (for example, video data of a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) may also receive the coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to restore video pictures, and display video pictures corresponding to the restored video data on an accessible display device.
[0034] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) can be implemented as a server, a personal computer, and a smartphone. However, it can be said that the scope of application of the fundamental principle of the present disclosure is not limited in this way. Embodiments of the present disclosure may be implemented with a desktop computer, a laptop computer, a tablet computer, a media player, a wearable computer, dedicated videoconferencing equipment, and / or the like. The network (350) represents any number and any type of network that transmits coded video data between the terminal devices (310), (320), (330), and (340), including, for example, a wired and / or wireless communication network. The communication network (350) may exchange data through a circuit-switched channel, a packet-switched channel, and / or other types of channels. Representative networks include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the description of this example, unless explicitly described in this application, the architecture and topology of the network (350) may not be important for the operation of the present disclosure.
[0035] FIG. 4 shows the arrangement of a video encoder and a video decoder in a video streaming environment as an example of applying to the disclosed protection target. The disclosed protection target may be similarly applicable to other application examples of video, including, for example, videoconferencing, digital television broadcasting, games, virtual reality, storage of compressed video on digital media such as CDs, DVDs, memory sticks, etc.
[0036] The video streaming system may include a video source (401), such as a digital camera, that generates a stream (402) of uncompressed video pictures or images. In one example, the stream (402) of video pictures includes samples recorded by the digital camera of the video source 401. The stream (402) of video pictures is shown in bold lines to emphasize the large amount of data when compared to the coded video data (404) (i.e., the encoded video bitstream), but the stream (402) of video pictures can be processed by an electronic device (420) that includes a video encoder (403) connected to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof that enables or implements the disclosed aspects of the subject matter to be described in detail hereinafter. The encoded video data (404) (i.e., the encoded video bitstream (404)) is shown in thin lines to emphasize the small amount of data when compared to the stream (402) of uncompressed video pictures, but this encoded video data (404) can be stored in the streaming server (405) for later use or directly stored in a downstream video device (not shown). One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to obtain copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) inputs and decodes a copy (407) of the encoded video data and generates an output stream (411) of video pictures that are not compressed and can be drawn on a display (412) (e.g., a display screen) or other drawing device (not shown). The video decoder 410 may be configured to perform some or all of the various functions described in this disclosure.In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to a specific video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used assuming VVC or other video coding standards.
[0037] Note that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and similarly, the electronic device (430) can include a video encoder (not shown).
[0038] FIG. 5 shows a block diagram of a video decoder (510) according to any one of the embodiments of the present disclosure described below. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used instead of the video decoder (410) in the example of FIG. 4.
[0039] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). It may decode one coded video sequence at a time in the same embodiment or in another embodiment, and the decoding of each coded video sequence is independent of other coded video sequences. Each video sequence may be related to a plurality of video frames or images. The coded video sequence may be received from a channel (501), and the channel (501) may be a storage device storing the encoded video data or hardware / software connected to a streaming source that sends the encoded video data. The receiver (531) may receive the encoded video data together with other data such as coded audio data and / or an attached data stream, and transfer the other data to a corresponding processing circuit (not shown). The receiver (531) may separate the coded video sequence from other data. To handle network jitter, a buffer memory (515) may be arranged between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as the "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may exist outside the video decoder (510) and be separable from the video decoder (510) (not shown). In still other applications, for example, to handle network jitter, a buffer memory (not shown) may exist outside the video decoder (510), and for example, to process the playback timing, another additional buffer memory (515) may exist inside the video decoder (510). If the receiver (531) receives data from a storage / transfer device or an isochronous network with sufficient bandwidth and controllability, the buffer memory (515) may not be required, or the buffer memory (515) may be small.When used in a best-effort packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and the size of the buffer memory (515) can be made relatively large. Such a buffer memory may be implemented in an optimal size, or may be at least partially implemented by an operating system or similar elements (not shown) outside the video decoder (510).
[0040] The video decoder (510) may include a parser (520) that reconstructs symbols (521) from the coded video sequence. Such symbol categories may include information used to manage the operation of the video decoder (510), and as shown in FIG. 5, information for controlling a rendering device such as a display (512) (e.g., a display screen) that may or may not be an integral part of the electronic device (530) and can be connected to the electronic device (530). The control information for one or more rendering devices may take the form of Supplemental Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) can parse / entropy decode the coded video sequence received by the parser (520). The entropy coding of the coded video sequence can conform to video coding techniques and standards, and can follow various principles including variable length coding, Huffman coding, and arithmetic coding with or without context dependence. The parser (520) may extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the subgroup. The subgroups can include Group of Picture (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The parser (520) may also perform extraction from coded video sequence information such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc.
[0041] The parser (520) may generate symbols (521) by performing entropy decoding / parser operations on the video sequence received from the buffer memory (515).
[0042] The reconstruction of the symbol (521) can involve multiple different processing sites or functional sites according to the type of the coded video picture or a portion thereof (e.g., between pictures and within a picture, between blocks and within a block) and other factors. The sites to be involved and how to involve the sites may be controlled by subgroup control information parsed by a parser (520) from the coded video sequence. For the sake of simplification, the flow of such subgroup control information between the parser (520) and a plurality of processing sites or functional sites described later is not shown.
[0043] Regarding the functional blocks already described, further speaking, the video decoder (510) can be conceptually subdivided into a plurality of functional sites as described later. In an actual implementation example operated under commercial constraints, many of such functional sites interact closely with each other and can be at least partially integrated with each other. On the other hand, in order to clarify and describe the various functions of the disclosed protection object, the following disclosure adopts the conceptually subdivision into functional sites.
[0044] The first site may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) may receive, as symbols (521) from the parser (520), quantized transform coefficients, information indicating which type of inverse transform should be used, block size, quantization factor / quantization parameter, quantization scaling matrix, and control information including a state. The scaler / inverse transform unit (551) can output a block having sample values that can be input to an assembler (555).
[0045] In some examples, the output samples of the scaler / inverse transform (551) may be suitable for an intra-coding block, i.e., a block that does not use predictive information obtained from a previously reconstructed picture, but can use predictive information obtained from a previously reconstructed portion of the current picture. Such predictive information can be provided by the intra-picture prediction unit (552). In some examples, the intra-picture prediction unit (552) may generate a block of the same size and shape as the block being reconstructed, using information from surrounding blocks that have already been reconstructed and stored in the current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (555) may, for each sample, add the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).
[0046] In another example, the output samples of the scaler / inverse transform unit (551) may be suitable for an inter-coding block and may, in some cases, be suitable for a motion-compensated block. In such an example, the motion compensation prediction unit (553) can access the reference picture memory (557) to obtain samples used for picture-to-picture prediction. Depending on the symbol (521) related to the block, after motion-compensating the obtained samples, the samples are added by an aggregator (555) to the output of the scaler / inverse transform unit (551) (the output of part 551 may be referred to as a residual sample or a residual signal) to generate output sample information. The address in the reference picture memory (557) from which the motion compensation prediction unit (553) obtains the prediction samples can be controlled by a motion vector, and the motion vector can be available to the motion compensation prediction unit (553) in the form of a symbol (521) that can have, for example, X, Y components (shifts) and a reference picture component (time). Motion compensation can also include interpolation of sample values obtained from the reference picture memory (557) when an accurate sub-sample motion vector is used, and can also be related to a motion vector prediction mechanism or the like.
[0047] The output samples of the aggregator (555) can undergo various loop filtering techniques in the loop filter unit (556). The video compression technology includes loop filter technology that is coded as a symbol (521) from the parser (520) in a video sequence (also referred to as a coded video bitstream) and is controlled by parameters that can be used in the loop filter unit (556). However, it is susceptible to the influence of meta-information obtained when decoding a coded picture or a previous part (in decoding order) of a coded video sequence, and may also be susceptible to the influence of sample values that have been previously reconstructed and passed through the loop filter. Several types of loop filters may be included as part of the loop filter unit 556 in various orders. This will be described in more detail later.
[0048] The output of the loop filter section (556) can be output not only to the device (512) for rendering but also can be a sample stream that can be stored in the reference picture memory (557) for use in subsequent inter-picture prediction.
[0049] Once some coded pictures are completely reconstructed, they can be used as reference pictures for subsequent inter-picture prediction. For example, when the coded picture corresponding to the current picture is completely reconstructed and the coded picture is specified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and before starting the reconstruction of the next coded picture, a new current picture buffer can be reallocated.
[0050] The video decoder (510) may perform a decoding operation according to a predetermined video compression technique adopted in a predetermined standard (e.g., ITU-T Recommendation H.265). In the sense that the coded video sequence complies with both the video compression technique and standard syntax and the profile described in the video compression technique and standard document, it can be said that the coded video sequence conforms to the syntax defined by the used video compression technique and standard. In particular, for the profile, a predetermined tool can be selected from all the tools available in the video compression technique and standard as the only tool that can be used in accordance with the profile. To comply with the standard, the complexity of the coded video sequence may be within the limits defined by the level of the video compression technique and standard. In some examples, depending on the level, the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured at 1 million samples per second), maximum reference picture size, etc. are limited. The limits provided by the level may be further restricted in some examples by the specifications of the Hypothetical Reference Decoder (HRD) and the metadata for HRD buffer management signaled in the coded video sequence.
[0051] In some embodiments, the receiver (531) may receive the coded video along with the accompanying (redundant) data. The accompanying data may be included as part of the coded video sequence. The accompanying data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. For example, the accompanying data can take forms such as a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slice, redundant picture, forward error correction code, etc.
[0052] FIG. 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (for example, a transmission circuit). The video encoder (603) can be used instead of the video encoder (403) in the example of FIG. 4.
[0053] The video encoder (603) may receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that can capture one or more video images to be coded by the video encoder (603). In another example, the video source (601) may be implemented as part of the electronic device (620).
[0054] The video source (601) may provide a source video sequence to be coded by the video encoder (603) as a digital video sample stream capable of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit,...), any color space (e.g., BT.601 YCrCb, RGB, XYZ,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media providing system, the video source (601) may be a storage device capable of storing pre-prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures or images that realize movement when viewed in order. The pictures may be arranged as a spatial array of pixels as they are, and each pixel may include one or more samples according to the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. In the following description, samples will be focused on.
[0055] According to some embodiments, the video encoder (603) may code and compress pictures of the source video sequence into the coded video sequence (643) in real time, or may code and compress while subject to any other time constraints required by the application. Enforcing an appropriate coding speed constitutes one of the functions of the controller (650). In some embodiments, the controller (650) may be functionally connected to other functional parts described below, and the other functional parts may be controlled by the controller (650). For simplicity, the connections are not shown. The parameters set by the controller (650) may include rate control related parameters (picture skip, quantizer, value of λ of rate-distortion optimization techniques, …), picture size, layout of group of picture (GOP), search range of the longest motion vector, etc. The controller (650) may be configured to have other appropriate functions related to the video encoder (603) optimized for a specific system design.
[0056] In some embodiments, the video encoder (603) may be configured to operate in a coding loop. As a maximally simplified explanation, in one example, the coding loop may include a source coder (630) (e.g., responsible for generating symbols such as a symbol stream based on an input picture to be coded and one or more reference pictures) and a decoder (633) incorporated in (and located at) the video encoder (603). The decoder (633) reconstructs the symbols and generates sample data as if it were generated by a decoder (in another location), even if the built-in decoder 633 processes a video stream coded by the source coder 630 without using entropy coding (any compression between the symbols and the coded video bitstream in entropy coding may be reversible compression included in the video compression technology under consideration in the disclosed subject matter). The reconstructed sample stream (sample data) is input to the reference picture memory (634). Since decoding the symbol stream results in an accurate bit result regardless of the location of the decoder (whether it is located here or elsewhere), the content of the reference picture memory (634) also has accurate bits in a comparison between the encoder located here and an encoder located elsewhere. In other words, the prediction part of the encoder "sees" exactly the same sample values as if the decoder "sees" when using prediction during decoding as reference picture samples. The principle of such reference picture simultaneity (and further, for example, the drift resulting when simultaneity cannot be maintained due to channel errors) is used to improve coding quality.
[0057] The operation of the "here" decoder (633) can be the same as that of a decoder such as the "elsewhere" video decoder (510) already described in detail above with reference to FIG. 5. However, referring somewhat to FIG. 5 as well, since the symbol is valid and the decoding of the symbol by the entropy encoder (645) and the parser (520) into the coded video sequence is reversible, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the here decoder (633) of the encoder.
[0058] What can be said at this point is that any decoder technology other than parse / entropy decoding that can only exist in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, for the disclosed subject matter, attention may be paid to the operation of the decoder related to the decoding part of the encoder as appropriate. Therefore, since the description of the encoder technology is the reverse of the decoder technology described in its entirety, it can be omitted. A more detailed description of the encoder will be given below only for specific parts or aspects.
[0059] During the operation of some embodiments, the source coder (630) may perform motion compensation predictive coding by predicting and coding an input picture with reference to one or more previously coded pictures obtained from the video sequence designated as "reference picture". In this way, the coding engine (632) codes the color channel difference (i.e., the residual value) between the pixel block of the input picture and the pixel blocks of one or more reference pictures that can be selected as one or more prediction references for the input picture. The terms "residue" and its adjective form "residual" may be used interchangeably.
[0060] The video decoder (633) located here may decode the coded video data of the picture that can be specified as the reference picture based on the symbol generated by the source coder (630). It may be effective that the operation of the coding engine (632) is an irreversible process. When the coded video data is decoded by some video decoder (not shown in FIG. 6), the reconstructed video sequence may generally be a copy of the source video sequence with some errors. The video decoder (633) located here may repeat the decoding process that may be performed by the video decoder for the reference picture so that the reconstructed reference picture is stored in the reference picture cache (634). In this way, the video encoder (603) can store a copy of the reconstructed reference picture having the same content as the reconstructed reference picture that will be obtained by a remote (located elsewhere) video decoder at this location (without transmission errors).
[0061] The predictor (635) may perform a prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may obtain sample data (sample data as a candidate reference pixel block), or search the reference picture memory (634) to obtain specific metadata such as a reference picture motion vector and a block shape that can be responsible for an appropriate prediction reference for the new picture. The predictor (635) may operate to use one sample block for each pixel block to detect an appropriate prediction reference. In some examples, when judging according to the search result obtained by the predictor (635), the input picture may include a prediction reference obtained from a plurality of reference pictures stored in the reference picture memory (634).
[0062] For example, the controller (650) may manage the coding operation of the source coder (630), including setting parameters and subgroup parameters used to encode video data.
[0063] The outputs of all the above functional parts may be subjected to entropy coding by an entropy encoder (645). The entropy encoder (645) converts symbols generated by various functional parts into a coded video sequence by reversibly compressing the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.
[0064] One or more coded video sequences generated by the entropy encoder (645) may be buffered for transmission by the transmitter (640) through the communication channel (660). The transmitter (640) may be hardware / software connected to a storage device that stores the encoded video data. The transmitter (640) may integrate other data to be transmitted, such as coded audio data and / or an attached data stream (the source is not shown), into the coded video data obtained from the video coder (603).
[0065] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a type of a specific coded picture that can affect the coding method applicable to each picture for each coded picture. For example, in many cases, a picture can be assigned as one of the following picture types.
[0066] An Intra Picture (I picture) can be said to be a picture that can be coded and decoded without using any other picture in the sequence as a prediction source. In some video codecs, various types of intra pictures, such as Independent Decoder Refresh (IDR) pictures, are possible. Those skilled in the art will conceive of such variations of I pictures and their corresponding application examples and features.
[0067] A predictive picture (P picture) is a picture that can be coded and decoded using intra prediction or inter prediction, with at most one motion vector and a reference index to predict the sample values of each block.
[0068] A bi-directionally predictive picture (B picture) is a picture that can be coded and decoded using intra prediction or inter prediction, with at most two motion vectors and a reference index to predict the sample values of each block. Similarly, in a multiple-predictive picture, three or more reference pictures and associated metadata can be used for the reconstruction of one block.
[0069] Generally, a source picture can be spatially subdivided into a plurality of sample coding blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each), and the blocks can be coded one by one. When determined by the coding assignment applied to the picture corresponding to each block, a block may be coded by referring to other (already coded) blocks for prediction. For example, a block of an I picture may be coded without using prediction, or may be coded using prediction by referring to an already coded block of the same picture (spatial prediction or intra prediction). A pixel block of a P picture may be coded using prediction by using spatial prediction or by referring to one previously coded reference picture for temporal prediction. A block of a B picture may be coded using prediction by using spatial prediction or by referring to one or two previously coded reference pictures for temporal prediction. The source picture or an intermediate processed picture may be subdivided into other types of blocks for other purposes. The division of coding blocks and other types of blocks may or may not follow the same way, which will be described in more detail later.
[0070] The video encoder (603) may perform a coding operation according to a predetermined video coding technique or standard (for example, ITU-T Recommendation H.265). In this operation, the video encoder (603) may perform various compression operations including a predictive coding operation that utilizes the temporal redundancy and spatial redundancy of the input video sequence. Therefore, the coded video data may conform to the syntax defined by the video coding technique or standard being used.
[0071] In some embodiments, the transmitter (640) may transmit the associated data together with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The associated data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures or redundant slices, SEI messages, VUI parameter set fragments, and the like.
[0072] The video may be captured in time order as a plurality of source pictures (video pictures). Intra prediction (often abbreviated as intra prediction) utilizes the spatial correlation within a given picture, and inter prediction utilizes the correlation between pictures (temporal correlation or other correlations). For example, a specific picture being encoded / decoded, referred to as the current picture, may be divided into blocks. If a block in the current picture is similar to a reference block in a reference picture that has been previously coded and is still buffered in the video, the block in the current picture may be coded by a vector referred to as a motion vector. The motion vector points towards the reference block in the reference picture, and when multiple reference pictures are used, the reference picture can be specified with a third dimension.
[0073] In some embodiments, a bi-prediction technique can be used for inter-picture prediction. According to this bi-prediction technique, two reference pictures such as a first reference picture and a second reference picture that both precede a current picture in the video in decoding order (although they may be in the past or future respectively in display order) are used. A block in the current picture can be coded by a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. These can be used in combination with each other depending on the combination of the first reference block and the second reference block to predict the block.
[0074] Furthermore, the coding efficiency can be improved by using a merge mode method for inter-picture prediction.
[0075] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, for compression, pictures in a series of video pictures are divided into coding tree units (CTUs), and the CTUs in a picture may have the same size (e.g., 64×64 pixels, 32×32 pixels, or 16×16 pixels). Generally, a CTU may include three parallel coding tree blocks (CTBs), i.e., one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a CTU of 64×64 pixels can be divided into one CU of 64×64 pixels or four CUs of 32×32 pixels. Each of one or more of the 32×32 blocks may be further divided into four CUs of 16×16 pixels. In some embodiments, each CU may be analyzed during encoding to determine the prediction type of the CU from various prediction types such as an inter prediction type or an intra prediction type. The CU may be divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes one luma prediction block (prediction block: PB) and two chroma PBs. In one embodiment, the prediction operation (encoding / decoding) of coding is performed in units of prediction blocks. The division of the CU into PUs (or PBs of different color channels) may be performed in various spatial patterns. The luma PB or chroma PB may include a matrix of sample values (e.g., luma values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 samples, etc.
[0076] FIG. 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) receives a processing block (e.g., a prediction block) of sample values in a current video picture in a series of video pictures, and is configured to encode the processing block into an encoded picture that is part of an encoded video sequence. An example of the video encoder (703) may be used instead of the video encoder (403) in the example of FIG. 4.
[0077] For example, a video encoder (703) receives a matrix of sample values of a processing block, such as a prediction block of 8×8 samples. Then, the video encoder (703) determines, for example, using rate-distortion optimization (RDO), whether the processing block is best coded using the intra mode, the inter mode, or the bi-prediction mode. If the processing block is determined to be coded in the intra mode, the video encoder (703) may use the intra prediction method to encode the processing block into a coded picture. If the processing block is determined to be coded in the inter mode or the bi-prediction mode, the video encoder (703) may use the inter prediction method or the bi-prediction method, respectively, to encode the processing block into a coded picture. In some embodiments, the merge mode may be used as a sub-mode of inter-picture prediction to obtain motion vectors from one or more motion vector predictors without assistance of coding motion vector components outside the predictor. In some other embodiments, there may be motion vector components applicable to the target block. Therefore, the video encoder (703) may include components not explicitly shown in FIG. 7, such as a mode determination module that determines the prediction mode of the processing block.
[0078] In the example of FIG. 7, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual value calculator (723), a switch (726), a residual value encoder (724), a general controller (721), and an entropy encoder (725), which are connected to each other as shown in the arrangement example of FIG. 7.
[0079] The inter encoder (730) receives samples of a current block (e.g., a processing block), compares the block with one or more reference blocks (e.g., blocks in the previous picture and the subsequent picture in display order) in a reference picture, generates inter prediction information (e.g., a description of redundant information according to an inter coding method, a motion vector, merge mode information), and is configured to calculate an inter prediction result (e.g., a predicted block) based on the inter prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information using a decoding unit 633 (shown as a residual decoder 728 in FIG. 7, which will be described in more detail later) incorporated in the example of the encoder 620 in FIG. 6.
[0080] The intra encoder (722) receives samples of a current block (e.g., a processing block), compares the block with already coded blocks in the same picture, generates quantized coefficients after transformation, and is optionally configured to also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). The intra encoder (722) may calculate an intra prediction result (e.g., a predicted block) based on the intra prediction information and reference blocks in the same picture.
[0081] The overall controller (721) may be configured to determine overall control data and control other components of the video encoder (703) based on the overall control data. In one example, the overall controller (721) determines the prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is the intra mode, the overall controller (721) controls the switch (726) to select the intra mode result for use by the residual value calculator (723), controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream, and when the predication mode of the block is the inter mode, the overall controller (721) controls the switch (726) to select the inter prediction result for use by the residual value calculator (723), and controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.
[0082] The residual value calculator (723) may be configured to calculate the difference (residual value data) between the received block and the prediction result for the block selected from the intra encoder (722) or the inter encoder (730). The residual value encoder (724) may be configured to encode the residual value data to generate transform coefficients. For example, the residual value encoder (724) may be configured to convert the residual value data from the spatial domain to the frequency domain to generate transform coefficients. Thereafter, the transform coefficients are subjected to quantization processing, and quantized transform coefficients are obtained. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation to generate decoded residual value data. The decoded residual value data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual value data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual value data and the intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.
[0083] The entropy encoder (725) may be configured to format the bitstream to include the encoded block and perform entropy coding. The entropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual value information, and other appropriate information in the bitstream. When coding a block in either the inter mode or the merge sub-mode of the bi-prediction mode, the residual value information may not be present.
[0084] FIG. 8 shows an example of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and decode the coded picture to generate a reconstructed picture. In one example, the video decoder (810) may be used instead of the video decoder (410) of the example of FIG. 4.
[0085] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) that are connected together as shown in the arrangement example of FIG. 8.
[0086] The entropy decoder (871) can be configured to reconstruct from the coded picture specific symbols representing the syntax elements that make up the coded picture. Such symbols can include, for example, the mode used for block coding (e.g., intra mode, inter mode, bi-predicted mode, merge sub-mode, or another sub-mode), prediction information (e.g., intra prediction information or inter prediction information) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880), and residual information, for example, in the form of quantized transform coefficients. In one example, when the prediction mode is an inter mode or a bi-predicted mode, the inter prediction information is provided to the inter decoder (880), and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). Inverse quantization can be performed on the residual information, which is provided to the residual decoder (873).
[0087] The inter decoder (880) may be configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.
[0088] The intra decoder (872) may be configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0089] The residual decoder (873) may be configured to perform inverse quantization to extract the inverse quantized transform coefficients, and process the inverse quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize specific control information (such as including the quantizer parameter (QP)), and this information may be provided by the entropy decoder (871) (since this is control information with only a small amount of data volume, the data path is not shown).
[0090] The reconstruction module (874) may be configured to combine, in the spatial domain, the residual output by the residual decoder (873) and the prediction result (output by the inter prediction module or the intra prediction module as appropriate) to form a reconstructed block that forms a part of the reconstructed picture as a part of the reconstructed video. Note that other appropriate operations such as a deblocking operation can be performed to improve the video quality.
[0091] Note that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using any appropriate method. In some embodiments, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) can be implemented using one or more processors that execute software instructions.
[0092] Focusing on the block partitioning used for coding and decoding, general partitioning can start from a base block and follow a set of predefined rules, a specific pattern, a partitioning tree, or some partitioning structure or method. The partitioning can be hierarchical or recursive. The base block can be divided, i.e., partitioned, according to an example of the partitioning procedure described later, another procedure, or a combination of these, and finally, the final set of partitions, i.e., coding blocks, can be obtained. Each of such partitions may be a partition at one of various partitioning levels of the partitioning hierarchy and may have various shapes. Each of the partitions may be referred to as a coding block (CB). In various partitioning implementation examples further described below, each obtained CB may be a CB with an allowed size and an allowed partitioning level. Such a partition can form a unit, and for the unit, some basic coding / decoding decisions can be made, and optimization, determination, and signaling of coding / decoding parameters can be performed in the encoded video bitstream. Therefore, such a partition is referred to as a coding block. The highest level or the deepest level of the final partition represents the depth of the coding block partitioning structure of the tree. The coding block may be a luma coding block or a chroma coding block. The CB tree structure of each color may be referred to as a coding block tree (CBT).
[0093] All coding blocks of all color channels may be collectively referred to as a coding unit (CU). The hierarchical structure for what is used for all color channels may be collectively referred to as a coding tree unit (CTU). The partitioning patterns or structures of various color channels in the CTU may be the same or different.
[0094] In some implementation examples, the splitting tree method or structure used for the luma channel and the chroma channel may not need to be the same. In other words, the luma channel and the chroma channel may have separate coding tree structures or patterns. Further, whether the same coding split tree structure or different coding split tree structures are used for the luma and chroma channels may depend on whether the slice during coding is a P slice, a B slice, or an I slice. For example, in an I slice, the chroma channel and the luma channel may have separate coding split tree structures or coding split tree structure modes, while in a P slice or a B slice, the luma channel and the chroma channel may share the same coding split tree method. When separate coding split tree structures or modes are applied, the luma channel can be split into CBs by a predetermined coding split tree structure, and the chroma channel can be split into chroma CBs by another coding split tree structure.
[0095] In some implementation examples, a default splitting pattern can be applied to a base block. As shown in FIG. 9, in an example of four splitting trees, the first default level (for example, a 64×64 block level or other size as the base block size) can be used as a starting point, and the base block can be split to lower the hierarchy to a default minimum level (for example, a 4×4 level). For example, four default splitting options, that is, splitting patterns, indicated by 902, 904, 906, and 908 may be applied to the base block. The splitting part indicated by R can be used for recursive splitting in that the same splitting options shown in FIG. 9 can be repeated at a smaller scale down to the minimum level (for example, a 4×4 level). In some implementation examples, another limitation may be applied to the splitting method of FIG. 9. In the implementation example of FIG. 9, a rectangular splitting part (for example, a 1:2 / 2:1 rectangular splitting) can be used, but the rectangular splitting part cannot be made recursive, while the square splitting part can be made recursive. By splitting according to FIG. 9 using recursion when necessary, a final set of coding blocks is generated. The coding tree depth may be further defined to indicate the splitting depth starting from the root node, that is, the root block. For example, the coding tree depth of the root node, that is, the root block (for example, a 64×64 block) can be set to 0, and further, after splitting the root block once according to FIG. 9, the coding tree depth increases by 1. The maximum level, that is, the deepest level, from a 64×64 base block to a 4×4 minimum splitting part is 4 (starting from level 0) in the above method. Such a splitting method can be applied to one or more of the color channels. Each color channel can be split individually according to the method of FIG. 9 (for example, the splitting pattern, that is, the option, may be determined individually from a default pattern for each color channel at each hierarchical level). Alternatively, two or more of the color channels may share the same hierarchical pattern tree of FIG. 9 (for example, the same splitting pattern, that is, the option, may be selected from a default pattern for two or more color channels at each hierarchical level).
[0096] FIG. 10 shows another example of a default partitioning pattern that enables the formation of a partitioning tree by recursive partitioning. As shown in FIG. 10, 10 different partitioning structures or pattern examples can be predefined. The root block can start from a default level (e.g., a base block at the 128×128 level or the 64×64 level). The example of the partitioning structure in FIG. 10 includes various 2:1 / 1:2 rectangular partitions and 4:1 / 1:4 rectangular partitions. The partitioning type with the three subdivisions 1002, 1004, 1006, and 1008 in the second row of FIG. 10 may be applied to a "T-type" partition. The "T-type" partitions 1002, 1004, 1006, and 1008 may be referred to as the left T-type, top T-type, right T-type, and bottom T-type, respectively. In some implementations, none of the rectangular partitions in FIG. 10 can be further subdivided. The coding tree depth may be further defined to indicate the partitioning depth starting from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., a 128×128 block) can be set to 0, and further, after the root block is partitioned once according to FIG. 10, the coding tree depth increases by only 1. In some implementations, for the recursive partitioning to the next level of the partitioning tree according to the pattern in FIG. 10, only the partitions that are all squares in 1010 can be used. In other words, the square partitions in the T-type patterns 1002, 1004, 1006, and 1008 cannot use recursive partitioning. By using the partitioning procedure according to FIG. 10 that uses recursion when necessary, the final set of coding blocks is generated. Such a method can be applied to one or more of the color channels. In some implementations, higher flexibility can be added for the use of partitions at the 8×8 level and below. For example, 2×2 chroma interpolation prediction can be used in some examples.
[0097] In other implementation examples of coding block splitting, a quadtree structure can be used to split a base block or an intermediate block into a quadtree splitting section. Such quadtree splitting can be applied hierarchically and recursively to any square-shaped splitting. According to various local features of the base block or the intermediate block / splitting section, it is possible to optimize whether to further perform quadtree splitting on the base block or the intermediate block, that is, the splitting section. Furthermore, quadtree splitting at the picture boundary can be optimized. For example, an implicit quadtree splitting may be performed at the picture boundary so that the block continues to be quadtree split until its size suits the picture boundary.
[0098] In other implementation examples, a hierarchical binary splitting starting from a base block can be used. In such a method, a base block or an intermediate level block can be split into two splitting sections. The binary splitting can be either horizontal or vertical. For example, a horizontal binary splitting can split a base block or an intermediate block into equal left and right splitting sections. Similarly, a vertical binary splitting can split a base block or an intermediate block into equal upper and lower splitting sections. Such binary splitting can also be hierarchical and recursive. For each of the base blocks or intermediate blocks, it is possible to determine whether to continue the binary splitting method. If this method is to be continued, it is possible to determine whether to use horizontal binary splitting or vertical binary splitting. In some implementation examples, further splitting can be stopped at a predetermined minimum splitting size (it can be stopped either at the minimum splitting size of one dimension or at the minimum splitting size of both dimensions). Instead of this, further splitting may be stopped when a predetermined splitting level, that is, depth, starting from the base block is reached. In some implementation examples, the aspect ratio of the splitting section can be restricted. For example, the aspect ratio of the splitting section cannot be made less than 1:4 (or more than 4:1). Therefore, for a vertically elongated splitting section with a vertical-to-horizontal aspect ratio of 4:1, it can only be further binary split vertically into upper and lower splitting sections, and the vertical-to-horizontal aspect ratio of each of the upper and lower splitting sections is 2:1.
[0099] In several other examples, a three-way splitting method as shown in FIG. 13 can be used to split the base block or any intermediate block. The triple pattern can be implemented vertically as shown at 1302 in FIG. 13 or horizontally as shown at 1304 in FIG. 13. The example of the splitting ratio in FIG. 13 is shown as 1:2:1 for both vertical and horizontal, but other ratios may be predefined. In some implementations, two or more different ratios may be predefined. Since the quadtree and the binary tree are always split along the block center to split an object into separate split parts, such a triple-tree split can capture an object at the block center of one continuous split part without a break, and thus such a three-way splitting method can be used to complement the quadtree or binary splitting structure in this regard. In some implementations, the width and height of the splitting parts of a typical triple tree are always a power of 2 to avoid further conversion.
[0100] The above-described splitting methods can be combined in various ways at various splitting levels. As an example, the above-mentioned quadtree splitting method and binary splitting method may be combined to split the base block into a quadtree-binary-tree (QTBT) structure. In such a method, the base block or intermediate block / splitting part may be either quadtree split or binary split according to a set of predefined conditions (if conditions are specified). A specific example is shown in FIG. 14. In the example of FIG. 14, first, the base block is quadtree split into four splitting parts shown as 1402, 1404, 1406, and 1408. Then, each of the obtained splitting parts is either quadtree split into four further splitting parts at the next level (for example, 1408), binary split into two further splitting parts (for example, split either horizontally or vertically as in 1402 or 1406, both of which are symmetric), or not split (for example, 1404). As shown in the overall typical splitting pattern of 1410 and the corresponding tree structure / representation of 1420, binary splitting or quadtree splitting can be recursively used for square-shaped splitting parts. In 1410 and 1420, solid lines represent quadtree splitting and dashed lines represent binary splitting. A flag can be used for each binary splitting node (non-leaf binary splitting) to indicate whether the binary splitting is horizontal or vertical. For example, as shown in 1420 corresponding to the splitting structure of 1410, it can be said that flag "0" represents horizontal binary splitting and flag "1" represents vertical binary splitting. For the splitting parts of quadtree splitting, in quadtree splitting, since the block or splitting part is always split both horizontally and vertically so that four sub-blocks / splitting parts of the same size are generated, there is no need to indicate the splitting type. In some embodiments, flag "1" may represent horizontal binary splitting and flag "0" may represent vertical binary splitting.
[0101] In some embodiments of QTBT, the set of rules for quadtree splitting and binary splitting can be represented by the following predefined parameters and the corresponding functions related thereto. CTU size: The size of the root node of the quadtree (the size of the base block) MinQTSize: Minimum allowable quadtree leaf node size MaxBTSize: Maximum allowable binary tree root node size MaxBTDepth: Maximum allowable binary tree depth MinBTSize: Minimum allowable binary tree leaf node size
[0102] In some implementation examples of the QTBT splitting structure, the CTU size can be set as a 128×128 luma sample having a corresponding 64×64 block of chroma samples (when assuming and using typical chroma subsampling), MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (MinBTSize for both width and height) can be set to 4×4, and MaxBTDepth can be set to 4. Quadtree splitting can be applied to the CTU so as to generate quadtree leaf nodes first. The quadtree leaf nodes can have a size ranging from the minimum size 16×16 (i.e., MinQTSize) that a quadtree leaf node can take to 128×128 (i.e., CTU size). When the node is 128×128, since its size exceeds MaxBTSize (i.e., 64×64), it is not split by the binary tree first. In other cases, nodes not exceeding MaxBTSize can be split by the binary tree. In the example of FIG. 14, the base block is 128×128. According to a given set of rules, only quadtree splitting can be performed on the base block. The splitting depth of the base block is zero. Each of the four obtained split parts is 64×64, does not exceed MaxBTSize, and each split part can be further split by quadtree or binary tree at level 1. The process proceeds. When the binary tree depth reaches MaxBTDepth (i.e., 4), it can be considered that no further splitting occurs. When the width of the binary tree node is equal to MinBTSize (i.e., 4), it can be considered that no further horizontal splitting occurs. Similarly, when the height of the binary tree node is equal to MinBTSize, it is considered that no further vertical splitting occurs.
[0103] In some implementation examples, the above QTBT method can be configured to support the flexibility of luma and chroma with the same QTBT structure, or can also be configured to support the flexibility of luma and chroma with separate QTBT structures. For example, in P slices and B slices, the luma CTB and chroma CTB in one CTU may share the same QTBT structure. On the other hand, in I slices, the luma CTB may be divided into CUs by the QTBT structure, and the chroma CTB may be divided into chroma CUs by another QTBT structure. This means that CUs can be used to apply to different color channels in the I slice. For example, the I slice may consist of a coding block of the luma component or coding blocks of two chroma components, and the CU of the P slice or B slice may consist of coding blocks of all three color components.
[0104] In other implementation examples, the QTBT method can be complemented using the above-mentioned ternary method. Such implementation examples may be referred to as a multi-type-tree (MTT) structure. For example, in addition to the binary division of nodes, one of the ternary division patterns in FIG. 13 can be selected. In some implementation examples, ternary division can be performed only on square nodes. Another flag can be used to indicate whether the ternary division is horizontal or vertical.
[0105] It can be said that the main reason for attempting to design two-level trees or multi-level trees such as QTBT implementation examples and QTBT implementation examples complemented by ternary division is to alleviate complexity. Theoretically, the complexity across the tree is T D . Here, T represents the number of division types, and D represents the depth of the tree. A trade-off can be achieved by using multiple types (T) while reducing the depth (D).
[0106] In some implementation examples, the CB can be further divided. For example, the CB can be further divided into a plurality of prediction blocks (PBs) for intra-frame prediction or inter-frame prediction during the coding process and the decoding process. In other words, the CB can be further divided into different subdivisions, and individual prediction judgments / configurations can be performed. In parallel with this, the CB can be further divided into a plurality of transform blocks (TBs) to determine the level at which the conversion or inverse conversion of video data is executed. The method of dividing the CB into PBs and TBs may or may not be the same. For example, each division method may be executed using a procedure unique to that division method based on various characteristics of the video data, for example. In some implementation examples, the PB division method and the TB division method may be independent of each other. In other implementation examples, the PB and TB division methods and the boundaries may be correlated. In one implementation example, for example, the TB may be divided after the PB division. In particular, after being determined according to the division of the coding block, each PB may be further divided into one or more TBs. For example, in some implementation examples, the PB may be divided into one TB, two TBs, four TBs, or other numbers of TBs.
[0107] In some implementation examples, when splitting a base block into coding blocks and further into prediction blocks and / or transform blocks, the luma channel and the chroma channel can be handled in different ways. For example, in some implementation examples, the splitting of a coding block into prediction blocks and / or transform blocks can be used for the luma channel, while the splitting of the coding block into prediction blocks and / or transform blocks cannot be used for one or more chroma channels. Therefore, in such implementation examples, only the conversion and / or prediction of luma blocks can be performed at the coding block level. In another example, the minimum transform block size of the luma channel and one or more chroma channels may be different. For example, it may be possible to split a coding block of the luma channel into smaller transform blocks and / or prediction blocks than those of the chroma channel. In yet another example, the maximum depth of splitting a coding block and / or a prediction block into transform blocks may be different between the luma channel and the chroma channel. For example, it may be possible to split a coding block of the luma channel into deeper transform blocks and / or prediction blocks than those of one or more chroma channels. In a specific example, a luma coding block may be split into transform blocks of multiple sizes, which can be expressed as recursive splitting down to a maximum of two levels, enabling transform block shapes such as square, 2:1 / 1:2, or 4:1 / 1:4, and transform block sizes from 4×4 to 64×64. In contrast, for chroma blocks, only the possible maximum transform blocks specified for luma blocks may be enabled.
[0108] In some implementation examples of splitting a coding block into PBs, the depth, shape, and / or other characteristics of the PB splitting may depend on whether the PB is intra-coded or inter-coded.
[0109] The splitting of a coding block (or prediction block) into transform blocks may be performed recursively or non-recursively in various typical ways including, but not limited to, quadtree splitting and a predefined pattern splitting, and may be further performed with further consideration of the transform blocks at the boundaries of the coding block or prediction block. Generally speaking, the obtained transform blocks may be transform blocks at different splitting levels, may not be transform blocks of the same size, and do not have to be square in shape (for example, it is possible to be a rectangle with several allowable sizes and aspect ratios). Further examples regarding FIGS. 15, 16, and 17 are described in more detail later.
[0110] For the above-mentioned manner, in other embodiments, the CB obtained through any of the above-mentioned splitting methods may be used as a basic coding block, i.e., a minimum coding block, for prediction and / or transformation. In other words, no further splitting is performed for the purpose of inter prediction / intra prediction and / or transformation. For example, the CB obtained from the above-mentioned QTBT method may be directly used as a unit for performing prediction. Specifically, in such a QTBT structure, the concept of a plurality of splitting types is excluded, that is, the separation of CU, PU, and TU is excluded, to correspond to the higher flexibility of the above-mentioned CU / CB splitting shape. In such a QTBT block structure, the CU / CB can have both a square shape and a rectangular shape. Such a leaf node of QTBT is used as a unit for prediction and transformation processing without any further splitting. This means that in such an example of the QTBT coding block structure, the block sizes of CU, PU, and TU are the same.
[0111] The above-mentioned various CB splitting methods may be combined with further splitting the CB into PB and / or TB (excluding PB / TB splitting) in any way. The following specific embodiments are provided as non-limiting examples.
[0112] The following describes specific implementation examples of coding block splitting and transform block splitting. In such implementation examples, a base block can be split into coding blocks using recursive quadtree splitting or the aforementioned default splitting patterns (such as the splitting patterns in FIGS. 9 and 10). For each level, it can be determined based on local video data characteristics whether a specific split part should be further split into a quadtree. The obtained CBs may be CBs at various quadtree splitting levels or CBs of various sizes. A determination can be made at the CB level as to whether a picture area should be coded using inter-picture (temporal) prediction or intra-picture (spatial) prediction (it can also be done at the CU level for all three color channels). Each CB can be further split into one PB, two PBs, four PBs, or other numbers of PBs according to a default PB splitting type. In one PB, the same prediction process can be applied, and relevant information can be sent to the decoder in PB units. After obtaining a residual block by applying a prediction process based on the PB splitting type, the CB can be split into TBs according to another quadtree structure similar to the coding tree of the CB. In this specific implementation example, the CB or TB may be limited to a square shape, but it is not necessarily so. Furthermore, in this specific example, in the case of inter prediction, the PB may be square or rectangular, and in the case of intra prediction, the PB may be limited to square. The coding block may be split into, for example, four square-shaped TBs. Furthermore, each TB may be recursively split into smaller TBs (using quadtree splitting), which is referred to as a Residual Quadtree (RQT).
[0113] Next, another implementation example of dividing the base block into CB, PB, and / or TB will be further described. For example, instead of using a plurality of division unit types such as those shown in FIG. 9 or FIG. 10, a quadtree (e.g., QTBT with the above-mentioned QTBT or QTBT with a ternary division) with a nested multi-type tree using a binary division structure and a ternary division structure can be used. The separation of CB, PB, and TB (i.e., the division of CB into PB and / or TB and the division of PB into TB) can be terminated except when necessary for a CB that is too large in size relative to the maximum conversion length (such a CB may need to be further divided). This typical division method can be designed to accommodate higher flexibility in the CB division shape so that both prediction and conversion can be performed at the CB level without further division. In such a coding tree structure, the CB can have both a square shape and a rectangular shape. In particular, the coding tree block (CTB) can be first divided by a quadtree structure. Thereafter, the quadtree leaf nodes can be further divided by a nested multi-type tree structure. An example of a nested multi-type tree structure using binary or ternary division is shown in FIG. 11. In particular, the typical multi-type tree structure of FIG. 11 includes four division types called vertical binary division (SPLIT_BT_VER) (1102), horizontal binary division (SPLIT_BT_HOR) (1104), vertical ternary division (SPLIT_TT_VER) (1106), and horizontal ternary division (SPLIT_TT_HOR) (1108). Thereafter, the CB corresponds to the leaf of the multi-type tree. In this implementation example, except when the CB is too large relative to the maximum conversion length, this segmentation is used for both the prediction process and the conversion process without any further division. This means that in most cases, the block sizes of CB, PB, and TB are the same in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum convertible length that can be accommodated is less than the width or height of the color component of the CB. In some implementation examples, in addition to binary or ternary division, the nested pattern of FIG. 11 can further include quadtree division.
[0114] One example of a quadtree (including a quadtree option, a binary split option, and a ternary split option) with a nested multi-type tree coding block structure for a single base block is shown in FIG. 12. More specifically, FIG. 12 shows how the base block 1200 is quad-tree divided into four square partitions 1202, 1204, 1206, and 1208. For each partition of the quadtree division, it is determined whether to further use the multi-type tree structure and the quadtree of FIG. 11 for further division. In the example of FIG. 12, the partition 1204 is not further divided. Another quadtree division is used for each of the partitions 1202 and 1208. In the partition 1202, third-level divisions of a quadtree, the horizontal binary split 1104 of FIG. 11, no split, and the horizontal ternary split 1108 of FIG. 11 are used for the upper-left, upper-right, lower-left, and lower-right partitions resulting from the second-level quadtree division, respectively. Another quadtree division is used for the partition 1208, and third-level divisions of the vertical ternary split 1106 of FIG. 11, no split, no split, and the horizontal binary split 1104 of FIG. 11 are used for the upper-left, upper-right, lower-left, and lower-right partitions resulting from the second-level quadtree division, respectively. Two subdivisions of the third-level upper-left partition of 1208 are further divided according to the horizontal binary split 1104 and the horizontal ternary split 1108 of FIG. 11, respectively. A second-level split pattern according to the vertical binary split 1102 of FIG. 11 is used for the partition 1206 to divide it into two partitions, and the two partitions are further divided at the third level according to the horizontal ternary split 1108 and the vertical binary split 1102 of FIG. 11. A fourth-level split is further applied to one of these according to the horizontal binary split 1104 of FIG. 11.
[0115] In the above specific example, the maximum luma transform size may be 64×64, and the maximum chroma transform size that can be supported may be different from that of luma (for example, 32×32). In the above typical CB in FIG. 12, it is usually not further divided into smaller PBs and / or TBs. However, if the width or height of the luma coding block or chroma coding block exceeds the maximum transform width or height, the luma coding block or chroma coding block can be automatically divided in the horizontal and / or vertical directions to meet the transform size limit in that direction.
[0116] In a specific example of dividing the above-described base block into CBs, as described above, the coding tree method can correspond to the fact that luma and chroma can have separate block tree structures. For example, in P slices and B slices, the luma CTB and chroma CTB in one CTU may share the same coding tree structure. In I slices, for example, luma and chroma may have separate coding block tree structures. When separate block tree structures are applied, the luma CTB can be divided into luma CBs by a predetermined coding tree structure, and the chroma CTB can be divided into chroma CBs by another coding tree structure. This means that the CU in the I slice can consist of a coding block of the luma component or coding blocks of two chroma components, and the CU in the P slice or B slice always consists of coding blocks of all three color components, except when the video is a monochrome video.
[0117] When a coding block is further divided into a plurality of transform blocks, the transform blocks in the coding block may be in the order in the bitstream following various orders, or may be in the order in the bitstream following the scanning method. Implementation examples of dividing a coding block or a prediction block into transform blocks and the coding order of the transform blocks are described in more detail later. In some implementation examples, as described above, the transform division can correspond to transform blocks of a plurality of shapes (for example, when the range of the transform block size is from 4×4 to 64×64, for example, 1:1 (square), 1:2 / 2:1, or 1:4 / 4:1). In some implementation examples, when the coding block is 64×64 or less, the transform block division can be applied only to the luma component so that the transform block size is the same as the coding block size in the chroma block. In other cases, when the width or height of the coding block exceeds 64, both the luma coding block and the chroma coding block can be implicitly divided into transform blocks that are multiples of min(W,64)×min(H,64) and transform blocks that are multiples of min(W,32)×min(H,32), respectively.
[0118] In some implementation examples of transform block division, in the intra-coding block and the inter-coding block, the coding block can be further divided into a plurality of transform blocks using a division depth up to a predetermined number of levels (for example, 2 levels). The division depth and the division size of the transform block can be associated. For some implementation examples, the correspondence from the transform size of the current depth to the transform size of the next depth is shown in Table 1 as follows.
Table 1
[0119] Based on the example of the correspondence in Table 1, for 1:1 square blocks, four 1:1 square sub-conversion blocks can be generated in the next-level conversion division part. The conversion division can be terminated, for example, at 4×4. Therefore, the conversion size at the current depth of 4×4 corresponds to the same 4×4 size at the next depth. In the example of Table 1, for non-square 1:2 / 2:1 blocks, two 1:1 square sub-conversion blocks can be generated in the next-level conversion division part, while for non-square 1:4 / 4:1 blocks, two 1:2 / 2:1 sub-conversion blocks can be generated in the next-level conversion division part.
[0120] In some embodiments, for the luma component of an intra-coding block, further different restrictions can be applied regarding the transform block splitting. For example, for each level of the transform splitting, all sub-transform blocks can be restricted to be of equal size. For example, for a 32×16 coding block, two 16×16+ sub-transform blocks are generated in the level 1 transform division part, and eight 8×8 sub-transform blocks are generated in the level 2 transform division part. In other words, the second-level splitting is always applied to all the first-level sub-blocks so as to keep the transform unit of the same size. An example of the transform block splitting for an intra-coding square block according to Table 1 is shown in FIG. 15 together with the coding order indicated by the arrows. Specifically, 1502 shows a square coding block. 1504 shows the result of performing the first-level splitting according to Table 1 to obtain four transform blocks of the same size, and the coding order is indicated by the arrows. 1506 shows the result of performing the second-level splitting on all the first-level blocks of the same size according to Table 1 to obtain 16 transform blocks of the same size, and the coding order is indicated by the arrows.
[0121] In some implementation examples, the luma component of the inter-coding block may not be subject to the above restrictions for intra-coding. For example, after the first-level transform split, any one of the sub-transform blocks may be further split by adding one more level independently of the others. Therefore, the resulting transform blocks may be transform blocks of the same size or may not be transform blocks of the same size. An example of splitting an inter-coding block into transform blocks is shown in FIG. 16 together with its coding order. In the example of FIG. 16, the inter-coding block 1602 is split into two-level transform blocks according to Table 1. At the first level, the inter-coding block is split into four transform blocks of the same size. Thereafter, only one of the four transform blocks (not all four) is further split into four sub-transform blocks, resulting in a total of seven transform blocks of two different sizes as shown in 1604. An example of the coding order of these seven transform blocks is indicated by arrows in 1604 of FIG. 16.
[0122] In various embodiments, during the video coding process, data about operations may be sent to the input section of an entropy encoder, such as the entropy encoder 645 shown in FIG. 6 or the entropy encoder 725 shown in FIG. 7. The entropy encoder may output a bitstream (i.e., the coded video sequence / bitstream), which may be transmitted via a transmission channel to another device. The transmission channel may include a wired connection, a wireless channel, or a combination of both, but is not limited thereto. In various embodiments, during the video decoding process, the coded video sequence / bitstream may be sent to the input section of an entropy decoder, such as the entropy decoder 871 of FIG. 8. The entropy decoder may output data about operations based on the coded video sequence / bitstream, which may include intra-prediction information, residual value information, and the like.
[0123] In various embodiments, the entropy encoder / decoder may encode / decrypt according to an arithmetic coding algorithm based on the occurrence probability of symbols (or notations) based on arithmetic coding. In some implementations, the occurrence probability of symbols (or notations) may be dynamically updated during the coding / decryption process. For example, if there are only two available symbols ("a" and "b"), and the occurrence probability of "a" is denoted as p_a and the occurrence probability of "b" is denoted as p_b, then p_a + p_b = 1 (or any other constant value). Therefore, when "a" appears in the coding / decryption process, p_a may be updated to a larger value, and p_b may be updated to a smaller value. This is because the sum of these values can be kept constant. This probability update process may be called a probability transition process in some implementations, or a probability state index update process in other implementations.
[0124] In some implementations, for example, in HEVC, the context adaptive binary arithmetic coding (CABAC) algorithm may be used in the entropy encoder / decoder. The CABAC algorithm may be a reversible compression algorithm, and in the CABAC algorithm, a probability transition process using a lookup table may be used among 64 different representative probability states.
[0125] In some implementations, binary arithmetic coding is applied to the syntax elements describing the video frame content to obtain a stream encoded as a binary bin stream. During CABAC, the initial interval [0, 1) may be stretched by an integer multiplier (for example, 512), and the probability of the disadvantaged symbol (PLPS) may be represented as the divisor of the integer obtained by rounding the quotient. Then, the interval division operation using typical arithmetic coding may be performed as an approximate calculation using integer operations with the specified resolution.
[0126] In some implementation examples, the updated section length corresponding to LPS (RLPS) may be calculated as RLPS = R * PLPS. Here, R is the value of the current section length. To shorten the time and improve the efficiency, the multiplication operation that uses a large amount of computing power as described above may be replaced with a look-up table into which the pre-calculated multiplication result is input. As a result, the updated section length (ivLpsRange) corresponding to LPS can be obtained by two indexes pStateIdx and qRangeIdx (that is, ivlLpsRange = rangeTabLps[pStateIdx][qRangeIdx]).
[0127] During encoding / decoding, every time a new value (binVal) of the bin to be encoded / decoded is obtained, the probability value PLPS may be recursively updated. For example, at the k-th step (that is, during the encoding or decoding of the k-th bin), when binVal is the value of LPS, a new value of PLPS may be calculated to be a larger value, or when binVal is the value of the dominant symbol (MPS), a new value of PLPS may be calculated to be a smaller value.
[0128] In some implementation examples, PLPS may be one of 64 possible values indexed by a 6-bit pStateIdx variable. The update of the probability value may be realized by updating the index pStateIdx, and this can be executed by searching for a value from a pre-calculated table to suppress the computing power and / or improve the efficiency.
[0129] In some embodiments, prior to calculating a new interval range, the range ivlCurrRange representing the state of the coding engine may be quantized into a set of four values. A transition state may be implemented using a table containing all 64×4 pre-calculated 8-bit values to approximate the value of ivlCurrRange*pLPS(pStateIdx). Also, a decoding decision may be implemented using one pre-calculated look-up table (LUT). After obtaining the initial ivlLpsRange using the LUT, the ivlCurrRange may be updated based on one LUT using the ivlLpsRange to calculate the output binVal.
[0130] In some embodiments, for example, with VCC, the probability may be linearly represented by the probability index pStateIdx. Thus, all calculations can be performed in an expression that does not use LUT operations. To improve the accuracy of probability estimation, a multiple hypothesis probability update model as shown in FIG. 17 may be used. The pStateIdx used for interval subdivision in the binary arithmetic coder is a combination of two probabilities pStateIdx0 and pStateIdx1. The two probabilities are associated with each context model and are updated individually at different adaptation rates. The adaptation rates of pStateIdx0 and pStateIdx1 for each context model are pre-trained based on the statistics of the associated bins. The probability estimate value pStateIdx is the average of the estimates obtained from two hypotheses.
[0131] FIG. 17 shows a flowchart used to decode one decision value (DecodeDecision), which includes a normalization process in an arithmetic decoding engine (RenomD, 1740). In some embodiments, the input to DecodeDecision may be a context table (ctxTable) and a context index (ctxIdx). The value of the variable ivlLpsRange is derived as in step 1710. Given the current value of ivlCurrRange, the variable qRangeIdx is derived as qRangeIdx = ivlCurrRange >> 5. Given qRangeIdx, pStateIdx0, and pStateIdx1 related to ctxTable and ctxIdx, valMps and ivlLpsRange are derived as pState = pStateIdx1 + 16 * pStateIdx0, valMps = pState >> 14, and ivlLpsRange = (qRangeIdx * ((valMps? 32767 - pState : pState) >> 9) >> 1) + 4. The variable ivlCurrRange is set to ivlCurrRange - ivlLpsRange. If ivlOffset is greater than or equal to ivlCurrRange, the variable binVal is set to 1 - valMps, ivlOffset is decreased by ivlCurrRange, and ivlCurrRange is set to ivlLpsRange. Otherwise, the variable binVal is set to valMps.
[0132] Regarding the probability update, during the state transition process, the inputs to this process are the current pStateIdx0 and pStateIdx1 and the decoded value binVal, and the outputs of this process are the updated pStateIdx0 and pStateIdx1 of the context variables related to ctxTable and ctxIdx. The variables shift0 and shift1 are derived from the shiftIdx value related to ctxTable and ctxIdx at 1730. That is, shift0 = (shiftIdx >> 2) + 2 and shift1 = (shiftIdx & 3) + 3 + shift0. Depending on the decoded value binVal, the updated values of the two variables pStateIdx0 and pStateIdx1 related to ctxTable and ctxIdx are obtained as pStateIdx0 = pStateIdx0 - (pStateIdx0 >> shift0) + (1023 * binVal >> shift0) and pStateIdx1 = pStateIdx1 - (pStateIdx1 >> shift1) + (16383 * binVal >> shift1).
[0133] In some implementation examples, for example, VVC CABAC may include an initialization process that depends on quantization parameters (QP) called at the beginning of each slice. When an initial value of the luma QP is given for a slice, the initial probability state of a context model denoted as preCtxState can be obtained by m = slopeIdx × 5 - 45, n = (offsetIdx << 3) + 7, and preCtxState = Clip3(1, 127, ((m × (QP - 32)) >> 4) + n). In some implementation examples, slopeIdx and offsetIdx are limited to 3 bits, and the entire initialization value is represented with 6-bit precision. The probability state preCtxState can directly represent the probability in the linear region. Therefore, preCtxState only requires an appropriate shift operation before being input to the arithmetic coding engine, and the mapping from the logarithmic region to the linear region and a 256-byte table can be stored / saved in memory by default. pStateIdx0 and pStateIdx1 can be obtained by pStateIdx0 = preCtxState << 3 and pStateIdx1 = preCtxState << 7.
[0134] In some implementation examples, a binary-based CABAC algorithm may be used, which includes two available symbols (for example, "0" and "1"). In an arithmetic coding algorithm using binary, the two available symbols may also be referred to as the less probable symbol (LPS) and the more probable symbol (MPS).
[0135] This disclosure describes various embodiments for performing multi-symbol arithmetic coding and / or decoding to solve at least one problem / issue related to a binary symbol arithmetic system, improve efficiency, and / or increase flexibility.
[0136] In some embodiments, an arithmetic algorithm using an M-ary base may be used in the entropy encoder or decoder, which includes M available symbols. M may be any integer value from 2 to 16. As a non-limiting example, when M = 5, the M-ary base includes five available symbols, which may be denoted as "0", "1", "2", "3", and "4".
[0137] An M-ary arithmetic coding engine is used to entropy code the syntax elements. Each syntax element is associated with an alphabet of M elements. As input to the encoder or decoder, the coding context may include a sequence of M-ary symbols along with a set of M probabilities. Each of the M probabilities may correspond to each of the M-ary symbols, and this may be represented by a cumulative distribution function (CDF). In some embodiments, the cumulative distribution function of the M-ary symbols is C = [c0, c1, …, c (M-2) , c (M-1) . In some embodiments, the cumulative distribution function of the M-ary symbols may be represented by an array of M 15-bit integers. Here, c (M-1) = 2 15 and c n / 32768 is the probability of symbols less than or equal to n, where n is an integer from 0 to M - 1.
[0138] In some embodiments, the M probabilities (i.e., the array of cumulative distribution functions) may be updated after coding / parsing each syntax element. In some embodiments, the M probabilities (i.e., the array of cumulative distribution functions) may be updated after coding / decoding each M-ary symbol. As a non-limiting example, since the array of cumulative distribution functions includes [c0, c1, …, c (M-2) , c (M-1) , when M = 4, the array of cumulative distribution functions includes [c0, c1, c2, c3].
[0139] In some embodiments, the update of the M probabilities (i.e., the array of cumulative distribution functions) may be performed using the following formula.
Number
[0140] Here, symbol is the M-ary symbol currently being coded / decoded, α is the probability update rate adapted based on the number of times the symbol has been coded or decoded (up to a maximum of 32 times), and m is the index of the element of the CDF. This adaptation of α enables rapid probability update at the start of the coding / parsing of the syntax element. As a non-limiting example, when M = 5 and the currently decoded M-ary symbol is "3", it can be said that m ∈ [0, symbol) becomes m ∈ [0, 3), where m includes any integer from 0 (including 0) to 3 (not including 3) (i.e., m includes 0, 1, and 2), and m ∈ [symbol, M - 1) becomes m ∈ [3, 4), where m includes any integer from 3 (including 3) to 4 (not including 4) (i.e., m includes only 3).
[0141] In some embodiments, the M-ary arithmetic coding process may follow the conventional arithmetic coding engine design. However, only the most significant 9 bits of the 15-bit probability value are input to the arithmetic encoder / decoder. The probability update rate α related to the symbol is calculated based on the number of occurrences of the related symbol when parsing the bit stream, and the value of α is reset at the start of the frame or tile using the following formula.
[0142] In some embodiments, as an example,[[]]
Number
[0143] In some embodiments, a single probability update model may be used for all contexts. In this case, the fact that the optimal design of the probability update model may vary for different syntaxes and / or the fact that a single probability update model may be sub - optimal may be overlooked. In the present disclosure, various embodiments applying syntax - dependent probability update models and multi - parameter probability update models are described, thereby enabling at least addressing the above - mentioned problems / issues and / or expecting an improvement in coding / decoding performance.
[0144] FIG. 18 shows a flowchart of a method example 1800 of multiple - symbol arithmetic coding during video decoding / encoding. The flow of the decoding method example starts from S1801 and ends at S1899. Method 1800 includes a step S1810 in which a device including a memory for storing instructions and a processor communicating with the memory receives a coded video bitstream, a step S1820 in which the device obtains an array of cumulative distribution functions and corresponding M symbols, where M is an integer greater than 1, step S1820, and / or a step S1830 in which the device performs arithmetic decoding on the coded video bitstream to extract at least one symbol based on the array of cumulative distribution functions, and for each extracted symbol, the device updates the array of cumulative distribution functions according to the extracted symbol based on at least one probability update rate, and the device continues to perform arithmetic decoding on the coded video bitstream to extract the next symbol based on the updated array of cumulative distribution functions. Some embodiments may include some or all of step S1830. In some embodiments, the array of cumulative distribution functions may include [c0, c1, …, c (M-2) , c (M-1) . By way of non - limiting example, when M = 4, the array of cumulative distribution functions may include [c0, c1, c2, c3], and each component of the array may be calculated as a 15 - bit integer.
[0145] In embodiments and implementations of the present disclosure, any number and order of any steps and / or operations may be combined or arranged as desired. Two or more of the steps and / or operations may be executed in parallel. The embodiments and implementations of the present disclosure may be used individually or combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. Embodiments of the present disclosure may be applied to luma blocks or chroma blocks. The term "block" may be interpreted as a prediction block, a coding block, or a coding unit (i.e., CU). The term "block" in the present disclosure may also be used to refer to a transform block. When explaining the block size for the following things, it may refer to the block width or height, or the maximum value of the width and height, or the minimum value of the width and height, or the area dimension (width × height), or the aspect ratio of the block (width:height or height:width).
[0146] In some implementations related to method 1800, at least one probability update rate has one probability update rate, and / or this probability update rate has a function f(N,M). Here, N is the number of occurrences of related symbols when parsing a coded video bitstream.
[0147] In some implementations related to method 1800, the function f(N,M) is
Equation
[0148] In some implementations related to method 1800, g(M,D) has min(log2(M),D).
[0149] In some implementations related to method 1800, the default function parameters are initialized to decode a tile or a frame.
[0150] In some implementations related to method 1800, the default function parameters are obtained from a high-level syntax comprising at least one of a video parameter set (VPS), a picture parameter set (PPS), a sequence parameter set (SPS), an adaptation parameter set (APS), a picture header, a frame header, a slice header, a tile header, or a coding tree unit (CTU) header.
[0151] In some implementations related to method 1800, the function f(N,M) is
Number
[0152] In some implementations related to method 1800, the function f(N,M) is
Number
[0153] In some implementations related to method 1800, at least one probability update rate comprises K probability update rates, where K is an integer greater than 1.
[0154] In some implementations related to method 1800, each probability update rate has a function f(N,M), where N is the number of occurrences of related symbols when parsing the coded video bitstream, and / or updating an array of cumulative distribution functions according to symbols extracted based on at least one probability update rate includes updating K arrays of cumulative distribution functions based on the corresponding K probability update rates, and / or updating the array of cumulative distribution functions as a weighted sum of the updated K arrays of cumulative distribution functions.
[0155] In some implementations related to method 1800, K is 2, and / or updating the array of cumulative distribution functions as a weighted sum of the updated K arrays of cumulative distribution functions includes updating the array of cumulative distribution functions as w1*p1 + w2*p2, where w1 is the first weight, p1 is the updated array of the first cumulative distribution function based on the first probability update rate, w2 is the second weight, and p2 is the updated array of the second cumulative distribution function based on the second probability update rate.
[0156] In some implementations related to method 1800, the first weight and the second weight are equal or different and are predefined.
[0157] In some implementations related to method 1800, one of the K probability update rates is
Number
[0158] In some implementations related to method 1800, one of the K probability update rates is
Number
Number
[0159] In some embodiments related to method 1800, the predefined function parameters are obtained from a high-level syntax comprising at least one of a video parameter set (VPS), a picture parameter set (PPS), a sequence parameter set (SPS), an adaptation parameter set (APS), a picture header, a frame header, a slice header, a tile header, or a coding tree unit (CTU) header.
[0160] As a non-limiting example, the probability update rate α may be a function f(count, M) of count and M, where count is the number of occurrences of the relevant symbol / context when parsing the bitstream during decoding, or the number of occurrences of the relevant symbol / context when coding the data corresponding to a series of syntax elements during encoding, and M represents the number of different symbol values of the relevant syntax / context. The function f(count, M) may be different for different syntax / contexts.
[0161] In some embodiments, the function [Number] where A, B, C, and D are default function parameters according to the selected syntax / context, and g(M, D) is a function of M with D as a function parameter. Examples of g(M, D) include, but are not limited to, min(log2(M), D). In some embodiments, A, B, C, and / or D are positive integers. In some embodiments, there may be limitations on the values of count, M, A, B, C, and D such that A + (count > B) + (count > C) + g(M, D) is less than and / or equal to 15. In some embodiments, the values of A, B, C, and D are initialized at the start of encoding of tiles or frames. In some embodiments, A, B, C, and D are positive integers.
[0162] In some embodiments,
Number
[0163] In some embodiments, the function
Number
[0164] In some implementation examples, the values of the function parameters defined as above may be specified in a high-level syntax / context including, but not limited to, VPS, PPS, SPS, APS, picture / frame header, slice header, tile header, or CTU header.
[0165] In some implementation examples, a probability may be updated as a weighted sum of a plurality of probabilities (e.g., N probabilities) derived using different probability update rates α i , i = 0, 1, …, N−1, while each probability update rate α i is a function f i (count, M) of count and M. Here, count is the number of occurrences of relevant symbols when parsing the bitstream, and M represents the number of different symbols in the relevant syntax / context.
[0166] In some implementation examples, the probability is updated as a weighted sum of N probabilities. As a non-limiting example, N may be equal to 2, and two probabilities (p0 and p1) may be derived using different probability update rates α0 and α1, respectively. In some implementation examples, the weights are equal. In some implementation examples, the weights are not equal and are predefined.
[0167] In some implementation examples, the probability is updated as a weighted sum of a plurality of probabilities (e.g., N probabilities), and one of the probabilities is
Number
[0168] In some implementation examples, the probability is updated as a weighted sum of a plurality of probabilities (e.g., N probabilities), and one of the probabilities is [Number] derived using a probability update rate that is, where another probability is [Number] derived using a probability update rate that is, where A’ is a given model parameter different from A. In some embodiments, another one (or more) of a plurality of probabilities (e.g., N probabilities) may be a function of any of the embodiments described above.
[0169] In some embodiments, the function f i (count,M) is different for different syntax / contexts.
[0170] In some embodiments, the values of the function parameters and / or weight parameters of the various embodiments described above may be specified in a high-level syntax / context including (but not limited to) VPS, PPS, SPS, APS, picture / frame headers, slice headers, tile headers, and CTU headers.
[0171] The various embodiments and / or examples described in this disclosure may be used individually or combined in any order. Further, parts, all or any partial combinations or all combinations of these embodiments and / or examples may be implemented as part of an encoder and / or decoder, and may also be implemented in hardware and / or software. For example, hard coding may be performed for the dedicated processing circuit (e.g., one or more integrated circuits) described above. In another example, the above may be implemented by one or more processors executing a program stored in a non-transitory computer-readable medium. Further, each of the method (or embodiment), encoder and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The embodiments of the present disclosure may be applied to luma blocks or chroma blocks. In the case of chroma blocks, the embodiments may be applied individually to one or more color components or together to one or more color components.
[0172] Using computer-readable instructions, the above-described techniques can be implemented as computer software physically stored on one or more computer-readable media. For example, FIG. 19 shows a computer system (2000) suitable for implementing some embodiments of the disclosed subject matter.
[0173] The computer software can be coded using any suitable machine language, i.e., computer language, which may be processed by an assembly, compilation, linking, or similar mechanism for generating code with instructions that can be executed directly or through interpretation, microcode execution, etc. by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.
[0174] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game consoles, Internet of Things devices, and the like.
[0175] The components shown in FIG. 20 of the computer system (2000) are of course typical, and no suggestion of any limitation is intended with respect to the use or function of the computer software implementing the embodiments of the present disclosure. Furthermore, the configuration of the components should not be construed as having any dependency or requirement related to any one or combination of the components shown in the typical embodiments of the computer system (2000).
[0176] The computer system (2000) may include several human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, through tactile input (keystrokes, swipes, movement of a data glove, etc.), voice input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). By using a human interface device, it is also possible to capture some media that is not necessarily directly related to conscious human input, such as voice (speech, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).
[0177] The input human interface device may include one or more of a keyboard (2001), a mouse (2002), a trackpad (2003), a touch screen (2010), a data glove (not shown), a joystick (2005), a microphone (2006), a scanner (2007), and a camera (2008) (only one of each is shown).
[0178] The computer system (2000) may also include several human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (2010), a data glove (not shown), or a joystick (2005), although there may also be tactile feedback devices that are not used as input devices), audio output devices (such as speakers (2009), headphones (not shown), etc.), visual output devices (screens including CRT screens, LCD screens, plasma screens, OLED screens (2010), etc., each of which may or may not have a touch screen input function and may or may not have a tactile feedback function - some of the above screens may be capable of outputting two-dimensional visual output or outputting three-dimensional or higher-dimensional output through means such as stereographic output; virtual reality glasses (not shown), hologram displays, and smoke tanks (not shown)), and printers (not shown).
[0179] The computer system (2000) can also include storage devices directly operable by humans and related media, such as optical media including CD / DVD ROM / RW (2020) using CD / DVD or similar media (2021), thumb-drives (2022), removable hard drives and solid-state drives (2023), legacy magnetic media such as tapes and floppy disks (not shown), and specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0180] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carrier waves, or other transient signals.
[0181] The computer system (2000) can also include an interface (2054) for one or more communication networks (2055). The network can be, for example, a wireless network, a wired network, or an optical network. Further, the network can be a local, wide area, metropolitan area, vehicle and industrial, real-time, delay-tolerant, etc. network. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital networks for television including cable television, satellite television, and terrestrial television, vehicle and industrial including CAN bus, etc. For some networks, it is common to require an external network interface adapter attached to some general-purpose data ports or peripheral buses (2049) (such as the USB port of the computer system (2000)), and for other networks, it is common to be incorporated into the core of the computer system (2000) by attaching to the system bus as described later (for example, an Ethernet interface is incorporated into a PC computer system or a cellular network interface is incorporated into a smartphone computer system). Using any of such networks, the computer system (2000) can communicate with a counterpart. Such communication can be unidirectional, reception-only (such as television broadcast), transmission-only unidirectional (such as from a CANbus to a specific CANbus device), or bidirectional, for example, bidirectional communication with other computer systems using a local or wide area digital network. Various protocols and protocol stacks can be used for each of the networks and network interfaces described above.
[0182] The above-described human interface device, storage device directly operable by a human, and network interface can be attached to the core (2040) of the computer system (2000).
[0183] The core (2040) can include one or more central processing units (CPUs) (2041), a graphics processing unit (GPU) (2042), a specialized programmable processing device in the form of a field programmable gate array (FPGA) (2043), a hardware accelerator (2044) for a specific task, a graphics adapter (2050), and the like. Such devices can be connected via a system bus (2048) together with a read-only memory (ROM) (2045), a random access memory (2046), an internal mass storage such as a hard drive or SSD (2047) that is internal and not directly accessible by the user. In some computer systems, the system bus (2048) can be in the form of one or more physical plugs that can be directly manipulated to allow expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core system bus (2048) or via a peripheral bus (2049). In one example, a screen (2010) can be connected to the graphics adapter (2050). Architectures of peripheral buses include PCI, USB, and the like.
[0184] The CPU (2041), GPU (2042), FPGA (2043), and accelerator (2044) can execute some instructions that can be combined to build the above-mentioned computer code. The computer code can be stored in the ROM (2045) or RAM (2046). Transient data can be stored in the RAM (2046), while immutable data can be stored, for example, in the internal mass storage (2047). By using a cache memory that can be closely associated with one or more CPUs (2041), GPUs (2042), mass storage (2047), ROM (2045), RAM (2046), etc., fast storage and reading to any of the memory devices can be enabled.
[0185] A computer-readable medium can carry computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and configured for this disclosure, or can be of the kind well known to and available from those skilled in the computer software arts.
[0186] As an example without limitation, a computer system having an architecture (2000), and in particular a core (2040), can provide functions as a result of one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software implemented on one or more tangible computer-readable media. Such a computer-readable media can be a media related to a large-capacity storage that can be directly operated by a user as described above, and further can be a specific storage having a non-transitory nature of the core, such as an on-core large-capacity storage (2047) or a ROM (2045). Software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (2040). The computer-readable media can include one or more memory devices or chips according to individual requirements. By the software, the core (2040), in particular the processors (including CPUs, GPUs, FPGAs, etc.) in the core (2040), can execute the specific process or a specific part of the specific process described above, including the definition of the data structure stored in the RAM (2046) and the modification of such a data structure according to the process defined by the software. In addition to or instead of this, logic can be hard-wired into a circuit (such as an accelerator (2044)) that operates instead of or in cooperation with the software to execute the specific process or a specific part of the specific process described in this application, and as a result of the logic being implemented in another way, the computer system can provide functions. If necessary, when described as "software", it may include logic, and vice versa. If necessary, when described as "computer-readable media", it may include a circuit (such as an integrated circuit (IC)) that stores the software to be executed, a circuit that implements the logic to be executed, or both. The present disclosure includes any suitable combination of hardware and software.
[0187] Although a particular invention has been described with reference to exemplary embodiments, this description is not intended to be limiting. Various modifications of the exemplary embodiments and further alternative embodiments of the invention will be apparent to those skilled in the art from this description. It will be readily understood by those skilled in the art that such modifications and various other modifications can be made to the exemplary embodiments illustrated and described in this disclosure without departing from the spirit and scope of the invention. Therefore, it is contemplated that such modifications may be covered by the appended claims and the embodiments may be changed. While certain ratios may be taken large in the examples, other ratios may be taken small. Therefore, this disclosure and the figures are to be considered illustrative rather than limiting.
Explanation of Reference Numerals
[0188] 300 Communication system 310 Terminal device 320 Terminal device 330 Terminal device 340 Terminal device 350 Communication network 400 Communication system 401 Video source 402 Stream 403 Video encoder 404 Encoded video bitstream 405 Streaming server 406 Client subsystem 407 Duplicate of encoded video data 408 Client subsystem 409 Duplicate of encoded video data 410 Video decoder 411 Output stream 412 Display 413 Video imaging subsystem 420 Electronic device 430 Electronic device 501 Channel 510 Video decoder 512 Display device 515 Buffer Memory 520 Parser 521 Symbol 530 Electronic Device 531 Receiver 551 Scaler / Inverse Converter 552 Intra Prediction Unit 553 Motion Compensation Prediction Unit 555 Assembler 556 Loop Filter Unit 557 Reference Picture Memory 558 Current Picture Buffer 601 Video Source 603 Video Encoder 620 Electronic Device 630 Source Coder 632 Coding Engine 633 Built-in Decoder 634 Reference Picture Memory 635 Predictor 640 Transmitter 643 Video Sequence 645 Entropy Encoder 650 Controller 660 Communication Channel 703 Video Encoder 721 Overall Controller 722 Intra Encoder 723 Residual Value Calculator 724 Residual Value Encoder 725 Entropy Encoder 726 Switch 728 Residual Decoder 730 Inter Encoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Reconstruction Module 880 Inter Decoder 2000 Computer System 2001 Keyboard 2002 Mouse 2003 Trackpad 2005 Joystick 2006 Microphone 2007 Scanner 2008 Camera 2009 Audio Output Device 2010 Touch Screen 2020 CD / DVD ROM / RW 2021 CD / DVD 2022 Thumb-drive 2023 Solid State Drive 2040 Core 2041 Central Processing Unit 2042 Graphics Processing Unit 2043 Field Programmable Gate Array 2044 Accelerator 2044 Hardware Accelerator 2045 Read-Only Memory 2046 Random Access Memory 2047 Mass Storage 2048 System Bus 2049 Peripheral Bus 2050 Graphics Adapter 2054 Interface 2055 Communication Network
Claims
1. A method for multiple symbol arithmetic coding in video decoding, the method comprising: a device comprising a memory for storing instructions and a processor communicating with the memory receiving a coded video bitstream; the device obtaining an array of cumulative distribution functions and corresponding M symbols, where M is an integer greater than 1; the device performing arithmetic decoding on the coded video bitstream to extract at least one symbol based on the array of cumulative distribution functions; for each extracted symbol, the device updating the array of cumulative distribution functions according to the extracted symbol based on at least one probability update rate; the device performing arithmetic decoding by continuing to perform arithmetic decoding on the coded video bitstream to extract the next symbol based on the updated array of cumulative distribution functions. A method comprising the above steps.
2. The at least one probability update rate comprises one probability update rate, the one probability update rate comprises a function f(N, M), where N is the number of occurrences of relevant symbols when parsing the coded video bitstream, The method according to claim 1.
3. The function f(N, M) comprises 【Number 1】 where A, B, C, and D are predefined function parameters, and g(M, D) is a function of M having D as a function parameter, The method according to claim 2.
4. g(M, D) has min(log 2 (M), D), The method according to claim 3.
5. The predefined function parameters are initialized for decoding a tile or a frame, The method according to claim 3.
6. The predefined function parameters are obtained from a high-level syntax comprising at least one of a video parameter set (VPS), a picture parameter set (PPS), a sequence parameter set (SPS), an adaptation parameter set (APS), a picture header, a frame header, a slice header, a tile header, or a coding tree unit (CTU) header, The method according to claim 3.
7. The function f(N, M) comprises 【Number 2】 where A, B, C, D, and E are predefined function parameters, and g(M, E) is a function of M having E as a function parameter, The method according to claim 2.
8. The function f(N, M) comprises [Number 3] comprising, where A and B are predefined function parameters, LUT(N) is a lookup table having N as an index, and g(M, B) is a function of M having B as a function parameter, The method according to claim 2.
9. wherein the at least one probability update rate comprises K probability update rates, and K is an integer greater than 1, each probability update rate comprises a function f(N, M), where N is the number of occurrences of a relevant symbol when parsing the coded video bitstream, updating the array of the cumulative distribution function according to the extracted symbol based on the at least one probability update rate comprises updating K arrays of the cumulative distribution function based on the corresponding K probability update rates, and updating the array of the cumulative distribution function as a weighted sum of the updated K arrays of the cumulative distribution function and comprising, The method according to claim 1.
10. K is 2, updating the array of the cumulative distribution function as a weighted sum of the updated K arrays of the cumulative distribution function comprises w 1 *p 1 +w 2 *p 2 and updating the array of the cumulative distribution function as, w 1 is the first weight, p 1 is the updated array of the first cumulative distribution function based on the first probability update rate, w 2 is the second weight, p 2 is the updated array of the second cumulative distribution function based on the second probability update rate The method according to claim 9.
11. the first weight and the second weight are equal, or the first weight and the second weight are different and predefined, The method according to claim 10.
12. one of the K probability update rates 【Number 4】 comprising, where A, B, C, and D are predefined model parameters, The method according to claim 9.
13. one of the K probability update rates [Number 5] comprising, another one of the K probability update rates 【Number 6】 comprising, where A, B, C, D, and E are predefined model parameters, and A is different from E, The method according to claim 9.
14. the predefined function parameters are obtained from a high-level syntax comprising at least one of a video parameter set (VPS), a picture parameter set (PPS), a sequence parameter set (SPS), an adaptation parameter set (APS), a picture header, a frame header, a slice header, a tile header, or a coding tree unit (CTU) header, The method according to claim 13.
15. An apparatus for multiple-symbol arithmetic coding in video decoding, the apparatus comprising a memory for storing instructions, and a processor communicating with the memory, wherein when the processor executes the instructions, the processor receives a coded video bitstream, Obtaining an array of cumulative distribution functions and corresponding M symbols, where M is an integer greater than 1, Performing arithmetic decoding on the coded video bitstream to extract at least one symbol based on the array of cumulative distribution functions, For each extracted symbol, updating the array of cumulative distribution functions according to the extracted symbol based on at least one probability update rate, Continuing to perform arithmetic decoding on the coded video bitstream to extract the next symbol based on the updated array of cumulative distribution functions, Performing arithmetic decoding by, A processor configured to cause the device to perform, A device comprising.
16. The at least one probability update rate comprises one probability update rate, The one probability update rate comprises a function f(N, M), where N is the number of occurrences of related symbols when parsing the coded video bitstream, The device according to claim 15.
17. The function f(N, M) comprises, 【Number 7】 where A, B, C, and D are predefined function parameters, and g(M, D) is a function of M having D as a function parameter, The device according to claim 16.
18. The function f(N, M) comprises, 【Number 8】 where A, B, C, D, and E are predefined function parameters, and g(M, E) is a function of M having E as a function parameter, The device according to claim 16.
19. The at least one probability update rate comprises K probability update rates, where K is an integer greater than 1, Each probability update rate comprises a function f(N, M), where N is the number of occurrences of related symbols when parsing the coded video bitstream, Updating the array of cumulative distribution functions according to the extracted symbol based on the at least one probability update rate comprises, Updating K arrays of cumulative distribution functions based on the corresponding K probability update rates, Updating the array of cumulative distribution functions as a weighted sum of the updated K arrays of cumulative distribution functions, Comprising, The device according to claim 15.
20. A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, cause the processor to, Receive a coded video bitstream, obtaining an array of cumulative distribution functions and corresponding M symbols, where M is an integer greater than 1, performing arithmetic decoding on the coded video bitstream to extract at least one symbol based on the array of cumulative distribution functions, updating the array of cumulative distribution functions according to the extracted symbols based on at least one probability update rate for each of the extracted symbols, continuing to perform arithmetic decoding on the coded video bitstream to extract the next symbol based on the updated array of cumulative distribution functions, performing arithmetic decoding by, a non-transitory computer-readable storage medium configured to cause the processor to perform the above. Claims 21 An apparatus for multi-symbol arithmetic coding in video decoding, the apparatus comprising: a memory for storing instructions; a processor in communication with the memory, the processor being configured to cause the apparatus to perform the method according to any one of claims 4, 5, 6, 8, 10, 11, 12, 13, and 14 when the processor executes the instructions, a processor. Claims 22 A non-transitory computer-readable storage medium storing instructions, the instructions being configured to cause a processor to perform the method according to any one of claims 2 to 14 when the instructions are executed by the processor.
Citation Information
Patent Citations
Parallel Entropy Coding
JP2024510268A
Efficient update of cumulative distribution functions for image compression
WO2022010531A1
Parallel entropy coding
WO2022231452A1