Method and apparatus for further improved context design for prediction mode and coded block flag (CBF)
Patent Information
- Application Number
- JP2024211961
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-04-24
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2039-11-27
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] [Related Applications] This application claims priority to U.S. Provisional Patent Application No. 62 / 777,041, filed December 7, 2018, and U.S. Patent Application No. 16 / 393,439, filed April 24, 2019, in the U.S. Patent and Trademark Office, which are hereby incorporated by reference in their entireties.
[0002] [Technical field] Methods and apparatus consistent with embodiments relate to video encoding, and in particular to methods and apparatus for improving prediction mode and coded block flag (CBF) context design. [Background technology]
[0003] FIG. 1A shows the intra prediction modes used in High Efficiency Video Coding (HEVC). There are a total of 35 intra prediction modes in HEVC, among which mode 10 (101) is the horizontal mode, mode 26 (102) is the vertical mode, and mode 2 (103), mode 18 (104), and mode 34 (105) are diagonal modes. The intra prediction modes are signaled by the three most probable modes (MPMs) and 32 remaining modes.
[0004] For Versatile Video Coding (VVC), the following is a partial coding unit syntax table. The flag pred_mode_flag is signaled if the slice type is not intra, skip mode is not selected, and only one context (e.g., the variable pred_mode_flag) is used to code this flag. The partial coding unit syntax table is as follows: [Table 1]
[0005] 1B, VVC has a total of 87 intra-prediction modes. Among them, mode 18 (106) is a horizontal mode, mode 50 (107) is a vertical mode, and mode 2 (108), mode 34 (109), and mode 66 (110) are diagonal modes. Modes -1 to -10 (111) and modes 67 to 76 (112) are Wide-Angle Intra Prediction (WAIP) modes.
[0006] The encoder selects the best chroma prediction mode from among five modes including planar mode (mode index 0), DC mode (mode index 1), horizontal mode (mode index 18), vertical mode (mode index 66), and diagonal mode (mode index 66) for the chroma components of an intra-coded block, and a direct copy of the intra prediction mode, i.e., DM mode, for the associated luma component. For chroma, the mapping between intra prediction direction and intra prediction mode number is shown in Table 1 below. [Table 2]
[0007] In order to avoid overlapping modes, the four modes other than the DM mode are assigned according to the intra prediction mode of the associated luma component. When the intra prediction mode number of the chroma component is 4, the intra prediction direction of the luma component is used to generate the intra prediction sample of the chroma component. When the intra prediction mode number of the chroma component is not 4 but is the same as the intra prediction mode number of the luma component, 66 intra prediction directions are used to generate the intra prediction sample of the chroma component.
[0008] Multi-hypothesis intra-inter prediction combines one intra prediction and one merge indexed prediction, i.e., intra-inter prediction mode. In a merge coding unit (CU), one flag is signaled for the merge mode, and when the flag is True, an intra mode is selected from the intra candidate list. For the luma component, the intra candidate list is derived from four intra prediction modes including DC, planar, horizontal, and vertical, and the size of the intra candidate list can be 3 or 4 depending on the block shape. When the CU width is greater than twice the CU height, the horizontal mode is removed from the intra candidate list, and when the CU height is greater than twice the CU width, the vertical mode is removed from the intra candidate list. The one intra prediction mode selected by the intra mode index and the one merge index prediction selected by the merge index are combined using a weighted average. For the chroma components, DM is always applied without additional signaling.
[0009] The weights for combining predictions are described below. When DC or planar mode is selected, or the coding block (CB) width or height is less than 4, equal weights are applied. For CBs with CB width or height equal to or greater than 4, when horizontal / vertical mode is selected, one CB is first divided horizontally / vertically into four equal area regions. Each weight set is divided into (w_intra i ,w_inter i), where i is 1 to 4, and (w_intra1,w_inter1)=(6,2), (w_intra2,w_inter2)=(5,3), (w_intra3,w_inter3)=(3,5), (w_intra4,w_inter4)=(2,6) are applied to the corresponding regions. (w_intra1,w_inter1) is for the region closest to the reference sample, and (w_intra4,w_inter4) is for the region furthest from the reference sample. The combined prediction can then be calculated by adding the two weighted predictions and shifting them right by 3 bits. Furthermore, the intra prediction modes of the intra multi-directional predictors can be preserved for subsequent intra mode coding of neighboring CBs if they are intra coded. Summary of the Invention
[0010] According to an embodiment, there is provided a method of video decoding or encoding, said method comprising the steps of: determining whether at least one of a plurality of neighboring blocks in the video sequence is coded using an intra-prediction mode; in response to determining that at least one of the plurality of neighboring blocks is coded in the intra prediction mode, entropy coding a prediction mode flag of the current block using a first context; in response to determining that none of the plurality of neighboring blocks are coded with at least the intra-prediction mode, entropy coding the prediction mode flag of the current block with a second context; There are methods including:
[0011] According to an embodiment, there is an apparatus for video encoding, decoding or encoding, the apparatus comprising at least one memory storing computer program code, and at least one processor configured to access the at least one memory and operate according to the computer program code, the computer program code comprising: a first decision code configured to cause the at least one processor to determine whether at least one of a plurality of neighboring blocks in a video sequence is coded using an intra-prediction mode; and executable code configured to cause the at least one processor to entropy encode a prediction mode flag of a current block using a first context in response to determining that at least one of the plurality of neighboring blocks is encoded using the intra prediction mode; and second execution code configured to cause the at least one processor to entropy code the prediction mode flag of the current block using a second context in response to determining that none of the plurality of neighboring blocks are coded using the intra prediction mode.
[0012] According to an embodiment, a non-transitory computer-readable storage medium storing instructions, the instructions being configured to cause at least one processor to: determining whether at least one of a plurality of neighboring blocks in the video sequence is coded using an intra-prediction mode; in response to determining that at least one of the plurality of neighboring blocks is coded in the intra prediction mode, entropy coding a prediction mode flag of a current block using a first context; in response to determining that none of the plurality of neighboring blocks are coded with at least the intra-prediction mode, entropy coding the prediction mode flag of the current block with a second context. There is a non-transitory computer readable storage medium.
[0013] According to an embodiment, the step of entropy coding the prediction mode flag of the current block comprises the step of coding only by the first context and the second context.
[0014] According to an embodiment, the method includes the steps of: determining whether at least one of the neighboring blocks is coded in an intra-inter prediction mode; in response to determining that at least one of the plurality of neighboring blocks is coded in the intra-inter prediction mode, entropy coding the prediction mode flag of the current block using the first context; in response to determining that none of the neighboring blocks are coded in either the intra prediction mode or the intra-inter prediction mode, entropy coding the prediction mode flag of the current block using the second context; Further includes:
[0015] According to an embodiment, in response to determining that at least one of the plurality of neighboring blocks is coded in the intra prediction mode, entropy coding a skip flag of the current block using a first skip context; In response to determining that none of the neighboring blocks are coded with at least the intra-prediction mode, entropy coding the skip flag of the current block with a second skip context.
[0016] According to an embodiment, in response to determining that at least one of the plurality of neighboring blocks is coded in the intra prediction mode, entropy coding an affine flag of the current block using a first affine context; In response to determining that none of the neighboring blocks are coded with at least the intra prediction mode, there is a step of entropy coding the affine flag of the current block with a second affine context.
[0017] According to an embodiment, in response to determining that at least one of the plurality of neighboring blocks is coded in the intra prediction mode, entropy coding a sub-block integration flag of the current block using a first sub-block integration context; In response to determining that none of the plurality of neighboring blocks are coded in at least the intra-prediction mode, there is a step of entropy coding the sub-block integration flag of the current block with a second sub-block integration context.
[0018] According to an embodiment, in response to determining that at least one of the plurality of neighboring blocks is coded by the intra prediction mode, entropy coding a CU split flag of the current block according to a first coding unit (CU) split context; In response to determining that none of the plurality of neighboring blocks are coded in at least the intra prediction mode, there is a step of entropy coding the CU" split flag of the current block using a second CU split context.
[0019] According to an embodiment, in response to determining that at least one of the plurality of neighboring blocks is coded in the intra prediction mode, entropy coding an AMVR flag of the current block using a first adaptive motion vector resolution (AMVR) context; In response to determining that none of the neighboring blocks are coded with at least the intra-prediction mode, entropy coding the AMVR flag of the current block with a second AMVR context.
[0020] According to an embodiment, in response to determining that at least one of the plurality of neighboring blocks is coded in the intra prediction mode, entropy coding an intra-inter mode flag of the current block using a first intra-inter prediction mode context; In response to determining that none of the plurality of neighboring blocks are coded with at least the intra prediction mode, there is a step of entropy coding the intra-inter mode flag of the current block with a second intra-inter mode context.
[0021] According to an embodiment, in response to determining that at least one of the plurality of neighboring blocks is coded in the intra prediction mode, entropy coding a triangular partition mode flag of the current block according to a first triangular partition mode context; In response to determining that none of the plurality of neighboring blocks are coded using at least the intra prediction mode, there is a step of entropy coding the triangular partition mode flag of the current block using a second triangular partition mode context.
[0022] According to an embodiment, in response to determining that at least one of the plurality of neighboring blocks is coded in the intra prediction mode, entropy coding a CBF of the current block according to a first coded block flag (CBF) context; In response to determining that none of the neighboring blocks are coded with at least the intra-prediction mode, entropy coding the CBF of the current block with a second CBF context. [Brief description of the drawings]
[0023] [Figure 1A]FIG. 2 is a diagram of intra prediction modes in HEVC.
[0024] [Figure 1B] FIG. 1 is a diagram showing intra prediction modes in VVC.
[0025] [Diagram 2] FIG. 1 is a simplified block diagram of a communication system according to one embodiment.
[0026] [Diagram 3] FIG. 2 is a diagram of an arrangement of a video encoder and a video decoder in a streaming environment according to one embodiment.
[0027] [Figure 4] FIG. 2 is a functional block diagram of a video decoder according to one embodiment.
[0028] [Diagram 5] FIG. 2 is a functional block diagram of a video encoder according to one embodiment.
[0029] [Figure 6] 2 is a diagram of a current block and neighboring blocks of the current block, according to one embodiment.
[0030] [Figure 7] 1 is a flowchart illustrating a method for controlling intra-inter prediction for decoding or encoding a video sequence according to one embodiment.
[0031] [Figure 8] 1 is a simplified block diagram of a device for controlling intra-inter prediction for decoding or encoding a video sequence according to one embodiment.
[0032] [Figure 9] FIG. 1 is a diagram of a computer system suitable for implementing embodiments.
[0033] [Figure 10] 4 is a flow chart illustrating a method for controlling the decoding or encoding of a video sequence according to one embodiment.
[0034] [Figure 11] 4 is a flow chart illustrating a method for controlling the decoding or encoding of a video sequence according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0035] 2 is a simplified block diagram of a communication system 200 according to one embodiment. The communication system 200 may include at least two terminals 210-220 interconnected via a network 250. In a unidirectional transmission of data, a first terminal 210 may locally encode video data for transmission to another terminal 220 via the network 250, and a second terminal 220 may receive the encoded video data of the other terminal from the network 250, decode the encoded data, and display the reconstructed video data. Unidirectional data transmission may be common in media serving applications, etc.
[0036] 2 shows a second pair of terminals (230, 240) adapted to support bidirectional transmission of encoded video, such as may occur during a video conference. In the bidirectional transmission of data, each terminal 220, 240 may encode locally captured video data for transmission to the other terminal over network 250. Each terminal 230, 240 may also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the recovered video data on a local display device.
[0037] In FIG. 2, the terminal devices 210-240 may be depicted as servers, personal computers, and smartphones, although the principles of the embodiments are not so limited. The embodiments have application with laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network 250 represents any number of networks that carry encoded video data between the terminal devices 210-240, including, for example, wired and / or wireless communication networks. The communication network 250 may exchange data over circuit-switched and / or packet-switched channels. Representative networks include electronic communication networks, local area networks, wide area networks, and / or the Internet. For purposes of the present discussion, the architecture and topology of the network 250 may not be important to the operation of the embodiments, unless otherwise noted below.
[0038] 3 is a diagram of an arrangement of video encoders and video decoders in a streaming environment 300, according to one embodiment. The disclosed subject matter is equally applicable to, for example, video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc., other video-enabled applications, and the like.
[0039] The streaming system may include a video source 301, a capture subsystem 313 which may include, for example, a digital camera, generating an uncompressed video sample stream 302. The sample stream 302 may be processed by an encoder 303, shown in bold to emphasize its high data volume when compared to an encoded video bitstream, coupled to the camera 301. The encoder 303 may include hardware, software, or a combination thereof, and may enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream 304, shown in thin to emphasize its low data volume when compared to the sample stream, may be stored on a streaming server 305 for future use. One or more streaming clients 306, 308 may access the streaming server 305 to retrieve copies 307, 309 of the encoded video bitstream 304. The client 306 may include a video decoder 310. The video decoder 310 decodes an incoming copy of the encoded bitstream 307 and generates an output video sample stream 311 that can be rendered on a display 312 or other rendering device (not shown). In some streaming systems, the video bitstreams 304, 307, 309 may be encoded according to a particular video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265. A video coding standard, VVC, is under development. The disclosed subject matter may be used in the context of VVC.
[0040] FIG. 4 is a functional block diagram 400 of a video decoder 310 according to one embodiment.
[0041] The receiver 410 may receive one or more coded video sequences to be encoded by the video decoder 310, one coded video sequence at a time in the same or another embodiment, where the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequences may be received from a channel 412, which may be a hardware / software link to a storage device that stores the coded video data. The receiver 410 may receive the coded video data along with other data, e.g., coded audio data and / or ancillary data streams, which may be forwarded to a respective using entity (not shown). The receiver 410 may separate the coded video sequences from the other data. To eliminate network jitter, a buffer memory 415 may be coupled between the receiver 410 and the entropy decoder / parser 420 (hereinafter "parser 420"). When the receiver 410 is receiving data controllably from a storage / forwarding device of sufficient bandwidth or from an isochronous network, the buffer 415 may not be needed or may be small. For use in a best effort packet network such as the Internet, buffer 415 may be necessary and may be relatively large, and may advantageously be adaptively sized.
[0042] The video decoder 310 may include a parser 420 to reconstruct symbols 421 from the coded video sequence. These categories of symbols include information used to manage the operation of the video decoder 310 and information for controlling a rendering device such as a display 312, which may possibly be coupled to the decoder, but is not an integral part of the decoder as shown in FIG. 4. The control information for the rendering device may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser 420 may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser 420 may extract a set of subgroup parameters from the coded video sequence for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the subgroup. Subgroups may include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter (QP) values, motion vectors, etc. from the coded video sequence.
[0043] Parser 420 may perform an entropy decoding / parsing operation on the video sequence received from buffer 415 to generate symbols 421. Parser 420 may receive the encoded data and selectively decode particular symbols 421. Additionally, parser 420 may determine whether a particular symbol 421 should be provided to motion compensated prediction unit 453, scaler / inverse transform unit 451, intra prediction unit 452, or loop filter unit 454.
[0044] The reconstruction of symbols 421 may include different units depending on the type of coded video picture or portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. Which units are included and how can be controlled by subgroup control information parsed from the coded video sequence by parser 420. The flow of such subgroup control information between parser 420 and the following units is not shown for clarity.
[0045] Beyond the functional blocks already mentioned, the decoder 310 can be conceptually subdivided into a number of functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is adequate.
[0046] The first unit is a scalar / inverse transform unit 451. The scalar / inverse transform unit 451 receives quantized transform coefficients and control information including which transform should be used, block size, quantization coefficients, quantization scaling matrix, etc. as symbols 421 from the parser 420. The scalar / inverse transform unit 451 can output a block containing sample values that can be input to an aggregator 455.
[0047] In some examples, the output samples of the scalar / inverse transform unit 451 may belong to intra-coded blocks, i.e. blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information may be provided by the intra-picture prediction unit 452. In some cases, the intra-picture prediction unit 452 generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the current (partially reconstructed) picture 456. The aggregator 455 adds the prediction information generated by the intra-prediction unit 452 to the output sample information provided by the scalar / inverse transform unit 451, in some cases, on a sample-by-sample basis.
[0048] In other cases, the output samples of the scalar / inverse transform unit 451 may relate to an inter-coded, possibly motion-compensated block. In such cases, the motion compensated prediction unit 453 may access the reference picture memory 457 to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols 421 related to the block, these samples may be added by the aggregator 455 to the output of the scalar / inverse transform unit to generate output sample information (in this case referred to as residual samples or residual signals). The addresses in the reference picture memory from which the motion compensated prediction unit fetches prediction samples may be controlled by the available motion vectors of the motion compensated prediction unit, for example in the form of symbols 421 that may have X, Y and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory when sub-sample accurate motion vectors are in use, motion vector prediction mechanisms, etc.
[0049] The output samples of aggregator 455 may be subjected to various loop filtering techniques in loop filter unit 454. The video compression techniques are controlled by parameters contained in the coded video bitstream and made available to loop filter unit 454 as symbols 421 from parser 420, but may include in-loop filter techniques that are also responsive to meta-information obtained during decoding of previous portions of the coded pictures or coded video sequences (in decoding order), and that may also be responsive to previously reconstructed loop filtered sample values.
[0050] The output of the loop filter unit 454 may be a sample stream that can be output to the render device 312 and stored in the reference picture memory 456 for use in future inter-picture prediction.
[0051] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 420), the current reference picture 456 can become part of reference picture memory 457, and fresh current picture memory can be reallocated before beginning reconstruction of a subsequent coded picture.
[0052] The video decoder 310 may perform decoding operations according to a given video compression technique defined in a standard such as ITU-T Rec. H.265. The encoded video sequence may comply with the syntax specified by the video compression technique or standard in use, in the sense that the encoded video sequence complies with the syntax of the video compression technique or standard specified in the video compression technique or standard, specifically in a profile document therein. Also, a requirement for compliance may be that the complexity of the encoded video sequence is within the limits defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples / second), maximum reference picture size, etc. The limits set by the level may in some cases be further restricted through a Hypothetical Reference Decoder (HRD) specification and metadata for HDR buffer management signaled in the encoded video sequence.
[0053] In one embodiment, the receiver 410 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (310) to correctly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0054] FIG. 5 is a functional block diagram 500 of the video encoder 303 according to one embodiment.
[0055] The encoder 303 may receive video samples from a video source 301 (not part of the encoder) that may capture video images to be encoded by the encoder 303 .
[0056] The video source 301 may provide a source video sequence to be encoded by the encoder 303 in the form of a digital video sample stream of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb4:2:0, YCrCb4:4:4). In a media presentation system, the video source (301) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of individual pictures that, when viewed in succession, give the appearance of motion. The pictures themselves may be organized as a spatial array of pixels. Each pixel may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily appreciate the relationship between pixels and samples. The following description focuses on samples.
[0057] According to an embodiment, the encoder 303 may encode and compress pictures of a source video sequence into an encoded video sequence 543 in real-time or under any other time constraint required by the application. Enforcing an appropriate encoding rate is one function of the controller 550. The controller controls and is functionally coupled to other functional units as described below. Couplings are not shown for clarity. Parameters set by the controller may include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller 550 as they may be relevant to a video encoder 303 optimized for a particular system design.
[0058] Some video encoders operate in what those skilled in the art immediately recognize as a "coding loop." As a very simplified explanation, the coding loop can include an encoder 530 (hereafter "source coder") (which generates symbols based on the input picture to be coded and the reference pictures) and a coding portion of a (local) decoder 533 that is embedded within the encoder 303 and reconstructs the symbols to generate sample data that a (remote) decoder can generate (when any compression between the symbols and the coded video bitstream is lossless among the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream is input to a reference picture memory 534. When the decoding of the symbol stream results in bit-exact results independent of the decoder location (local or remote), the contents of the reference picture buffer are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" exactly the same sample values as the decoder "sees" when using prediction during decoding as reference picture samples. This basic principle of reference picture synchrony (and the resulting drift if synchrony cannot be maintained, for example due to channel errors) is well known to those skilled in the art.
[0059] The operation of the "local" decoder 533 may be the same as that of the "remote" decoder 310, detailed above in connection with Figure 4. Referring also briefly to Figure 4, however, the entropy decoding portion of the decoder 310, including the channel 412, receiver 410, buffer 415, and parser 420, may not be fully implemented in the local decoder 533, since symbols are available and the encoding / decoding of the symbols into the encoded video sequence by the entropy coder 545 and parser 420 may be lossless.
[0060] An observation to be made at this point is that any decoder techniques, other than parsing / entropy decoding, present in the decoder must also be present in substantially the same functional form as in the corresponding encoder. Descriptions of the encoder techniques can be omitted since they are the inverse of the decoder techniques which are described generically. Only in certain areas are more detailed descriptions necessary and are provided below.
[0061] In operation, in some examples, the source coder 530 may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence designated as "reference frames." In this method, the coding engine 532 codes the difference between pixel blocks of the input frame and pixel blocks of reference frames that may be selected as prediction references for the input frame.
[0062] The local video decoder 533 may decode the encoded video data of frames, which may be designated as reference frames, based on the symbols generated by the source coder 530. The operation of the encoding engine 532 may advantageously be a lossy process. When the encoded video data may be decoded in a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder 533 may replicate the decoding process that may be performed by the video decoder on the reference frames, resulting in reconstructed reference frames to be stored in the reference picture cache 534. In this way, the encoder 303 may locally store copies of reconstructed reference frames that have common content with the reconstructed reference frames that would be obtained by the far-end video decoder (in the absence of transmission errors).
[0063] The predictor 535 may perform a prediction search for the coding engine 532. That is, for a new frame to be coded, the predictor 535 may search the reference picture memory 534 for sample data (such as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that may serve as suitable prediction references for the new picture. The predictor (535) may operate on a sample block-pixel block basis to find a suitable prediction reference. In some examples, the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory 534, as determined by the search results obtained by the predictor 535.
[0064] Control unit 550 may manage the encoding operations of video coder 530, including, for example, setting parameters and subgroup parameters used for encoding the video data.
[0065] The output of all the aforementioned functional units may undergo entropy coding in entropy coder 545. Entropy coder 545 converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques well known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0066] The transmitter 540 may buffer the encoded video sequence generated by the entropy coder 545 to prepare it for transmission over a communication channel 560, which may be a hardware / software link to a storage device that may store the encoded video data. The transmitter 540 may merge the encoded video data from the video coder 530 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).
[0067] A control unit 550 may manage the operation of the encoder 303. During encoding, the control unit 550 may assign a particular encoding picture type to each encoded picture, which may affect the encoding technique that may be applied to the respective picture. For example, pictures may often be assigned as one of the following picture types:
[0068] An intra picture (I-picture) may be a picture that can be coded and decoded without using any other picture in a sequence as a source of prediction. Some video codecs allow different kinds of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will recognize the variations of I-pictures and their respective applications and characteristics.
[0069] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra- or inter-prediction, in most cases using one motion vector and reference index to predict the sample values of each block.
[0070] A bidirectionally predicted picture (B-picture) may be a picture that can be coded and decoded using intra- or inter-prediction with up to two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predictive picture can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0071] A source picture may commonly be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks determined by a coding assignment applied to the respective picture of the block. For example, blocks of an I-picture may be non-predictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). Pixel blocks of a P-picture may be predictively coded, via spatial prediction or via temporal prediction, with reference to one previously coded reference picture. Blocks of a B-picture may be non-predictively coded, via spatial prediction or via temporal prediction, with reference to one or two previously coded reference pictures.
[0072] The video decoder 303 may perform encoding operations according to a given video encoding technique or standard, such as ITU-T Rec. H.265. In its operations, the video coder 303 may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. The encoded video data may therefore conform to a syntax specified by the video encoding technique or standard being used.
[0073] In one embodiment, the transmitter 540 may transmit additional data along with the encoded video. The video coder 530 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.
[0074] In the related art, for the flag pred_mode_flag indicating whether a block is intra or inter predicted, only one context is used, and the value of the flag applied to the neighboring block is not used. Furthermore, when a neighboring block is coded with an intra-inter prediction mode, it is predicted using a mixture of intra and inter prediction modes, and therefore, for the context design signaling the flag pred_mode_flag, it may be more efficient to consider whether the neighboring block is coded using an intra-inter prediction mode.
[0075] The embodiments described herein may be used separately or combined in any order.In the following description, the flag pred_mode_flag indicates whether the current block is intra or inter coded.
[0076] FIG. 6 is a diagram 600 of a current block and neighboring blocks of the current block, according to one embodiment.
[0077] 6, a current block 610 is shown along with a neighboring block 620 above and a neighboring block 630 to the left of the current block 610. The neighboring block 620 above and the neighboring block 630 to the left may each have a width of four and a height of four.
[0078] In an embodiment, information that neighboring blocks (e.g., the upper neighboring block 620 and the left neighboring block 630) are coded with an intra prediction mode, an inter prediction mode, or an intra-inter prediction mode is used to derive a context value used to entropy code the flag pred_mode_flag of a current block (e.g., the current block 610). In particular, when a neighboring block is coded with an intra-inter prediction mode, the associated intra prediction mode is used for intra mode coding and / or MPM derivation of the current block, but the neighboring block is considered as an inter coded block when deriving a context value for entropy coding the flag pred_mode_flag of the current block, even if intra prediction is used for the neighboring block.
[0079] In one example, the associated intra prediction mode of an intra-inter prediction mode is always planar.
[0080] In another example, the associated intra prediction mode of an intra-inter prediction mode is always DC.
[0081] In yet another example, the associated intra prediction mode is aligned with the intra prediction mode applied in the intra-inter prediction mode.
[0082] In one embodiment, when neighboring blocks (e.g., the upper neighboring block 620 and the left neighboring block 630) are coded using an intra-inter prediction mode, the associated intra prediction mode is used for intra mode coding and / or MPM derivation of the current block (e.g., the current block 610), but the neighboring blocks are also considered as intra-coded blocks when deriving context values for entropy coding the flag pred_mode_flag of the current block.
[0083] In one example, the associated intra prediction mode of an intra-inter prediction mode is always planar.
[0084] In another example, the associated intra prediction mode of an intra-inter prediction mode is always DC.
[0085] In yet another example, the associated intra prediction mode is aligned with the intra prediction mode applied in the intra-inter prediction mode.
[0086] In one embodiment, the context index or value is incremented by 2, 0, and 1 when the neighboring block is coded with intra prediction mode, inter prediction mode, and inter-intra prediction mode, respectively.
[0087] In another embodiment, the context index or value is incremented by 1, 0, and 0.5 when neighboring blocks are coded in intra prediction modes, inter prediction modes, and inter-intra prediction modes, respectively, and the final context index is rounded to the nearest integer.
[0088] After the context index or value is incremented for all of the neighboring blocks of the current block and a final context index is determined, an average context index may be determined based on the determined final context index divided by the number of neighboring blocks rounded to the nearest integer. A flag pred_mode_flag may be set to indicate that the current block is intra-coded or inter-coded based on the determined average context index. For example, if the determined average context index is 1, the flag pred_mode_flag may be set to indicate that the current block is intra-coded, and if the determined average context index is 0, the flag pred_mode_flag may be set to indicate that the current block is inter-coded.
[0089] In an embodiment, information about whether a current block (e.g., current block 610) is coded in an intra prediction mode, an inter prediction mode, or an inter-intra prediction mode is used to derive one or more context values for entropy coding the CBF of the current block.
[0090] In one embodiment, three separate contexts (e.g., variables) are used to entropy code the CBF: one used when the current block is coded in an intra prediction mode, one used when the current block is coded in an inter prediction mode, and one used when the current block is coded in an intra-inter prediction mode. The three separate contexts may be applied only for coding the luma CBF, only for coding the chroma CBF, or only for coding both the luma and chroma CBFs.
[0091] In another embodiment, two separate contexts (e.g., variables) are used to entropy code the CBF: one is used when the current block is coded with an intra prediction mode and one is used when the current block is coded with an inter prediction mode or an intra-inter prediction mode. The two separate contexts may be applied only for coding the luma CBF, only for coding the chroma CBF, or only for coding both the luma and chroma CBFs.
[0092] In yet another embodiment, two separate contexts (e.g., variables) are used to entropy code the CBF: one is used when the current block is coded in an intra- or inter-prediction mode, and one is used when the current block is coded in an intra-inter prediction mode. The two separate contexts may be applied only for coding the luma CBF, only for coding the chroma CBF, or only for coding both the luma and chroma CBFs.
[0093] 7 is a flow chart illustrating a method 700 for controlling intra-inter prediction for decoding or encoding a video sequence, according to one embodiment. In some implementations, one or more of the processing blocks in FIG. 7 may be performed by the decoder 310. In some implementations, one or more of the processing blocks in FIG. 7 may be performed by another device or group of devices separate from or including the decoder 310, such as the encoder 303.
[0094] 7, in a first block 710, the method 700 includes determining whether a neighboring block of the current block is coded in an intra-inter prediction mode. Based on determining that the neighboring block is not coded in an intra-inter prediction mode (No in 710), the method 700 ends.
[0095] Based on determining that the neighboring block is coded using an intra-inter prediction mode (Yes at 710), in a second block 720, the method 700 includes a step of performing intra-mode coding of the current block using an intra prediction mode related to the intra-inter prediction mode.
[0096] In a third block 730, the method 700 includes setting a prediction mode flag indicating whether the current block is intra-coded or inter-coded, such that the prediction mode flag indicates that the current block is inter-coded.
[0097] Method 700 may further include, based on determining that the neighboring block is coded using an intra-inter prediction mode (Yes at 710), performing MPM derivation for the current block using an intra prediction mode related to the intra-inter prediction mode.
[0098] An intra prediction mode related to the intra-inter prediction mode may be a planar mode, a DC mode, or an intra prediction mode applied in the intra-inter prediction mode.
[0099] The method 700 includes the steps of determining whether a neighboring block is coded in an intra prediction mode, an inter prediction mode, or an intra-inter prediction mode; incrementing a context index of a prediction mode flag by 2 based on determining that the neighboring block is coded using an intra prediction mode; incrementing a context index by 0 based on determining that the neighboring block is coded in an inter prediction mode; incrementing a context index by one based on determining that the neighboring block is coded in an intra-inter prediction mode; determining an average context index based on the incremented context index and a number of neighboring blocks of the current block; Setting a prediction mode flag based on the determined average context index may further be included.
[0100] The method includes the steps of determining whether a neighboring block is coded in an intra prediction mode, an inter prediction mode, or an intra-inter prediction mode; incrementing a context index of a prediction mode flag by one based on determining that the neighboring block is coded using an intra prediction mode; incrementing a context index by 0 based on determining that the neighboring block is coded in an inter prediction mode; incrementing a context index by 0.5 based on determining that the neighboring block is coded in an intra-inter prediction mode; determining an average context index based on the incremented context index and a number of neighboring blocks of the current block; Setting a prediction mode flag based on the determined average context index may further be included.
[0101] Although Figure 7 illustrates example blocks of method 700, method 700 may, in some implementations, include more, fewer, or a different arrangement of blocks than those illustrated in Figure 7. Additionally or alternatively, two or more of the blocks of method 700 may be performed in parallel.
[0102] Additionally, the proposed methods may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium to perform one or more of the proposed methods.
[0103] FIG. 8 is a simplified block diagram of a device 800 for controlling intra-inter prediction for decoding or encoding a video sequence, according to one embodiment.
[0104] 8, the device 800 includes a first decision code 810, an execution code 820, and a setting code 830. The device 800 may further include an increment code 840 and a second decision code 850.
[0105] The first decision code 810 is configured to cause the at least one processor to determine whether a neighboring block of the current block is coded in an intra-inter prediction mode.
[0106] The execution code 820 is configured to cause at least one processor to perform intra-mode encoding of the current block using an intra-prediction mode associated with the intra-inter prediction mode based on determining that the neighboring block is encoded in the intra-inter prediction mode.
[0107] The setting code 830 is configured to cause the at least one processor to set a prediction mode flag indicating whether the current block is intra-coded or inter-coded based on determining that the neighboring block is coded in an intra-inter prediction mode, such that the prediction mode flag indicates that the current block is inter-coded.
[0108] The execution code 820 is configured to cause at least one processor to perform a Most Probable Mode (MPM) derivation for the current block using an intra prediction mode related to the intra-inter prediction mode based on determining that the neighboring block is encoded in the intra-inter prediction mode.
[0109] An intra prediction mode related to the intra-inter prediction mode may be a planar mode, a DC mode, or an intra prediction mode applied in the intra-inter prediction mode.
[0110] The first decision code 810 may be further configured to cause the at least one processor to determine whether the neighboring block is coded in an intra prediction mode, an inter prediction mode, or an intra-inter prediction mode. The increment code 840 may be configured to cause the at least one processor to increment a context index of the prediction mode flag by 2 based on determining that the neighboring block is coded in an intra prediction mode, to increment the context index by 0 based on determining that the neighboring block is coded in an inter prediction mode, and to increment the context index by 1 based on determining that the neighboring block is coded in an intra-inter prediction mode.
[0111] The second determining code 850 may be configured to cause the at least one processor to determine an average context index based on the incremented context index and the number of neighboring blocks of the current block. The setting code 830 may be further configured to cause the at least one processor to set a prediction mode flag based on the determined average context index.
[0112] The first decision code 810 may be further configured to cause the at least one processor to determine whether the neighboring block is coded in an intra prediction mode, an inter prediction mode, or an intra-inter prediction mode. The increment code 840 may be configured to cause the at least one processor to increment a context index of the prediction mode flag by 1 based on determining that the neighboring block is coded in an intra prediction mode, to increment the context index by 0 based on determining that the neighboring block is coded in an inter prediction mode, and to increment the context index by 0.5 based on determining that the neighboring block is coded in an intra-inter prediction mode. The second decision code 850 may be configured to cause the at least one processor to determine an average context index based on the incremented context index and the number of neighboring blocks of the current block. The setting code 830 may be further configured to cause the at least one processor to set the prediction mode flag based on the determined average context index.
[0113] The techniques described above may be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. Figure 9 is a diagram of a computer system 900 suitable for implementing embodiments.
[0114] Computer software can be encoded using any suitable machine code or computer language that can be processed by mechanisms such as assembly, compilation, linking, etc. to generate code including instructions that can be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly or through interpretation, microcode execution, etc.
[0115] The instructions may be executed in a variety of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.
[0116] 9 of the computer system 900 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments. Furthermore, the arrangement of components should not be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of the computer system (900).
[0117] The computer system 900 may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users, for example, through sensory input (e.g., keystrokes, swipes, data grabbing actions), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as sound (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a digital camera), and video (including, for example, two-dimensional video, three-dimensional video, stereoscopic video).
[0118] The input human interface devices may include one or more of a keyboard 901, a mouse 902, a trackpad 903, a touch screen 910, a data glove 904, a joystick 905, a microphone 906, a scanner 907, and a camera 908 (only one of which is shown).
[0119] The computer system 900 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through sensory output, sound, light, and smell / taste. Such human interface output devices may include sensory output devices (e.g., sensory feedback via touch screen 910, data glove 904, or joystick 905, although there may also be sensory feedback devices that do not function as input devices), audio output devices (e.g., speakers 909, headphones (not shown), visual output devices (e.g., screen 910, including cathode ray tube (CRT) screen, liquid crystal display (LCD) screen, plasma screen, organic light emitting diode (OLED) screen, each with or without touch screen input capability, each with or without sensory feedback capability, some of which may be capable of outputting more than one output, for example, stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke generator tanks (not shown), and printers (not shown)).
[0120] The computer system 900 may also include human accessible storage and associated media such as optical media such as CD / DVD ROM / RW 920 with media 921 such as CDs / DVDs, thumb drives 922, removable hard drives or solid state drives 923, legacy magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), etc.
[0121] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0122] The computer system 900 may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan area, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, global systems for mobile communications (GSM), cellular networks including 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, vehicular and industrial including CANBus, etc. Particular networks generally require an external network interface that is attached to a particular general purpose data port or peripheral bus 949 (e.g., a USB port on the computer system 900). Others are generally integrated into the core of the computer system 900 by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Using these networks, the computer system (900) can communicate with other entities. Such communications may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANBus to a particular CANbus device), or bidirectional to other computer systems, for example, using local or wide area digital networks. Specific protocols and protocol stacks may be used with each of the above-mentioned networks and network interfaces.
[0123] The aforementioned human interface devices, human accessible storage, and network interfaces may be attached to a core 940 of the computer system 900 .
[0124] The core 940 may include one or more central processing units (CPUs) 941, graphics processing units (GPUs) 942, dedicated programmable processing units in the form of GPGAs 943, hardware accelerators for specific tasks 944, etc. These devices may be connected through a system bus 948, along with read only memory (ROM) 945, random access memory (RAM) 946, internal mass storage device 947 such as an internal non-user accessible hard drive, SSD, etc. In some computer systems, the system bus 948 is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripherals can be attached directly to the core's system bus 948 or through a peripheral bus 949. Peripheral bus architectures include PCI, USB, etc.
[0125] The CPU 941, GPU 942, FPGA 943, and accelerator 944 may execute certain instructions that may be combined to generate the aforementioned computer code. The computer code may be stored in ROM 945 or RAM 946. Temporary data may also be stored in RAM 946, while permanent data may be stored, for example, in an internal mass storage device 947. Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU (941), GPU (942), mass storage device (947), ROM (945), RAM (946), etc.
[0126] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the embodiments, or they may be of the kind well known and available to those skilled in the computer software arts.
[0127] As an example and not by way of limitation, computer system 900 having the architecture, and core 940 in particular, can provide functionality as a result of processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be specific storage of core 940 of a non-transitory nature, such as core internal mass storage 947 or ROM 945, as well as media associated with user-accessible mass storage devices as described above. Software implementing various embodiments of the present disclosure can be stored in such devices and executed by core 940. Computer-readable media can include one or more memory devices or chips, according to specific needs. The software can cause core (940) and specifically the processors therein (including CPUs, GPUs, FPGAs, etc.) to perform specific operations or specific portions of specific operations described herein, including defining and modifying data structures stored in RAM (946) according to software-defined operations. Additionally or alternatively, the computer system may provide functionality as a result of implementation in hardwired or other circuitry (e.g., accelerator (944)) that can operate in conjunction with or in place of software to perform certain processes or certain portions of certain processes described herein. Reference to software includes logic, and vice versa, where appropriate. Reference to computer-readable medium may include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that implements logic for execution, or both, where appropriate. Embodiments include any appropriate combination of hardware and software.
[0128] The inventors have found that pred_mode_flag has some correlation with the current block and its neighbors, the neighboring blocks. If the neighbors use pred_mode_flag equal to 1, the current block may or should also use pred_mode_flag equal to 1. In such a case, improved context efficiency can be achieved in arithmetic encoding / decoding thanks to the correlation. The neighboring block information can be used as context to entropy encode / decode the current pred_mode_flag for improved encoding efficiency.
[0129] Additionally, in VVC, there may be a special prediction mode, intra-inter prediction mode, which is a mixture of inter prediction and intra prediction.
[0130] Therefore, if it is desired to add a neighboring block as a context for entropy encoding / decoding of pred_mode_flag, it may be necessary to determine whether the neighboring block is considered as an intra or inter mode for deriving the context when the neighboring block is coded with an intra-inter prediction mode. Aspects of the present application relate to the design of the context for entropy encoding / decoding of pred_mode_flag, for which there may be different designs.
[0131] As mentioned above, when multiple neighboring blocks are identified, three contexts may be used depending on how many neighboring blocks are intra-coded. For example, if there are no intra-coded neighboring blocks, context index 0 may be used. If there is one intra-coded neighboring block, there may be two neighboring blocks, context index 1 may be used, otherwise context index 2 may be used. Thus, depending on the number of neighboring blocks coded by intra-coding, three contexts may exist.
[0132] However, further improvement may be achieved by reducing the number of contexts from three to two. For example, if none of the neighboring blocks are coded by intra coding, the first context may be used, otherwise, if any of the neighboring blocks use intra coding, another context may be used. Such a design may be adopted as part of VVC.
[0133] FIG. 10 illustrates an exemplary flowchart 1000 according to an embodiment, which is similar to the above-described embodiment except for the differences described herein.
[0134] In step 1001, a check is made to determine one or more neighboring blocks of the current block, and in step 1002, it may be determined whether one or more of these neighboring blocks are intra-coded. If not, in step 1003, a first context may be used for the current block, and if yes, in step 1004, another context may be used for the current block.
[0135] According to an example embodiment, not only are one or more neighboring blocks checked to determine whether intra-coding is used, but also whether intra-inter coding is used may be checked.
[0136] FIG. 11 illustrates an exemplary flowchart 1100 according to an embodiment, similar to the above-described embodiment except for the differences described herein.
[0137] In step 1101, a check is made to determine one or more neighboring blocks of the current block, and in step 1102, it may be determined whether one or more of these neighboring blocks are intra-coded. If not, in step 1103, it may be checked whether one or more neighboring blocks use intra-inter coding, and if yes, in step 1104, a first context may be used for the current block. Otherwise, in step 1102, intra-coding is determined, or in step 1103, another context may be used for the current block in step 1105.
[0138] In addition to FIG. 10, and similarly in FIG. 11, pred_mode_flag was discussed, but further adoption by VVC includes other syntax elements, such as skip flags, affine flags, and sub-block merge flags, as further described by the syntax below.
[0139] [Table 3]
[0140] [Table 4]
[0141] [Table 5] TIFF2025026534000007.tif20193
[0142] That is, in an exemplary embodiment, several neighboring blocks are first identified and two contexts are used to entropy code the pred_mode_flag of the current block, and when none of the identified neighboring blocks are coded in an intra prediction mode, the first context may be used, and in other cases the other context may be used.
[0143] Also, some neighboring blocks may be identified first, and two contexts may be used to entropy code the pred_mode_flag of the current block. When none of the identified neighboring blocks are coded with intra prediction mode or intra-inter prediction, the first context may be used, otherwise the second context may be used.
[0144] In another example, several neighboring blocks may be first identified, and two contexts may be used to entropy code the CBF flag of the current block. When none of the identified neighboring blocks are coded by an intra-prediction mode, the first context may be used, otherwise the second context may be used.
[0145] In another example, several neighboring blocks may be first identified, and two contexts may be used to entropy code the CBF flag of the current block. When none of the identified neighboring blocks are coded in an intra-prediction mode or an intra-inter prediction mode, the first context may be used, otherwise the second context may be used.
[0146] For entropy coding syntax elements such as skip flag (cu_skip_flag), affine flag (inter_affine_flag), subblock merge flag (merge_subblock_flag), CU split flag (qt_split_cu_flag, mtt_split_cu_flag, mtt_split_cu_vertical_flag, mtt_split_cu_binary_flag), IMV flag (amvr_mode), intra-inter mode flag, triangular partitioning flag, according to the above embodiment, we propose to use two contexts depending on the corresponding flag value used for the neighboring block. Examples of such flags are introduced in the above table. The meaning of these flags may suggest different respective modes such as one or another of skip mode, affine mode, subblock merge mode, etc.
[0147] Further, according to an exemplary embodiment, some neighboring blocks are first identified. When entropy coding the above flags, if none of the identified neighboring blocks are coded by the corresponding mode (meaning that the associated flag value is signaled with a value indicating that the corresponding mode is possible), the first context may be used, otherwise the second context may be used. Additionally, as can be understood by those skilled in the art in view of the present disclosure, any of the embodiments referring to Figures 10 and 11 may be used with these additional flags.
[0148] As mentioned above, in accordance with an example embodiment, the number of contexts may be advantageously reduced to two contexts for flags and such predictions.
[0149] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents, which are encompassed within the scope of this disclosure. It will be apparent to those skilled in the art that numerous systems and methods can be devised that, although not explicitly shown or described herein, embody the principles of this disclosure and thus are encompassed within the spirit and scope of this disclosure.
Claims
A method for entropy encoding a first prediction mode flag of a current block by determining a context of the first prediction mode flag of the current block based on a first prediction mode flag of an adjacent block, wherein the adjacent block is encoded in any one of an intra mode, an inter mode, or an intra-inter mode, and the intra-inter prediction mode is different from both the inter mode and the intra mode, the neighboring block is at least one of the left block and the upper block of the current block, and the method includes: when it is determined that none of the adjacent blocks is encoded in the intra mode, entropy encoding the first prediction mode flag of the current block with a first context; when it is determined that the neighboring block is encoded in the intra mode, entropy encoding the first prediction mode flag of the current block with a second context; including wherein the first prediction mode flag is pred_mode_flag. A method. The method according to claim 1, wherein the second context is 1 and the first context is different from the second context. A method for entropy decoding a first prediction mode flag of a current block by determining a context of the first prediction mode flag of the current block based on a first prediction mode flag of an adjacent block, wherein the adjacent block is decoded in any one of an intra mode, an inter mode, or an intra-inter mode, and the intra-inter prediction mode is different from both the inter mode and the intra mode, the neighboring block is at least one of the left block and the upper block of the current block, and the method includes: when it is determined that none of the adjacent blocks is decoded in the intra mode, entropy decoding the first prediction mode flag of the current block with a first context; when it is determined that the neighboring block is decoded in the intra mode, entropy decoding the first prediction mode flag of the current block with a second context; including wherein the first prediction mode flag is pred_mode_flag. A method. A method for entropy encoding a first prediction mode flag of a current block into encoded data by determining a context of the first prediction mode flag of the current block based on a first prediction mode flag of an adjacent block, wherein the adjacent block is encoded in any of an intra mode, an inter mode, or an intra-inter mode, and the intra-inter prediction mode is different from both the inter mode and the intra mode, the neighboring block is at least one of a left block and an upper block of the current block, and the method includes generating an encoded data bitstream, and transmitting the bitstream, and the method includes when it is determined that none of the adjacent blocks is encoded in the intra mode, entropy encoding the first prediction mode flag of the current block with a first context; and when it is determined that the neighboring block is encoded in the intra mode, entropy encoding the first prediction mode flag of the current block with a second context, wherein the first prediction mode flag is pred_mode_flag, the method. A device having at least one memory configured to store computer program code, and at least one processor configured to access the at least one memory and operate in accordance with the computer program code, wherein the computer program code causes the at least one processor to perform the method according to any one of claims 1 to 4. A computer program for causing at least one processor to perform the method according to any one of claims 1 to 4.