Video bitstream encoding method, apparatus, and program
By determining context values based on adjacent block encoding modes and applying weighted averaging, the method optimizes entropy coding for intra-inter prediction, addressing inefficiencies in existing video coding technologies and enhancing compression efficiency.
Patent Information
- Application Number
- JP2024137497
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-27
- Filing Date
- 2024-08-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2039-11-12
AI Technical Summary
Existing video coding technologies, such as HEVC and VVC, face inefficiencies in entropy coding of prediction modes and coding block flags due to the lack of context-based signaling for intra-inter prediction modes, leading to redundant mode signaling and suboptimal compression efficiency.
Implement a method for controlling intra-inter prediction by determining the context value based on the encoding mode of adjacent blocks, using specific intra prediction modes for entropy encoding, and applying weighted averaging for combined predictions to optimize mode signaling.
Improves video coding efficiency by reducing redundant mode signaling and enhancing compression performance through context-aware entropy encoding.
Smart Images

Figure 0007701528000003 
Figure 0007701528000004 
Figure 0007701528000005
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Provisional Patent Application No. 62 / 767,473, filed on November 14, 2018, and U.S. Application No. 16 / 454,294, filed on June 27, 2019, with the United States Patent and Trademark Office, the entire contents of which are incorporated herein by reference.
[0002] The methods and apparatuses corresponding to the embodiments relate to video coding, and in particular, to methods and apparatuses for improved context design for entropy coding of prediction modes and coding block flags (CBFs).
Background Art
[0003] FIG. 1A shows the intra - prediction modes used in High Efficiency Video Coding (HEVC). In HEVC, there are a total of 35 intra - prediction modes. In these intra - prediction modes, mode 10 (101) is the horizontal mode, mode 26 (102) is the vertical mode, and modes 2 (103), 18 (104), and 34 (105) are the diagonal modes. These intra - prediction modes are signaled by three most probable modes (MPMs) and the remaining 32 modes.
[0004] Regarding Versatile Video Coding (VVC), some coding unit syntax tables are shown below. When the slice type is not intra and the skip mode is not selected, the flag pred_mode_flag is signaled and the flag is encoded using only one context (e.g., the variable pred_mode_flag). The syntax tables of some coding units are as follows.
Table 1
[0005] Referring to FIG. 1B, in VVC, there are a total of 87 intra prediction modes. In these intra prediction modes, mode 18 (106) is the horizontal mode, mode 50 (107) is the vertical mode, and modes 2 (108), 34 (109), and 66 (110) are the diagonal modes. Modes -1 to -10 (111), and modes 67 to 76 (112) are called Wide - Angle Intra Prediction (WAIP) modes.
[0006] For the chrominance component of the intra - coding block, the coder selects the optimal chrominance prediction mode among five modes including the planar mode (mode index 0), the DC mode (mode index 1), the horizontal mode (mode index 18), the vertical mode (mode index 50), and the diagonal mode (mode index 66), and selects the direct copy of the intra - prediction mode of the related luminance component, that is, the DM mode. Table 1 below shows the mapping between the intra - prediction direction of chrominance and the number of the intra - prediction mode. [Table 2] [Summary of the Invention] [Problems to be Solved by the Invention]
[0007] To avoid duplicate modes, the four modes other than the DM mode are assigned based on the intra - prediction mode of the related luminance component. When the number of the intra - prediction mode of the chrominance component is 4, the intra - prediction direction of the luminance component is applied to generate the intra - prediction samples of the chrominance component. When the number of the intra - prediction mode of the chrominance component is not 4 and is the same as the number of the intra - prediction mode of the luminance component, the intra - prediction direction 66 is applied to generate the intra - prediction samples of the chrominance component.
[0008] Multi-hypothesis inter-intra prediction combines one intra prediction and one prediction with a merge index, i.e., it enters the inter-intra prediction mode. In a merge coding unit (CU), for the merge mode, by signaling one flag, if the flag is true, an intra mode is selected from the intra candidate list. For the luminance component, the intra candidate list is obtained from four intra prediction modes including the DC mode, the planar mode, the horizontal mode, and the vertical mode. Depending on the block shape, the size of the intra candidate list may be 3 or 4. If the width of the CU is greater than twice the height of the CU, the horizontal mode is removed from the intra candidate list. If the height of the CU is greater than twice the width of the CU, the vertical mode is removed from the intra candidate list. Using weighted average, one intra prediction mode selected by the intra mode index and one prediction with a merge index selected by the merge index are combined. For the chrominance component, no additional signaling is required and DM is always used.
[0009] The weights for combining the predictions are described as follows. Equal weights are applied if the DC mode, or the planar mode is selected, or if the width or height of the coding block (CB) is less than 4. For these CBs where the width or height of the CB is 4 or more, if the horizontal / vertical mode is selected, first, one CB is divided vertically / horizontally into four equal-area regions. For each region, the corresponding (w_intra i , w_inter i) Apply the weight set shown as i ranges from 1 to 4, where (w_intra1, w_inter1) = (6, 2), (w_intra2, w_inter2) = (5, 3), (w_intra3, w_inter3) = (3, 5), and (w_intra4, w_inter4) = (2, 6). (w_intra1, w_inter1) corresponds to the region closest to the reference sample, and (w_intra4, w_inter4) corresponds to the region farthest from the reference sample. Then, sum the two weighted predictions and right-shift by 3 bits to calculate the combined prediction. Also, the intra prediction mode of the predictor's intra hypothesis can be saved to perform intra mode encoding on these CBs when the subsequent adjacent CBs are intra-coded.
Means for Solving the Problem
[0010] According to an embodiment, a method for controlling intra-inter prediction for decoding or encoding a video sequence is executed by at least one processor. The method includes determining whether adjacent blocks of a current block are encoded in an intra-inter prediction mode, and based on determining that the adjacent blocks are encoded in the intra-inter prediction mode, using an intra prediction mode associated with the intra-inter prediction mode to perform intra mode encoding of the current block, setting a prediction mode flag associated with the adjacent blocks, obtaining a context value based on the set prediction mode flag associated with the adjacent blocks, and performing entropy encoding on a prediction mode flag associated with the current block indicating that the current block is intra-encoded using the obtained context value.
[0011] According to an embodiment, an apparatus for controlling intra-inter prediction for decoding or encoding a video sequence includes at least one memory arranged to store computer program code, and at least one processor arranged to access the at least one memory and operate based on the computer program code. The computer program code is arranged to cause the at least one processor to: a first determination code for determining whether an adjacent block of a current block is encoded in an intra-inter prediction mode; an execution code for causing the at least one processor to perform intra-mode encoding of the current block using an intra prediction mode associated with the intra-inter prediction mode based on a determination that the adjacent block is determined to be encoded in the intra-inter prediction mode; and a setting code for causing the at least one processor to set a prediction mode flag associated with the adjacent block based on a determination that the adjacent block is determined to be encoded in the intra-inter prediction mode, obtain a context value based on the set prediction mode flag associated with the adjacent block, and perform an operation of performing entropy encoding on a prediction mode flag associated with the current block indicating that the current block is intra-encoded using the obtained context value.
[0012] According to an embodiment, a non-transitory computer-readable storage medium storing instructions, the instructions causing at least one processor to perform steps of determining whether adjacent blocks of a current block are encoded in an intra-inter prediction mode, performing intra-mode encoding of the current block using an intra prediction mode associated with the intra-inter prediction mode based on a determination that the adjacent blocks are encoded in the intra-inter prediction mode, setting a prediction mode flag associated with the adjacent blocks, obtaining a context value based on the set prediction mode flag associated with the adjacent blocks, and performing entropy encoding on a prediction mode flag associated with the current block indicating that the current block is intra-encoded using the obtained context value.
Brief Description of the Drawings
[0013]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Best Mode for Carrying Out the Invention
[0014] FIG. 2 is a simplified block diagram of a communication system (200) according to an embodiment. The communication system (200) may include at least two terminals (210-220) interconnected via a network (250). In the case of unidirectional data transmission, the first terminal (210) can transmit video data at a local location, encoded, via the network (250) to another terminal (220). The second terminal (220) can receive the encoded video data of another terminal from the network (250), decode the encoded video data, and display the restored video data. Unidirectional data transmission may be common in media service applications and the like.
[0015] FIG. 2 shows a second pair of terminals (230, 240) provided to support bidirectional transmission of encoded video that may occur, for example, during a video conference. In the case of bidirectional data transmission, each terminal (230, 240) can transmit video data captured at a local location, encoded, via the network (250) to another terminal. Each terminal (230, 240) can also receive the encoded video data transmitted from another terminal, decode the encoded video data, and display the restored video data on a local display device.
[0016] In FIG. 2, the terminals (210-240) are shown as a server, a personal computer, and a smartphone, but the principles of the embodiments are not limited thereto. The embodiments are applicable to laptop computers, tablets, media players, and / or dedicated video conferencing equipment. The network (250) represents any number of networks, including, for example, wired and / or wireless communication networks, for transmitting encoded video data between the terminals (210-240). The communication network (250) can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the examination of this application, the architecture and topology of the network (250) may not be important for the operation of the embodiments, unless otherwise described herein below.
[0017] FIG. 3 is a diagram of the arrangement of a video encoder and a video decoder in a streaming environment according to an embodiment. The disclosed theme can equivalently be applied to other applications for supporting video, including, for example, video conferencing, digital TV, storage on digital media such as CDs, DVDs, memory sticks, etc. that contain compressed video.
[0018] The streaming system may include a capture subsystem (313), and the capture subsystem may include, for example, a video source (301) (such as a digital camera) for creating an uncompressed video sample stream (302). Compared with the encoded video bitstream, the sample stream (302) is drawn as a thick line to emphasize that it has a large data volume, and the sample stream (302) can be processed by an encoder (303) connected to the imaging device (301). The encoder (303) may include hardware, software, or a combination thereof to implement or carry out each aspect of the disclosed theme described in more detail below. Compared with the sample stream, the encoded video bitstream (304) is drawn as a thin line to emphasize that it has a small data volume, and the encoded video bitstream (304) can be stored in a streaming server (305) for future use. One or more streaming clients (306, 308) can access the streaming server (305) to retrieve replicas (307, 309) of the encoded video bitstream (304). The client (306) can include a video decoder (310) for decoding the incoming replica (307) of the encoded video data and creating an outgoing video sample stream (311) to be rendered on a display (312) or other rendering device (not shown). In some streaming systems, the video bitstreams (304, 307, 309) can be encoded based on some video encoding / compression standards. Examples of these standards include the ITU-T H.265 proposal. A video encoding standard, informally called VVC, is under development. The disclosed theme may be applicable in the context of VVC.
[0019] FIG. 4 is a functional block diagram of a video decoder (310) according to an embodiment.
[0020] The receiver (410) can receive one or more codec video sequences decoded by the decoder (310), and in the same or other embodiments, receives one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. It can receive the encoded video sequence from the channel (412), and the channel may be a hardware / software link to a storage device for storing the encoded video data. The receiver (410) can receive the encoded video data and other data, for example, the encoded audio data and / or the auxiliary data stream that can be transferred to their respective usage entities (not shown). The receiver (410) can separate the encoded video sequence from other data. To handle network jitter, a buffer memory (415) can be connected between the receiver (410) and the entropy decoder / parser (hereinafter referred to as "parser"). The receiver (410) may not require the buffer memory (415) when receiving data from a storage / transfer device with sufficient bandwidth and controllability, or an isochronous real-time network, or the buffer (615) may be small. To make the best use of packet networks such as the Internet, the buffer memory (415) may be required, and the buffer memory may be relatively large and advantageously have an adaptable size.
[0021] The video decoder (310) may include a parser (420) to reconstruct symbols (421) based on an entropy - encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (310) and potential information for controlling a rendering device such as a display (312). As shown in FIG. 4, although the rendering device is not a component of the decoder, it can be connected to the decoder. The control information for the rendering device may be in the form of Supplemental Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (420) performs parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can be based on a video encoding technology or standard and can follow principles well - known to those skilled in the art, including variable - length coding, Huffman coding, and arithmetic coding with or without context dependence. The parser extracts a sub - group parameter set for at least one sub - group of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The groups include Groups of Picture (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The entropy decoder / parser may further extract information such as transform coefficients, quantizer parameter (QP) values, motion vectors, etc. from the encoded video sequence.
[0022] The parser (420) can create a symbol (421) by performing an entropy decoding / analysis operation on the video sequence received from the buffer memory (415). The parser may receive the encoded data and selectively decode a specific symbol (421). Also, the parser can determine whether a specific symbol (421) is provided to the motion compensation prediction unit (453), the scaler / inverse transform unit (451), the intra prediction unit (452), or the loop filter (454).
[0023] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbol (421) can involve multiple different units. The units involved and the way of involvement are controlled by the subgroup control information parsed by the parser (420) from the encoded video sequence. For the sake of brevity, the flow of such subgroup control information between the parser (420) and the following multiple units is not described.
[0024] In addition to the function blocks already mentioned, the decoder (310) can conceptually be subdivided into multiple functional units described below. In the actual implementation mode executed under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of explaining the disclosed theme, it is appropriate to conceptually subdivide into the following functional units.
[0025] The first unit is the scaler / inverse transform unit (451). The scaler / inverse transform unit (451) receives from the parser (420) the quantized transform coefficients and control information as a symbol (421), including the transform method to be used, the block size, the quantization factor, the quantization scaling matrix, etc. It can output a block including sample values that can be input to the aggregator (455).
[0026] In some cases, the output samples of the scaler / inverse transform (451) can belong to an intra-coding block, i.e., a block that can use prediction information from a previously reconstructed part of the current picture instead of using prediction information from a previously reconstructed picture. Such prediction information can be provided by the intra prediction unit (452). In some cases, the intra prediction unit (452) uses the information around the currently (partially reconstructed) picture (456) that has already been reconstructed to generate a block of the same size and shape as the block being reconstructed. In some cases, the aggregator (455) adds the prediction information generated by the intra prediction unit (452) to the output sample information provided by the scaler / inverse transform unit (451) based on each sample.
[0027] In other cases, the output samples of the scaler / inverse transform unit (451) can belong to an inter-coded and potentially motion-compensated block. In such cases, the motion compensation prediction unit (453) can access the reference picture buffer (457) to obtain samples for prediction. Based on the symbols (421) belonging to the block, after performing motion compensation on the obtained samples, these samples can be added by the aggregator (455) to the output of the scaler / inverse transform unit (in this case, called the residual samples or residual signal) to generate output sample information. The address in the reference picture memory where the motion compensation prediction unit extracts the prediction samples may be controlled by the motion vector, and the motion vector can be in the form of symbols (421) and used by the motion compensation prediction unit. The symbols (421) may have, for example, X, Y, and reference picture components. Motion compensation may further include interpolation of the sample values obtained from the reference picture memory when an accurate motion vector of sub-samples is used, a motion vector prediction mechanism, etc.
[0028] The output samples of the aggregator (455) may be processed in the loop filter unit (454) by various loop filtering techniques. The video compression technology may include in-loop filter technology, and the in-loop filter technology is controlled by parameters included in the encoded video bitstream, and the parameters can be applied to the loop filter unit (454) as symbols (421) from the parser (420). However, the video compression technology may further respond to meta-information obtained during the period of decoding the previous part (in the decoding order) of the encoded picture or the encoded video sequence, or may respond to previously constructed loop filter processed sample values.
[0029] The output of the loop filter unit (454) may be a sample stream, and the sample stream may be output to the rendering device (312) and stored in the reference picture buffer (456) for use in future inter-picture prediction.
[0030] An encoded picture is used as a reference picture for future prediction when it is completely reconstructed. When the encoded picture is completely reconstructed and the encoded picture is recognized as a reference picture (e.g., by the parser (420)), the current reference picture (456) becomes part of the reference picture buffer (457), and a new current picture memory can be reallocated before starting the reconstruction of subsequent encoded pictures.
[0031] The video decoder (310) may perform a decoding operation based on a predetermined video compression technique recorded in, for example, the ITU-T H.265 proposal standard. The encoded video sequence may be used in accordance with the syntax of the video compression technique or standard clearly specified in its profile in, for example, a video compression technique document or standard. In terms of compliance, it is also required that the complexity of the encoded video sequence be within the range limited by the level of the video compression technique or standard. In some cases, the level limits the size of the maximum picture, the maximum frame rate, the maximum reconstructed sample rate (measured, for example, in megasamples per second), the maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the specifications of the Hypothetical Reference Decoder (HRD) and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0032] In an embodiment, the receiver (410) can receive additional (redundant) data together with the encoded video. The additional data is included as part of the encoded video sequence. The additional data is utilized by the video decoder (310) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.
[0033] FIG. 5 may be a functional block diagram of a video encoder (303) according to an embodiment.
[0034] The encoder (303) can receive video samples from a video source (301) (which is not part of the encoder), and the video source can capture the video image to be encoded by the encoder (303).
[0035] The video source (301) may provide a source video sequence in the form of a digital video sample stream, which is encoded by an encoder (303). The digital video sample stream may include any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling configuration (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device for storing previously prepared videos. In a video conferencing system, the video source (301) may be a photographing device for capturing local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. Each picture itself may be organized as a spatial pixel array, and depending on the sampling configuration, color space, etc. used, each pixel may include one or more samples. The relationship between pixels and samples is easily understandable to those skilled in the art. The following description focuses on samples.
[0036] According to an embodiment, the encoder (303) encodes the pictures of the source video sequence in real time or with any other time constraints required by the application, and is compressed as an encoded video sequence (543). Executing at an appropriate encoding speed is one of the functions of the controller (550). The controller controls the other functional units described below and is functionally coupled to these functional units. For the sake of brevity, the coupling is not shown. The parameters set by the controller may include rate control related parameters (such as picture skip, quantizer, λ value of rate distortion optimization technology, etc.), the size of the picture, the arrangement of the group of pictures (GOP), the search range of the maximum motion vector, etc. Those skilled in the art can easily recognize the other functions of the controller (550) as being related to a video encoder (303) optimized for a specific system design.
[0037] Some video encoders operate in an "encoding loop" that is readily understandable to those skilled in the art. As a very simplified explanation, the encoding loop may include an encoding portion of an encoder (530) (hereinafter referred to as the "source encoder") that is responsible for creating symbols based on the input picture and the reference picture to be encoded, and a (local) decoder (533) embedded in the encoder (303). The decoder (533) reconstructs the symbols to create the sample data that the (remote) decoder would also attempt to create (because in the video compression techniques contemplated by the theme of the present disclosure, any compression between the symbols and the encoded video bitstream is reversible). The reconstructed sample stream is input into the reference picture memory (534). Since the decoding of the symbol stream results in bits that are accurate regardless of the decoder location (local or remote), the contents of the reference picture buffer are bit-accurate between the local encoder and the remote encoder. That is, the reference picture samples "seen" from the prediction portion of the encoder are exactly the same as the sample values "seen" when the decoder utilizes the prediction during decoding. Such a basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, for example due to channel errors) is well known to those skilled in the art.
[0038] The operation of the "local" decoder (533) may be the same as the operation of the "remote" decoder (310) described in detail above with reference to FIG. 4. However, referring briefly to FIG. 4, if the symbols are available and the entropy encoder (545) and the parser (420) can encode / decode the symbols into the encoded video sequence losslessly, it is not necessary for the decoder (533) to fully implement the entropy decoding portion of the decoder (310) that includes the channel (412), the receiver (410), the buffer memory (415), and the parser.
[0039] In this case, it can be observed that any decoder technology other than parsing / entropy decoding existing in the decoder will necessarily exist in the corresponding encoder in basically the same functional form. Since the encoder technology and the fully described decoder technology are inverses of each other, the description of the encoder technology can be simplified. A more detailed description is only necessary for some parts and is provided below.
[0040] As part of the operation of the source encoder (530), the source encoder (530) can perform motion-compensated predictive coding, which performs predictive coding on an input frame by referring to one or more previously encoded frames designated as "reference frames" from a video sequence. In this way, the encoding engine (532) may encode the difference between a pixel block of the input frame and a pixel block that can be selected as a reference frame for predictive reference of the input frame.
[0041] The decoder (533) decodes previously encoded video data of a frame designated as a reference frame based on the code created by the source encoder (530). The operation of the encoding engine (532) may advantageously be an irreversible process. When the encoded video data is decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may generally be a replica of the source video sequence with some errors. The decoder (533) can copy the decoding process that can be performed on the reference frame by the video decoder and store the reconstructed reference frame in the reference picture cache (534). In this way, the video encoder (303) can locally store replicas of the reconstructed reference frames, and these replicas have common content with the reconstructed reference frames obtained by the remote video decoder (in the absence of transmission errors).
[0042] The predictor (535) performs a prediction search for the encoding engine (532). That is, for a new frame to be encoded, the predictor (535) may search the reference picture memory (534) for sample data (as candidate reference pixel blocks) that can be used as appropriate prediction references for the new picture, or specific metadata such as motion vectors and block shapes of the reference pictures. The predictor (535) can find appropriate prediction references by operating on a per-pixel-block basis based on the sample blocks. In some cases, the input picture may have prediction references obtained from a plurality of reference pictures stored in the reference picture memory (534), as determined from the search results obtained by the predictor (535).
[0043] The controller (550) can manage the encoding operations of the source coder (530), including setting parameters and subgroup parameters for encoding video data, for example.
[0044] In the entropy coder (545), entropy encoding may be performed on the outputs of all the above functional units. The entropy coder performs lossless compression on the symbols generated by the various functional units based on techniques known to those skilled in the art (such as Huffman coding, variable-length coding, arithmetic coding, etc.) to convert these symbols into an encoded video sequence.
[0045] The transmitter (540) can buffer the encoded video sequence created by the entropy coder (545) to prepare for transmission via the communication channel (560), and the communication channel may be a hardware / software link to a storage device for storing the encoded video data. The transmitter (540) can merge the encoded video data from the source coder (530) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (the source is not shown).
[0046] The controller (550) can manage the operation of the coder (303). During encoding, the controller (550) may assign a specific encoded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, generally, a picture is assigned as one of the following frame types.
[0047] An intra picture (I picture) may be a picture that is encoded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, for example, including Independent Decoder Refresh pictures. Those skilled in the art are aware of those variations of I pictures and their corresponding uses and characteristics.
[0048] A predicted picture (P picture) may be a picture that is encoded and decoded using intra prediction or inter prediction to predict the sample values of each block using at most one motion vector and a reference index.
[0049] A bi - directionally predictive picture (B picture) may be a picture that is encoded and decoded using intra prediction or inter prediction to predict the sample values of each block using at most two motion vectors and reference indices. Similarly, multiple predicted pictures can use more than two reference images and associated metadata for the reconstruction of a single block.
[0050] The source picture is generally spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and may be encoded block by block. The blocks may be predictedly encoded by referring to other (already encoded) blocks determined by the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture may be encoded non-predictively or may be predictedly encoded by referring to the encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be encoded non-predictively via spatial prediction or temporal prediction by referring to one previously encoded reference picture. Blocks of a B picture may be encoded non-predictively via spatial prediction or temporal prediction by referring to one or two previously encoded reference pictures non-predictively.
[0051] The video encoder (303) can perform an encoding operation, for example, based on a predetermined video encoding technique or standard of the ITU-T H.265 proposal. During the operation of the video encoder (303), the video encoder (303) can perform various compression operations including a predictive encoding operation due to temporal and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard being used.
[0052] In an embodiment, the transmitter (540) can transmit additional data together with the encoded video. The source encoder (530) may include such data as part of the encoded video sequence. The additional data may include a temporal / spatial / SNR enhancement layer, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, video user information (VUI) parameter set segments, and the like.
[0053] In the prior art, in order to encode a flag pred_mode_flag that indicates whether a block is intra-coded or inter-coded, only one context is used instead of the value of the flag applied to adjacent blocks. Also, when an adjacent block is coded by an intra-inter prediction mode, the combination of the intra prediction mode and the inter prediction mode is used to predict the adjacent block, and thus, it may be more effective to consider whether to code the adjacent block by the intra-inter prediction mode in the context design for signaling the flag pred_mode_flag.
[0054] The embodiments described in this specification may be used alone or in combination in any order. The following is that the flag pred_mode_flag indicates whether the current block is intra-coded or inter-coded.
[0055] FIG. 6 is a diagram of a current block and adjacent blocks of the current block according to an embodiment.
[0056] Referring to FIG. 6, a current block (610) and a top adjacent block (620) and a left adjacent block (630) of the current block (610) are shown. The width of each of the top adjacent block (620) and the left adjacent block (630) is 4, and the height is 4.
[0057] In an embodiment, information on whether adjacent blocks (e.g., the top adjacent block (620) and the left adjacent block (630)) are encoded in any of the intra prediction mode, the inter prediction mode, or the intra - inter prediction mode is used to obtain a context value for entropy - encoding the flag pred_mode_flag of the current block (e.g., the current block (610)). Specifically, when an adjacent block is encoded in the intra - inter prediction mode, the associated intra prediction mode is applied to the intra - mode encoding and / or MPM derivation of the current block. However, when deriving a context value for entropy - encoding the flag pred_mode_flag of the current block, the adjacent block is regarded as an inter - coding block even though intra prediction is utilized for the adjacent block.
[0058] In one example, the associated intra prediction mode of the intra - inter prediction mode is always the planar mode.
[0059] In other examples, the associated intra prediction mode of the intra - inter prediction mode is always the DC mode.
[0060] In still other examples, the associated intra prediction mode aligns with the intra prediction mode applied in the intra - inter prediction mode.
[0061] In an embodiment, when encoding adjacent blocks (e.g., the top adjacent block (620) and the left adjacent block (630)) in the intra - inter prediction mode, the associated intra prediction mode is applied to the intra - mode encoding and / or MPM derivation of the current block (e.g., the current block (610)). When deriving a context value for entropy - encoding the flag pred_mode_flag of the current block, the adjacent block is also regarded as an intra - coding block.
[0062] In one example, the associated intra prediction mode of the inter-intra prediction mode is always the planar mode.
[0063] In other examples, the associated intra prediction mode of the inter-intra prediction mode is always the DC mode.
[0064] In still other examples, the associated intra prediction mode aligns with the intra prediction mode applied in the inter-intra prediction mode.
[0065] In one embodiment, when adjacent blocks are respectively encoded by the intra prediction mode, the inter prediction mode, and the inter-intra prediction mode, the context index or value is incremented by 2, 0, and 1 respectively.
[0066] In other embodiments, when adjacent blocks are respectively encoded by the intra prediction mode, the inter prediction mode, and the inter-intra prediction mode, the context index or value is incremented by 1, 0, and 0.5 respectively, and the final context index is rounded to the nearest integer.
[0067] After incrementing the context index or value for all adjacent blocks of the current block to determine the final context index, the determined final context index is divided by the number of adjacent blocks and rounded to the nearest integer to determine the average context index. Based on the determined average context index, the flag pred_mode_flag is set to indicate whether the current block is intra-encoded or inter-encoded, and arithmetic coding is performed to encode the pred_mode_flag of the current block.
[0068] In an embodiment, information indicating whether the current block (e.g., the current block (610)) is encoded in an intra prediction mode, an inter prediction mode, or an inter-intra prediction mode is used to obtain one or more context values for entropy encoding the CBF of the current block.
[0069] In one embodiment, three separate contexts (e.g., variables) are used to entropy encode the CBF. One context is used when the current block is encoded in an intra prediction mode, one context is used when the current block is encoded in an inter prediction mode, and one context is used when the current block is encoded in an intra-inter prediction mode. The three separate contexts may be applied only to encode the luminance CBF, only to encode the chrominance CBF, or to encode both the luminance CBF and the chrominance CBF.
[0070] In other embodiments, two separate contexts (e.g., variables) are used to entropy encode the CBF. One context is used when the current block is encoded in an intra prediction mode, and one context is used when the current block is encoded in an inter prediction mode or an intra-inter prediction mode. The two separate contexts may be applied only to encode the luminance CBF, only to encode the chrominance CBF, or to encode both the luminance CBF and the chrominance CBF.
[0071] In yet other embodiments, two separate contexts (e.g., variables) are used to entropy encode the CBF, one context is used when the current block is encoded by an intra-prediction mode or an intra-inter prediction mode, and one context is used when the current block is encoded by an inter-prediction mode. The two separate contexts may be applied only to entropy encode the luma CBF, only to entropy encode the chroma CBF, or only to entropy encode both the luma and chroma CBFs.
[0072] FIG. 7 is a flowchart showing a method (700) for controlling intra-inter prediction for decoding or encoding a video sequence according to an embodiment. In some implementations, one or more processing blocks of FIG. 7 may be performed by a decoder (310). In some implementations, one or more processing blocks of FIG. 7 may be performed by another device separate from the decoder (310), or another device or group of devices (e.g., an encoder (303)) including the decoder (310).
[0073] Referring to FIG. 7, in a first block (710), the method (700) includes determining whether an adjacent block of the current block is encoded by an intra-inter prediction mode. Based on a determination that the adjacent block is not encoded by an intra-inter prediction mode (710 - NO), the method (700) ends.
[0074] Based on a determination that the adjacent block is encoded by an intra-inter prediction mode (710 - YES), in a second block (720), the method (700) includes performing intra-mode encoding of the current block using an intra-prediction mode associated with the intra-inter prediction mode.
[0075] In a third block (730), the method (700) includes setting a prediction mode flag associated with the adjacent block.
[0076] In the fourth block (740), the method (700) includes obtaining a context value based on a set prediction mode flag associated with an adjacent block.
[0077] In the fifth block (750), the method (700) includes performing entropy coding of a prediction mode flag associated with a current block, which indicates that the current block is intra-coded, using the obtained context value.
[0078] The method (700) further includes performing derivation of the MPM of the current block using an intra prediction mode associated with an intra-inter prediction mode, based on a determination (710 - YES) that an adjacent block is coded by an intra-inter prediction mode.
[0079] The intra prediction mode associated with the intra-inter prediction mode may be a planar mode, a DC mode, or an intra prediction mode applied in the intra-inter prediction mode.
[0080] Setting the prediction mode flag associated with an adjacent block may include setting the prediction mode flag associated with the adjacent block to indicate that the adjacent block is intra-coded.
[0081] Setting the prediction mode flag associated with an adjacent block may include setting the prediction mode flag associated with the adjacent block to indicate that the adjacent block is inter-coded.
[0082] The method (700) further includes determining whether an adjacent block is encoded by an intra prediction mode, an inter prediction mode, or an intra-inter prediction mode, and incrementing the context index of the prediction mode flag associated with the current block by 2 based on a determination that the adjacent block is encoded by the intra prediction mode, incrementing the context index by 0 based on a determination that the adjacent block is encoded by the inter prediction mode, and incrementing the context index by 1 based on a determination that the adjacent block is encoded by the intra-inter prediction mode, determining an average context index based on the incremented context index and the number of adjacent blocks of the current block, and setting the prediction mode flag associated with the current block based on the determined average context index.
[0083] The method may further include determining whether an adjacent block is encoded by an intra prediction mode, an inter prediction mode, or an intra-inter prediction mode, incrementing the context index of the prediction mode flag associated with the current block by 1 based on a determination that the adjacent block is encoded by the intra prediction mode, incrementing the context index by 0 based on a determination that the adjacent block is encoded by the inter prediction mode, and incrementing the context index by 0.5 based on a determination that the adjacent block is encoded by the intra-inter prediction mode, determining an average context index based on the incremented context index and the number of adjacent blocks of the current block, and setting the prediction mode flag associated with the current block based on the determined average context index.
[0084] FIG. 7 shows a block example of method (700). However, in some implementation manners, method (700) may include additional blocks, fewer blocks, different blocks, or blocks in a different arrangement than these blocks depicted in FIG. 7. Additionally or alternatively, two or more of the blocks of method (700) may be executed in parallel.
[0085] Also, the proposed method may be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, one or more processors execute a program for executing one or more of the proposed methods stored in a non-transitory computer-readable medium.
[0086] FIG. 8 is a simplified block diagram of an apparatus (800) for controlling in-line inter prediction for decoding or encoding a video sequence according to an embodiment.
[0087] Referring to FIG. 8, apparatus (800) includes a first determination code (810), an execution code (820), and a setting code (830). Apparatus (800) may further include an increment code (840) and a second determination code (850).
[0088] The first determination code (810) is arranged to cause at least one processor to determine whether adjacent blocks of the current block are encoded in an in-line inter prediction mode.
[0089] The execution code (820) is arranged to cause at least one processor to execute intra-mode encoding of the current block using an intra prediction mode associated with the in-line inter prediction mode based on a determination that the adjacent blocks are encoded in the in-line inter prediction mode.
[0090] The setting code (830) is arranged to cause at least one processor to operate as follows based on a determination that an adjacent block is encoded in an intra-inter prediction mode, that is, to set a prediction mode flag associated with the adjacent block, to obtain a context value based on the set prediction mode flag associated with the adjacent block, to perform entropy encoding of the prediction mode flag associated with the current block using the obtained context value, and the prediction mode flag indicates that the current block is intra-encoded.
[0091] The execution code (820) may further be arranged to cause at least one processor to derive the most probable mode (MPM) of the current block using an intra prediction mode associated with the intra-inter prediction mode based on a determination that an adjacent block is encoded in the intra-inter prediction mode.
[0092] The intra prediction mode associated with the intra-inter prediction mode may be a planar mode, a DC mode, or an intra prediction mode applied in the intra-inter prediction mode.
[0093] The setting code (830) may further be arranged to cause at least one processor to set a prediction mode flag associated with an adjacent block so as to indicate that the adjacent block is intra-encoded.
[0094] The setting code (830) may further be arranged to cause at least one processor to set a prediction mode flag associated with an adjacent block so as to indicate that the adjacent block is inter-encoded.
[0095] The first decision code (810) may further be arranged to cause at least one processor to determine whether an adjacent block is encoded by any of an intra prediction mode, an inter prediction mode, or an intra-inter prediction mode. The increment code (840) may be arranged to cause at least one processor to increment the context index of a prediction mode flag associated with the current block by 2 based on a determination that an adjacent block is encoded by the intra prediction mode, to increment the context index by 0 based on a determination that the adjacent block is encoded by the inter prediction mode, and to increment the context index by 1 based on a determination that the adjacent block is encoded by the intra-inter prediction mode. The second decision code (850) may further be arranged to cause at least one processor to determine an average context index based on the incremented context index and the number of adjacent blocks of the current block. The setting code (830) may further be arranged to cause at least one processor to set a prediction mode flag associated with the current block based on the determined average context index.
[0096] The first decision code (810) may further be arranged to cause at least one processor to determine whether an adjacent block is encoded by an intra prediction mode, an inter prediction mode, or an intra - inter prediction mode. The increment code (840) causes at least one processor to increment the context index of a prediction mode flag associated with the current block by 1 based on a determination that an adjacent block is encoded by an intra prediction mode, increment the context index by 0 based on a determination that an adjacent block is encoded by an inter prediction mode, and increment the context index by 0.5 based on a determination that an adjacent block is encoded by an intra - inter prediction mode. The second decision code (850) may be arranged to cause at least one processor to determine an average context index based on the incremented context index and the number of adjacent blocks of the current block. The setting code (830) may further be arranged to cause at least one processor to set a prediction mode flag associated with the current block based on the determined average context index.
[0097] The above - described technology may be implemented as computer software using computer - readable instructions and may be physically stored on one or more computer - readable media.
[0098] FIG. 9 is a diagram of a computer system (900) suitable for implementing an embodiment.
[0099] Computer software can be encoded in any suitable machine code or computer language, and for the machine code or computer language, by executing mechanisms such as assembly, compilation, and linking, code containing instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed by interpretation, microcode, etc., can be created.
[0100] Instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0101] The components for the computer system (900) shown in FIG. 9 are essentially exemplary and are not intended to limit the scope of use or functionality of the computer software for implementing each embodiment. The arrangement of the components should not be construed as having any dependencies or requirements related to any of the components shown in the exemplary embodiments of the computer system (900), or combinations thereof.
[0102] The computer system (900) may include several human interface input devices. Such human interface input devices can respond to input by one or more human users, such as tactile input (e.g., keystrokes, slides, data glove movements), audio input (e.g., voice, hand clapping sounds), visual input (e.g., gestures), olfactory input (not shown), etc. The human interface devices can also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, environmental sounds), images (e.g., scanned images, photographic images obtained from a still image capture device), videos (e.g., 2D videos, 3D videos including stereoscopic videos), etc.
[0103] The human interface input device may include one or more of a keyboard (901), a mouse (902), a touch pad (903), a touch panel (910), a data glove (904), a joystick (905), a microphone (906), a scanner (907), and a photographing device (908) (each is shown only once).
[0104] The computer system (900) may further include several human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, via tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (for example, there are tactile feedback devices by a touch panel (910), a data glove (904), or a joystick (905), and there are also tactile feedback devices not used as input devices), audio output devices (for example, speakers (909), headphones (not shown)), visual output devices (for example, a screen (910), including a cathode ray tube (CRT) screen, a liquid crystal display (LCD) screen, a plasma screen, an organic light emitting diode (OLED) screen, each may or may not have touch screen input ability and tactile feedback ability, and some of the screens may be capable of outputting two-dimensional visual output or output of three dimensions or more by means such as stereoscopic graphics output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown).
[0105] The computer system (900) may further include a memory device and its associated media accessible by a human, for example, an optical medium including a CD / DVD ROM / RW (920) having a medium (921) such as a CD / DVD, a thumb drive (922), a removable hard drive or solid state drive (923), conventional magnetic media such as magnetic tapes and floppy disks (not shown), devices based on dedicated ROM / ASIC / PLD such as a security dongle (not shown), and the like.
[0106] Also, one of ordinary skill in the art should understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.
[0107] The computer system (900) may further include an interface to one or more communication networks. The network may be, for example, wireless, wired, optical, etc. The network may further be a local area, wide area, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks (including Global System for Mobile Communications (GSM), Third Generation (3G), Fourth Generation (4G), Fifth Generation (5G), Long Term Evolution (LTE), etc.), television wired or wireless wide area digital networks (including cable television, satellite television, and terrestrial television), vehicle and industrial (including CANBus), etc. Some networks generally require an external network interface adapter connected to a specific general-purpose data port or peripheral bus (949) (e.g., the Universal Serial Bus (USB) port of the computer system (900)), and other networks are generally integrated into the core of the computer system (900) by being connected to the system bus described below (e.g., integrated into the Ethernet interface of a PC computer system or the cellular network interface of a smartphone computer system). Using any of these networks, the computer system (900) can communicate with other entities. Such communication may be one-way reception only (e.g., broadcast television), one-way transmission only (e.g., CANbus to a certain CANbus device), or two-way (e.g., reaching another computer system via a local area or wide area digital network). Specific protocols and protocol stacks can be used for each of these networks and network interfaces as described above.
[0108] The above human interface devices, storage devices accessible by humans, and network interfaces may be connected to the core (940) of the computer system (900).
[0109] The core (940) includes one or more central processing units (CPUs) (941), a graphics processing unit (GPU) (942), a specialized programmable processing unit in the form of a field-programmable gate array (FPGA) (943), hardware accelerators (944) for some tasks, etc. These devices are connected via a system bus (948) together with a read-only memory (ROM) (945), a random access memory (RAM) (946), and an internal mass storage device (e.g., an internal hard disk drive that cannot be accessed by users, a solid-state drive (SSD), etc.) (947). In some computer systems, expansion by additional CPUs, GPUs, etc. can be enabled by accessing the system bus (948) in the form of one or more physical plugs. Peripheral devices can be connected directly or via a peripheral bus (949) to the system bus (948) of the core. The architecture of the peripheral bus includes peripheral component interconnect (PCI), USB, etc.
[0110] The CPU (941), GPU (942), FPGA (943), and accelerator (944) can execute some instructions, and by combining these instructions, the above computer code can be configured. The computer code may be stored in the ROM (945) or RAM (946). Temporary data is also stored in the RAM (946), and permanent data may be stored, for example, in the internal mass storage device (947). High-speed storage and retrieval to any of the storage devices can be achieved by a cache memory, and the cache memory can be closely related to one or more CPUs (941), GPUs (942), mass storage devices (947), ROM (945), RAM (946), etc.
[0111] A computer-readable medium can have computer code thereon for performing various operations implemented by a computer. The medium and the computer code may be media and computer code specially designed and constructed for the purposes of the embodiments, or may be of the type well-known and available to those of ordinary skill in the computer software arts.
[0112] By way of example and not limitation, a computer system (900) having an architecture, and in particular a core (940), can provide functionality by a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software implemented on one or more tangible computer-readable media. Such computer-readable media can be media related to the mass storage devices accessible to the user as introduced above, and some storage devices of the core (940) having a non-transitory nature such as the on-core mass storage device (947) or ROM (945). Software for implementing various embodiments is stored in such devices and executed by the core (940). Depending on specific needs, the computer-readable media may include one or more memory devices or chips. The software causes the core (940), and in particular the processor (including a CPU, GPU, FPGA, etc.) therein, to execute a specific process or a specific part of a specific process that includes limiting the data configurations stored in the RAM (946) as described herein and modifying these data configurations based on the processes defined by the software. Further or alternatively, the computer system can provide functionality by logic implemented in a circuit (e.g., an accelerator (944)) in a hardwired or other manner, and the logic can execute the specific process or a specific part of a specific process described herein by operating in place of the software or together with the software. Where appropriate, references to software may include logic, and conversely, references to logic may include software. Where appropriate, references to the computer-readable media may include a circuit (e.g., an integrated circuit (IC)) in which the software for execution is stored, a circuit embodying the logic for execution, or both. Embodiments include any suitable combination of hardware and software.
[0113] Although several exemplary embodiments have already been described in this disclosure, there are changes, replacements, and various alternative equivalents that are included within the scope of this disclosure. Accordingly, it should be understood by those skilled in the art that, although not explicitly shown or described herein, many systems and methods that embody the principles of this disclosure can be devised and are within its spirit and scope.
Claims
1. 1. A method for encoding a video bitstream executed by at least one processor, comprising: setting a first prediction mode flag pred_mode_flag associated with an upper or left neighboring block of a current block to indicate whether the neighboring block is in inter or intra prediction mode; encoding the first prediction mode flag pred_mode_flag into a video bitstream; encoding a decision flag into the video bitstream, the decision flag being used to finally decide whether the neighboring block above or to the left of the current block is coded in an inter prediction mode or an intra-inter prediction mode; deriving a context value for a second prediction mode flag pred_mode_flag associated with the current block based on the first prediction mode flag pred_mode_flag associated with the neighboring block; encoding the second prediction mode flag pred_mode_flag associated with the current block into the video bitstream using the derived context value; A method comprising:
2. performing intra-predictive coding of the current block using an intra-prediction mode associated with the intra-inter prediction mode based on a final determination that the neighboring block is coded using the intra-inter prediction mode; The method of claim 1 further comprising:
3. The method of claim 1 or 2, further comprising: based on a final determination that the neighboring block is coded by the intra-inter prediction mode, performing a derivation of a most probable mode (MPM) of the current block using an intra prediction mode associated with the intra-inter prediction mode.
4. The method according to claim 2 or 3, wherein an intra prediction mode associated with the intra-inter prediction mode is a planar mode.
5. The method according to claim 2 or 3, wherein an intra prediction mode associated with the intra-inter prediction mode is a direct current DC mode.
6. The method according to claim 2 or 3, wherein an intra prediction mode associated with the intra-inter prediction mode is an intra prediction mode applied in the intra-inter prediction mode.
7. The step of deriving the context value comprises: incrementing the context value by 2 based on determining that the neighboring block is coded using an intra prediction mode; incrementing the context value by 0 based on determining that the neighboring block is coded using an inter prediction mode; incrementing the context value by one based on determining that the neighboring block is coded using an intra-inter prediction mode; The method further comprises: determining an average context index based on the incremented context value and a number of neighboring blocks of the current block; setting the second prediction mode flag based on the determined average context index; The method of claim 1 , further comprising:
8. The step of deriving the context value related to the second prediction mode flag comprises: incrementing the context value by one based on determining that the neighboring block is coded using an intra prediction mode; incrementing the context value by 0 based on determining that the neighboring block is coded using an inter prediction mode; incrementing the context value by 0.5 based on determining that the neighboring block is coded using an intra-inter prediction mode; The method further comprises: determining an average context index based on the incremented context value and a number of neighboring blocks of the current block; setting the second prediction mode flag based on the determined average context index; The method of claim 1 , further comprising:
9. 9. A method according to claim 1, wherein when the decision flag is true, the prediction mode actually used for the neighboring block is the intra-inter prediction mode, which is different from the inter prediction mode indicated by the first prediction mode flag pred_mode_flag.
10. at least one memory storing a program; at least one processor coupled to said memory; Including, The program is configured to cause the at least one processor to execute a method according to any one of claims 1 to 9. Video bitstream coding device.
11. A program for causing a computer to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method and apparatus of video coding
US20170251213A1
Method for processing image based on joint inter-intra prediction mode and apparatus therefor
US20180249156A1
Video encoding method and apparatus, and video decoding method and apparatus
US20180288410A1