Video coding and decoding method and device and storage medium
By combining intra-inter prediction modes and combining geometric segmentation and inter-prediction of affine motion, the problem of insufficient compression efficiency in the prior art is solved, and more efficient video encoding and decoding and more accurate motion vector estimation are achieved.
Patent Information
- Application Number
- CN202510034989.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-26
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-29
AI Technical Summary
The existing video encoding technology is insufficient in compression when dealing with stationary and fast moving scenes, making it difficult to adapt to multiple scenario needs, and the inter-frame prediction mode limits the accuracy of the motion vector.
The combined intra-inter prediction (CIIP) mode is used, and the inter prediction mode based on geometric segmentation, affine motion and temporal interpolation is combined to improve the accuracy of the motion vector.
It improves the compression efficiency and decoding accuracy of video encoding and decoding, adapts to video content in more scenarios, reduces the amount of data transmission, and maintains video quality.
Smart Images

Figure CN120390082A_ABST
Abstract
Description
Related Applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 625,705, entitled "Combined Intra and Inter Prediction Mode", filed on January 26, 2024, which is hereby incorporated by reference in its entirety. Technical Field
[0002] This application relates to the field of video coding and decoding technologies, and in particular, to a video coding and decoding method, apparatus, and storage medium. Background Art
[0003] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital imaging devices, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise transmit digital video data across communication networks and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding can be used to compress video data according to one or more video coding standards before transmitting or storing the video data. Video coding can be performed by hardware and / or software on an electronic / client device or server providing cloud services.
[0004] Video coding typically uses prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality. A variety of video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard designed to succeed HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verification version 1.0.0 with Specification Errata 1 was released.
[0005] In existing video coding technologies, intra-frame prediction or inter-frame prediction is generally adopted. Intra-frame prediction is a spatial prediction technology that uses the relationship between pixels within a single frame (picture) to predict data, and it is particularly effective for static or slowly moving scenes. Inter-frame prediction is a temporal prediction technology that uses the temporal correlation between different frames (pictures) in a video sequence to predict data, and it can utilize motion information to reduce the amount of data that needs to be transmitted, being particularly effective for fast-moving scenes.
[0006] Therefore, there is a need for a video encoding and decoding method, apparatus, and storage medium that can adapt to more scenarios and improve compression efficiency. Summary of the Invention
[0007] Among other aspects, the present disclosure describes a set of methods for video (image) compression, more specifically relating to a combined intra-frame and inter-frame prediction (CIIP) mode that combines intra-frame prediction and inter-frame prediction to generate a final predicted block. For example, any intra-frame prediction mode can be used to obtain the intra-frame prediction of the CIIP mode. The present disclosure describes techniques in which the inter-frame prediction of the CIIP mode can be an inter-frame prediction mode based on geometric segmentation, an inter-frame prediction mode based on affine / distorted motion, or a temporal interpolation prediction mode. Compared with restricting inter-frame prediction to a translational mode, the advantage of making inter-frame prediction one of these modes is that a more accurate motion vector (MV) can be obtained, which improves video decoding accuracy.
[0008] According to some embodiments, a method for video decoding includes: (i) receiving a video bitstream (e.g., a source video sequence) (e.g., corresponding to one or more pictures) including a plurality of blocks, the plurality of blocks including a first block, wherein the first block is encoded using a combined intra-frame and inter-frame prediction (CIIP) mode; (ii) identifying an intra-frame prediction mode for the first block; (iii) identifying an inter-frame prediction mode for the first block, the identified inter-frame prediction mode being one of: (a) an inter-frame prediction mode based on geometric segmentation, (b) an inter-frame prediction mode based on affine motion, and (c) a temporal interpolation prediction mode; and (iv) decoding the first block using the identified intra-frame prediction mode and the identified inter-frame prediction mode.
[0009] According to some embodiments, a method of video encoding includes: (i) receiving video data (e.g., a source video sequence) (e.g., corresponding to one or more pictures) including a plurality of blocks, the plurality of blocks including a first block, wherein the first block is to be encoded using a combined intra-inter prediction (CIIP) mode; (ii) identifying an intra prediction mode for the first block; (iii) identifying an inter prediction mode for the first block, the identified inter prediction mode being one of: (a) an inter prediction mode based on geometric segmentation, (b) an inter prediction mode based on affine motion, and (c) a temporal interpolation prediction mode; and (iv) encoding the first block using the identified intra prediction mode and the identified inter prediction mode.
[0010] According to some embodiments, a method of processing visual media data includes: (i) obtaining a source video sequence including a plurality of frames; and (ii) performing a conversion between the source video sequence and a video bitstream of visual media data according to format rules, (a) wherein the video bitstream includes a plurality of blocks, the plurality of blocks including a first block, wherein the first block is encoded using a combined intra-inter prediction (CIIP) mode; and (b) wherein the format rules specify that the inter prediction mode of the CIIP mode is one of: an inter prediction mode based on geometric segmentation, an inter prediction mode based on affine motion, and a temporal interpolation prediction mode.
[0011] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic device. The computing system includes control circuitry and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).
[0012] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions for execution by a computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.
[0013] Accordingly, apparatuses and systems that utilize methods for encoding and decoding video are disclosed. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, and systems for video encoding / decoding.
[0014] In this application, a combined intra - inter prediction CIIP mode is proposed. The prediction of a block depends not only on the pixels within the same frame (intra - prediction), but also on the pixels in other frames (inter - prediction). This combined method can adapt to more scenarios, whether static or fast - moving, because it takes into account both spatial and temporal correlations. The CIIP mode can improve the compression efficiency because it can more accurately predict and encode complex content in the video, thereby reducing the amount of data to be transmitted while maintaining the video quality and improving the video decoding accuracy.
[0015] The features and advantages described in the specification are not necessarily all included, and in particular, given the drawings, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. In addition, it should be noted that the language used in this specification is mainly selected for readability and guidance purposes and is not necessarily selected to depict or limit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To understand the present disclosure in more detail, a more specific description can be made by referring to the features of various embodiments, and some of the features of various embodiments are shown in the drawings. However, the drawings only show the relevant features of the present disclosure and are therefore not necessarily considered restrictive, as the specification may allow other valid features that those skilled in the art will understand when reading the present disclosure.
[0017] Figure 1 is a block diagram showing an example communication system according to some embodiments.
[0018] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.
[0019] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.
[0020] Figure 3 is a block diagram showing an example server system according to some embodiments.
[0021] Figure 4A shows an example motion sample for obtaining model parameters of a block using local warped motion prediction according to some embodiments.
[0022] Figure 4B shows the motion vectors in a block using the warped extension mode according to some embodiments.
[0023] Figure 5AShows an example of a segmentation-based prediction mode according to some embodiments.
[0024] Figures 5B to 5C Shows an example segmentation mode mix according to some embodiments.
[0025] Figure 5D Shows an example of deriving sub-block motion vectors according to some embodiments.
[0026] Figure 5E Shows an example of decoder-side motion vector refinement according to some embodiments.
[0027] Figure 6A Shows an example video decoding process according to some embodiments.
[0028] Figure 6B Shows an example video encoding process according to some embodiments.
[0029] By convention, the various features shown in the drawings are not necessarily drawn to scale, and throughout the specification and drawings, like reference numerals may be used to represent like features. Detailed Description
[0030] The present disclosure describes video / image compression techniques including encoding / decoding video blocks using a combined intra-inter prediction (CIIP) mode, in which an intra mode prediction is combined with an inter mode prediction to generate a final prediction block. In some conventional systems, a translational inter prediction mode is used to generate the inter mode prediction of the CIIP mode. In the present disclosure, systems and methods are described in which the inter mode prediction can be obtained from one of a geometric segmentation-based inter prediction mode, an affine motion-based inter prediction mode, and a temporal interpolation prediction mode. Using a non-translational inter prediction mode allows for more accurate motion vectors to be used during the CIIP encoding / decoding process, which improves the accuracy of video coding and decoding. Example Systems and Devices
[0031] Figure 1 Is a block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0032] The source device 102 includes a video source 104 (e.g., a camera device component or a media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera device (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams based on the video stream. The video stream from the video source 104 can be of high data volume compared to the encoded video bitstreams 108 generated by the encoder component 106. Since the encoded video bitstreams 108 are of lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less bandwidth to transmit and less storage space to store. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video to the network 110).
[0033] One or more networks 110 represent any number of networks for transmitting information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 can exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0034] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as the encoded video stream from the source device 102). The server system 112 includes a codec component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec component 114 includes an encoder component and / or a decoder component. In various embodiments, the codec component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the codec component 114 is configured to decode the encoded video bitstreams 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings based on the encoded video bitstreams 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 can be configured to trim the encoded video bitstreams 108 to customize potentially different bitstreams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.
[0035] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an outgoing video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0036] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, one or more of the electronic devices 120 and / or the source device 102 are examples of server systems, personal computers, portable devices (e.g., smart phones, tablet computers, or laptop computers), wearable devices, video conferencing devices, and / or other types of electronic devices.
[0037] In an example operation of the communication system 100, the source device 102 transmits the encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may decode and / or encode the encoded video bitstream 108 using the codec component 114. For example, the server system 112 may apply an encoding that is more optimized for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.
[0038] Figure 2AFIG. 0 is a block diagram showing example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be organized as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. A person of ordinary skill in the art can easily understand the relationship between pixels and samples.
[0039] The encoder component 106 is configured to encode and / or compress the pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, the encoder component 106 is configured to perform a conversion between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ value of rate distortion optimization techniques), picture size, Group Of Pictures (GOP) layout, maximum motion vector search range, etc. A person of ordinary skill in the art can easily identify other functions of the controller 204, as such functions may belong to the encoder component 106 optimized for a specific system design.
[0040] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols such as a symbol stream based on an input picture and reference pictures to be encoded) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data (in the case where the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results regardless of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the same sample values as the sample values that the decoder will interpret when using prediction during decoding as reference picture samples.
[0041] The operation of the decoder 210 can be the same as that of a remote decoder such as the decoder component 122 described in detail below in connection with Figure 2B However, briefly referring to Figure 2B , since the symbols are available and the encoding of the symbols into an encoded video sequence by the entropy encoder 214 and the decoding of the symbols by the parser 254 can be lossless, the entropy decoding part including the buffer memory 252 and the parser 254 of the decoder component 122 may not be fully implemented in the local decoder 210.
[0042] Except for parsing / entropy decoding, the decoder techniques described herein can exist in a corresponding encoder in a form with substantially the same functionality. For this reason, the disclosed subject matter focuses on decoder operations. Additionally, the description of encoder techniques can be simplified because encoder techniques can be inverse to decoder techniques.
[0043] As part of the operation of the source encoder 202, the source encoder 202 can perform motion-compensated predictive coding that predictively encodes an input frame by referring to one or more previously encoded frames designated as reference frames from a video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of the reference frame, and the reference frame can be selected as the prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and sub-group parameters for encoding video data.
[0044] The decoder 210 decodes the encoded video data of a frame that can be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 can advantageously be a lossy process. When the encoded video data is in a video decoder ( Figure 2AWhen decoded at the (not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that can be performed by the remote video decoder on the reference frames, and can cause the reconstructed reference frames to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frames, which has the same content (no transmission errors) as the reconstructed reference frames that will be obtained by the remote video decoder.
[0045] The predictor 206 can perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search in the reference picture memory 208 for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that can be used as an appropriate prediction reference for the new picture. The predictor 206 can operate on a per-pixel block basis of the sample blocks to find an appropriate prediction reference. As determined by the search results obtained by the predictor 206, the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory 208.
[0046] The outputs of all the above-mentioned functional units can be subjected to entropy coding in the entropy encoder 214. The entropy encoder 214 converts the symbols into an encoded video sequence by losslessly compressing the symbols generated by the various functional units according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).
[0047] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer the encoded video sequence created by the entropy encoder 214 in preparation for transmission via the communication channel 218, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter can be configured to merge the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter can transmit additional data along with the encoded video. The source encoder 202 can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.
[0048] The controller 204 can manage the operation of the encoder component 106. During encoding, the controller 204 can assign a certain type of encoded picture to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture can be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of those variants of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. Predictive pictures can be encoded and decoded using inter prediction or intra prediction that uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures can be encoded and decoded using inter prediction or intra prediction that uses at most two motion vectors and a reference index to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0049] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture can be non-predictively encoded, or can be predictively encoded (spatial prediction or intra prediction) with reference to already encoded blocks of the same picture. Pixel blocks of a P picture can be non-predictively encoded with reference to one previously encoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture can be non-predictively encoded with reference to one or two previously encoded reference pictures via spatial prediction or via temporal prediction.
[0050] Video can be captured as a sequence of multiple source pictures (video pictures) in time series. Intra picture prediction (commonly abbreviated as intra prediction) exploits the spatial correlation within a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In an example, a specific picture in encoding / decoding, which is called the current picture, is segmented into blocks. In the case where a block in the current picture is similar to a reference block in a previously encoded and still buffered reference picture in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0051] The encoder component 106 may perform encoding operations according to any of the predetermined video encoding techniques or standards such as those described herein. In operation of the encoder component 106, the encoder component 106 may perform various compression operations, including predictive encoding operations that utilize temporal redundancy and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.
[0052] Figure 2B FIG. 4 is a block diagram illustrating example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in FIG. 4 is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to a loop filter 256 and configured to transmit data (e.g., via a wired connection or a wireless connection) to the display 124.
[0053] In some embodiments, the decoder component 122 includes a receiver coupled to the channel 218 and configured to receive data (e.g., via a wired connection or a wireless connection) from the channel 218. The receiver may be configured to receive one or more encoded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from the channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of the (one or more) encoded video sequences. The additional data may be used by the decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0054] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. The decoder component 122 may be implemented at least in part in software.
[0055] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to counter network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 inside the decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided outside the decoder component 122 (e.g., to counter network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer memory 252 may not be needed, or the buffer memory 252 may be smaller. To make the best use of packet networks such as the Internet, the buffer memory 252 may be needed, the buffer memory 252 may be relatively large and / or have an adaptive size, and may be implemented at least partially in the operating system or a similar element outside the decoder component 122.
[0056] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. The symbols may include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device may be in the form of, for example, a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be according to a video coding technique or standard, and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser 254 may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0057] Depending on the type of the encoded video picture or a part thereof (e.g., inter picture and intra picture, inter block and intra block) and other factors, the reconstruction of the symbol 270 may involve multiple different units. Which units are involved and the manner of involvement may be controlled by the parser 254 through subgroup control information parsed from the encoded video sequence. For the sake of brevity, this subgroup control information flow between the parser 254 and the multiple units below is not depicted.
[0058] The decoder component 122 can conceptually be subdivided into multiple functional units, and in some implementations, these units interact closely with each other and can be at least partially integrated with each other. However, for the sake of brevity, the conceptual subdivision of the functional units is maintained herein.
[0059] The scaler / inverse transform unit 258 receives, from the parser 254, the quantized transform coefficients as symbols 270 and control information (such as which transform to use, block size, quantization factor, and / or quantization scaling matrix). The scaler / inverse transform unit 258 can output a block including sample values, and the sample values can be input into the aggregator 268. In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is: blocks that do not use predictive information from previously reconstructed pictures, but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 can generate a block of the same size and shape as the block being reconstructed using the surrounding reconstructed information obtained from the current (partially reconstructed) picture in the current picture memory 264. The aggregator 268 can add the prediction information generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 based on each sample.
[0060] In other cases, the output samples of the scaler / inverse transform unit 258 belong to inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit 260 can access the reference picture memory 266 to obtain samples for prediction. After motion-compensating the obtained samples according to the symbols 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The address in the reference picture memory 266 from which the motion compensation prediction unit 260 obtains the prediction samples can be controlled by a motion vector. The motion vector can be available to the motion compensation prediction unit 260 in the form of symbols 270, which can have, for example, an X component, a Y component, and a reference picture component. Motion compensation can also include, for example, interpolation of sample values obtained from the reference picture memory 266 when using sub-sampled accurate motion vectors, a motion vector prediction mechanism.
[0061] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and available to loop filter unit 256 as symbols 270 from parser 254, but video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to sample values of previously reconstructed and loop-filtered samples. The output of loop filter unit 256 can be a sample stream that can be output to a rendering device such as display 124, and stored in reference picture memory 266 for use in future inter-picture prediction.
[0062] Once reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture has been reconstructed and the encoded picture has been identified (e.g., by parser 254) as a reference picture, the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0063] Decoder component 122 can perform decoding operations according to a predetermined video compression technique that can be recorded in a standard such as any of the standards described herein. An encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard. In addition, in order to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.
[0064] Figure 3is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes control circuitry 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuitry includes a field programmable gate array, a hardware accelerator, and / or an integrated circuit (e.g., an application specific integrated circuit).
[0065] The network interface 304 may be configured to interface with one or more communication networks (e.g., a wireless network, a wired network, and / or an optical network). The communication network may be local, wide area, metropolitan area, vehicular and industrial, real-time, delay tolerant, etc. Examples of communication networks include: local area networks such as Ethernet, wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus, etc. Such communication may be only one-way receiving (e.g., broadcast TV), only one-way transmitting (e.g., CANBus to certain CANBus devices), or two-way (e.g., to other computer systems using local digital networks or wide area digital networks). Such communication may include communication to one or more cloud computing networks.
[0066] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 may include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera device, etc. The output device 308 may include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays), etc.
[0067] The memory 314 may include high-speed random access memory (e.g., DRAM, SRAM, DDR RAM, and / or other random access solid state memory devices) and / or non-volatile memory (e.g., one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid state storage devices). The memory 314 optionally includes one or more storage devices remote from the control circuitry 302. The memory 314, or alternatively, the non-volatile solid state memory device within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: · An operating system 316, which includes procedures for handling various basic system services and for performing hardware-related tasks; · A network communication module 318, which is used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via a wired connection and / or a wireless connection); · A codec module 320, which is used to perform various functions regarding encoding and / or decoding data such as video data. In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following: ○ A decoding module 322, which is used to perform various functions regarding decoding encoded data, such as those functions previously described regarding the decoder component 122; and ○ An encoding module 340, which is used to perform various functions regarding encoding data, such as those functions previously described regarding the encoder component 106; and · A picture memory 352, such as for storing pictures and picture data for use with the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.
[0068] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described regarding the parser 254), a transformation module 326 (e.g., configured to perform various functions previously described regarding the scaler / inverse transformation unit 258), a prediction module 328 (e.g., configured to perform various functions previously described regarding the motion compensation prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described regarding the loop filter 256).
[0069] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions previously described regarding the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described regarding the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 a subset of the modules shown. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.
[0070] Each of the modules identified above that are stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The modules identified above (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but instead uses the same set of modules to perform two sets of functions. In some embodiments, the memory 314 stores a subset of the modules and data structures identified above. In some embodiments, the memory 314 stores additional modules and data structures not described above.
[0071] Although Figure 3 FIG. 112 shows a server system 112 in accordance with some embodiments, Figure 3 it is intended more as a functional description of the various features that may exist in one or more server systems rather than as a structural schematic of the embodiments described herein. In practice, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in FIG. may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among the servers will vary with the implementation and, optionally, will depend in part on the amount of data traffic processed by the server system during peak usage periods as well as during average usage periods. Example Encoding Techniques
[0072] The encoding processes and techniques described below may be performed at the devices and systems described above (e.g., source device 102, server system 112, and / or electronic device 120). In the following, methods for combining non-translational inter prediction modes and intra prediction modes are described.
[0073] As described in detail below, the combined intra and inter prediction (CIIP) mode combines intra prediction and inter prediction to generate a final prediction block. Any intra prediction mode such as DC prediction, direction / angle intra prediction mode, or an intra prediction mode that models smooth transitions of texture values can be used to perform the intra prediction. The intra prediction and inter prediction can be combined using a wedge mask that divides the block into two parts at various angles. One part can be filled with intra prediction samples and the other part can be filled with inter prediction samples. Alternatively, a predefined set of weights can be used to combine the intra prediction and inter prediction, where the predefined set of weights gradually decreases the intra prediction weight along its prediction direction. In the CIIP mode, the prediction block is derived as a combination of the intra prediction and the inter prediction. The intra prediction block can be derived using an intra prediction that is one of the allowed intra prediction modes, and the inter prediction block can be derived using a single reference inter prediction with translational motion. Depending on how the weights are derived for the intra prediction samples and the inter prediction samples, different composite inter-intra prediction modes can be applied, including a conventional inter-intra prediction mode and a wedge-based inter-intra prediction mode.
[0074] As used herein, the smooth mode refers to an intra prediction mode that models smooth transitions of texture values (e.g., the SMOOTH mode defined in AV1 or the Planar mode defined in HEVC / VVC).
[0075] Motion estimation involves determining a motion vector that describes the transformation from one image (picture) to another image (picture). The reference image (or block) can be from an adjacent frame in a video sequence. The motion vector can be associated with the entire image (global motion estimation) or a specific block. Additionally, the motion vector can correspond to a translational or warping model of approximate motion (e.g., three-dimensional rotation, translation, and scaling). By further dividing the block, motion estimation can be improved in some cases (e.g., for more complex video objects).
[0076] Regarding the inter prediction mode based on affine motion, motion compensation typically assumes a translational motion model between the reference block and the target block. However, warping motion utilizes an affine model. The affine motion model can be represented by Equation 1 below. where [x, y] are the coordinates of the original pixel and [x′, y′] are the warped coordinates of the reference block. According to Equation 1, up to six parameters are required to specify the warping motion: a3 and b3 specify the translational MV; a1 and b2 specify the scaling along the MV; and a2 and b1 specify the rotation.
[0077] In global warping motion compensation, global motion information is signaled for each inter-frame reference frame, which includes the global motion type and a number of motion parameters. After signaling the reference frame index, if global motion is selected, the global motion type and parameters associated with the given reference frame are used for the current coding block.
[0078] In local warping motion compensation, local warping motion for an inter-frame coded block may be allowed when the following conditions are met. First, the current block must use single reference prediction. The width or height of the coded block must be greater than or equal to eight. Finally, at least one of the adjacent neighboring blocks must use the same reference frame as the current block.
[0079] If local warping motion is used for the current block, the affine model parameters are estimated by minimizing the mean square of the difference between the reference projection and the modeled projection based on the MVs of the current block and its adjacent neighboring blocks. To estimate the parameters of local warping motion, if a neighboring block uses the same reference frame as the current block, a pair of projected samples of the central sample in the neighboring block and its corresponding sample in the reference frame is obtained. Subsequently, three additional samples are created by shifting the central position by a quarter sample in one or two dimensions. These additional samples can also be considered as pairs of projected samples to ensure the stability of the model parameter estimation process.
[0080] The MVs of the neighboring blocks used to derive the motion parameters are referred to as motion samples. Motion samples are selected from neighboring blocks that use the same reference frame as the current block. Note that the warping motion prediction mode is only enabled for blocks that use a single reference frame.
[0081] Figure 4A An example of motion samples for deriving the model parameters of a block using local warping motion prediction according to some embodiments is shown. As Figure 4A shown, the MVs of neighboring blocks B0, B1, and B2 are referred to as MV0, MV1, and MV2, respectively. The current block is predicted using unidirectional prediction with reference frame Ref0. For example, neighboring block B0 is predicted using composite prediction with reference frames Ref0 and Ref1; neighboring block B1 is predicted using unidirectional prediction with reference frame Ref0; and neighboring block B2 is predicted using composite prediction with reference frames Ref0 and Ref2. The motion vector MV0 Ref0 of B0, the motion vector MV1 Ref0 of B1, and the motion vector MV2 Ref0 of B2 can be used as motion samples for deriving the affine motion parameters of the current block.
[0082] As mentioned above, two types of warping motion models can be supported: a global warping model and a local warping model. For example, the global warping model is associated with each reference frame, where each of the four non-translation parameters has 12-bit precision, and the translational motion vector is encoded with 15-bit precision. The coded block can choose to use it directly (provided that the reference frame index). The global warping model captures frame-level scaling and rotation. Thus, the global warping model mainly focuses on the rigid motion across the entire frame. A local warping model at the coded block level is also supported. In the local warping mode (also known as WARPED_CAUSAL), the warping parameters of the current block are derived by fitting the model to the nearby motion vectors using least squares.
[0083] In the warping motion mode WARP_EXTEND, the motion of neighboring blocks is smoothly extended into the current block, but to some extent, the warping parameters can be modified. This enables complex warping motions to be represented and propagated across multiple blocks while minimizing block artifacts. To achieve this, the WARP_EXTEND mode applied to the NEWMV block constructs a new warping model based on two constraints: the per-pixel motion vector generated by the new warping model should be continuous with the per-pixel motion vector in the neighboring blocks, and the pixel at the center of the current block should have a per-pixel motion vector that matches the motion vector signaled for the entire block. Figure 4B Shows the motion vectors in a block using the warping extension mode according to some embodiments. As Figure 4B shown, for example, if the neighboring block to the left of the current block is warped, a model fitting the Figure 4B shown motion vectors is used as the warping model.
[0084] The two constraints for constructing the new warping model imply certain equations involving the warping parameters of the neighboring blocks and the current block. Then, these equations can be solved to calculate the warping model of the current block. For example, if (A,..., F) represents the neighboring warping model and (A',..., F') represents the new warping model, the first constraint is as follows, at each point along the common edge: Equation 2 - The first constraint for warping modeling
[0085] Note that the points along the edge have different y values, but they all have the same x value. This means that the coefficients of y must be the same on both sides (e.g., B' = B and D' = D). At the same time, the x coefficients provide equations related to the other coefficients, defined by Equation 3 below: B′ = B D′ = D A′x + E′ = Ax + E C′x + F′ = Cx + F Equation 3 - x coefficient for warping modeling Wherein, in Equation 3, x is the horizontal position of the vertical column of pixels and is thus actually a constant. The second constraint specifies that the motion vector at the center of the block must be equal to the motion vector signaled using the NEWMV mechanism. This provides two additional equations, resulting in a system of six equations with six variables having a unique solution. These equations can be solved efficiently in both software and hardware. The solution can be obtained using basic addition, subtraction, multiplication, and division by powers of 2. Thus, this mode is much simpler than the least squares-based local warping mode.
[0086] Note that there may be multiple neighboring blocks from which to extend. Thus, a method for choosing from which block to extend is needed. This problem is similarly encountered in motion vector prediction. Specifically, there may be several possible motion vectors from nearby blocks, and one should be chosen as the basis for NEWMV coding. The solution to this can be extended to handle the requirements of WARP_EXTEND. This is done by locating the source of each motion vector prediction. Then, WARP_EXTEND is only enabled if the selected motion vector prediction is taken from a directly neighboring block. Then, this block is used as a single "neighboring block" for the remainder of the algorithm.
[0087] Note that sometimes the warping model of the neighbor will be very good without any additional modification. To reduce the coding cost in such cases, WARP_EXTEND can be used for NEARMV blocks. The neighbor selection is the same as NEWMV, except that the neighbor needs to be warped (not just translated via translational motion) for the selection in NEWMV. But if this is true and WARP_EXTEND is selected, then the warping model parameters of the neighbor are copied to the current block.
[0088] In some embodiments, the motion mode WARP_DELTA can be used. In this mode, the warping model of the block is encoded as the increment from a predicted warping model, similar to how the motion vector is encoded as the increment from a predicted motion vector. The prediction can be derived from a global motion model (if any) or a neighboring block.
[0089] To avoid encoding the same prediction distortion model in multiple ways, restrictions can be applied. For example, if the mode is NEARMV or NEWMV, the same neighbor selection logic as described for WARP_EXTEND is used. If this results in a warped neighbor block, the model of that neighbor block (without applying the rest of the WARP_EXTEND logic) is used as the prediction. Otherwise, the global distortion model is used as the basis. Other restrictions can be applied. This example is not intended to limit the scope of the implementation. Then, the increment for each of the non-translation parameters can be encoded. Finally, the translation part of the model is adjusted so that the per-pixel motion vector at the center of the block matches the overall motion vector of the block.
[0090] Since this tool (WARP_DELTA) involves explicitly encoding the increment for each distortion parameter, it uses more bits for encoding than other distortion modes. Therefore, for blocks smaller than 16×16, WARP_DELTA can be disabled. However, the decoding logic is very simple and can thus represent more complex motions that other distortion modes cannot represent.
[0091] In the example single prediction mode (e.g., WARPMV mode), the MV is derived from the warped model of the WRL list. The precision of the derived MV is set to 1 / 8 pixel precision. In the WARPMV mode, the Dynamic Reference List (DRL) is not used, and thus, ref_mv_idx is not signaled.
[0092] The wedge or Geometric Partition Mode (GPM) focuses on the CUs for inter-picture prediction. When the wedge or GPM is applied to a CU, the CU is divided into two parts by a (straight) partitioning boundary. The position of the partitioning boundary can be mathematically defined by an angle parameter and an offset parameter ρ or a predefined lookup table. These parameters can be quantized and combined into a predefined partitioning index lookup table. The wedge / GPM partitioning index of the current CU can be encoded into the bitstream. The two GPM / wedge partitions contain the respective (e.g., separate, distinct, and / or independent) motion information for predicting the corresponding parts in the current CU. In one example, for each part of the GPM / wedge, only unidirectional motion compensation prediction is allowed, such that the memory bandwidth required for motion compensation prediction in the GPM / wedge is equal to the memory bandwidth required for conventional bidirectional motion compensation prediction.
[0093] After splitting, two GPM parts (partitions) contain separate motion information that can be used to predict the corresponding parts in the current block. In some embodiments, for each part of the GPM, only unidirectional motion-compensated prediction (MCP) is allowed, such that the memory bandwidth required for MCP under GPM is equal to the memory bandwidth required for conventional bidirectional MCP. To simplify motion information coding and reduce the possible combinations of GPM, the motion information can be coded in a merge mode. The GPM merge candidate list can be derived from the merge candidate list to ensure that only unidirectional motion information is included.
[0094] Figure 5A Shows a prediction process of GPM according to some embodiments. The current block 510 is split into a right part and a left part via splitting 516. The right prediction part of the current block 510 (e.g., CU) of the current picture 502 (e.g., having a size of w×h) is predicted by MV0 of the reference block 512 from the reference picture 504, while the left part is predicted by MV1 of the reference block 514 from the reference picture 506.
[0095] Figure 5B Shows an example mixing matrix for splitting (e.g., splitting 516) according to some embodiments. In this example, the final GPM prediction (PG) is generated by performing a mixing process using integer mixing matrices W0 and W1 (e.g., weights in the value range from 0 to 8). This can be expressed as: PG = (W0 P0 + W1 P1 + 4) >> 3 where W0 + W1 = 8J Equation 4 - Hybrid Prediction In Equation 4, J is a matrix of all 1s, with a size of w×h. The weights of the mixing matrix can depend on the displacement between the sample position and the splitting boundary. The computational complexity of mixing matrix derivation can be very low, such that these matrices can be dynamically generated at the decoder side. Note that >> in Equation 4 indicates a right shift operation.
[0096] Then, the generated GPM prediction (PG) can be subtracted from the original signal to generate a residual. The residual can be transformed, quantized, and coded into the bitstream, e.g., using a conventional VVC transformation, quantization, and entropy coding engine. At the decoder side, the signal is reconstructed by adding the residual to the GPM prediction PG. GPM can also support a skip mode, e.g., when the residual can be ignored. For example, the residual drops due to the encoder, and the GPM prediction PG is directly used by the decoder as the reconstructed signal. GPM can be further enhanced, e.g., by GPM+TM (bilateral matching), GPM+MMVD (merge mode with motion vector difference), and inter+intra GPM. As Figure 5CAs shown, for all different contents, the mixing intensity or the mixing region width θ can be fixed.
[0097] Decoder-side motion vector refinement is a method of refining motion vectors using decoder-side information (e.g., reconstructed samples). Bilateral matching is a commonly used refinement matching method. In decoder-side motion vector refinement based on bilateral matching, a distortion metric such as the sum of absolute differences can be used to directly compare two candidate predictor blocks. The candidate block with the lowest distortion can be used as the refined block. Decoder-side motion vector refinement can be performed at the block level or the sub-block level. For example, in the sub-block level, the current block is divided into multiple sub-blocks, and bilateral matching is performed independently for each sub-block. Sub-block-based decoder-side motion vector refinement can achieve motion vector refinement with a finer granularity (e.g., sub-block level), but the matching cost may also be less accurate because fewer samples are used to infer the matching cost.
[0098] As discussed above, some codecs (e.g., AV1 and VVC) operate on pixel blocks. Each pixel block can be processed in a predictive transform coding scheme, where prediction is obtained using reference pixels and / or motion compensation. For an inter-predicted block, motion parameters such as motion vectors, reference picture indices, reference picture list indices, and / or required additional information can be used for inter-predicted sample generation. The motion parameters can be signaled in an explicit or implicit manner. As discussed above, inter-predicted blocks can use temporal motion vectors and / or spatial motion vectors. Additionally, sub-block level motion vector refinement can be applied to extend block-level TMVP.
[0099] Figure 5D An example of deriving sub-block motion vectors according to some embodiments is shown. In Figure 5D the current picture 402 includes a current block 403 composed of sub-blocks 405-1 to 405-16. Figure 5D The number and size of sub-blocks and blocks in Figure 5D are merely examples, and in other embodiments, different numbers and sizes of blocks and sub-blocks are used. Figure 5D A reference picture 406 is also shown, and the reference picture 406 has a reference block 407 corresponding to the current block 403. In some embodiments, the distances of the reference pictures 406 and 404 from the current picture 402 are different. In the example of Figure 5D the motion shift derived from the motion in block A1 is used to identify the reference block 407.
[0100] Therefore, Figure 5DAn example of Subblock-based TMVP (SbTMVP) is shown. SbTMVP can predict the motion vectors of subblocks in a current block in two steps. In the first step, spatial neighbors (denoted as A1 in Figure 5D are identified. If A1 has a motion vector that uses a collocated picture as its reference picture, that motion vector is selected as the motion shift (or displacement vector) to be applied. If no such motion is identified, the motion shift can be set to (0, 0). Figure 5D The example in
[0101] uses a motion shift based on the motion vector from block A1. Figure 5D In the second step, the motion shift identified in the first step (e.g., added to the coordinates of the current block) is applied to obtain subblock-level motion information (motion vectors and reference indices) from the collocated picture as shown in
[0102] For each subblock, the motion information of the subblock is then derived using the motion information of its corresponding block (e.g., the smallest motion grid covering the central sample) in the collocated picture. After the motion information of the collocated subblocks is identified, it is converted into the motion vectors and reference indices of the current subblocks in a manner similar to TMVP processing, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current block.
[0103] In this way, SbTMVP uses the sports field in the collocated picture to improve the motion vector prediction and merge mode of the coded block in the current picture. The same collocated picture used by TMVP can be used for SbTMVP. SbTMVP differs from TMVP in that TMVP predicts motion at the coded block level, while SbTMVP predicts motion at the sub-coded block level. Additionally, TMVP obtains the temporal motion vector from the collocated block in the collocated picture (e.g., the collocated block is the bottom-right or center block relative to the current CU), while SbTMVP applies a motion shift before obtaining the temporal motion information from the collocated picture. The motion shift can be obtained from the motion vector of one of the spatial neighboring blocks of the current coded block.
[0104] Figure 5E An example of decoder-side motion vector refinement according to some embodiments is shown. In Figure 5E , a set of reference pictures 404 and 406 are used to derive the refined motion vectors (refined MV0 and refined MV1). The initial motion vectors MV0 and MV1 can be used to identify the initial reference blocks in each reference picture. The motion differences (indicated by the arrows 416 and 418 in Figure 5E ) are applied to each motion vector to derive the refined motion vectors. Thus, Figure 5E shows an example of decoder-side motion vector refinement (Decoder Side Motion Vector Refinement, DMVR) applied to a coded block (e.g., in the merge mode). The MV pair obtained from the regular merge candidates can be used as the input for the DMVR process. DMVR applies bilateral matching (BM) to refine the input MV pair {mvL0, mvL1}, and the refined MV pair is used for motion compensation prediction (e.g., motion compensation prediction for both the luminance component and the chrominance component). The output MV (refined MV pair) of DMVR is defined in Equation Set 5: mv 细化L0 =mv L0 +Δmv mv 细化L1 =mv L1 -Δmv Equation Set 5 - Refined motion vector pair
[0105] In Equation Set 5, the motion vector difference Δmv is applied to the input MV pair to use the MVD (Motion Vector Differenc e) The mirroring attribute obtains a refined MV pair (e.g., because the input MV pair points to two different reference pictures, the two different reference pictures have an equal picture order count (POC) difference from the current picture, and the two reference pictures are in different temporal directions).
[0106] Integer sample offset search can be performed in DMVR. In an example implementation, the search space includes MV pair candidates (e.g., 25 pairs of candidates), as shown in Equation Set 6: mv L0(i,j) = mv L0(0,0) +(i, j) mv L1(i,j) = mv L1(0,0) -(i, j) Equation Set 6 - Search space for MV pair candidates where (i, j) represents the coordinates of the search points around the initial MV pair, and i and j are integer values between -2 and 2 (including -2 and 2). The sum of absolute differences (Sum of Absolute e Differences e , SAD) is calculated as shown in Equation Set 7 below: diff m,n = abs(P0 i,j [m + i, 2n + j] - P1 i,j [m - i, 2n - j]) Equation Set 7 - SAD calculation where W and H are the width and height of the sub - block. If the SAD of the initial MV pair is less than the threshold, the integer sample stage of DMVR is terminated. Otherwise, the SADs of the remaining 24 points are calculated and checked in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search stage. In some embodiments, for example, to reduce the penalty for the uncertainty of DMVR refinement, the SAD between the reference blocks referred to by the initial MV candidate reduces the SAD value by 1 / 4.
[0107] In some embodiments, the candidate MV pairs selected in the integer sample offset search step are further refined. For example, fractional sample refinement can be derived by using parametric error surfaces (e.g., to save computational complexity), rather than performing additional searches using SAD comparisons. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. For example, fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. As an example, when the integer sample search phase is terminated in the case of the center with the minimum SAD in the first or second iteration search, fractional sample refinement is further applied.
[0108] BDOF (Bi-Directional Optical Flow) can be used to refine the bi-prediction signals of coded blocks (e.g., CUs). For example, BDOF can be performed at the 4×4 sub-block level. BDOF can be applied to a coded block if the coded block satisfies at least one subset of the following conditions: (i) the coded block is coded using a bi-prediction mode in which one of the two reference pictures is before the current picture in display order and the other is after the current picture in display order, (ii) the distances from the two reference pictures to the current picture (e.g., POC difference) are the same, (iii) both reference pictures are short-term reference pictures, (iv) the coded block is not coded using the affine mode or the SbTVMP merge mode, (v) the coding unit has more than 64 luma samples, (vi) the coding unit height and width are greater than or equal to 8 luma samples, (vii) the BCW (Bi-Prediction with CU-level Weight) weight index indicates equal weights, (viii) WP (Weighting Prediction) is not enabled for the current coded block, and (ix) the CIIP mode is not used for the current coded block. In some implementations, BDOF is applied only to the luma component.
[0109] The BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each sub-block (e.g., 4×4 sub-block), the motion refinement (v x , v y ) can be calculated by minimizing the difference between the L0 prediction samples and the L1 prediction samples. Then, the motion refinement can be used to adjust the bi-prediction sample values in the sub-block.
[0110] In some implementations, multiple passes of decoder-side motion vector refinement are applied. For example, in the first pass, bilateral matching (BM) is applied to the coding block, in the second pass, BM is applied to each 16×16 sub-block within the coding block, and in the third pass, the MVs in each 8×8 sub-block are refined by applying bidirectional optical flow (BDOF). Then, the refined MVs can be stored for subsequent spatial and / or temporal motion vector prediction.
[0111] As discussed above, decoder-side motion vector refinement uses existing decoder-side information such as reconstructed samples to refine the motion vectors. Additionally, an optical flow approach can be applied to formulate a least squares problem, from which the refined motion can be derived from the gradients of the composite inter-predicted samples. Using these refined motions, the MVs can be refined in each sub-block within the prediction block, thus enhancing the inter-prediction quality. This implementation is an extension of BDOF as it supports MV refinement when the two reference blocks have an arbitrary temporal distance to the current block. The gradients of the current whole-block predictor can be pre-computed and then the sub-block refined MVs can be calculated according to the optical flow model.
[0112] Bilateral matching can be used for MV refinement. In decoder-side motion vector refinement based on bilateral matching, a distortion metric such as sum of absolute differences (SAD) can be used to directly compare two candidate predictor blocks. Then, the candidate block with the lowest distortion can be used as the refined block. Decoder-side motion vector refinement can be performed at the block level or sub-block level. In the sub-block level, the current block is divided into multiple sub-blocks and bilateral matching is performed independently for each sub-block. Sub-block based decoder-side motion vector refinement can achieve motion vector refinement at a finer granularity (e.g., sub-block level), but the matching cost may also be less accurate as fewer samples are used to infer the matching cost.
[0113] The prefetch samples mentioned herein refer to the limited maximum number of samples in the reference picture that can be used for bilateral matching processing. The prefetch region mentioned herein refers to the region occupied by the prefetch samples. When performing bilateral matching processing, the more the number of prefetch samples, the greater the memory bandwidth required.
[0114] In some embodiments, if a block is encoded in a composite mode with bi - directional reference frames, the MV of the block is refined before applying optical flow refinement. In some embodiments, if the block size is greater than 16×16, the block is divided into multiple 16×16 sub - blocks. If the block size is less than or equal to 16×16, the refinement is performed on the entire block basis. For each sub - block, an offset motion vector (e.g., referred to as ΔMV) can be derived. The offset motion vector of the sub - block can be found by searching the neighborhood of the initial motion vectors MV0 and MV1. For example, the decoder searches a predefined 5×5 region centered on the initial motion vector and selects the offset that produces the minimum sum of absolute differences (SAD) between P0 and P1. In some embodiments, only integer offsets are searched.
[0115] In some embodiments, only integer offsets are searched and a bilinear interpolation filter is used during the search. In some embodiments, after calculating the SAD for each offset, the SAD value is checked with a threshold. If the SAD is less than (width * height), the search terminates. In some embodiments, MV refinement is applied to both the chrominance and luminance channels. In some embodiments, the search is only performed on the luminance channel and the refined MV values are stored for reuse by the chrominance channel.
[0116] As used herein, temporal interpolation prediction (sometimes also referred to as intra - temporal interpolation prediction) refers to an inter - frame prediction mode for deriving an inter - frame predicted block using a block (or reference frame) generated by interpolating the current block using coded pictures and motion vectors.
[0117] Some embodiments include using a composite inter - frame mode. The composite inter - frame mode creates a prediction of a block by combining two hypotheses from two different reference frames. In these modes, two motion information components (e.g., motion vectors) can be sent in the bitstream. Although motion vectors can be well predicted using predictors from spatial and temporal neighbors or historical motion vectors, the bytes for motion information can still be quite substantial for many content and applications.
[0118] In some embodiments, the inter-frame prediction mode is based on temporal interpolation techniques. For example, by utilizing the already available motion vector fields of the forward and backward reference frames, an intermediate frame can be generated by interpolation. This interpolated frame can have a high correlation with the current encoded frame. Thus, it can be used as an additional reference frame for the current frame, or can be directly associated with the current frame as the output of the decoder, in which case no additional encoding steps need to be performed. The temporal interpolation technique can be based on the direct projection of the existing motion vectors in the available reference frames without performing any additional motion estimation. In some embodiments, the Temporal Interpolated Prediction (TIP) mode can be enabled only in the random access mode.
[0119] The TIP process can include combining the information in two reference frames and using interpolation processing to project to the same temporal instance as the current frame. In an example TIP mode, the interpolated frame (TIP frame) can be used as an additional reference frame. The encoded blocks of the current frame can directly refer to the interpolated frame and utilize the information from two different references, while only requiring the overhead cost of a single inter-frame prediction mode. In another example TIP mode, the interpolated frame can be directly specified as the output of the decoding process of the current frame (e.g., skipping other traditional encoding steps).
[0120] In an example process, the TIP frame is first generated. Then, the TIP frame is used as an additional reference frame for the current frame, or is directly specified as the reconstructed output of the decoder for the current frame. The TIP mode used can be indicated using a syntax element, for example, named tip_frame_mode.
[0121] Figure 6A FIG. 600 is a flowchart showing a method 600 for decoding video according to some embodiments. Method 600 can be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 600 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.
[0122] Some embodiments disclose an adaptive frame filling method for reconstructing pictures in a video bitstream. The system receives (602) a video bitstream (e.g., an encoded video sequence) including a plurality of blocks (e.g., corresponding to a group of pictures), the plurality of blocks including a first block encoded using a combined intra - inter prediction (CIIP) mode. The system identifies (604) the intra - prediction mode of the first block. The system identifies (606) the inter - prediction mode of the first block, and the identified inter - prediction mode is one of: an inter - prediction mode based on geometric segmentation, an inter - prediction mode based on affine motion, and a temporal interpolation prediction mode. The system decodes (608) the first block using the identified intra - prediction mode and the identified inter - prediction mode. In this way, the inter - prediction part in the combined intra - and inter - prediction mode can employ a wedge - shaped (or geometric - segmentation - based) inter - prediction or an inter - prediction based on warped motion (or affine motion).
[0123] In some embodiments, when a wedge - shaped or geometric - segmentation - based or warped - motion - based method is used for the inter - prediction of the combined intra - and inter - prediction mode, only a predefined set of weights can be used to blend intra - prediction samples and inter - prediction samples and generate a final prediction sample. In this way, in this case, a wedge - mask - based weighting factor is not allowed to blend intra - prediction samples and inter - prediction samples.
[0124] In some embodiments, the predefined weighting factor can be a) the same for all samples within the prediction block; or b) gradually decrease the intra - prediction weight along its prediction direction.
[0125] In some embodiments, when a wedge - shaped or geometric - segmentation - based method is used for the inter - prediction of the combined intra - and inter - prediction mode, the wedge mask used for generating the inter - prediction can also be used to combine the intra - prediction samples and the inter - prediction samples.
[0126] In some embodiments, when the intra - prediction mode is from a predefined set of intra - prediction modes, a wedge - shaped or geometric - segmentation - based method or a warped - motion - based method can be used for the inter - prediction of the combined intra - and inter - prediction mode. For example, the predefined set of intra - prediction modes includes only DC and / or smooth modes. In another example, the predefined set of intra - prediction modes includes only DC, smooth modes, angular intra - (e.g., vertical intra - prediction or horizontal intra - prediction) prediction modes.
[0127] In some embodiments, when the warping motion is used for inter - prediction in the combined intra - and inter - prediction mode, the warping motion mode is set to a predefined warping motion mode. In some embodiments, the warping motion mode is set to a warping model that derives model parameters using a linear method, and the warping model parameters are not transmitted to the bitstream, such as the WARP_CAUSAL mode. In some embodiments, the warping motion mode is set to a warping model that extends warping model parameters from neighboring blocks to the current block, such as the WARP_EXTENDED mode. In some embodiments, the predefined warping motion mode can be signaled to the bitstream via sequence / frame / slice high - level syntax.
[0128] In some embodiments, when the warping motion is used for inter - prediction in the combined intra - and inter - prediction mode, the warping motion mode can only be selected from a subset of the allowed warping motion modes. For example, the warping motion mode can be only the WARP_CAUSAL mode or the WARP_EXTENDED mode.
[0129] In some embodiments, one of the high - level syntaxes at the sequence, frame, or slice level is parsed into the bitstream to indicate whether the wedge (or geometry - based partition) inter - prediction or warping motion can be applied to the combined intra - and inter - prediction mode.
[0130] In some embodiments, a decoder - side motion refinement method is applied to the combined intra - and inter - prediction mode to refine the motion vectors used for the inter - prediction part. In some embodiments, decoder - side motion vector refinement based on bilateral matching is used for the inter - prediction part of the combined intra - and inter - prediction mode. In some embodiments, a distortion metric such as the sum of absolute differences is used to compare the inter - prediction samples of the current block with the intra - prediction samples of the current block. The offset (or delta) motion vector can be found by searching the neighboring region of the decoded motion vector M0 used for the inter - prediction part. In some embodiments, after applying the bilateral - matching - based method to the combined intra - and inter - prediction mode, the intra - prediction part remains unchanged, and the motion vectors of the inter - prediction part can be refined.
[0131] In some embodiments, an optical - flow - based motion vector refinement method or an optical - flow - based prediction refinement method is used for the inter - prediction part of the combined intra - and inter - prediction mode. In some embodiments, the refinement is applied to the de - meaned intra - and inter - prediction mode.
[0132] In some embodiments, one of the high - level syntaxes at the sequence, frame, or slice level is parsed into the bitstream to indicate whether the decoder - side motion refinement method can be applied to the combined intra - and inter - prediction mode.
[0133] In some embodiments, a temporal interpolation prediction mode is applied to derive an inter prediction block in a combined intra and inter prediction mode.
[0134] In some embodiments, the weight between an inter prediction block and an intra prediction block in a combined inter and intra prediction mode depends on which inter prediction mode is used to derive the inter prediction mode and / or which intra prediction mode is used to derive the intra prediction mode.
[0135] Figure 6B FIG. 650 is a flow diagram illustrating a method 650 for encoding video according to some embodiments. Method 650 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 650 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system. In some embodiments, method 650 is performed by the same system as method 600 described above.
[0136] The system receives (652) video data (e.g., a source video sequence) including a plurality of blocks (e.g., corresponding to a set of pictures), the plurality of blocks including a first block to be encoded using a combined intra-inter prediction (CIIP) mode. The system identifies (654) an intra prediction mode of the first block. The system identifies (656) an inter prediction mode of the first block, the identified inter prediction mode being one of: an inter prediction mode based on geometric partitioning, an inter prediction mode based on affine motion, and a temporal interpolation prediction mode. The system encodes (658) the first block using the identified intra prediction mode and the identified inter prediction mode. As previously described, the encoding process may mirror the decoding process described herein (e.g., applying the CIIP mode). For the sake of brevity, these details are not repeated here.
[0137] Although Figure 6A and Figure 6B The various logical stages are shown in a particular order, but stages that are not order dependent may be reordered and other stages may be combined or split. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, and thus the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that the stages may be implemented in hardware, firmware, software, or any combination thereof.
[0138] Turning now to some example embodiments.
[0139] (A1)In one aspect, some embodiments include a method of video decoding (e.g., method 600). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). In some embodiments, the method is performed at a source coding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving a video bitstream (e.g., an encoded video sequence) including a first block, wherein the first block is encoded using a combined intra-inter prediction (CIIP) mode; (ii) identifying an intra prediction mode for the first block; (iii) identifying an inter prediction mode for the first block, the identified inter prediction mode being one of: (a) an inter prediction mode based on geometric partitioning, (b) an inter prediction mode based on affine motion, and (c) a temporal interpolation prediction mode; and (iv) decoding the first block using the identified intra prediction mode and the identified inter prediction mode. For example, the inter prediction portion under the CIIP mode may employ a wedge (or other geometric partitioning-based) inter prediction, an inter prediction based on warped motion (or affine motion), or a temporal interpolation prediction mode. In some embodiments, the inter prediction mode is selected from the group including: an inter prediction mode based on geometric partitioning, an inter prediction mode based on affine / warped motion, a temporal interpolation prediction mode, and a translational inter prediction mode. In some embodiments, the inter prediction mode is selected from the group including: a temporal interpolation prediction mode and a translational inter prediction mode. In some embodiments, the inter prediction mode is selected from the group including: an inter prediction mode based on geometric partitioning, an inter prediction mode based on affine / warped motion, and a translational inter prediction mode.
[0140] (A2)In some embodiments of A1, decoding the first block using the identified intra prediction mode and the identified inter prediction mode includes using a weighted combination of the identified intra prediction mode and the identified inter prediction mode, wherein the weighted combination uses a predefined set of weights. For example, when a wedge or geometric partitioning-based method or a warped motion-based method is used for inter prediction, only the predefined set of weights is allowed to mix intra prediction samples and inter prediction samples and generate a final prediction sample. In this way, a wedge mask-based weighting factor is not allowed to mix intra prediction samples and inter prediction samples.
[0141] (A3)In some embodiments of A2, the predefined set of weights includes the same weight applied to each sample of the first block.
[0142] (A4) In some embodiments of A2, the predefined set of weights has intra prediction weights that decrease according to the intra prediction angle. For example, the predefined weighting factor can be the same for all samples within the prediction block, or the intra prediction weights can gradually decrease along its prediction direction.
[0143] (A5) In some embodiments of any one of A2 to A4, a predefined set of weights is identified based on at least one of the identified intra prediction mode and the identified inter prediction mode. For example, the weights between the inter prediction block and the intra prediction block in the combined inter and intra prediction mode can depend on which inter prediction mode is used to derive the inter prediction mode and / or which intra prediction mode is used to derive the intra prediction mode.
[0144] (A6) In some embodiments of any one of A1 to A5, (i) the identified inter prediction mode is a geometric segmentation-based inter prediction mode using a wedge mask; and (ii) decoding the first block using the identified intra prediction mode and the identified inter prediction mode includes using the wedge mask to combine the inter prediction samples and the intra prediction samples. For example, when a wedge or geometric segmentation-based method is used for inter prediction, the wedge mask used to generate the inter prediction can also be used to combine the intra prediction samples and the inter prediction samples.
[0145] (A7) In some embodiments of any one of A1 to A6, when the identified intra prediction mode is one of a predefined set of intra prediction modes, the inter prediction mode for the first block is selected from the group including: a geometric segmentation-based inter prediction mode, an affine motion-based inter prediction mode, and a temporal interpolation prediction mode. For example, when the intra prediction mode is from a predefined set of intra prediction modes, a wedge or geometric segmentation-based method or a warped motion-based method can be used for the inter prediction in the combined intra and inter prediction mode. In some embodiments, the inter prediction mode for the first block is selected from the group including: a geometric segmentation-based inter prediction mode, an affine motion-based inter prediction mode, and a temporal interpolation prediction mode. In some embodiments, the inter prediction mode for the first block is selected from the group including: a geometric segmentation-based inter prediction mode, an affine motion-based inter prediction mode, a temporal interpolation prediction mode, and a translational inter prediction mode. In some embodiments, based on the determination that the identified intra prediction mode is one of a predefined set of intra prediction modes, the inter prediction mode for the first block is selected from the group including: a geometric segmentation-based inter prediction mode, an affine motion-based inter prediction mode, and a temporal interpolation prediction mode.
[0146] (A8)In some embodiments of A7, the set of predefined intra prediction modes includes a smooth intra prediction mode and a direct current (DC) intra prediction mode. For example, the set of predefined intra prediction modes includes only the DC mode and / or the smooth mode.
[0147] (A9)In some embodiments of A7, the set of predefined intra prediction modes includes a smooth intra prediction mode, a direct current (DC) intra prediction mode, and a set of angular intra prediction modes. For example, the set of predefined intra prediction modes includes only the DC prediction mode, the smooth prediction mode, and angular intra (e.g., vertical and horizontal intra prediction) prediction modes.
[0148] (A10)In some embodiments of any one of A1 to A5 and A6 to A9, identifying an inter prediction mode includes identifying a predefined warped motion mode as an inter prediction mode based on affine motion. For example, when warped motion is used for inter prediction of the CIIP mode, the warped motion mode is set to a predefined warped motion mode.
[0149] (A11)In some embodiments of A10, the video bitstream does not include warping model parameters for a first block. For example, the warped motion mode is set to a warping model that derives model parameters in a linear fashion, and the warping model parameters are not transmitted to the bitstream. The warped motion mode is, for example, the WARP_CAUSAL mode.
[0150] (A12)In some embodiments of A10 or A11, the predefined warped motion mode uses a warping model that has warping model parameters from one or more neighboring blocks of a first block. For example, the warped motion mode is set to extend warping model parameters from neighboring blocks to the warping model of the current block. The warped motion mode is, for example, the WARP_EXTENDED mode.
[0151] (A13)In some embodiments of A10, the predefined warped motion mode is signaled in the video bitstream. For example, the predefined warped motion mode can be signaled into the bitstream via sequence / frame / slice high-level syntax. In some embodiments, one or more parameters for the predefined warped motion mode are signaled in the video bitstream. For example, the warped motion mode can be an incremental warping mode, and a set of incremental values can be signaled for the first block.
[0152] (A14)In some embodiments of any one of A10 to A13, the predefined warping motion pattern is selected from a subset of the allowed warping motion patterns. For example, when warping motion is used for inter-frame prediction in the CIIP mode, the warping motion pattern is only selected from a subset of the allowed warping motion patterns. In some embodiments, the subset of the allowed warping motion patterns includes the WARP_CAUSAL pattern and the WARP_EXTENDED pattern. For example, the warping motion pattern can be only the WARP_CAUSAL pattern or the WARP_EXTENDED pattern.
[0153] (A15)In some embodiments of any one of A1 to A14, an inter-frame prediction mode is identified for a first block based on an indicator in the video bitstream. For example, the high-level syntax in the bitstream can be parsed to determine whether wedge (or geometry-based segmentation) inter-frame prediction or warping motion inter-frame prediction can be applied to the CIIP mode.
[0154] (A16)In some embodiments of any one of A1 to A15, motion vector refinement is further applied to the motion vector of the identified inter-frame prediction mode. For example, a decoder-side motion refinement method can be applied to the CIIP mode to refine the motion vector for the inter-frame prediction part. In some embodiments, motion vector refinement is applied when an indicator in the video bitstream indicates that motion vector refinement is to be applied. For example, the high-level syntax is parsed from the bitstream to determine whether the decoder-side motion refinement method can be applied to the CIIP mode. In some embodiments, applying motion vector refinement includes applying the refinement to de-meaned intra- and inter-frame prediction modes. For example, the refinement can be applied to de-meaned intra- and inter-frame prediction modes. In some embodiments, applying motion vector refinement includes using bilateral matching. For example, decoder-side motion vector refinement based on bilateral matching is used for the inter-frame prediction part of the CIIP mode. In some embodiments, after applying the bilateral matching-based method to the CIIP mode, the intra-frame prediction part remains unchanged, and the motion vector of the inter-frame prediction part is refined.
[0155] (A17)In some embodiments of A16, applying motion vector refinement includes comparing the inter-frame prediction samples for the first block with the intra-frame prediction samples for the first block. For example, the inter-frame prediction samples of the current block are compared with the intra-frame prediction samples of the current block using a distortion metric such as the sum of absolute differences. The offset (or delta) motion vector can be found by searching a neighboring region of the decoded motion vector M0 for the inter-frame prediction part.
[0156] (A18)In some embodiments of A16, the motion vector refinement is optical-flow-based motion vector refinement. For example, an optical-flow-based motion vector refinement method or an optical-flow-based prediction refinement method is used for the inter-frame prediction part of the CIIP mode.
[0157] (B1)In another aspect, some embodiments include a method of video encoding (e.g., method 650). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) receiving video data (e.g., a source video sequence) including a plurality of blocks, the plurality of blocks including a first block, wherein the first block is to be encoded using a combined intra-inter prediction (CIIP) mode; (ii) identifying an intra prediction mode for the first block; (iii) identifying an inter prediction mode for the first block, the identified inter prediction mode being one of: (a) an inter prediction mode based on geometric partitioning, (b) an inter prediction mode based on affine motion, and (c) a temporal interpolation prediction mode; and (iv) encoding the first block using the identified intra prediction mode and the identified inter prediction mode.
[0158] (C1)In another aspect, some embodiments include a method of visual media data processing. In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) obtaining a source video sequence including a plurality of frames; and (ii) performing a conversion between the source video sequence and a video bitstream of visual media data according to format rules. The video bitstream includes a plurality of blocks, the plurality of blocks including a first block, wherein the first block is encoded using a combined intra-inter prediction (CIIP) mode. The format rules specify that the inter prediction mode of the CIIP mode is one of: an inter prediction mode based on geometric partitioning, an inter prediction mode based on affine motion, and a temporal interpolation prediction mode.
[0159] (D1) In one aspect, some embodiments include a method of video decoding (e.g., method 600). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). In some embodiments, the method is performed at a source coding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving a video bitstream (e.g., an encoded video sequence) including a first block, wherein the first block is encoded using a combined intra-inter prediction (CIIP) mode; (ii) identifying an intra prediction mode for the first block; (iii) identifying an inter prediction mode for the first block; (iv) obtaining a motion vector for the first block using the inter prediction mode; (v) refining the motion vector for the first block using a motion vector refinement technique; and (vi) decoding the first block using the identified intra prediction mode and the refined motion vector. Refining the motion vector of the CIIP mode can improve the accuracy of the motion vector, which in turn improves the encoding accuracy compared to the CIIP mode without motion vector refinement.
[0160] (D2) In some embodiments of D1, the motion vector refinement technique includes at least one of decoder-side motion vector refinement and optical flow-based motion vector refinement.
[0161] (D3) In some embodiments of D1 or D2, the CIIP mode includes a non-translational inter prediction mode as described in any one of A1 to A18.
[0162] (E1) In one aspect, some embodiments include a method of video decoding (e.g., method 600). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). In some embodiments, the method is performed at a source coding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving a video bitstream (e.g., an encoded video sequence) including a first block, wherein the first block is encoded using a combined intra-inter prediction (CIIP) mode; (ii) identifying an intra prediction mode for the first block; (iii) identifying a temporal interpolation prediction mode for the first block; and (iv) decoding the first block using the identified intra prediction mode and the identified temporal interpolation prediction mode.
[0163] (E2) In some embodiments of E1, the CIIP mode includes a non-translational inter-frame prediction mode as described in any one of A1 to A18.
[0164] (F1) In one aspect, some embodiments include a method of video decoding (e.g., method 600). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). In some embodiments, the method is performed at a source coding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving a video bitstream (e.g., an encoded video sequence) including a first block, wherein the first block is encoded using a combined intra-inter prediction (CIIP) mode; (ii) identifying an intra prediction mode for the first block; (iii) identifying an inter prediction mode for the first block; (iv) identifying a set of weights based on one or more of the identified intra prediction mode and the identified inter prediction mode; and (v) decoding the first block by combining a first sample obtained using the identified intra prediction mode and a second sample obtained using the identified inter prediction mode using the set of weights. Adjusting the weights for the combination of intra prediction and inter prediction allows for higher coding accuracy (e.g., by weighting more likely predictions higher and / or reducing the weight assigned to less likely predictions) compared to combining the CIIP mode using preset weights.
[0165] (F2) In some embodiments of F1, the CIIP mode includes a non-translational inter-frame prediction mode as described in any one of A1 to A18.
[0166] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more sets of instructions configured to be executed by the control circuitry, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A18, B1, C1, D1 to D3, E1 to E2, and F1 to F2 above).
[0167] In another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by a control circuitry of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A18, B1, C1, D1 to D3, E1 to E2, and F1 to F2 above).
[0168] Unless otherwise specified, any syntactic element (e.g., indicator) described herein may be High-Level Syntax (HLS). As used herein, HLS is signaled at a level higher than the block level. For example, HLS may correspond to the sequence level, frame level, slice level, or tile level. As another example, HLS elements may be signaled in a Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptation Parameter Set (APS), slice header, picture header, tile header, and / or CTU header.
[0169] It will be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are also intended to include the plural forms. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0170] As used herein, the term "if" may be construed, depending on the context, as "when the precondition is true" or "after the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "in response to detecting that the precondition is true". Similarly, the phrase "if it is determined [that the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" may be construed, depending on the context, as "after determining that the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "after detecting that the precondition is true" or "in response to detecting that the precondition is true".
[0171] For purposes of illustration, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best illustrate the principles of operation and practical application, thereby enabling others skilled in the art to implement them.
Claims
1. A video decoding method, characterized in that, The method includes: Receiving a video bitstream including a plurality of blocks, the plurality of blocks including a first block, wherein the first block is encoded using a combined intra - inter prediction (CIIP) mode; Identifying an intra - prediction mode for the first block; Identifying an inter - prediction mode for the first block, the identified inter - prediction mode being one of the following: An inter - prediction mode based on geometric segmentation, An inter - prediction mode based on affine motion, and A temporal interpolation prediction mode; and Decoding the first block using the identified intra - prediction mode and the identified inter - prediction mode.
2. The method according to claim 1, wherein Decoding the first block using the identified intra - prediction mode and the identified inter - prediction mode includes using a weighted combination of the identified intra - prediction mode and the identified inter - prediction mode, wherein the weighted combination uses a predefined set of weights.
3. The method according to claim 2, wherein, The predefined set of weights includes the same weight applied to each sample of the first block; or The predefined set of weights has a decreasing intra - prediction weight according to the intra - prediction angle.
4. The method according to claim 2, characterized in that, The predefined set of weights is identified based on at least one of the identified intra - prediction mode and the identified inter - prediction mode.
5. The method according to claim 1, wherein, The identified inter - prediction mode is the inter - prediction mode based on geometric segmentation using a wedge - shaped mask; And Decoding the first block using the identified intra - prediction mode and the identified inter - prediction mode includes using the wedge - shaped mask to combine inter - prediction samples and intra - prediction samples.
6. The method according to claim 1, characterized in that When the identified intra - prediction mode is one of a predefined set of intra - prediction modes, the inter - prediction mode for the first block is selected from the group including the following: the inter - prediction mode based on geometric segmentation, the inter - prediction mode based on affine motion, and the temporal interpolation prediction mode.
7. The method according to claim 6, wherein, The predefined set of intra - prediction modes includes a smooth intra - prediction mode and a direct current (DC) intra - prediction mode; or The predefined set of intra - prediction modes includes a smooth intra - prediction mode, a direct current (DC) intra - prediction mode, and a set of angular intra - prediction modes.
8. The method according to claim 1, wherein Identifying the inter - prediction mode includes identifying a predefined warped motion mode as the inter - prediction mode based on affine motion.
9. The method according to claim 8, characterized in that, The video bitstream does not include warping model parameters for the first block.
10. The method according to claim 8, wherein The predefined warped motion mode uses a warping model having warping model parameters from one or more neighboring blocks of the first block.
11. The method according to claim 8, characterized in that The predefined warped motion mode is signaled in the video bitstream.
12. The method according to claim 8, characterized in that The predefined warped motion mode is selected from a subset of allowed warped motion modes.
13. The method according to claim 1, wherein The inter - prediction mode for the first block is identified based on an indicator in the video bitstream.
14. The method according to claim 1, wherein The method further includes applying motion vector refinement to the motion vector of the identified inter - prediction mode.
15. The method according to claim 14, wherein Applying the motion vector refinement includes comparing an inter - frame prediction sample for the first block with an intra - frame prediction sample for the first block.
16. The method according to claim 14, wherein, The motion vector refinement is an optical - flow - based motion vector refinement.
17. A video encoding method, characterized in that, The method includes: Generating a video bitstream including a plurality of blocks, the plurality of blocks including a first block; Determining an intra - frame prediction mode for the first block; Determining an inter - frame prediction mode for the first block, the determined inter - frame prediction mode being one of the following: An inter - frame prediction mode based on geometric segmentation, An inter - frame prediction mode based on affine motion, and A temporal interpolation prediction mode; and Encoding the first block using a combined intra - inter prediction (CIIP) mode, including: encoding the first block based on the determined intra - frame prediction mode and the determined inter - frame prediction mode.
18. A video decoding device, characterized in that, The apparatus includes: At least one memory configured to store computer program code; and At least one processor configured to read the computer program code and perform the method according to any one of claims 1 to 15 as indicated by the computer program code.
19. A video encoding device, characterized in that, The apparatus includes: At least one memory configured to store computer program code; and At least one processor configured to read the computer program code and perform the method according to claim 17 as indicated by the computer program code.
20. A non-transitory computer-readable storage medium, characterized in that, The storage medium stores one or more sets of instructions configured for execution by a computing device having control circuitry and a memory of the method according to any one of claims 1 to 17.