Improvements in motion vector redundancy and similarity check for merge mode
By using a processing circuit system in a video decoder, performing verification of motion and non-motion inheritable parameters, and building an optimized candidate list, the problem of inefficient management of motion vector and non-motion vector redundant information in the prior art is solved, and efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202480004317.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-20
- Filing Date
- 2024-04-19
- Publication Date
- 2025-05-30
AI Technical Summary
When existing video encoding and decoding technologies process video data, it is difficult to effectively manage the redundant information of motion vectors and the redundant information of non-motion vectors, resulting in ineffective encoding and decoding.
By introducing a processing circuit system into the video decoder, checking based on motion and non-motion inheritable parameters is performed, redundant information between potential candidates and existing candidates is determined, and an optimized candidate list is constructed to improve encoding and decoding efficiency.
It realizes efficient encoding and decoding of video data, reduces the storage and transmission of redundant information, and improves the overall performance of the video processing system.
Smart Images

Figure CN120077642A_ABST
Abstract
Description
Related Applications
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 460,888, filed on April 20, 2023, entitled "IMPROVEMENT OF MOTION VECTOR REDUNDANCY AND SIMILARITY CHECK FOR MERGE MODE", the entire content of which is incorporated herein by reference. Technical Field
[0002] The present disclosure generally describes embodiments related to video coding and decoding. Background Art
[0003] The background art description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work described in this background art section, the work of the presently named inventors, and aspects that may not otherwise be qualified as prior art at the time of filing are neither expressly nor implicitly admitted as prior art against the present disclosure.
[0004] Image / video compression can help transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial redundancy and temporal redundancy. In an example, a video codec can use a technique called intra prediction, which can compress an image based on spatial redundancy. For example, intra prediction can use reference data from the current picture in reconstruction for sample prediction. In another example, a video codec can use a technique called inter prediction, which can compress an image based on temporal redundancy. For example, inter prediction can utilize motion compensation to predict samples in the current picture based on a previously reconstructed picture. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention
[0005] Aspects of the present disclosure include methods and apparatuses for video encoding / decoding. In some examples, an apparatus for video decoding includes processing circuitry.
[0006] Some aspects of the present disclosure provide a method for processing visual media data. The method includes processing a bitstream of visual media data according to format rules. The bitstream includes an index pointing to a specific candidate in a candidate list for prediction of a current block in a current picture. The format rules specify obtaining a first check result by performing a first check of motion-based inheritable parameters between a potential candidate and a first existing candidate among one or more existing candidates in the candidate list, the first check result indicating whether the potential candidate has redundant information of motion-based inheritable parameters relative to the first existing candidate. The format rules also specify obtaining a second check result by performing a second check of non-motion-based inheritable parameters between the potential candidate and the first existing candidate, the second check result indicating whether the potential candidate has redundant information of non-motion-based inheritable parameters relative to the first existing candidate. The format rules further specify that when the first check result indicates that the potential candidate has redundant information of motion-based inheritable parameters relative to the first existing candidate and the second check result indicates that the potential candidate has redundant information of non-motion-based inheritable parameters relative to the first existing candidate, the potential candidate is redundant for the first existing candidate.
[0007] In some examples, when the current block is in an inter prediction mode, the motion-based inheritable parameters include a motion vector (MV), and the non-motion-based inheritable parameters include bidirectional prediction with a block coding unit level weight (BCW) index.
[0008] In some examples, when the current block is in an intra block copy (IBC) mode, the motion-based inheritable parameters include a block vector (BV), and the non-motion-based inheritable parameters include a reconstruction reordering type for IBC for the IBC mode.
[0009] Some aspects of the present disclosure provide a video encoding method. The method includes: determining a candidate list for constructing a prediction for a current block in a current picture; determining a potential candidate for addition to the candidate list, the candidate list including one or more existing candidates that have been added to the candidate list; and performing a first check of motion-based inheritable parameters between the potential candidate and a first existing candidate among the one or more existing candidates to obtain a first check result. The first check result indicates whether the potential candidate has redundant information of motion-based inheritable parameters relative to the first existing candidate. The method further includes performing at least a second check of non-motion-based inheritable parameters between the potential candidate and the first existing candidate to obtain a second check result, the second check result indicating whether the potential candidate has redundant information of non-motion-based inheritable parameters relative to the first existing candidate. The method further includes determining that the potential candidate is redundant for the first existing candidate when the first check result indicates that the potential candidate has redundant information of motion-based inheritable parameters relative to the first existing candidate and the second check result indicates that the potential candidate has redundant information of non-motion-based inheritable parameters relative to the first existing candidate.
[0010] In some examples, when the current block is in an inter prediction mode, the motion-based inheritable parameters include a motion vector (MV). In an example, the non-motion-based inheritable parameters include bidirectional prediction with a coding and decoding unit-level weight (BCW) index.
[0011] In some examples, when the current block is in an intra block copy (IBC) mode, the motion-based inheritable parameters include a block vector (BV). In an example, the non-motion-based inheritable parameters include a reconstruction reordering type of IBC for the IBC mode. For example, the non-motion-based inheritable parameters include one of a vertical flip and a horizontal flip.
[0012] In some examples, the second check result indicates whether the non-motion-based inheritable parameters of the potential candidate are exactly the same as those of the first existing candidate.
[0013] In some examples, the second check result indicates whether the difference in non-motion-based inheritable parameters between the potential candidate and the first existing candidate is within a certain range.
[0014] Some aspects of the present disclosure provide an apparatus for video decoding. The apparatus includes processing circuitry configured to receive a coded video bitstream including coded information of one or more pictures, and construct a candidate list for a current block in a current picture, the candidate list including at least a first candidate and a second candidate, the first candidate and the second candidate having redundant information based on motion-based inheritable parameters and non-redundant information based on at least non-motion-based inheritable parameters. The processing circuitry further selects a specific candidate from the candidate list and reconstructs the current block based on the specific candidate.
[0015] In some examples, the processing circuitry determines a potential candidate to be added to the candidate list and performs a first check on motion-based inheritable parameters between the potential candidate and a first candidate that is an existing candidate in the candidate list to obtain a first check result, the first check result indicating whether the potential candidate has redundant information of motion-based inheritable parameters identical to those of the first candidate.
[0016] In some examples, the processing circuitry performs at least a second check on non-motion-based inheritable parameters between the potential candidate and the first candidate to obtain a second check result, the second check result indicating whether the potential candidate has redundant information of non-motion-based inheritable parameters identical to those of the first candidate. When the first check result indicates that the potential candidate has redundant information of motion-based inheritable parameters identical to those of the first candidate and the second check result indicates that the potential candidate has redundant information of non-motion-based inheritable parameters identical to those of the first candidate, the processing circuitry determines that the potential candidate is redundant with respect to the first candidate.
[0017] In some examples, when the current block is in an inter prediction mode, the motion-based inheritable parameters include a motion vector (MV). In an example, the non-motion-based inheritable parameters include bi-prediction with a coded unit-level weight (BCW) index.
[0018] In some examples, when the current block is in an intra block copy (IBC) mode, the motion-based inheritable parameters include a block vector (BV). The non-motion-based inheritable parameters include a reconstruction reordering type for IBC for the IBC mode. For example, the non-motion-based inheritable parameters include at least one of a vertical flip and a horizontal flip.
[0019] In some examples, when the non-motion-based inheritable parameters of the potential candidate are exactly the same as those of the first candidate, the processing circuitry determines that the potential candidate has redundant information of non-motion-based inheritable parameters identical to those of the first candidate.
[0020] In some examples, when the difference in non-motion-based inheritable parameters between a potential candidate and a first candidate is within a certain range, the processing circuitry determines that the potential candidate has redundant information of non-motion-based inheritable parameters that is the same as that of the first candidate.
[0021] According to another aspect of the present disclosure, there is provided an apparatus. The apparatus includes processing circuitry. The processing circuitry may be configured to perform any of the methods described for video decoding / encoding.
[0022] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the methods described for video decoding / encoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0024] Figure 1 is a schematic illustration of an exemplary block diagram of a communication system (100).
[0025] Figure 2 is a schematic illustration of an exemplary block diagram of a decoder.
[0026] Figure 3 is a schematic illustration of an exemplary block diagram of an encoder.
[0027] Figure 4 shows the positions of spatial merge candidates according to an embodiment of the present disclosure.
[0028] Figure 5 shows candidate pairs considered for redundancy checking of spatial merge candidates according to an embodiment of the present disclosure.
[0029] Figure 6 shows motion vector scaling for temporal merge candidates in some examples.
[0030] Figure 7 shows the candidate positions of temporal merge candidates of a current block in some examples.
[0031] Figure 8 shows an example of template matching.
[0032] Figure 9 shows a graph of spatial merge candidates of a current block in some examples.
[0033] Figure 10 shows a graph of a template and reference samples of a block with sub-block motion for using motion information of sub-blocks of a current block in some examples.
[0034] Figure 11 A diagram showing an adaptive reordering (ARMC) reordering process for merge candidates in some examples.
[0035] Figure 12A and Figure 12B A diagram showing block vector adjustment in some examples;
[0036] Figure 13 A flowchart outlining an encoding process according to some embodiments of the present disclosure.
[0037] Figure 14 A flowchart outlining a decoding process according to some embodiments of the present disclosure.
[0038] Figure 15 A schematic illustration of a computer system according to an embodiment. Detailed description
[0039] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example for the application of the disclosed subject matter, video encoders, and video decoders in a streaming environment. The disclosed subject matter can be similarly applied to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0040] The video processing system (100) includes a capture subsystem (113), which can include a video source (101), such as a digital camera device, that creates, for example, an uncompressed video picture stream (102). In an example, the video picture stream (102) includes samples taken by the digital camera device. The video picture stream (102) is depicted as a thick line to emphasize the high data volume when compared to the encoded video data (104) (or the codec video bitstream), and the video picture stream (102) can be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) can include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or the encoded video bitstream) is depicted as a thin line to emphasize the low data volume when compared to the video picture stream (102), and the encoded video data (104) can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1The client subsystems (106) and (108) therein can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an outgoing video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or another rendering device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., video bitstreams) can be encoded according to certain video codec / compression standards. Examples of such standards include ITU-T Recommendation H.265. In an example, a video codec standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC.
[0041] Note that the electronic devices (120) and (130) can include other components (not shown). For example, the electronic device (120) can include a video decoder (not shown), and the electronic device (130) can also include a video encoder (not shown).
[0042] Figure 2 An exemplary block diagram of a video decoder (210) is shown. The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., receiving circuitry). The video decoder (210) can be used to replace Figure 1 the video decoder (110) in the example.
[0043] A receiver (231) may receive one or more coded video sequences, such as included in a bitstream, to be decoded by a video decoder (210). In an embodiment, one coded video sequence is received at a time, where the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data as well as other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not depicted). The receiver (231) may separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and an entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be external to the video decoder (210) (not depicted). In still other applications, a buffer memory (not depicted) may exist external to the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may additionally exist inside the video decoder (210) to, for example, handle playout timing. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer memory (215) may not be needed, or the buffer memory (215) may be small. To make the best use of packet networks such as the Internet, a buffer memory (215) may be needed, which may be relatively large and may advantageously have an adaptive size, and may be implemented at least partially in an operating system or a similar element (not depicted) external to the video decoder (210).
[0044] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the coded video sequences. The categories of these symbols include information for managing the operation of the video decoder (210), and potential information for controlling a presentation device such as a presentation device (212) (e.g., a display screen), which is not part of the electronic device (230) but may be coupled to the electronic device (230), as Figure 2As shown. The control information for presenting the device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set segment (not depicted). The parser (220) may perform parsing / entropy decoding on the received coded video sequence. The coding / decoding of the coded video sequence may be performed according to a video coding / decoding technology or standard and may follow various principles, including variable length coding / decoding, Huffman coding / decoding, arithmetic coding / decoding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set regarding at least one of a subgroup of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Picture (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The parser (220) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.
[0045] The parser (220) may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (215) to create symbols (221).
[0046] Depending on the type of the coded video picture or a part of the coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (221) may involve multiple different units. Which units are involved and the way they are involved may be controlled by the subgroup control information parsed by the parser (220) from the coded video sequence. For the sake of brevity, this subgroup control information flow between the parser (220) and the multiple units below is not depicted.
[0047] In addition to the functional blocks already mentioned, the video decoder (210) may conceptually be subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.
[0048] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives the quantized transform coefficients as symbols (221) and control information from the parser (220), including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) may output a block including sample values, and the block may be input into the aggregator (255).
[0049] In some cases, the output samples of the scaler / inverse transform unit (251) can belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but can use predictive information from a previously reconstructed part of the current picture. Such predictive information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses the reconstructed surrounding information obtained from the current picture buffer (258) to generate a block having the same size and shape as the size and shape of the block being reconstructed. For example, the current picture buffer (258) caches a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the predictive information already generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.
[0050] In other cases, the output samples of the scaler / inverse transform unit (251) can belong to an inter-coded block and potentially to a motion-compensated block. In such a case, the motion-compensation prediction unit (253) can access the reference picture memory (257) to obtain samples for prediction. After motion-compensating the obtained samples according to the symbol (221) related to the block, these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (in this case referred to as residual samples or a residual signal) to generate output sample information. The address within the reference picture memory (257) from which the motion-compensation prediction unit (253) obtains the prediction samples can be controlled by a motion vector, and the motion vector can be obtained by the motion-compensation prediction unit (253) in the form of a symbol (221), and the symbol (221) can have, for example, an X component, a Y component, and a reference picture component. Motion compensation can also include interpolation of sample values obtained from the reference picture memory (257) when using sub-sampled accurate motion vectors, a motion vector prediction mechanism, etc.
[0051] The output samples of the aggregator (255) can undergo various loop-filtering techniques in the loop filter unit (256). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the coded video sequence (also referred to as a coded video bitstream), and the parameters are available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression can also be responsive to meta-information obtained during decoding of a previous (in decoding order) part of the coded picture or coded video sequence, and responsive to previously reconstructed and loop-filtered sample values.
[0052] The output of the loop filter unit (256) can be a sample stream that can be output to a rendering device (212) and stored in a reference picture memory (257) for future inter-picture prediction.
[0053] Once fully reconstructed, some decoded pictures can be used as reference pictures for future prediction. For example, once the decoded picture corresponding to the current picture is fully reconstructed and the decoded picture has been identified as a reference picture (e.g., by a parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting to reconstruct subsequent decoded pictures.
[0054] The video decoder (210) can perform decoding operations according to standards such as the ITU-T H.265 recommendation or a predetermined video compression technique. In the sense that the encoded video sequence follows both the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence can conform to the syntax specified by the used video compression technique or standard. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as tools only available under that profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.
[0055] In an embodiment, the receiver (231) can receive additional (redundant) data along with the encoded video. The additional data can be included as part of the encoded video sequence. The video decoder (210) can use the additional data to decode the data appropriately and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0056] Figure 3 An exemplary block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit system). The video encoder (303) can be used instead of Figure 1The video encoder (103) in the example.
[0057] The video encoder (303) can receive video samples from a video source (301) (not Figure 3 part of the electronic device (320) in the example) that can capture video images to be encoded and decoded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).
[0058] The video source (301) can provide a source video sequence in the form of a stream of digital video samples to be encoded and decoded by the video encoder (303), and the digital video sample stream can have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 YCrCb, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) can be a camera device that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves can be organized as a spatial pixel array, where each pixel can include one or more samples depending on the sampling structure, color space, etc. used. The following description focuses on samples.
[0059] According to an embodiment, the video encoder (303) can encode, decode, and compress pictures of the source video sequence into a coded video sequence (343) in real time or under any other time constraints required. Implementing an appropriate encoding and decoding speed is a function of the controller (350). In some embodiments, the controller (350) controls other functional units as described below and is functionally coupled to other functional units. For the sake of brevity, the couplings are not depicted. The parameters set by the controller (350) can include rate control related parameters (picture skip, quantizer, λ value of rate distortion optimization techniques,...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions belonging to the video encoder (303) optimized for a specific system design.
[0060] In some embodiments, the video encoder (303) is configured to operate in a codec loop. For simplicity of description, in the example, the codec loop may include a source codec (330) (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be coded and decoded and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to the way a (remote) decoder would also create. The reconstructed sample stream (sample data) is input to the reference picture memory (334). Since the decoding of the symbol stream produces bit-exact results regardless of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact between the local encoder and the remote encoder. In other words, as reference picture samples, the sample values that the prediction part of the encoder "sees" are exactly the same as the sample values that the decoder will "see" when using prediction during decoding. The basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, e.g., due to channel errors) is also used in some related technologies.
[0061] The operation of the "local" decoder (333) can be the same as that of a "remote" decoder such as the video decoder (210) that has been described in detail above in conjunction with Figure 2 However, also briefly referring to Figure 2 , since the symbols are available and the encoding of the symbols into a coded video sequence by the entropy codec (345) and the decoding of the symbols by the parser (220) can be lossless, the entropy decoding part including the buffer memory (215) and the parser (220) of the video decoder (210) may not be fully implemented in the local decoder (333).
[0062] In an embodiment, decoder techniques other than the parsing / entropy decoding present in the decoder exist in the corresponding encoder in the same or substantially the same functional form. Thus, the disclosed subject matter focuses on decoder operations. The description of encoder techniques can be simplified because the encoder techniques are contrary to the decoder techniques described comprehensively. More detailed descriptions are provided in some parts below.
[0063] In some examples, during operation, the source codec (330) may perform motion-compensated predictive coding, which predictive-codes an input picture by referring to one or more previously coded pictures designated as "reference pictures" from a video sequence. In such a manner, the codec engine (332) codes the difference between a pixel block of the input picture and a pixel block of the reference picture, and the reference picture may be selected as the prediction reference for the input picture.
[0064] The local video decoder (333) can decode the coded video data of pictures that can be designated as reference pictures based on the symbols created by the source codec (330). The operation of the codec engine (332) can advantageously be lossy processing. When the coded video data can be decoded at the video decoder ( Figure 3 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that the video decoder performs on the reference pictures and can store the reconstructed reference pictures in the reference picture memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference pictures, which has the same content (no transmission errors) as the reconstructed reference pictures to be obtained by the remote video decoder.
[0065] The predictor (335) can perform a prediction search for the codec engine (332). That is, for a new picture to be coded, the predictor (335) can search the reference picture memory (334) for sample data (as a candidate reference pixel block) or specific metadata such as reference picture motion vectors, block shapes, etc. that can be used as an appropriate prediction reference for the new picture. The predictor (335) can operate on a per-pixel block basis of the sample blocks to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (335), the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory (334).
[0066] The controller (350) can manage the coding and decoding operations of the source codec (330), including, for example, setting parameters and subgroup parameters for encoding the video data.
[0067] The outputs of all the above-mentioned functional units can be entropy-coded in the entropy codec (345). The entropy codec (345) converts these symbols into a coded video sequence by losslessly compressing the symbols generated by various functional units according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.
[0068] The transmitter (340) can cache the coded video sequence created by the entropy codec (345) in preparation for transmission via the communication channel (360), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown).
[0069] The controller (350) may manage the operation of the video encoder (303). During encoding and decoding, the controller (350) may assign a certain encoding and decoding picture type to each encoded and decoded picture, which may affect the encoding and decoding techniques that can be applied to the corresponding picture. For example, pictures can generally be assigned to one of the following picture types:
[0070] An intra picture (I picture) can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh (IDR) pictures.
[0071] A predictive picture (P picture) can be encoded and decoded using inter prediction or intra prediction that utilizes motion vectors and reference indices to predict the sample values of each block.
[0072] A bi-predictive picture (B picture) can be encoded and decoded using inter prediction or intra prediction that utilizes two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predictive picture can use more than two reference pictures and associated metadata for reconstructing a single block.
[0073] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and encoded and decoded on a block-by-block basis. These blocks can be predictively encoded and decoded with reference to other (already encoded and decoded) blocks, which are determined by the encoding and decoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture can be non-predictively encoded, or blocks of an I picture can be predictively encoded (spatial prediction or intra prediction) with reference to the encoded and decoded blocks of the same picture. Pixel blocks of a P picture can be predictively encoded and decoded with reference to one previous encoded and decoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture can be predictively encoded and decoded with reference to one or two previous encoded and decoded reference pictures via spatial prediction or via temporal prediction.
[0074] The video encoder (303) may perform encoding and decoding operations according to a standard such as the ITU-T H.265 recommendation or a predetermined video encoding and decoding technique. In the operation of the video encoder (303), the video encoder (303) may perform various compression operations, including predictive encoding and decoding operations that utilize temporal redundancy and spatial redundancy in the input video sequence. Thus, the encoded and decoded video data may conform to the syntax specified by the used video encoding and decoding technique or standard.
[0075] In an embodiment, the transmitter (340) may send additional data along with the encoded video. The source codec (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set segments, etc.
[0076] Video may be captured as a plurality of source pictures (video pictures) in a time series. Intra picture prediction (commonly abbreviated as intra prediction) utilizes the spatial correlation within a given picture, while inter picture prediction utilizes the (temporal or other) correlation between pictures. In an example, a particular picture in encoding / decoding, which is referred to as the current picture, is partitioned into blocks. In the case where a block in the current picture is similar to a reference block in a reference picture that has been previously coded and decoded and is still cached in the video, the block in the current picture may be coded and decoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector may have a third dimension identifying the reference picture.
[0077] In some embodiments, bidirectional prediction techniques may be used in inter picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). The block in the current picture may be coded and decoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.
[0078] In addition, merge mode techniques may be used in inter picture prediction to improve coding and decoding efficiency.
[0079] According to some embodiments of the present disclosure, predictions such as inter - picture prediction and intra - picture prediction are performed on a block - by - block basis. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into Coding Tree Units (CTUs) for compression, and CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three Coding Tree Blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more Coding Units (CUs) in a quadtree. For example, a 64×64 - pixel CTU can be divided into one 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. In an example, each CU is analyzed to determine the prediction type for that CU, such as an inter - prediction type or an intra - prediction type. According to temporal and / or spatial predictability, a CU is divided into one or more Prediction Units (PUs). Generally, each PU includes a luminance Prediction Block (PB) and two chrominance PBs. In an embodiment, the prediction operation in coding / decoding (encoding / decoding) is performed on a prediction - block basis. Taking the luminance prediction block as an example of the prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0080] Note that any suitable technology can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In an embodiment, one or more integrated circuits can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210).
[0081] Various inter - frame prediction modes can be used in video coding and decoding. For example, in VVC, for an inter - frame prediction CU, the motion parameters can include an MV, one or more reference picture indices, a reference picture list usage index, and additional information of certain coding and decoding features to be used for generating inter - frame prediction samples. The motion parameters can be signaled explicitly or implicitly. When coding a CU in skip mode, the CU can be associated with a PU and does not have valid residual coefficients, no coding motion vector change or MV difference (e.g., MVD) or reference picture index. A merge mode can be specified, in which the motion parameters of the current CU are obtained from adjacent CUs, including spatial candidates and / or temporal candidates, and optionally including additional information such as that introduced in VVC. The merge mode can be applied to CUs for inter - frame prediction, not only for the skip mode. In an example, an alternative to the merge mode is the explicit transmission of motion parameters, in which the MV, the corresponding reference picture indices for each reference picture list, and the reference picture list usage flag and other information are signaled explicitly for the CU.
[0082] In an embodiment, for example, in VVC, the VVC Test Model (VTM) reference software includes one or more improved inter-prediction codec tools, and the improved inter-prediction codec tools include: extended merge prediction, Merge Motion Vector Difference (MMVD) mode, Adaptive Motion Vector Prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, Subblock-based Temporal Motion Vector Prediction (SbTMVP), Adaptive Motion Vector Resolution (AMVR), motion field storage (1 / 16th luminance sample MV storage and 8×8 motion field compression), Bi-prediction with CU-level Weight (BCW), Bi-Directional Optical Flow (BDOF), Prediction Refinement using Optical Flow (PROF), Decoder side Motion Vector Refinement (DMVR), Combined Inter and Intra Prediction (CIIP), Geometric Partitioning Mode (GPM), etc. Inter-prediction and related methods are described in detail below.
[0083] In some examples, extended merge prediction can be used. In an example, for example, in VTM4, the merge candidate list is constructed by sequentially including the following five types of candidates: spatial motion vector predictor (MVP) from spatially adjacent CUs, temporal MVP from co-located CUs, history-based MVP (HMVP) from a First-In-First-Out (FIFO) table, pairwise-averaged MVP, and zero MV.
[0084] The size of the merge candidate list can be signaled in the slice header. In the example, in VTM4, the maximum allowed size of the merge candidate list is 6. For each CU coded in merge mode, a truncated unary binarization (TU) can be used to code the index of the best merge candidate (e.g., the merge index). The first bin of the merge index can be coded with context (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other bins.
[0085] Some examples of the generation process of merge candidates for each category are provided below. In an embodiment, spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC can be the same as the derivation of spatial merge candidates in HEVC. In the example, up to four merge candidates are selected from the candidates at the Figure 4 depicted positions.
[0086] Figure 4 The positions of spatial merge candidates according to an embodiment of the present disclosure are shown. Referring to Figure 4 , the order of derivation is B1, A1, B0, A0, and B2. Position B2 is considered only when any of the CUs at positions A0, B0, B1, and A1 are unavailable (e.g., because the CU belongs to another slice or another tile) or are intra-coded. After adding the candidate at position A1, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.
[0087] To reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the Figure 5 pairs linked by arrows in
[0088] Figure 5 are considered, and a candidate is added to the candidate list only if the corresponding candidates used for the redundancy check do not have the same motion information. Figure 5 The candidate pairs considered for the redundancy check of spatial merge candidates according to an embodiment of the present disclosure are shown. Referring to Figure 5 , the pairs linked by the corresponding arrows include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Thus, the candidates at positions B1, A0, and / or B2 can be compared with the candidate at position A1, and the candidates at positions B0 and / or B2 can be compared with the candidate at position B1.
[0089] In an embodiment, temporal candidates are derived as follows. In the example, only one temporal merge candidate is added to the candidate list. Figure 6Shows motion vector scaling for temporal merge candidates in some examples. To derive the temporal merge candidates for the current CU (611) in the current picture (601), a scaled MV (621) can be derived based on the collocated CU (612) belonging to the collocated reference picture (604) (e.g., Figure 6 as indicated by the dashed line in Figure 6 . The reference picture list for deriving the collocated CU (612) can be explicitly signaled in the slice header. The scaled MV (621) for the temporal merge candidates can be obtained, as
[0090] Figure 7 indicated by the dashed line in
[0091] . The scaled MV (621) can be scaled from the MV of the collocated CU (612) using the picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (602) of the current picture (601) and the current picture (601). The POC distance td can be defined as the POC difference between the collocated reference picture (604) of the collocated picture (603) and the collocated picture (603). The reference picture index of the temporal merge candidates can be set to zero. The collocated picture is a reference picture that serves as a source picture for temporal motion information derivation. The collocated picture can be identified in one of two lists called list 0 or list 1. In some examples, the encoder can determine the collocated picture and signal the collocated picture using appropriate syntax techniques.
[0092] In some examples, template matching (TM) techniques for refining motion at the decoder side can be used in video / image coding and decoding (e.g., VVC, ECM, etc.) to further improve compression efficiency. In TM mode, the motion vector (MV) can be refined by constructing a template (e.g., the current template) of a block (e.g., the current block) in the current picture and determining the closest match between the template of the block in the current picture and multiple possible templates (e.g., multiple possible reference templates) in the reference picture. In an embodiment, the template of the block in the current picture may include the left adjacent reconstructed samples and the upper adjacent reconstructed samples of the block.
[0093] Figure 8 An example of template matching (800) is shown. TM can be used to derive the motion information of the current coding unit (CU) (e.g., the current block) (801) by determining the closest match between a template (e.g., the current template) (821) of the current CU (801) in the current picture (810) and a template (e.g., the reference template) (e.g., one of the multiple possible templates is the template (825)) in the reference picture (811). The template (821) of the current CU (801) can have any suitable shape and any suitable size.
[0094] In an embodiment, the template (821) of the current CU (801) includes a top template (822) and a left template (823). Each of the top template (822) and the left template (823) can have any suitable shape and any suitable size.
[0095] The top template (822) may include samples in one or more top adjacent blocks of the current CU (801). In an example, the top template (822) includes one or more rows of samples above the current CU (801). The left template (823) may include samples in one or more left adjacent blocks of the current CU (801). In an example, the left template (823) includes one or more columns of samples to the left of the current CU (801).
[0096] Each possible template (e.g., template (825)) among multiple possible templates in the reference picture (811) corresponds to the template (821) in the current picture (810). In an embodiment, the initial MV (802) points from the current CU (801) to the reference block (803) in the reference picture (811). Each possible template (e.g., template (825)) among multiple possible templates in the reference picture (811) and the template (821) in the current picture (810) may have the same shape and the same size. For example, the template (825) of the reference block (803) includes the top template (826) in the reference picture (811) and the left template (827) in the reference picture (811). The top template (826) may include samples above the reference block (803). The left template (827) may include samples to the left of the reference block (803).
[0097] The TM cost can be determined based on a pair of templates, e.g., the template (e.g., the current template) (821) and the template (e.g., the reference template) (825). The TM cost can indicate the match between the template (821) and the template (825). The optimized MV (or the final MV) can be determined based on a search performed around the initial MV (802) of the current CU (801) within the search range (815). The search range (815) can have any suitable shape and any suitable number of reference samples. In an example, the search range (815) in the reference picture (811) includes a [-L, L]-pixel range, where L is a positive integer such as 8 (e.g., 8 samples). For example, a difference (e.g., [0, 1]) is determined based on the search range (815), and the intermediate MV is determined by the sum of the initial MV (802) and the difference (e.g., [0, 1]). The corresponding template and the intermediate reference block in the reference picture (811) can be determined based on the intermediate MV. The TM cost can be determined based on the template (821) and the intermediate template in the reference picture (811). The TM cost can correspond to the difference determined based on the search range (815) (e.g., [0, 0], [0, 1], etc. corresponding to the initial MV (802)). In an example, the difference corresponding to the minimum TM cost is selected, and the optimized MV is the sum of the difference corresponding to the minimum TM cost and the initial MV (802). As described above, TM can obtain the final motion information (e.g., the optimized MV) based on the initial motion information (e.g., the initial MV 802).
[0098] In Figure 8In an example, a better MV can be searched within a search range, such as [-8 pixels, +8 pixels], around the initial motion vector of the current CU. In some examples (such as ECM), template matching with several modifications is also adopted. In an example, the search step size is determined by the AMVR mode. In another example, TM can be cascaded with bilateral matching processing. In another example, template matching is also used to reorder the indices of candidates in the merge candidate list and the AMVP candidate list.
[0099] In some examples, the merge candidate list may include one or more non-neighboring spatial merge candidates. In an example, the non-neighboring spatial merge candidates are inserted after the TMVP in the merge candidate list.
[0100] Figure 9 A diagram showing the spatial merge candidates of the current block (901) in some examples is presented. In Figure 9 an example, positions 1-5 are the neighboring spatial merge candidate positions of the current block (901) (also referred to as the current coding / decoding block), and positions 6-23 can be the non-neighboring spatial merge candidate positions of the current block (901). For example, the motion information at one or more of positions 6-23 can be inserted as one or more non-neighboring spatial merge candidates into the merge candidate list. In some examples, the distance between the non-neighboring spatial candidates and the current coding / decoding block (901) is based on the width and height of the current block (the current coding / decoding block) (901).
[0101] In some examples, a technique called adaptive reordering of merge candidates using template matching can be used. In an example, the candidate list is constructed by adding candidates in order, such as a regular merge candidate list, a template matching (TM) merge candidate list, a bilateral matching (BM) merge candidate list, etc. For example, candidates can be added in the following order: the spatial MVP (SMVP) from spatially neighboring adjacent CUs (e.g., Figure 9 one or more of the neighboring spatial merge candidates at positions 1 to 5 in ), the temporal MVP Figure 7 (TMVP) from the co-located CU (e.g., Figure 9 the temporal MVP at C0 or C1 in ), the non-adjacent MVP (NA-MVP) from spatially non-neighboring CUs (e.g.,
[0102] In some examples, candidates in a candidate list are reordered according to the technique of an Adaptive Reordering of Merge Candidates (ARMC) method that utilizes template matching. ARMC can reorder MV candidates in the candidate list based on TM cost. In an example, a candidate list that may include various candidate types such as TMVP type, NA-SMVP type, etc. is reordered in ascending order by using the TM cost. In another example, candidates of a specific type such as TMVP type or NA-SMVP type in the candidate list can be reordered in ascending order using the TM cost. This reordering method can be applied not only to a regular merge mode and a template matching (TM) merge mode, but also to an affine merge mode (excluding SbTMVP candidates). In an example, for the TM merge mode, merge candidates are reordered before the refinement process.
[0103] In some examples, after constructing a merge candidate list, the merge candidates in the merge candidate list are divided into several subgroups. For example, the subgroup size is set to 5. Then, based on the cost value based on template matching, the merge candidates in each subgroup are reordered in ascending order. For simplicity, in the example, the merge candidates in the last subgroup rather than the first subgroup are not reordered.
[0104] In some examples, for sub-block-based merge candidates with a sub-block size equal to Wsub×Hsub, the upper template includes several sub-templates with a size of Wsub×1, and the left template includes several sub-templates with a size of 1×Hsub.
[0105] Figure 10 A diagram showing a template of a block with sub-block motion and a reference sample of the template using the motion information of sub-blocks of a current block in some examples. The motion information of the sub-blocks in the first row and the first column of the current block is used to obtain the reference sample of each sub-template.
[0106] In Figure 10 an example, the template (T) of the current block (1010) includes the row (1011) above the current block (1010) and the column (1012) to the left of the current block (1010). The row (1011) is also referred to as the upper template, and the column (1012) is also referred to as the left template. The current block (1010) includes multiple sub-blocks, such as Figure 10 16 sub-blocks in the example. The sub-blocks in the first row of the current block (1010) are shown as A, B, C, and D, and the sub-blocks in the first column of the current block (1010) are shown as A, E, F, and G.
[0107] In Figure 10In this example, the reference picture includes a co-located block (1020) of the current block (1010). The motion information of sub-blocks A, B, C, D, E, F, and G can be used to derive a reference sub-block in the reference picture, for example, by Figure 10 A ref, B ref, C ref, Dref, Eref, F ref and G ref are shown in FIG. The upper sub-template of A ref, B ref, C ref and D ref can constitute the upper reference template of the upper template (1011), and the left sub-template of Aref, Eref, F ref and G ref can constitute the left reference template of the left template (1012).
[0108] In some examples, in order to improve coding efficiency, MV candidate reordering is applied not only to a single TMVP type or a single NA-MVP type, but also to a mixed candidate type including NA-MVP, HMVP and PAMVP types. In some examples, the merged candidate list includes several repeated zero MV candidates. According to aspects of the present disclosure, several repeated zero MV candidates (for example, for all available reference lists, the motion vector is (0,0)) can be sorted to an early position in the candidate list. In some examples, all zero MV candidates are excluded from the ARMC reordering process. After ARMC reordering, repeated zero MVs are filled at the end of the merged candidate list, so the zero MV candidate stays at a later position in the candidate list.
[0109] Figure 11 A diagram showing the ARMC reordering process in some examples. Figure 11 In the example, all zero MV candidates are excluded from the ARMC reordering process. For example, the candidate list (1110) before ARMC reordering includes candidates in the order of Cand0, Cand1, Cand2, Cand3, Cand4, Cand5, Cand6, Cand7, Cand8, and Cand9. Candidates Cand2 to Cand6 are zero MV candidates and are excluded from the ARMC reordering. ARMC reordering is performed on Cand0, Cand1, Cand7, Cand8, and Cand9 to obtain the first 5 candidates in the reordered candidate list (1120), for example, as shown by NewCand0, NewCand1, NewCand2, NewCand3, and NewCand4. Then, zero MV candidates are filled at the end of the reordered candidate list (1120), for example, as shown by NewCand5, NewCand6, NewCand7, NewCand8, and NewCand9.
[0110] Note that block-based compensation can be used for both inter prediction and intra prediction. For inter prediction, block-based compensation from different pictures is referred to as motion compensation. For example, in intra prediction, block-based compensation can also be done based on previously reconstructed regions within the same picture. Block-based compensation from reconstructed regions within the same picture is referred to as intra picture block compensation, Current Picture Referencing (CPR), or Intra Block Copy (IBC). A displacement vector indicating the offset between the current block and a reference block (also referred to as a prediction block) in the same picture is called a Block Vector (BV), based on which the current block can be encoded / decoded. Different from a motion vector in motion compensation which can be any value (positive or negative in the x or y direction), the BV has some constraints to ensure that the reference block is available and has been reconstructed. Additionally, in some examples, for parallel processing considerations, some reference regions that are tile boundaries, slice boundaries, or wavefront trapezoid boundaries are excluded.
[0111] In some examples, the IBC mode can be used to significantly improve the encoding and decoding efficiency of screen content material. Generally, the IBC mode can be implemented as a block-level encoding and decoding mode. On the encoder side, the encoder can perform block matching (BM) to find the best block vector for each CU. In some examples, on the encoder side, hash-based motion estimation (also referred to as hash-based search) is performed for CUs in the IBC mode. The encoder can perform rate distortion (RD) checking on blocks with a width or height no greater than 16 luma samples. For non-merge modes, a hash-based search is first used to perform the block vector search. If the hash search does not return a valid candidate, a block-matching based local search can be performed.
[0112] In some examples, in the hash-based search, the hash key match (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes in the current picture. The hash key calculation for each position in the current picture can be based on 4x4 sub-blocks. For a current block of a larger size, when all hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions, it can be determined that the hash key matches the hash key of the reference block. If it is found that the hash keys of multiple reference blocks match the hash key of the current block, the block vector cost for each of the matching reference blocks can be calculated, and then the one matching reference block with the minimum cost is selected as the result of the hash-based search.
[0113] In some examples, the block matching search searches a region local to the current block. For example, in the block matching search, the search range is set to cover both the previous CTU and the current CTU.
[0114] The encoding and decoding of block vectors can be explicit or implicit. In the explicit mode, the BV difference between the block vector and its predictor is signaled. In the implicit mode, the block vector is recovered from a predictor (referred to as the block vector predictor) in a manner similar to the motion vector in the merge mode, without using the BV difference. In some examples, the explicit mode may be referred to as the non-merge BV prediction mode or the IBC AMVP mode. In some examples, the implicit mode may be referred to as the merge BV prediction mode, the IBC merge mode, or the IBC skip mode.
[0115] There may be variations for the IBC mode. In an example, the IBC mode is regarded as a third mode different from the intra prediction mode and the inter prediction mode. Thus, the BV prediction in the implicit mode (or IBC merge mode) and the explicit mode (IBC AMVP mode) is separated from the regular inter mode. In some examples, a separate merge candidate list may be defined for the IBC mode, and in the IBC mode, the entries in the separate merge candidate list are BVs. Similarly, in an example, the BV prediction candidate list in the IBC explicit mode (IBC AMVP mode) only includes BVs. The general rule applied to the two lists (i.e., the separate merge candidate list for the IBC merge mode and the BV prediction candidate list for the IBC AMVP mode) is that, in terms of the candidate derivation process, these two lists may follow the same logic as the merge candidate list used in the regular merge mode (for inter prediction) or the advanced motion vector prediction (AMVP) predictor list used in the regular AMVP mode (for inter prediction). For example, five spatially adjacent positions (e.g., Figure 2 A0, A1 and B0, B1, B2 in
[0116] are accessed for the IBC merge mode, such as the HEVC or VVC inter merge mode, to obtain the separate merge candidate list for the IBC merge mode.
[0117] In some examples, for the IBC skip / merge mode, a merge candidate index may be signaled to indicate which one of the block vectors from adjacent candidate IBC encoded / decoded blocks in the merge candidate list is used as the BV predictor to predict the current block. In some examples, the merge candidate list may include spatially-based, history-based motion vector prediction (HMVP), and paired candidates.
[0118] In some examples, a technique called Reconstruction Reordering IBC (RR-IBC) or RR-IBC mode is allowed to be used for IBC encoding and decoding blocks. When applying RR-IBC, samples in the reconstructed block are flipped according to the flip type of the current block. In an example, on the encoder side, the original block is flipped before motion search and residual calculation, while the prediction block is obtained without flipping. On the decoder side, the reconstructed block is flipped back to restore the original block. In some examples, the flip type can be horizontal flip or vertical flip.
[0119] In some examples, to better utilize symmetry, a flip-aware block vector (BV) adjustment method is applied to refine block vector candidates.
[0120] Figure 12A and Figure 12B A diagram showing BV adjustment in some examples.
[0121] Respectively, in Figure 12A and Figure 12B , (x n , y n ) represents the coordinates of the central sample of an adjacent block, and (x c , y c ) represents the coordinates of the central sample of the current block. BV n represents the BV of the adjacent block, and BV c represents the BV of the current block.
[0122] In Figure 12A 's example, instead of directly inheriting the BV from the adjacent block, the adjacent block is encoded and decoded with a horizontal flip, and the horizontal component of BV n (represented by BV n h ) is used to calculate the horizontal component of BV c (represented by BV c h ) by adding a motion shift to it because the adjacent block is encoded and decoded with a horizontal flip. For example, the calculation can be represented by BV c h = 2(x n - x c ) + BV n h.
[0123] In Figure 12B 's example, the adjacent block is encoded and decoded with a vertical flip. The vertical component of BV n (BV n v ) is used to calculate the vertical component of BV c (represented by BV c vrepresented), since neighboring blocks are encoded and decoded with a vertical flip. For example, the calculation can be performed by BV c v = 2(y n - y c ) + BV n v represented.
[0124] In some examples, before adding a potential candidate to the merge candidate list, a redundancy check or a similarity check (e.g., using "mvdSimilarityThresh") is applied to determine whether the potential candidate has the same or similar motion information as the existing candidates in the merge candidate list. The potential candidates having the same or similar motion information as the existing candidates can be excluded from the merge candidate list.
[0125] In some examples, the current block inherits additional information other than the motion information from the candidates in the candidate list. The information other than the motion information that can be passed from the candidates in the candidate list to the current block is referred to as other inherited information. For example, a Local Lumination Compensation (LIC) flag can be passed from the candidates in the candidate list to the current block, and the LIC flag can be included as a condition in the redundancy and similarity checks. In the example, the LIC flag is other inherited information.
[0126] In some examples, the inherited information refers to the part of the information of the candidate (for the current block) that is passed to the current block when the candidate is selected for prediction of the current block, so the current block inherits that part of the information of the candidate. In some examples, the inherited information is also referred to as inheritable parameters, and includes motion-based inheritable parameters and non-motion-based inheritable parameters. In some examples, the motion-based inheritable parameters include motion vectors, reference lists, reference indices, block vectors, etc.; the non-motion-based inheritable parameters can include BCW indices, interpolation filter selections, LIC flags, etc.
[0127] Some aspects of the present disclosure provide techniques for redundancy and similarity checking. In some examples, redundancy and similarity checking may include conditional checking of the current codec block (e.g., a potential candidate) and other inherited information (e.g., non-motion-based inheritable parameters) of each existing candidate in the candidate list. For example, to determine whether a potential candidate is redundant and should not be included in the candidate list, not only the similarity of the (multiple) MVs between the potential candidate and the existing candidates in the candidate list is checked, but also the similarity of other inherited information (e.g., non-motion-based inheritable parameters) between the potential candidate and the existing candidates in the candidate list is checked. In some examples, when both of the above two checks (1. similarity of MVs and 2. similarity of other inherited information (e.g., non-motion-based inheritable parameters)) return true (e.g., both the MV and other inherited information of the potential candidate are similar to the existing candidates in the candidate list), the potential candidate will not be filled into the candidate list. In an example, when the potential candidate has the same MV as an existing candidate but different other inherited information from the existing candidate, the potential candidate can still be added to the candidate list.
[0128] For example, an encoder / decoder may build a candidate list for a current block in a current picture, the candidate list including at least a first candidate and a second candidate, the first candidate and the second candidate having redundant information based on motion-based inheritable parameters and non-redundant information based at least on non-motion-based inheritable parameters.
[0129] In some embodiments, the inherited information includes a BCW index for regular merge candidate list construction.
[0130] In some examples, bidirectional prediction with CU-level weights (BCW) can be used to weight predictions from different reference pictures differently. The BCW technique is designed to predict a block by taking a weighted average of two motion-compensated prediction blocks. BCW is different from a technique called Weighting Prediction (WP), which indicates weights at the slice level. BCW can signal weight information at the CU level by using an index denoted as bcwIdx. The index can point to a weight selected from a list of predefined candidate weights. In some examples, the list includes 5 predefined candidate weights to be selected for reference pictures in reference list 1 (also known as reference picture list 1), such as {-2, 3, 4, 5, 10} / 8, where -2 / 8 and 10 / 8 can be used to reduce the negative correlation noise between prediction blocks of bidirectional prediction. When using forward and backward reference pictures in two reference lists, the list of predefined candidate weights can be reduced to {3, 4, 5} / 8 to achieve a better trade-off between performance and complexity. In some examples, due to the application of the unity gain constraint, once the weight (denoted as w) pointed to by bcwIdx corresponding to reference table 1 is determined, the weight corresponding to other reference tables can be calculated as 1–w. In an example, each luma / chroma prediction sample of BCW is calculated as Equation (1): P BCW =(8(1 - w)×P 0 +8w×P 1 +4)>>3 Equation (1) where P BCW is the final prediction of the current block sample (the sample in the current block), and P 0 and P 1 are the prediction samples pointed to by the motion vectors from the reference pictures in list 0 (also known as reference picture list 0) and list 1 (also known as reference picture list 1), respectively. In some examples, BCW is enabled only for bidirectional prediction CUs having at least 256 luma samples and with WP turned off. BCW is also extended to the affine AMVP mode.
[0131] In some examples, the subsequent CU buffer bcwIdx in the same frame can be used to perform spatial motion merging for the regular merge mode or for the affine merge mode. In an example, when a spatially adjacent merge candidate is bi-predicted and the current CU selects this candidate (this spatially adjacent merge candidate), all reference indexes and motion vectors including its bcwIdx (or in the case of inheriting the affine merge mode, the Control Point Motion Vector (CPMV)) are inherited by the current CU. In some examples, the only exception where the bcwIdx is not inherited occurs when the current CU enables the CIIP flag. In the case of constructing the affine merge mode, the bcwIdx is inherited from one associated with the top-left control point motion vector (or when not using the top-left control point motion vector, inherited from one associated with the top-right control point motion vector). It should be noted that when the inferred bcwIdx points to a non-0.5 weight, both decoder-side motion vector refinement (DMVR) and bidirectional optical flow (BDOF) are turned off.
[0132] In some examples, the BCW index is in the inheritance information, and the similarity and redundancy checks include conditions on the BCW index. Specifically, when the first check on the motion information between a potential candidate and an existing candidate in the candidate list returns true (e.g., the potential candidate has the same MV as the existing candidate), but the second check on the BCW index between the potential candidate and the existing candidate returns false (e.g., the BCW index of the potential candidate is different from that of the existing candidate), the potential candidate is not considered redundant relative to the existing candidate. In an example, when the potential candidate is not redundant with respect to any existing candidate in the candidate list, the potential candidate can be appropriately inserted into the candidate list.
[0133] In some embodiments, the inheritance information includes the reconstruction reordering type (e.g., flip type) for IBC merge candidate list construction for IBC. In some examples, the similarity and redundancy checks include conditions on the reconstruction reordering type. Specifically, when the first check on the motion information between a potential candidate and an existing candidate in the candidate list returns true (e.g., the potential candidate has the same MV as the existing candidate), but the second check on the reconstruction reordering type between the potential candidate and the existing candidate returns false (e.g., one of the potential candidate and the existing candidate has a vertical flip and the other has a horizontal flip), the potential candidate is not considered redundant relative to the existing candidate. In an example, when the potential candidate is not redundant with respect to any existing candidate in the candidate list, the potential candidate can be appropriately inserted into the candidate list.
[0134] In some embodiments, for the above-mentioned second verification (similarity of other inheritance information between a potential candidate and an existing candidate), it is determined that the second verification returns true only when the inheritance information between the potential candidate and the existing candidate is exactly the same. In the example, when the inheritance information includes a BCW index and the BCW index of the potential candidate is exactly the same as that of the existing candidate, the second verification returns true.
[0135] In some embodiments, for the above-mentioned second verification (similarity of other inheritance information between a potential candidate and an existing candidate), it is determined that the second verification returns true only when the difference in the inheritance information between the potential candidate and the existing candidate is within a threshold. In the example, when the inheritance information includes a BCW index and the difference in the BCW index between the potential candidate and the existing candidate is within a range such as [-1, 1], the second verification returns true.
[0136] Figure 13 A flowchart outlining a process (1300) according to embodiments of the present disclosure is shown. The process (1300) can be used in a video encoder. In various embodiments, the process (1300) is executed by processing circuitry, such as processing circuitry that performs the functions of video encoder (103), processing circuitry that performs the functions of video encoder (303), and so on. In some embodiments, the process (1300) is implemented as software instructions, so when the processing circuitry executes the software instructions, the processing circuitry executes the process (1300). The process starts at (S1301) and proceeds to (S1310).
[0137] At (S1310), it is determined to construct a candidate list for prediction of a current block in a current picture.
[0138] At (S1320), potential candidates to be added to the candidate list are determined, and the candidate list includes one or more existing candidates that have already been added to the candidate list.
[0139] At (S1330), a first verification of motion-based inheritable parameters between the potential candidate and a first existing candidate among the one or more existing candidates is performed to obtain a first verification result, and the first verification result indicates whether the potential candidate has redundant information of motion-based inheritable parameters relative to the first existing candidate.
[0140] At (S1340), a second verification of non-motion-based inheritable parameters between the potential candidate and the first existing candidate is performed to obtain a second verification result, and the second verification result indicates whether the potential candidate has redundant information of non-motion-based inheritable parameters relative to the first existing candidate.
[0141] At (S1350), when the first verification result indicates that the potential candidate has redundant information of motion-based inheritable parameters relative to the first existing candidate and the second verification result indicates that the potential candidate has redundant information of non-motion-based inheritable parameters relative to the first existing candidate, it is determined that the potential candidate is redundant for the first existing candidate.
[0142] In some examples, when the current block is in the inter prediction mode, the motion-based inheritable parameters include motion vectors (MVs). The non-motion-based inheritable parameters include bi-prediction with block coding unit level weights (BCW) indices.
[0143] In some examples, when the current block is in the intra block copy (IBC) mode, the motion-based inheritable parameters include block vectors (BVs). The non-motion-based inheritable parameters include the reconstruction reordering type for IBC for the IBC mode. In an example, the non-motion-based inheritable parameters include one of vertical flipping and horizontal flipping.
[0144] In some examples, the second verification result indicates whether the non-motion-based inheritable parameters of the potential candidate are exactly the same as those of the first existing candidate.
[0145] In some examples, the second verification result indicates whether the difference between the non-motion-based inheritable parameters between the potential candidate and the first existing candidate is within a certain range.
[0146] Then, the process proceeds to (S1399) and terminates.
[0147] The process (1300) can be appropriately adjusted. The steps in the process (1300) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.
[0148] Figure 14 A flowchart outlining the process (1400) according to an embodiment of the present disclosure is shown. The process (1400) can be used in a video decoder. In various embodiments, the process (1400) is executed by a processing circuitry, such as a processing circuitry that executes the functions of the video decoder (110), a processing circuitry that executes the functions of the video decoder (210), etc. In some embodiments, the process (1400) is implemented by software instructions, so when the processing circuitry executes the software instructions, the processing circuitry executes the process (1400). The process starts at (S1401) and proceeds to (S1410).
[0149] At (S1410), a coded video bitstream including coded information of one or more pictures is received.
[0150] At (S1420), a candidate list for the current block in the current picture is constructed. The candidate list includes at least a first candidate and a second candidate. The first candidate and the second candidate have redundant information based on motion-based inheritable parameters and non-redundant information based on non-motion-based inheritable parameters.
[0151] At (S1430), a specific candidate is selected from the candidate list.
[0152] At (S1440), the current block is reconstructed based on the specific candidate.
[0153] In some examples, potential candidates to be added to the candidate list are determined. A first check of the motion-based inheritable parameters between the potential candidate and the first candidate, which is an existing candidate in the candidate list, is performed to obtain a first check result. The first check result indicates whether the potential candidate has redundant information of the motion-based inheritable parameters the same as that of the first candidate. At least a second check of the non-motion-based inheritable parameters between the potential candidate and the first candidate is performed to obtain a second check result. The second check result indicates whether the potential candidate has redundant information of the non-motion-based inheritable parameters the same as that of the first candidate. When the first check result indicates that the potential candidate has redundant information of the motion-based inheritable parameters the same as that of the first candidate and the second check result indicates that the potential candidate has redundant information of the non-motion-based inheritable parameters the same as that of the first candidate, it is determined that the potential candidate is redundant with respect to the first candidate.
[0154] In some examples, when the current block is in the inter-frame prediction mode, the motion-based inheritable parameters include a motion vector (MV). The non-motion-based inheritable parameters include bidirectional prediction with a coding / decoding unit-level weight (BCW) index.
[0155] When the current block is in the intra-block copy (IBC) mode, the motion-based inheritable parameters include a block vector (BV). The non-motion-based inheritable parameters include a reconstruction reordering type for IBC in the IBC mode. In an example, the non-motion-based inheritable parameters include at least one of a vertical flip and a horizontal flip.
[0156] In some examples, when the non-motion-based inheritable parameters of the potential candidate are exactly the same as those of the first candidate, it is determined that the potential candidate has redundant information of the non-motion-based inheritable parameters the same as that of the first candidate.
[0157] In some examples, when the difference in the non-motion-based inheritable parameters between the potential candidate and the first candidate is within a certain range, it is determined that the potential candidate has redundant information of the non-motion-based inheritable parameters the same as that of the first candidate.
[0158] Then, the process proceeds to (S1499) and terminates.
[0159] The process (1400) can be adjusted appropriately. Steps in the process (1400) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.
[0160] According to aspects of the present disclosure, a method for processing visual media data is provided. In this method, a bitstream of visual media data is processed according to format rules. For example, the bitstream can be a bitstream decoded / encoded by any one of the decoding and / or encoding methods described herein. The format rules can specify one or more constraints of the bitstream and / or one or more processes to be performed by a decoder and / or an encoder.
[0161] In an example, the bitstream includes an index pointing to a specific candidate in a candidate list for prediction of a current block in a current picture. The format rules specify that a first check result is obtained by performing a first check of motion-based inheritable parameters between a potential candidate and a first existing candidate among one or more existing candidates in the candidate list, the first check result indicating whether the potential candidate has redundant information of motion-based inheritable parameters relative to the first existing candidate, a second check result is obtained by performing a second check of non-motion-based inheritable parameters between the potential candidate and the first existing candidate, the second check result indicating whether the potential candidate has redundant information of non-motion-based inheritable parameters relative to the first existing candidate, and when the first check result indicates that the potential candidate has redundant information of motion-based inheritable parameters relative to the first existing candidate and the second check result indicates that the potential candidate has redundant information of non-motion-based inheritable parameters relative to the first existing candidate, the potential candidate is redundant for the first existing candidate.
[0162] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 15 A computer system (1500) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0163] The computer software can be encoded and decoded using any suitable machine code or computer language, which can be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. or executed through interpretation, microcode execution, etc.
[0164] The instructions can be executed on various types of computers or their components, including for example personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0165] Figure 15 The components shown for the computer system (1500) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the present disclosure. The configuration of the components should also not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (1500).
[0166] The computer system (1500) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs made by one or more human users through, for example, tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, taps), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media not necessarily directly related to conscious human input, such as, for example, audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera device), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0167] The human-machine interface input devices may include one or more of the following (only one of each is depicted): keyboard (1501), mouse (1502), touchpad (1503), touch screen (1510), data glove (not shown), joystick (1505), microphone (1506), scanner (1507), camera device (1508).
[0168] The computer system (1500) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through the touch screen (1510), data glove (not shown), or joystick (1505), but there may also be tactile feedback devices that do not function as input devices); audio output devices (e.g., speakers (1509), headphones (not depicted)); visual output devices (e.g., screen (1510), including CRT screen, LCD screen, plasma screen, OLED screen, each with or without touch screen input capability, each with or without tactile feedback capability - some of which may be able to output two-dimensional visual output or more than three-dimensional output, such as through stereoscopic image output; virtual reality glasses (not depicted); holographic displays and smokeboxes (not depicted)); and printers (not depicted).
[0169] The computer system (1500) may also include human-accessible storage devices and their associated media, e.g., including CD / DVD ROM / RW with media such as CD / DVD (1521). (1520), thumb drives (1522), removable hard disk drives or solid state drives (1523), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc.
[0170] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0171] The computer system (1500) may also include an interface (1554) to one or more communication networks (1555). The network may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan area, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include: local area networks such as Ethernet; wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial including CANBus, etc. Certain networks typically require an external network interface adapter attached to certain common data ports or peripheral buses (1549) (such as, for example, the USB port of the computer system (1500)); other networks are typically integrated into the core of the computer system (1500) by attaching to a system bus as described below (e.g., to an Ethernet interface in a PC computer system or to a cellular network interface in a smart phone computer system). Using any of these networks, the computer system (1500) can communicate with other entities. Such communication can be only one-way receiving (e.g., broadcast TV), only one-way sending (e.g., CAN bus to certain CAN bus devices), or two-way, e.g., to other computer systems using local digital networks or wide area digital networks. Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.
[0172] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1540) of the computer system (1500).
[0173] The core (1540) may include one or more central processing units (CPUs) (1541), a graphics processing unit (GPU) (1542), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1543), a hardware accelerator (1544) for certain tasks, a graphics adapter (1550), etc. These devices, together with a read-only memory (ROM) (1545), a random access memory (1546), an internal mass storage device (1547) such as an internal non-user-accessible hard disk drive, SSD, etc., can be connected via a system bus (1548). In some computer systems, the system bus (1548) may be accessible in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices can be attached directly or via a peripheral bus (1549) to the system bus (1548) of the core. In an example, a screen (1510) can be connected to the graphics adapter (1550). The architecture of the peripheral bus includes PCI, USB, etc.
[0174] The CPU (1541), GPU (1542), FPGA (1543), and accelerator (1544) can execute certain instructions that, when combined, can constitute the computer code mentioned above. The computer code can be stored in the ROM (1545) or the RAM (1546). Transient data can also be stored in the RAM (1546), while permanent data can be stored in, for example, the internal mass storage device (1547). Fast storage and retrieval of any memory device in the memory devices can be achieved by using a cache memory that can be closely associated with one or more CPUs (1541), GPU (1542), mass storage device (1547), ROM (1545), RAM (1546), etc.
[0175] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the medium and the computer code can be of the type well-known and available to those skilled in the field of computer software.
[0176] By way of example and not limitation, a computer system (1500) having an architecture and in particular a core (1540) can provide functionality due to software executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.) contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as introduced above, as well as certain storage devices of the core (1540) having non-transitoriness, such as, for example, an on-core mass storage device (1547) or a ROM (1545). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1540). Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core (1540) and in particular the processors therein (including a CPU, GPU, FPGA, etc.) to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in a RAM (1546) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of being logically hardwired or otherwise embodied in a circuit (e.g., an accelerator (1544)), which can operate in place of or in conjunction with the software to perform specific processes or specific portions of specific processes described herein. In appropriate cases, references to software can include logic, and references to logic can also include software. In appropriate cases, references to computer-readable media can include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits implementing logic for execution, or both of the above. The present disclosure encompasses any suitable combination of hardware and software.
[0177] The use of "at least one of... " or "one of... " in the present disclosure is intended to include any one or combination of the recited elements. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include A or B or (A and B). The use of "one of... " does not exclude any combination of the recited elements in cases where, for example, the elements are not mutually exclusive.
[0178] Although the present disclosure has described several exemplary embodiments, there are variations, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be recognized that those skilled in the art will be able to conceive of many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.
Claims
1. A method for processing visual media data, the method comprising: The bitstream of visual media data including the current picture is processed according to the format rules, wherein: The bitstream includes an index pointing to a particular candidate in a candidate list for prediction of a current block in the current picture; and The format rules specify: Obtaining a first verification result by performing a first verification of the motion-based inheritable parameters between the potential candidate and a first existing candidate among the one or more existing candidates in the candidate list, the first verification result indicating whether the potential candidate has redundant information of the motion-based inheritable parameters relative to the first existing candidate; Obtaining a second verification result by performing a second verification of the non-motion-based inheritable parameter between the potential candidate and the first existing candidate, the second verification result indicating whether the potential candidate has redundant information of the non-motion-based inheritable parameter relative to the first existing candidate; and When the first verification result indicates that the potential candidate has redundant information of the motion-based inheritable parameter relative to the first existing candidate and the second verification result indicates that the potential candidate has redundant information of the non-motion-based inheritable parameter relative to the first existing candidate, the potential candidate is redundant for the first existing candidate.
2. The method according to claim 1, wherein: When the current block is in inter prediction mode, the motion-based inheritable parameters include a motion vector (MV), and The non-motion based inheritable parameters include bi-directional prediction with codec level weight (BCW) index.
3. The method according to claim 1, wherein: When the current block is in intra block copy (IBC) mode, the motion-based inheritable parameters include a block vector (BV), and The non-motion based inheritable parameters include a reconstruction reordering type of IBC for the IBC mode.
4. A video encoding method, comprising: Determining to construct a candidate list for prediction of a current block in a current picture; determining potential candidates for addition to the candidate list, the candidate list comprising one or more existing candidates that have been added to the candidate list; performing a first check of the motion-based inheritable parameters between the potential candidate and a first existing candidate among the one or more existing candidates to obtain a first check result, the first check result indicating whether the potential candidate has redundant information of the motion-based inheritable parameters relative to the first existing candidate; performing at least a second check of the non-motion-based inheritable parameter between the potential candidate and the first existing candidate to obtain a second check result, the second check result indicating whether the potential candidate has redundant information of the non-motion-based inheritable parameter relative to the first existing candidate; as well as When the first verification result indicates that the potential candidate has redundant information of the motion-based inheritable parameter relative to the first existing candidate and the second verification result indicates that the potential candidate has redundant information of the non-motion-based inheritable parameter relative to the first existing candidate, it is determined that the potential candidate is redundant with respect to the first existing candidate.
5. The method according to claim 4, wherein: When the current block is in inter prediction mode, the motion-based inheritable parameters include a motion vector (MV).
6. The method according to claim 5, wherein: The non-motion based inheritable parameters include bi-directional prediction with codec level weight (BCW) index.
7. An apparatus for video decoding, comprising a processing circuit system, the processing circuit system being configured to: receiving a codec video bitstream including codec information of one or more pictures; Constructing a candidate list for a current block in a current picture, the candidate list comprising at least a first candidate and a second candidate, the first candidate and the second candidate having redundant information based on inheritable parameters of motion and non-redundant information based on inheritable parameters of non-motion; selecting a particular candidate from the candidate list; as well as The current block is reconstructed based on the specific candidate.
8. The device according to claim 7, wherein: The processing circuit system is configured to: determining potential candidates for addition to the candidate list; performing a first check of the motion-based inheritable parameters between the potential candidate and the first candidate as an existing candidate in the candidate list to obtain a first check result, the first check result indicating whether the potential candidate has the same redundant information of the motion-based inheritable parameters as the first candidate; performing at least a second check of the non-motion-based inheritable parameter between the potential candidate and the first candidate to obtain a second check result, the second check result indicating whether the potential candidate has the same redundant information of the non-motion-based inheritable parameter as the first candidate; as well as When the first verification result indicates that the potential candidate has the same redundant information of the inheritable motion-based parameters as the first candidate and the second verification result indicates that the potential candidate has the same redundant information of the inheritable non-motion-based parameters as the first candidate, it is determined that the potential candidate is redundant with respect to the first candidate.
9. The device according to claim 8, wherein: When the current block is in inter prediction mode, the motion-based inheritable parameters include a motion vector (MV).
10. The device according to claim 9, wherein: The non-motion based inheritable parameters include bi-directional prediction with codec level weight (BCW) index.
11. The device according to claim 8, wherein: When the current block is in intra block copy (IBC) mode, the motion-based inheritable parameters include a block vector (BV).
12. The device according to claim 11, wherein The non-motion based inheritable parameters include a reconstruction reordering type of IBC for the IBC mode.
13. The device according to claim 12, wherein: The non-motion based inheritable parameter includes at least one of a vertical flip and a horizontal flip.
14. The device according to any one of claims 8 to 13, wherein: The processing circuit system is configured to: When the non-motion-based inheritable parameter of the potential candidate is completely the same as that of the first candidate, it is determined that the potential candidate has the same redundant information of the non-motion-based inheritable parameter as that of the first candidate.
15. The device according to any one of claims 8 to 14, wherein: The processing circuit system is configured to: When the difference of the non-motion-based inheritable parameter between the potential candidate and the first candidate is within a certain range, it is determined that the potential candidate has the same redundant information of the non-motion-based inheritable parameter as the first candidate.