Method and device for video processing and medium

By introducing airspace proximity information into the motion compensation time domain filter of video encoding, the problems of motion domain inconsistency and block boundary artifacts in the prior art are solved, and the encoding and decoding performance is improved.

CN120036002APending Publication Date: 2025-05-23DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101051.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-13
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

There are several problems in the design of motion compensation time domain filter (MCTF) in the existing video encoding technology, including not taking into account information of the current block and adjacent blocks, resulting in motion domain inconsistency and block boundary artifacts, and filtering on non-overlapping blocks, affecting performance.

Method used

A spatial adjacent information assisted motion compensation time domain filter (SNIMCTF) method is proposed, and the selection of motion vectors and the setting of filter parameters are optimized by introducing information of neighboring blocks during motion estimation and filtering.

Benefits of technology

By introducing airspace proximity information, the accuracy of motion estimation and consistency of the filtering process are improved, the encoding and decoding performance is enhanced, and block boundary artifacts and motion domain inconsistencies are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120036002A_ABST
    Figure CN120036002A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises: during a transition between a target block of a video and a bitstream of the target block, determining a target motion vector from a set of candidate motion vectors based on information of neighboring blocks associated with the target block; executing motion estimation of a filtering process based on the target motion vector; and performing the conversion according to the motion estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to video coding and decoding technology, and more particularly, to a motion compensated temporal filter (MCTF) design in video coding / decoding. Background Art

[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Versatile Video Codec (VVC) standard. However, the encoding and decoding efficiency of conventional video encoding and decoding technologies is generally low, which is undesirable. Summary of the invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method comprises: during conversion between a target block of a video and a bitstream of the target block, determining a target motion vector from a set of candidate motion vectors based on information of neighboring blocks associated with the target block; performing motion estimation of a filtering process based on the target motion vector; and performing conversion according to the motion estimation. Compared with conventional schemes, the proposed method can advantageously improve coding efficiency and performance.

[0005] In a second aspect, another method for video processing is proposed. The method includes: determining an error during conversion between a target block of a video and a bitstream of the target block, the error including neighboring information of the target block; performing a filtering process based on the error; and performing conversion according to the filtering process. Compared with conventional schemes, the proposed method can advantageously improve coding efficiency and performance.

[0006] In a third aspect, another method for video processing is proposed. The method includes: performing a filtering process on a set of overlapping blocks associated with the target block during conversion between a target block of the video and a bitstream of the target block; and performing conversion according to the filtering process. Compared with conventional schemes, the proposed method can advantageously improve coding efficiency and performance.

[0007] In a fourth aspect, another video processing method is proposed. The method includes: during conversion between a target block of a video and a bitstream of the target block, determining an encoding mode of a frame based on whether a filtering process is applied to a frame associated with the target block; and performing conversion based on the determination. Compared with conventional schemes, the proposed method can advantageously improve encoding and decoding efficiency and performance.

[0008] In a fifth aspect, a device for video processing is provided. The device for video processing includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform a method according to any one of the first, second, third or fourth aspects.

[0009] In a sixth aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to any one of the first, second, third or fourth aspects.

[0010] In a seventh aspect, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by a video processing device. The method includes: determining a target motion vector from a set of candidate motion vectors based on information of neighboring blocks associated with a target block of the video; performing motion estimation of a filtering process based on the target motion vector; and generating a bitstream of the target block according to the motion estimation.

[0011] In an eighth aspect, another method for storing a bitstream of a video is proposed. The method includes: determining a target motion vector from a set of candidate motion vectors based on information of neighboring blocks associated with a target block of the video; performing motion estimation of a filtering process based on the target motion vector; generating a bitstream of the target block according to the motion estimation; and storing the bitstream in a non-transitory computer-readable recording medium.

[0012] In a ninth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by a video processing apparatus. The method includes: determining an error, the error including neighboring information of a target block of the video; performing a filtering process based on the error; and generating a bitstream of the target block according to the filtering process.

[0013] In a tenth aspect, another method for storing a bitstream of a video is provided. The method includes: determining an error, the error including neighboring information of a target block of the video; performing a filtering process based on the error; generating a bitstream of the target block according to the filtering process; and storing the bitstream in a non-transitory computer-readable recording medium.

[0014] In an eleventh aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by a video processing apparatus. The method includes: performing a filtering process on a set of overlapping blocks associated with a target block of the video; and generating a bitstream of the target block according to the filtering process.

[0015] In a twelfth aspect, another method for storing a bitstream of a video is provided, the method comprising: performing a filtering process on a set of overlapping blocks associated with a target block of the video; generating a bitstream of the target block according to the filtering process; and storing the bitstream in a non-transitory computer-readable recording medium.

[0016] In a thirteenth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by a video processing device. The method includes: determining an encoding mode of a frame based on whether a filtering process is applied to a frame associated with a target block of the video; and generating a bitstream of the target block based on the determination.

[0017] In a fourteenth aspect, another method for storing a bitstream of a video is provided. The method includes: determining an encoding mode of a frame based on whether a filtering process is applied to a frame associated with a target block of the video; generating a bitstream of the target block based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0018] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0020] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;

[0021] Figure 2 A block diagram illustrating a first example video encoder is shown according to some embodiments of the present disclosure;

[0022] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;

[0023] Figure 4 It is an overview of the VVC standard;

[0024] Figure 5 A schematic diagram showing the different layers of hierarchical motion estimation;

[0025] Figure 6 A schematic diagram showing the decoding process using ACT;

[0026] Figure 7 An example of a block encoded and decoded in palette mode is shown;

[0027] Figure 8 A schematic diagram according to an embodiment of the present disclosure is shown;

[0028] Figure 9a The optimal MV obtained by conventional ME is shown in terms of exercise intensity and Figure 9b shows the motion intensity of the optimal MV according to an embodiment of the present disclosure;

[0029] Fig.10a The results of the distribution of the error in the spatial domain according to conventional filtering are shown and Fig.10b The result of the distribution of errors in the spatial domain according to the embodiment of the present disclosure is shown;

[0030] Fig.11 A flowchart showing a method according to some embodiments of the present disclosure is shown;

[0031] Fig.12 A flowchart showing a method according to some embodiments of the present disclosure is shown;

[0032] Fig.13 A flowchart showing a method according to some embodiments of the present disclosure is shown;

[0033] Fig.14 A flowchart showing a method according to some embodiments of the present disclosure; and

[0034] Fig.15 A block diagram of a computing device is shown in which various embodiments of the present disclosure may be implemented.

[0035] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION

[0036] The principle of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, without implying any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.

[0037] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0038] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art to affect correlation with other embodiments.

[0039] It should be understood that, although the terms "first" and "second" etc. may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another element. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0040] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "include", "comprises", "has", "has", "includes" and / or "comprising" are used herein to indicate the presence of the features, elements and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Example Environment

[0041] Figure 1 1 is a block diagram illustrating an example video codec system 100 that may utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0042] The video source 112 may include a source such as a video acquisition device. Examples of a video acquisition device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0043] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a bit sequence that forms a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of pictures. The associated data may include sequence parameter sets, picture parameter sets, and other grammatical structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.

[0044] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120, or may be outside the destination device 120, which is configured to be connected to an external display device interface.

[0045] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.

[0046] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of a video encoder 114 in the system 100 is shown.

[0047] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0048] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.

[0049] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture in which the current video block is located.

[0050] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail below. Figure 2 are shown separately in the example.

[0051] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0052] The mode selection unit 203 may select one of a plurality of encoding modes (intra-frame encoding or inter-frame encoding), for example, based on the error result, and provide the generated intra-frame encoded block or inter-frame encoded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode selection unit 203 may select an intra-frame inter-frame joint prediction (CIIP) mode, in which the prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 may also select a resolution for a motion vector (e.g., sub-pixel precision or integer pixel precision) for the block.

[0053] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.

[0054] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block. For example, depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" may refer to a part of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" may refer to parts of a picture composed of macroblocks independent of the macroblocks in the same picture.

[0055] In some examples, the motion estimation unit 204 can perform uni-directional prediction on the current video block, and the motion estimation unit 204 can search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0056] Alternatively, in other examples, the motion estimation unit 204 can perform bi-directional prediction on the current video block. The motion estimation unit 204 can search the reference pictures in list 0 to find one reference video block for the current video block, and can also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 can then generate multiple reference indexes and multiple motion vectors, where the multiple reference indexes indicate the multiple reference pictures in list 0 and list 1 that contain the multiple reference video blocks, and the multiple motion vectors indicate the multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 can output the multiple reference indexes and the multiple motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0057] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.

[0058] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block, and this value indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0059] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0060] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0061] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0062] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0063] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0064] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0065] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0066] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0067] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0068] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0069] Figure 3 is a block diagram showing an example of a video decoder 300 according to some embodiments of the present disclosure, which may be Figure 1 An example of a video decoder 124 in the system 100 is shown.

[0070] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of , video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0071] exist Figure 3 In the example of , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.

[0072] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video data, and the motion compensation unit 302 can determine motion information from the entropy decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data and reference pictures from adjacent PBs. The motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction areas in B strips, also includes an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatial neighboring blocks or temporal neighboring blocks.

[0073] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter.Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.

[0074] The motion compensation unit 302 may calculate interpolated values ​​for sub-integer pixels of a reference block using interpolation filters used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 based on received syntax information, and the motion compensation unit 302 may use the interpolation filters to generate a prediction block.

[0075] The motion compensation unit 302 may use at least part of the syntax information to determine the size of blocks used to encode (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice may be an entire picture, or it may be a region of a picture.

[0076] The intra prediction unit 303 may use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0077] The reconstruction unit 306 may obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and the buffer 307 also generates the decoded video for presentation on a display device.

[0078] Some example embodiments of the present disclosure are described in detail below. It should be noted that the section titles used in this document are for ease of understanding, and are not intended to limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a versatile video codec or other specific video codec, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps of de-encoding will be implemented by a decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1 Overview Embodiments of the present disclosure relate to video coding techniques. Specifically, it relates to a motion compensated temporal filter (MCTF) design in video coding. It can be applied to existing video encoders, such as VTM, x264, x265, HM, VVenC and others. It can also be applied to future video codec encoders or video codecs. 2 Introduction 2.1 Versatile Video Codec (VVC) Standard Figure 4 The functional diagram of a typical hybrid VVC encoder is shown, including block partitioning of video pictures into CTUs. For each CTU, quadtree, ternary tree, and binary tree structures are used to partition it into several blocks, called codec units. For each codec unit, block-based intra-frame or inter-frame prediction is performed, and then the generated residual is transformed and quantized. Finally, context-adaptive binary arithmetic coding and decoding (CABAC) entropy codec is used for bitstream generation. 2.2 Introduction to MCTF MCTF is a pre-filtering process for better compression efficiency. Several encoders, such as VVC Test Model (VTM) and HEVC Test Model (HM) support MCTF. And MCTF is applied before the encoding process. In MCTF, when the reference frame is ready, a hierarchical motion estimation scheme (ME) is used to find the best motion vector for each 8x8 block. Figure 5 As shown, three layers are used in the hierarchical motion estimation scheme. Each sub-sampling layer is half the width and half the height of the lower layer, and sub-sampling is done by calculating the rounded average of four corresponding sample values ​​from the lower layer. Different sub-sampling rates and sub-sampling filters can be applied. The ME process is described as follows. First, motion estimation is performed for each 16x16 block in L2. The ME difference (e.g., the sum of squared differences) is calculated for each selected motion vector, and the motion vector corresponding to the minimum matching difference is selected. The selected motion vector is then used as an initial value when estimating the motion in L1. The same operation is then performed for estimating the motion in L0. As a final step, one or more integer precision motions and fractional precision motions are estimated for each 8x8 block. Motion compensation is applied to the pictures before and after the current picture according to the best matching motion for each 8×8 block to align the sample coordinates of each block in the current picture with the best matching coordinates in the reference picture. During the filtering process, MCTF is performed on each 8x8 block. The samples of the current picture are then filtered separately for the luma and chroma channels to produce a filtered picture as follows. The filtered sample values ​​In of the current picture are calculated using the following formula: Among them I o is the original sample value, I r (i) is the predicted sample value motion compensated from picture i, and w r (i, a) is the weight of motion compensated picture i for a given value of a. If there is no reference frame after the current frame, α is set equal to 1, otherwise a is equal to 0. For the samples in the brightness channel, the weight w r (i, a) is calculated as follows: in s l =0.4, i and a for the remaining cases are: s r (i, a) = 0.3, as well as σ l (QP) = 3*(QP-10) ΔI(i)=I r ( / )-I o . Adjustment factor w aand σ w is calculated to calculate w r (i, a), as follows: where min(error) is the minimum error in the same position of all motion compensated pictures. The noise and error values ​​are calculated at a block granularity of 8×8 for luma and 4×4 for chroma as follows: in bsX and bsY represent the width and height of the block respectively. For the chroma channel, the weight w r (i, a) is calculated as follows: Among them, s c =0.55 and σ c =30. 2.3 Transform Skip Mode The residual of the block can be encoded and decoded using the transform skip mode that completely skips the transform process of the block. In addition, in VVC, for transform skip blocks, the minimum allowed quantization parameter (QP) signaled in the SPS is used, which is set equal to 6×(internalBitDepth-inputBitDepth)+4 in VTM. 2.4 Adaptive Color Transformation (ACT) Figure 6 FIG. 4 shows a decoding flow chart of VVC with ACT applied. Figure 6 As shown, the color space conversion is performed in the residual domain. Specifically, an additional decoding module, namely inverse ACT, is introduced after the inverse transform to convert the residual from the YCgCo domain back to the original domain. In VVC, unless the maximum transform size is less than the width or height of a coding unit (CU), a CU leaf node is also used as the unit of transform processing. Therefore, in the proposed implementation, an ACT flag is signaled for a CU to select the color space used to encode and decode its residual. In addition, following the HEVC ACT design, for inter and IBCCU, ACT is enabled only when there is at least one non-zero coefficient in the CU. For intra CU, ACT is enabled only when the chroma component selects the same intra prediction mode of the luminance component (i.e., DM mode). The core transform for color space conversion remains the same as that used for HEVC. In addition, similar to the ACT design in HEVC, a QP adjustment of (5, 5, -3) is applied to the transform residual to compensate for the dynamic range variation of the residual signal before and after the color transform. On the other hand, Figure 6 As shown, the forward and inverse color transforms require access to the residuals of all three components. Accordingly, in the proposed implementation, ACT is disabled in the following two cases, where not all residuals of the three components are available. 1. Split tree partitioning: When split tree is applied, the luma and chroma samples within a CTU are partitioned by different structures. This results in the CU in the luma tree containing only the luma component, and the CU in the chroma tree containing only two chroma components. 2. Intra-sub-partition prediction (ISP): ISP sub-partitioning is only applied to luma, while chroma signals are encoded and decoded without partitioning. In the current ISP design, except for the last ISP sub-partitioning, other sub-partitionings contain only luma components. 2.5 Block-based delta pulse codec modulation (BDPCM) In JVET-M0413, block-based delta pulse codec modulation (BDPCM) was proposed to efficiently encode and decode screen content, and then adopted into VVC. The prediction direction used in BDPCM can be vertical and horizontal prediction mode. The whole block is intra-predicted by copying the samples in the prediction direction (horizontal or vertical prediction) as intra-prediction. The residual is quantized and the delta between the quantized residual and its predicted value (horizontal or vertical) is encoded and decoded. This can be described as follows: For a block of size M (rows) × N (columns), let r i,j , 0≤i≤M-1, 0≤j≤N-1 is the prediction residual after performing intra prediction horizontally (copying left neighboring pixel values ​​across the predicted block line by line) or vertically (copying the top neighboring line to each line in the predicted block) using unfiltered samples from the upper or left block boundary samples. Let Q(r i,j ), 0≤i≤M-1, 0≤j≤N-1 represents the residual r i,j The quantized version of , where the residual is the difference between the original block and the predicted block value. Block DPCM is then applied to the quantized residual samples, resulting in a quantized residual with elements The modified M×N array When vertical BDPCM is transmitted by signal: For horizontal prediction, similar rules apply and the residual quantization samples are obtained as follows: Residual quantization samples is sent to the decoder. On the decoder side, the above calculation is inverted to produce Q(r i,j ), 0≤i≤M-1, 0≤j≤N-1. For the vertical prediction case, For the horizontal case, The inverse quantized residual Q -1 (Q(r i,j )) is added to the intra block prediction value to produce the reconstructed sample value. The main benefit of this approach is that inverse BDPCM can be done simply on the fly during coefficient parsing by adding prediction values ​​as the coefficients are parsed, or inverse BDPCM can be performed after parsing. In VTM-7.0, BDPCM can also be applied to chroma blocks, and the chroma BDPCM has a flag and BDPCM direction separated from the luma BDPCM mode. 2.6 Palette Mode The basic idea behind palette mode is that pixels in a CU are represented by a small set of representative color values. This set is called the palette. And it is also possible to indicate samples outside the palette by signaling an escape symbol followed by a (possibly quantized) component value. Such pixels are called escape pixels. Figure 7 As shown in Figure 7 As shown in , for each pixel having three color components (a luma component and two chroma components), an index into the palette is found and the block can be reconstructed based on the found values ​​in the palette. 3 Questions The current MCTF design has the following problems: 1. It does not include the information of neighboring blocks of the current block in the ME process, so the MCTF performance is limited due to the introduced motion domain inconsistency. 2. It does not include the information of neighboring blocks of the current block in the filtering process, and the MCTF performance is therefore limited due to the introduced boundary artifacts of adjacent blocks. 3. Performing MCTF filters on non-overlapping blocks, which may affect MCTF performance because the reference block can span two filtering blocks during the encoding process. 4. The reference block is filtered as the current block in the frame filtered by MCTF, and the frame needs to be processed carefully to avoid removing the reference block component from the frame. Existing codec solutions do not consider this aspect. 4. Embodiments of the present disclosure In order to solve the above problems and some other unmentioned problems, the following method is disclosed. The embodiments of the present disclosure should be considered as examples to explain the general concept and should not be interpreted in a narrow manner. In addition, these inventions can be applied individually or in combination in any way. It should be noted that “MCTF” may refer to a design in the prior art, or alternatively, it may refer to any variant of a MCTF design or other types of time-domain filtering methods in the prior art. MCTF ME processing including neighbor information Let C be the current block. Let F i is the difference metric, which corresponds to the motion vector MV associated with C i , where 1≤i≤L. Let R be the value corresponding to MV i Let CN j is the jth neighboring block of C. RN j is the jth neighboring block of the reference block, where 1≤j≤S. Let T be the cost between C and R. Let K j For CN j and RN j The cost between. Let mctf_frame be a frame with MCTF applied, and let non_mctf_frame be a frame without MCTF applied. 1. The decision of the best motion vector in the MCTF ME process may depend on the information of neighboring blocks, such as the cost of neighboring blocks. a. In one example, the cost of the neighboring blocks may depend on the motion vector to be checked during the ME process of the current block. b. In one example, the final cost of the motion vector to be checked for the current block may be calculated using a linear function of the costs associated with the current block and neighboring blocks. c. In one example, a final cost of a motion vector to be checked for the current block may be calculated using a non-linear function of the costs associated with the current block and neighboring blocks. d. In the ME process in MCTF, the ME difference (as described in Section 2) can include neighbor information. e. In one example, F i Can include T and / or K j . i. In one example, F i Can be evaluated as: 1. In one example, W 0 , W1 ....W s Can have the same or different values. 2. In one example, W 1 , W 2 ....W s can have the same value and value as W 0 The values ​​of are different. ii. In one example, T and / or K j Distortion metrics can be used to calculate the error, such as the sum of absolute differences (SAD), the sum of squared errors (SSE), or the mean sum of squared errors. (MSE). f. In one example, CN j and / or RN j At least one of top, bottom, left, right, upper-left, upper-right, lower-left and / or lower-right neighboring blocks may be included. i. In one example, CN j and / or RN j Top, bottom, left and / or right neighboring blocks may be included. ii. In one example, CN j and / or RN j Can include top, left, upper left, and / Or the upper right neighboring block. iii. In one example, CN j and / or RN j Can include bottom, right, lower right, and / or Or the lower left neighboring block. iv. In one example, CN j and / or RN j Top and / or left neighboring blocks may be included. v. In one example, CN j and / or RN j The upper left, upper right, lower left and / or lower right neighboring blocks may be included. vi. In one example, different block sizes may be used for different neighboring blocks. vii. In one example, CN is compared to C and R. j and / or RN j The block sizes can be the same or different. viii. In one example, CN j and / or RN j The dimensions can be W×H. ix. In one example, W 0 , W 1 ....W sOne or more of may be determined based on a block size of one or more neighboring blocks. g. In one example, different methods of introducing neighbor information can be adopted for different layers in the hierarchical ME scheme. i. In one example, W 0 , W 1 ....W s It can be the same or different values ​​for different layers in a layered ME. ii. In one example, S may be different for different layers in a layered ME. iii. In one example, ME with neighbor information is only applied to L1 and L0 layers in hierarchical ME. h. In one example, the above items can be applied to one or all layers in the layered ME process in MCTF. i. In one example, the above items may be applied or not applied according to different sizes of C in the hierarchical ME process in MCTF. j. In one example, the above items can be Figure 8 Shown. MCTF filtering process including neighbor information In the following project, let MV k is the optimal motion vector for block C. Let R be the motion vector corresponding to MV k Let CN j is the jth neighboring block of C. Let RN j is the jth neighboring block of the reference block, where 1≤j≤S. Let T be the cost between C and R. Let K j For CN j and RN j The cost between. 2. In the filtering process in MCTF, the error derived for each filtered block (eg, as mentioned in Section 2) may include neighbor information. a. In one example, the proximity information can be expressed as: i. In one example, W 0 , W 1 ....W s Can have the same or different values. ii. In one example, W 1 , W 2 ....W s can have the same value and value as W 0 The values ​​of are different. iii. In one example, K j Can be calculated by distortion metrics such as SAD, SSE or MSE. Overlapping MCTF filtering process 3. The filtering process in MCTF can be performed on overlapping blocks. a. In one example, a width step WS and a height step HS may be used, and they may not be equal to the size B×B of the filter block. i. In one example, WS and / or HS may be less than B. ii. In one example, after a block with position (X, Y) is filtered, the next block to be filtered may be at (X+WS, Y). 1. In one example, after all blocks with vertical position Y are filtered, the next block to be filtered may be at (X, Y+WS). b. In one example, the size of the block to be filtered can be B×B, WS×B, B× HS, WS×HS. c. In one example, the error and / or noise for the overlapping region may be determined by the neighboring blocks involved. i. In one example, the error and / or noise may be calculated by weighting or averaging the errors and / or noise of some or all of the involved neighboring blocks. d. In one example, the error and / or noise of a neighboring block may be used for the error and / or noise of the overlapping area. MCTF frame processing during encoding 4. How a frame is encoded may depend on whether MCTF is applied to the frame. a. In one example, a frame after MCTF filtering may be processed differently than a frame without MCTF filtering during encoding. b. In one example, the slice / CTU / CU / block level QP of mctf_frame can be decreased or increased. i. In one example, the above changes are applied only to luma QP. ii. Alternatively, in one example, the above changes are applied only to chroma QP. iii. Alternatively, in one example, the above changes are applied to luma and chroma QPs Both. c. In one example, the intra cost of some / all blocks in the mctf frame may be reduced by Q. d. In one example, the skip cost of some / all blocks in the mctf frame may be increased by V. e. In one example, the codec information F of one or more blocks may be determined differently for mctf_frame and non_mctf_frame. i. In one example, F may represent a prediction mode. ii. Alternatively, in one example, F may represent an intra prediction mode. iii. Alternatively, in one example, F may represent a quadtree partition flag. iv. Alternatively, in one example, F may represent a binary / ternary tree partition type. v. Alternatively, in one example, F may represent a motion vector. vi. Alternatively, in one example, F may represent a Merge flag. vii. Alternatively, in one example, F may represent a Merge index. f. In one example, whether and / or how blocks / regions / CTUs are partitioned may be different for mctf_frame and non_mctf_frame. i. In one example, the maximum depth of a CU in mctf_frame may be increased. g. In one example, different motion search methods can be used for mctf_frame and non_mctf_frame. h. In one example, different fast intra mode algorithms can be used for mctf_frame and non_mctf_frame. i. In one example, screen content codecs (eg, palette mode, IBC mode, BDPCM, ACT and / or transform skip mode) may not be allowed to be used to encode mctf_frame. j. In one example, the difference between the MCTF filtered block and the original block can be used as a metric to determine if the block needs to be processed differently in encoding. 5. The above items may be applied under certain conditions. a. In one example, the condition is the distortion of the original pixel and the filtered pixel (including SAD, SSE or MSE) exceeds a threshold X at the CTU / CU / block level. b. In one example, the condition is that the distortion of the filtered current pixel and the filtered neighboring pixels, including SAD, SSE or MSE, exceeds a threshold U at the CTU / CU / block level. c. In one example, the condition is that one of the values in the average motion vector exceeds a threshold at the slice / CTU / CU / block level. General 6. The above items can be applied regardless of the current block size used in MCTF. a. In one example, W and / or D can be greater than or equal to 4. b. In one example, W and / or D can be less than or equal to 64. c. In one example, W and / or D can be equal to 8. 7. In the above items, W, H, WS, HS, B, P, Q, V, X, Y, and / or Z are integers (e.g., 0 or 1), and can depend on: a. Slice / tile group type and / or picture type b. Color component (e.g., can be applied only to Cb or Cr), c. Temporal layer ID, d. Layer ID in pyramid ME search, e. Standard profile / level / layer 8. The above items can be applied to MCTF-related variants and other filtering methods such as bilateral filters, low-pass filters, and high-pass filters. 9. The above items can be applied to loop filters. 5. Embodiments MCTF is based on independent blocks with a fixed size during ME and filtering processes. Although the independent processes between blocks are convenient and efficient, the ME process is prone to early termination at local optimal MVs and during the filtering process, resulting in large inconsistent regions and causing block boundary artifacts after filtering. The processing method of independent blocks will affect the quality of the filtered frames and the coding / decoding efficiency after filtering. Therefore, a Spatially Neighboring Information-aided Motion Compensation Temporal Filter (SNIMCTF) method is proposed to improve the performance of MCTF, including the ME and filtering processes. 5.1 Spatially Neighboring Information-aided Motion Compensation Temporal Filter The ME process of conventional MCTF uses the SSE of the current block C and the reference block R from its reference picture for motion estimation. This estimation process can effectively and accurately match the minimum distortion of the reference block with the current block, but only the information of the current block is considered in the estimation process, and the validity of the current block and the neighboring blocks as a whole is not considered. In the encoding process, the frame filtered by MCTF is referenced by a larger block when referenced by subsequent frames, and the filtered frame is also encoded in a larger block, and the size of the larger block is usually larger than the current block. If only the optimal reference of the current block is considered in ME, it is likely to fall into a local optimal solution, resulting in a reduction in the reference of subsequent frames in the large block and a decrease in the encoding and decoding efficiency of the current filtered frame. Therefore, a spatial neighbor information assisted motion estimation (SNIME) method is proposed to solve this problem. like Figure 8 As shown in Figure 1, when SNIME performs motion estimation on the best reference block of the current block, neighboring information is introduced into the estimation process in a weighted manner and is calculated as: Among them, w c is the weight of the current block, w i is the weight of its neighboring blocks, CN i and RN i Represents the i-th spatial neighboring block of the current block and its corresponding reference block respectively. At different resolutions, the spatial distribution of pixels is different, so the correlation between the current block and the neighboring blocks is different. For example, in 1080p resolution, the current block and the surrounding blocks are more closely related, but in 480p resolution, the current block and the surrounding blocks may not have such a strong correlation at all [8] The correlation between the current block and the neighboring blocks is related to the size of the current block. For example, when the current block size is 8×8, it can be associated with another neighboring information, but when the current block size is 16×16, the range of neighboring information related to the current block will be reduced. Therefore, in the hierarchical ME process of MCTF, motion estimation of different layers with different resolutions will adopt different weighting parameters w c and w i , with different sizes of CN i and RN i . During the filtering process, the filtering parameters are dynamically set by the implicit information of the current block. The independent setting of the block-level filtering parameters does not consider the correlation between the current block and the neighboring information, resulting in inconsistent filtering between blocks, thereby reducing the filtering effect. Therefore, the present disclosure proposes a spatial neighbor information assisted block-level filtering (SNIBF) scheme. The filtering process of MCTF is expressed by equation (2-2), where w a and σ wDetermined by the error between C and R, the calculations are as follows: Equation (2-3), Equation (2-4), Equation (2-5), and Equation (2-6). In order to further reconcile with SNIME, SNIBF replaces the SSE in Equation (2-5) with Equation (5-1). After the replacement, neighboring information is considered to be the key factor in filtering, and the block-level filtering process is more relevant to improve the overall filtering effect. 5.2 Analysis of SNIME and SNIBF SNIMCTF mainly optimizes the ME and filtering process of MCTF by introducing spatial neighbor information. The following discussion will further demonstrate the impact of SNIMCTF in the visual mode. Figure 9a and 9b The figure shows the visualization of the motion intensity of motion estimation of POC0 and POC2 in 8×8 blocks with coordinates (128, 256) and size (512, 512) of BacketalDrive under QP15. Figure 9a and 9b As shown, the motion intensity comparison of the optimal MV obtained by conventional ME and SNIME is shown, corresponding to Figure 9a and 9b The intensity of motion in the figure is represented by the absolute value of the maximum value of the motion vector. As the reader can observe, Figure 9a There are many regions with strong motion intensity changes in the proposed method, but after the spatial neighborhood information is introduced, such as Figure 9b As shown in Figure 1, the motion intensity variation becomes relatively smooth in the spatial domain. This is mainly because the neighboring blocks are considered when estimating the MV. In the current design, the MV only considers the information of the current block and falls into the local optimal area. However, the proposed method estimates the MV including more useful information, so it can achieve a better MV and then enhance the encoding and decoding performance. Fig.10a and Fig.10b A visualization of the error in motion estimation of POC0 and POC2 in an 8×8 block on a region with coordinates (128, 256) and size (512, 512) of BacketalDrive under QP15 is shown. Fig.10a and 10b This is the result of the distribution of the error in the spatial domain obtained under conventional filtering and SNIBF. From the visualization results, after adopting SNIBF, the error distribution in the spatial domain also becomes more uniform. This error is used to determine the filter coefficients to make the filtering process more consistent between independent blocks.

[0079] Embodiments of the present disclosure relate to prediction from multiple component mixes in image / video codecs.

[0080] As used herein, the term “video unit” or “codec unit” or “block” used herein may refer to one or more of the following: a color component, a sub-picture, a slice, a slice, a codec tree unit (CTU), a CTU row, a CTU group, a codec unit (CU), a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec block (CB), a prediction block (PB), a transform block (TB), a block, a sub-block of a block, a sub-region within a block, or a region including more than one sample or pixel.

[0081] In the present disclosure, with respect to “blocks encoded and decoded using mode N”, the term “mode N” may be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a coding technology (e.g., AMVP, Merge, SMVD, BDOF, PROF, DMVR, AMVR, TM, Affine, CIIP, GPM, MMVD, BCW, HMVP, SbTMVP, etc.).

[0082] In this context, let C be the current block. Let F i is the difference metric, which corresponds to the motion vector MV associated with C i , where 1≤i≤L. Let R be the value corresponding to MV i Let CN j is the jth neighboring block of C. Let RN j Let C be the jth neighbor of the reference block, where 1≤j≤S. Let T be the cost between C and R. Let K j For CN j and RN j Let mctf_frame be the frame with MCTF applied, and let non_mctf_frame be the frame without MCTF applied. In this context, let MV k is the optimal motion vector for block C. Let R be the motion vector corresponding to MV k of the reference block.

[0083] Note that the terms mentioned below are not limited to the specific terms defined in existing standards. Any variants of the codec tool are also applicable.

[0084] Fig.11 A flow chart of a method 1100 for video processing according to some embodiments of the present disclosure is shown. The method 1100 may be implemented during conversion between a target block of a video and a bitstream of the target block.

[0085] like Fig.11As shown in, at box 1110, during the conversion between the target block of the video and the bitstream of the target block, the target motion vector is determined from a set of candidate motion vectors based on the information of the adjacent blocks associated with the target block. In some embodiments, the information of the adjacent blocks may include the cost of the adjacent blocks. In some embodiments, the cost of the adjacent blocks may depend on the candidate motion vectors in the motion estimation of the target block. In one example, the final cost of the candidate motion vector to be checked for the target block can be determined using a linear function of the cost associated with the target block and the adjacent blocks. In another example, the final cost of the candidate motion vector to be checked for the target block can be determined using a nonlinear function of the cost associated with the target block and the adjacent blocks.

[0086] At block 1120, motion estimation of the filtering process is performed based on the target motion vector. In some embodiments, in the motion estimation of the filtering process, the motion estimation difference may include neighbor information.

[0087] At block 1130, conversion is performed based on motion estimation. In some embodiments, conversion may include encoding the target block into a bitstream. Alternatively, conversion may include decoding the target block from the bitstream. Compared with conventional schemes, inconsistencies in the motion domain can be avoided and the performance of the filtering process can be improved.

[0088] Implementations of the present disclosure may be described in view of the following items, and features of these items may be combined in any reasonable manner.

[0089] In some embodiments, the difference measure of the candidate motion vector may include at least one of the following: a first cost between the target block and a reference block corresponding to the candidate motion vector; or a second cost between the jth neighboring block of the target block and the jth neighboring block of the reference block. j may be an integer. For example, F i Can include T and / or K j .

[0090] In some embodiments, the difference metric F i Can be evaluated as: Where W 0 represents the initial value, W j represents the jth initial value, T represents the first cost, K j represents the second cost, and S represents the total number of neighboring blocks.

[0091] In some embodiments, W 0 , W 1 ....W s Can have the same value. In some embodiments, W 0 , W 1 ....W sCan have different values. In some embodiments, W 1 , W 2 ....W s can have the same value and the value is the same as W 0 The values ​​of are different.

[0092] In some embodiments, at least one of the first cost or the second cost is determined using a distortion metric. For example, the distortion metric may include at least one of the following: a sum of absolute differences (SAD), a sum of squared errors (SSE), or a mean sum of squared errors (MSE). In one example, T and / or K j The distortion may be calculated using a distortion metric such as the sum of absolute differences (SAD), the sum of squared errors (SSE), or the mean sum of squared errors (MSE).

[0093] In some embodiments, the neighboring blocks of the target block may include at least one of the following: a top neighboring block of the target block, a bottom neighboring block of the target block, a left neighboring block of the target block, a right neighboring block of the target block, an upper left neighboring block of the target block, an upper right neighboring block of the target block, a lower left neighboring block of the target block, or a lower right neighboring block of the target block. In some embodiments, the neighboring blocks of the reference block associated with the target block may include at least one of the following: a top neighboring block of the reference block, a bottom neighboring block of the reference block, a left neighboring block of the reference block, a right neighboring block of the reference block, an upper left neighboring block of the reference block, an upper right neighboring block of the reference block, a lower left neighboring block of the reference block, or a lower right neighboring block of the reference block. For example, in one example, CN j (For example, Figure 8 The CN shown in 1 , CN 2 ,......,CN 8 ) and / or RN j (For example, Figure 8 RN shown 1 , R.N. 2 ,......,R.N. 8 ) may include top, bottom, left and / or right neighboring blocks. In one example, CN j and / or RN j It may include top, left, upper left, and / or upper right neighboring blocks. In one example, CN j and / or RN j The bottom, right, lower right and / or lower left neighboring blocks may be included. In one example, CN j and / or RN j Can include top and / or left neighboring blocks. In one example, CN j and / or RN j The upper left, upper right, lower left and / or lower right neighboring blocks may be included.

[0094] In some embodiments, different block sizes can be used for different adjacent blocks. In some embodiments, the first block size of the adjacent block can be identical with the second block size of the target block. In some embodiments, the first block size of the adjacent block can be different from the second block size of the target block.

[0095] In some embodiments, the third block size of the neighboring block of the reference block may be the same as the fourth block size of the reference block. Alternatively, the third block size of the neighboring block of the reference block may be different from the fourth block size of the reference block. j and / or RN j The block sizes can be the same or different.

[0096] In some embodiments, the size of the neighboring block may be W×H. In some embodiments, the size of the neighboring block of the reference block may be W×H. In this case, W represents the width of the target block and H represents the height of the target block. j and / or RN j The dimensions of may be W×H. In some embodiments, W 0 , W 1 ....W s At least one of may be determined based on a block size of one or more neighboring blocks.

[0097] In some embodiments, different neighbor information can be used for different layers in the hierarchical motion estimation scheme. In one example, different methods of introducing neighbor information can be used for different layers in the hierarchical motion estimation scheme. 0 , W 1 ....W s It can have the same value for different layers in a hierarchical motion estimation scheme. 0 , W 1 ....W s There may be different values ​​for different layers in a hierarchical motion estimation scheme.

[0098] In some embodiments, the total number of neighboring blocks may be different for different layers in a hierarchical motion estimation scheme.In one example, S may be different for different layers in a hierarchical motion estimation scheme.

[0099] In some embodiments, motion estimation with neighbor information may be applied to L1 layer and L0 layer in a hierarchical motion estimation scheme. In one example, ME with neighbor information is only applied to L1 and L0 layers in a hierarchical ME.

[0100] In some embodiments, determining the target motion vector based on information of neighboring blocks may be applied to at least one layer in a hierarchical motion estimation scheme. For example, the above method or embodiment may be applied to one or all layers in a hierarchical ME process in MCTF.

[0101] In some embodiments, determining whether the target motion vector is applied based on information of neighboring blocks may be based on different sizes of target blocks in the hierarchical motion estimation scheme in the filtering process. In one example, the above items may be applied or not applied based on different sizes of C in the hierarchical ME process in MCTF. In some embodiments, the above methods or embodiments may be as follows: Figure 8 For example, Figure 8 As shown, the current block 810 may include a neighboring block CN 1 , CN 2 , CN 3 , CN 4 , CN 5 , CN 6 , CN 7 and CN 8 The reference block 820 of the current block 810 may include a neighboring block RN 1 , RN 2 , RN 3 , RN 4 , RN 5 , RN 6 , RN 7 and RN 8 .

[0102] In some embodiments, a target motion vector from a set of candidate motion vectors can be determined based on information of neighboring blocks associated with the target block of the video. In some embodiments, motion estimation of the filtering process is performed based on the target motion vector. In some embodiments, a bitstream for the target block is generated based on the motion estimation.

[0103] In some embodiments, a target motion vector from a set of candidate motion vectors can be determined based on information of neighboring blocks associated with the target block of the video. In some embodiments, motion estimation of the filtering process is performed based on the target motion vector. In some embodiments, a bitstream of the target block is generated based on motion estimation. In some embodiments, the bitstream is stored in a non-transient computer readable recording medium.

[0104] Fig.12 A flow chart of a method 1200 for video processing according to some embodiments of the present disclosure is shown. The method 1200 may be implemented during conversion between a target block and a bitstream of the target block.

[0105] like Fig.12As shown, at block 1210, during conversion between a target block of a video and a bitstream of the target block, an error is determined, the error including neighboring information of the target block. For example, during filtering in MCTF, the error derived for each filtered block (e.g., as mentioned in Section 2) may include neighboring information.

[0106] At box 1220, the filtering process is performed based on the error. At box 1230, the conversion is performed according to the filtering process. In some embodiments, the conversion may include encoding the target block into a bitstream. Alternatively, the conversion may include decoding the target block from the bitstream. Compared with conventional schemes, boundary artifacts of adjacent blocks can be avoided, and the filtering process performance can be improved.

[0107] Implementations of the present disclosure may be described in view of the following items, and features of these items may be combined in any reasonable manner.

[0108] In some embodiments, the proximity information may be expressed as: In this case, W 0 represents the initial value, W j represents the jth value, T represents the first cost between the target block and the reference block corresponding to the candidate motion vector, K j represents the second cost between the j-th neighboring block of the target block and the j-th neighboring block of the reference block, S represents the total number of neighboring blocks, j can be an integer and 1≤j≤S.

[0109] In some embodiments, W 0 , W 1 ....W s can have the same value. Alternatively, W 0 , W 1 ....W s Can have different values. In some embodiments, W 1 , W 2 ....W s can have the same value and value as W 0 The values ​​of are different.

[0110] In some embodiments, the second cost may be determined using a distortion metric. For example, the distortion metric may include at least one of the following: sum of absolute differences (SAD), sum of squared errors (SSE), or mean sum of squared errors (MSE). In one example, K j It can be calculated by a distortion metric such as SAD, SSE or MSE.

[0111] In some embodiments, an error is determined, the error comprising neighboring information of a target block of the video. In some embodiments, a filtering process is performed based on the error. In some embodiments, a bitstream of the target block is generated according to the filtering process.

[0112] In some embodiments, an error is determined, and the error includes neighboring information of the target block of the video. In some embodiments, the filtering process is performed based on the error. In some embodiments, the bitstream of the target block is generated according to the filtering process. In some embodiments, the bitstream is stored in a non-transient computer readable recording medium.

[0113] Fig.13 A flow chart of a method 1300 for video processing according to some embodiments of the present disclosure is shown. The method 1300 may be implemented during conversion between a target block and a bitstream of the target block.

[0114] like Fig.13 As shown, at block 1310, during conversion between a target block of a video and a bitstream of the target block, a filtering process is performed on a set of overlapping blocks associated with the target block. For example, the filtering process in MCTF can be performed on overlapping blocks.

[0115] At box 1320, conversion is performed according to the filtering process. In some embodiments, conversion may include encoding the target block into a bitstream. Alternatively, conversion may include decoding the target block from the bitstream. Compared with conventional schemes, the performance of the filtering process can be improved. For example, if the reference block can span two filtering blocks during the encoding process, the filter process performance can still be guaranteed.

[0116] Implementations of the present disclosure may be described in view of the following items, and features of these items may be combined in any reasonable manner.

[0117] In some embodiments, a width step and a height step may be used. For example, the width step and the height step may be different from the size of the filter block. In one example, a width step WS and a height step HS may be used, and they may not be equal to the size B×B of the filter block.

[0118] In some embodiments, at least one of the width step or the height step may be smaller than the size of the filter block. In one example, WS and / or HS may be smaller than B.

[0119] In some embodiments, after a block with position (X, Y) is filtered, the next block to be filtered may be at (X+WS, Y). In this case, X represents the horizontal position, Y represents the vertical position, and WS represents the width step.

[0120] In some embodiments, after all blocks with vertical position Y are filtered, the next block to be filtered may be at (X, Y+WS). In this case, X represents the horizontal position, Y represents the vertical position, and WS represents the width step.

[0121] In some embodiments, the size of the block to be filtered may be one of: B×B, WS×B, B×HS, or WS×HS, where B represents the size of the filter block, WS represents the width step and HS represents the height step.

[0122] In some embodiments, at least one of the error or noise for a group of overlapping blocks can be determined based on adjacent blocks. In one example, the error and / or noise for the overlapping area can be determined by the adjacent blocks involved. For example, the error for a group of overlapping blocks can be determined by weighting the error of a part of the adjacent blocks or the error of all adjacent blocks. Alternatively, the error for a group of overlapping blocks can be determined by averaging the error of a part of the adjacent blocks or the error of all adjacent blocks. In some embodiments, the noise for a group of overlapping blocks can be determined by weighting the noise of a part of the adjacent blocks or the noise of all adjacent blocks. Alternatively, the noise for a group of overlapping blocks can be determined by averaging the noise of a part of the adjacent blocks or the noise of all adjacent blocks. In one example, the error and / or noise can be calculated by weighting or averaging the error and / or noise of some or all adjacent blocks involved.

[0123] In some embodiments, the error of the adjacent blocks can be used as the error for a group of overlapping blocks. Alternatively, the noise of the adjacent blocks can be used as the noise for a group of overlapping blocks. In one example, the error and / or noise for the overlapping area can use the error and / or noise of one adjacent block.

[0124] In some embodiments, the filtering process is performed on a set of overlapping blocks associated with a target block of the video. In some embodiments, a bitstream for the target block is generated according to the filtering process.

[0125] In some embodiments, the filtering process is performed on a set of overlapping blocks associated with a target block of the video. In some embodiments, a bitstream of the target block is generated according to the filtering process. In some embodiments, the bitstream is stored in a non-transitory computer-readable recording medium.

[0126] Fig.14 Flowchart illustrating a method 1400 for video processing according to some embodiments of the present disclosure. The method 1400 may be implemented during conversion between a target block and a bitstream of the target block.

[0127] like Fig.14As shown, at block 1410, during the conversion between the target block of the video and the bitstream of the target block, the encoding manner of the frame associated with the target block is determined based on whether the filtering process is applied to the frame. In other words, how to encode a frame may depend on whether MCTF can be applied to the frame.

[0128] At block 1420, conversion is performed based on the determination. In some embodiments, the conversion may include encoding the target block into a bitstream. Alternatively, the conversion may include decoding the target block from the bitstream. Compared to conventional schemes, it is possible to avoid removing reference block components from the frame.

[0129] Implementations of the present disclosure may be described in view of the following items, and features of these items may be combined in any reasonable manner.

[0130] In some embodiments, a frame after a filtering process may be processed differently than another frame without a filtering process. In some embodiments, at least one of the following quantization parameters (QP) of a frame using an applied filtering process may be in a change that reduces the P value or increases the P value, at a slice level, a codec tree unit (CTU) level, a codec unit (CU) level, or a block level. In one example, the slice / CTU / CU / block level QP of mctf_frame may be reduced or increased. In this case, P may be any suitable value. For example, P may be an integer or a non-integer. In some embodiments, a reduced change or an increased change may be applied to a luma QP. In some embodiments, a reduced change or an increased change may be applied to a chroma QP. In some embodiments, a reduced change or an increased change may be applied to both a luma QP and a chroma QP.

[0131] In some embodiments, the intra-frame cost of some / all blocks in the frame using the applied filtering process can be reduced by Q. In this case, Q can be any suitable value. For example, Q can be an integer or a non-integer. In some embodiments, the skip cost of some / all blocks in the frame using the applied filtering process can be increased by V. In this case, V can be any suitable value. For example, V can be an integer or a non-integer.

[0132] In some embodiments, the codec information of at least one block may be determined differently for a frame using the applied filtering process and a frame not using the applied filtering process. In some embodiments, the codec information may include at least one of the following: a prediction mode, an intra-frame prediction mode, a quadtree partition flag, a binary tree partition type, a ternary tree partition type, a motion vector, a Merge flag, or a Merge index.

[0133] In some embodiments, whether and / or how to segment at least one of the following may be different for frames with the applied filtering process and frames without the applied filtering process: block, region, or CTU. In one example, whether and / or how to segment blocks / regions / CTUs may be different for mctf frames and non mctf frames. In some embodiments, the maximum depth of a CU in a frame with the applied filtering process may be increased.

[0134] In some embodiments, different motion search methods can be utilized for frames with and without the applied filtering process. In some embodiments, different fast intra mode algorithms can be utilized for frames with and without the applied filtering process.

[0135] In some embodiments, the screen content codec tool may not be allowed to encode frames that utilize the applied filtering process. For example, the screen content codec tool may include at least one of the following: palette mode, intra-block copy (IBC) mode, block-based delta pulse code modulation (BDPCM), adaptive color transform (ACT), or transform skip mode.

[0136] In some embodiments, the difference between the block with the filtering process applied and the original block may be used as a metric to determine if the block needs to be processed differently in the conversion.

[0137] In some embodiments, determining the encoding method of the frame can be applied in a condition. For example, the condition may be that the distortion of the original pixel and the filtered pixel exceeds a first threshold at one of the following: CTU level, CU level, or block level. In some embodiments, the condition may be that the distortion of the filtered current pixel and the filtered neighboring pixel exceeds a second threshold at one of the following: CTU level, CU level, or block level. In some embodiments, the distortion may include one of the following: SAD, SSE, or MSE. In some embodiments, the condition may be that one of the values ​​in the average motion vector exceeds a third threshold at one of the following: CTU level, CU level, or block level.

[0138] In some embodiments, the encoding manner of a frame associated with the target block of the video is determined based on whether a filtering process is applied to the frame. In some embodiments, a bitstream for the target block is generated based on the determination.

[0139] In some embodiments, the encoding mode of the frame associated with the target block of the video is determined based on whether the filtering process is applied to the frame. In some embodiments, the bitstream of the target block is generated based on the determination. In some embodiments, the bitstream is stored in a non-transitory computer-readable recording medium.

[0140] The embodiments of the present disclosure may be implemented individually. Alternatively, the embodiments of the present disclosure may be implemented in any appropriate combination. The implementation of the present disclosure may be described in view of the following items, and the features of these items may be combined in any reasonable manner.

[0141] Item 1. A method for video processing, comprising: during a conversion between a target block of a video and a bitstream of the target block, determining a target motion vector from a set of candidate motion vectors based on information of neighboring blocks associated with the target block; performing motion estimation of a filtering process based on the target motion vector; and performing the conversion according to the motion estimation.

[0142] Item 2. The method of Item 1, wherein the information of the neighboring blocks comprises a cost of the neighboring blocks.

[0143] Item 3. A method according to item 2, wherein the cost of the neighboring block depends on a candidate motion vector in the motion estimation of the target block.

[0144] Item 4. A method according to Item 2, wherein the final cost of the candidate motion vector to be checked for the target block is determined using a linear function of the costs associated with the target block and the neighboring blocks, or wherein the final cost of the candidate motion vector to be checked for the target block is determined using a non-linear function of the costs associated with the target block and the neighboring blocks.

[0145] Item 5. A method according to item 1, wherein in the motion estimation of the filtering process, the motion estimation difference includes neighboring information.

[0146] Item 6. A method according to Item 1, wherein the difference measure of the candidate motion vector includes at least one of the following: a first cost between the target block and a reference block corresponding to the candidate motion vector; or a second cost between the jth neighboring block of the target block and the jth neighboring block of the reference block, and wherein j is an integer.

[0147] Clause 7. The method of clause 6, wherein the difference metric is evaluated as: Where W 0 represents the initial value, W j represents the jth initial value, T represents the first cost, K j denotes the second cost, S denotes the total number of neighboring blocks, and j is an integer.

[0148] Item 8. The method according to Item 7, wherein W 0 , W 1 ....Ws have the same value; or W 0 , W 1 ....W s have different values.

[0149] Item 9. The method according to Item 7, wherein W 1 , W 2 ....W s have the same value and the value is the same as W o The values ​​of are different.

[0150] Clause 10. The method of clause 6, wherein at least one of the first cost or the second cost is determined using a distortion metric.

[0151] Item 11. The method of Item 10, wherein the distortion metric comprises at least one of the following: sum of absolute differences (SAD), sum of squared errors (SSE), or mean sum of squared errors (MSE).

[0152] Item 12. A method according to Item 1, wherein the neighboring blocks of the target block include at least one of the following: a top neighboring block of the target block, a bottom neighboring block of the target block, a left neighboring block of the target block, a right neighboring block of the target block, an upper left neighboring block of the target block, an upper right neighboring block of the target block, a lower left neighboring block of the target block, or a lower right neighboring block of the target block.

[0153] Item 13. A method according to Item 1, wherein the neighboring blocks of the reference block associated with the target block include at least one of the following: a top neighboring block of the reference block, a bottom neighboring block of the reference block, a left neighboring block of the reference block, a right neighboring block of the reference block, an upper left neighboring block of the reference block, an upper right neighboring block of the reference block, a lower left neighboring block of the reference block, or a lower right neighboring block of the reference block.

[0154] Item 14. A method according to item 12 or 13, wherein different block sizes are used for different adjacent blocks.

[0155] Item 15. The method of Item 12, wherein the first block size of the neighboring block is the same as the second block size of the target block, or wherein the first block size of the neighboring block is different from the second block size of the target block.

[0156] Item 16. A method according to Item 13, wherein the third block size of the neighboring block of the reference block is the same as the fourth block size of the reference block, or wherein the third block size of the neighboring block of the reference block is different from the fourth block size of the reference block.

[0157] Item 17. The method according to Item 12, wherein the size of the neighboring block is W×H, where W represents the width of the target block and H represents the height of the target block.

[0158] Item 18. The method according to Item 13, wherein the size of the neighboring block of the reference block is W×H, where W represents the width of the target block and H represents the height of the target block.

[0159] Item 19. The method according to Item 1, wherein W 0 , W 1 ....W s is determined based on the block sizes of one or more neighboring blocks.

[0160] Item 20. The method according to Item 1, wherein different neighboring information is adopted for different layers in the hierarchical motion estimation scheme.

[0161] Item 21. The method according to Item 20, wherein W 0 , W 1 ....W s has the same value for different layers in the hierarchical motion estimation scheme, or W 0 , W 1 ....W s has different values for different layers in the hierarchical motion estimation scheme.

[0162] Item 22. The method according to Item 20, wherein the total number of neighboring blocks is different for different layers in the hierarchical motion estimation scheme.

[0163] Item 23. The method according to Item 20, wherein the motion estimation with the neighboring information is applied to layer L1 and layer L0 in the hierarchical motion estimation scheme.

[0164] Item 24. The method according to any one of Items 1 to 23, wherein determining the target motion vector based on the information of the neighboring blocks is applied to at least one layer in the hierarchical motion estimation scheme.

[0165] Item 25. The method according to any one of Items 1 to 23, wherein whether to apply determining the target motion vector based on the information of the neighboring blocks is according to the different sizes of the target blocks in the hierarchical motion estimation scheme in the filtering process.

[0166] Item 26. A method of video processing, comprising: determining an error during a conversion between a target block of a video and a bitstream of the target block, the error comprising neighborhood information of the target block; performing a filtering process based on the error; and performing the conversion according to the filtering process.

[0167] Clause 27. The method of clause 26, wherein the proximity information is expressed as: Where W 0 represents the initial value, W j represents the jth initial value, T represents the first cost between the target block and the reference block corresponding to the candidate motion vector, K j represents a second cost between the j-th neighboring block of the target block and the j-th neighboring block of the reference block, S represents the total number of neighboring blocks, j is an integer and 1≤j≤S.

[0168] Item 28. The method according to Item 27, wherein W 0 , W 1 ....W s have the same value; or where W 0 , W 1 ....W s have different values.

[0169] Item 29. The method according to Item 27, wherein W 1 , W 2 ....W s have the same value and the value is the same as W 0 The values ​​of are different.

[0170] Clause 30. The method of clause 27, wherein the second cost is determined using a distortion metric.

[0171] Item 31. A method according to Item 30, wherein the distortion metric includes at least one of the following: sum of absolute differences (SAD), sum of squared errors (SSE), or mean sum of squared errors (MSE).

[0172] Item 32. A method of video processing, comprising: performing a filtering process on a set of overlapping blocks associated with a target block of a video during a conversion between the target block and a bitstream of the target block; and performing the conversion according to the filtering process.

[0173] Item 33. A method according to Item 32, wherein a width step and a height step are used, and wherein the width step and the height step are different from the size of the filter block.

[0174] Item 34. A method according to item 33, wherein at least one of the width step or the height step is smaller than the size of the filter block.

[0175] Item 35. A method according to Item 33, wherein after a block with position (X, Y) is filtered, the next block to be filtered is at (X+WS, Y), where X represents the horizontal position, Y represents the vertical position, and WS represents the width step.

[0176] Item 36. A method according to Item 33, wherein after all blocks with vertical position Y are filtered, the next block to be filtered is at (X, Y + WS), where X represents the horizontal position, Y represents the vertical position, and WS represents the width step.

[0177] Item 37. A method according to Item 32, wherein the size of the block to be filtered is one of: B×B, WS×B, B×HS, or WS×HS, where B represents the size of the filter block, WS represents the width step and HS represents the height step.

[0178] Item 38. The method of Item 32, wherein at least one of error or noise for the set of overlapping blocks is determined based on neighboring blocks.

[0179] Item 39. A method according to Item 38, wherein the error for the set of overlapping blocks is determined by weighting the errors of a portion of the adjacent blocks or the errors of all adjacent blocks, or wherein the error for the set of overlapping blocks is determined by averaging the errors of the portion of the adjacent blocks or the errors of all adjacent blocks.

[0180] Item 40. A method according to Item 38, wherein the noise for the set of overlapping blocks is determined by weighting the noise of a portion of the adjacent blocks or the noise of all adjacent blocks, or wherein the noise for the set of overlapping blocks is determined by averaging the noise of the portion of the adjacent blocks or the noise of all adjacent blocks.

[0181] Item 41. A method according to Item 32, wherein an error of a neighboring block is used as an error for the set of overlapping blocks, or wherein noise of the neighboring block is used as noise for the set of overlapping blocks.

[0182] Item 42. A method of video processing, comprising: during conversion between a target block of a video and a bitstream of the target block, determining an encoding method for the frame based on whether a filtering process is applied to a frame associated with the target block; and performing the conversion based on the determination.

[0183] Item 43. A method according to Item 42, wherein a frame after the filtering process is processed differently than another frame without the filtering process.

[0184] Item 44. A method according to Item 42, wherein at least one of the following quantization parameters (QP) of the frame to which the filtering process is applied is in a change that decreases the P value or increases the P value, at a slice level, a codec tree unit (CTU) level, a codec unit (CU) level, or a block level.

[0185] Item 45. A method according to Item 44, wherein the reduced change or the increased change is applied to the luma QP, or wherein the reduced change or the increased change is applied to the chroma QP, or wherein the reduced change or the increased change is applied to both the luma QP and the chroma QP.

[0186] Item 46. The method of Item 42, wherein the intra-frame cost of some / all blocks in a frame with the filtering process applied is reduced by a Q value.

[0187] Item 47. The method of Item 42, wherein the skip cost of some / all blocks in a frame with the filtering process applied is increased by a value V.

[0188] Item 48. The method of Item 42, wherein codec information of at least one block is determined differently for frames with the filtering process applied and for frames without the filtering process applied.

[0189] Item 49. The method according to Item 48, wherein the encoding and decoding information includes at least one of the following: prediction mode, intra-frame prediction mode, quadtree partition flag, binary tree partition type, ternary tree partition type, motion vector, Merge flag, or Merge index.

[0190] Item 50. A method according to item 42, wherein whether and / or how at least one of the following is segmented is different for frames with the filtering process applied and frames without the filtering process applied: blocks, regions, or CTUs.

[0191] Item 51. A method according to item 42, wherein the maximum depth of a CU in a frame with the filtering process applied is increased.

[0192] Item 52. The method of Item 42, wherein different motion search methods are utilized for frames with the filtering process applied and frames without the filtering process applied.

[0193] Item 53. The method of Item 42, wherein different fast intra mode algorithms are utilized for frames with the filtering process applied and frames without the filtering process applied.

[0194] Item 54. A method according to Item 42, wherein a screen content codec tool is not allowed to be used to encode and decode frames that utilize the applied filtering process.

[0195] Item 55. A method according to Item 54, wherein the screen content codec tool includes at least one of the following: palette mode, intra block copy (IBC) mode, block-based delta pulse code modulation (BDPCM), adaptive color transform (ACT), or transform skip mode.

[0196] Item 56. The method of Item 42, wherein the difference between a block with the filtering process applied and an original block is used as a metric to determine whether the block needs to be processed differently in the conversion.

[0197] Item 57. A method according to any one of Items 42 to 56, wherein determining the encoding method of the frame is applied in a condition.

[0198] Item 58. A method according to item 57, wherein the condition is that the distortion of the original pixel and the filtered pixel exceeds a first threshold at one of the following: CTU level, CU level or block level.

[0199] Item 59. A method according to Item 57, wherein the condition is that the distortion of the filtered current pixel and the filtered neighboring pixels exceeds a second threshold at one of the following: CTU level, CU level or block level.

[0200] Item 60. A method according to item 58 or 59, wherein the distortion comprises one of the following: SAD, SSE, or MSE.

[0201] Item 61. A method according to item 57, wherein the condition is that one of the values ​​in the average motion vector exceeds a third threshold at one of the following: CTU level, CU level or block level.

[0202] Item 62. A method according to any one of items 1 to 61, wherein the block size of the target block used in the filtering process is not taken into account.

[0203] Item 63. A method according to Item 62, wherein at least one of the width or the height of the target block is greater than or equal to 4, or wherein at least one of the width or the height of the target block is less than or equal to 64, or wherein at least one of the width or the height of the target block is equal to 8.

[0204] Item 64. The method according to any one of Items 1 to 61, wherein at least one of the width of the target block, the height of the target block, the width step, the height step, the size of the filter block, P, Q, V, X, U, or Z is an integer and depends on: the slice group type, the picture group type, the picture type, the color component, the temporal layer identifier, the layer identifier in the pyramid motion estimation search; the profile of the standard, the level of the standard, or the layer of the standard.

[0205] Item 65. The method according to any one of Items 1 to 61, wherein the filtering process includes at least one of the following: motion compensated temporal filter (MCTF), MCTF related variance, bilateral filter, low pass filter, high pass filter, or loop filter.

[0206] Item 66. The method according to any one of Items 1 to 65, wherein the transformation includes encoding the target block into the bitstream.

[0207] Item 67. The method according to any one of Items 1 to 65, wherein the transformation includes decoding the target block from the bitstream.

[0208] Item 68. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 67.

[0209] Item 69. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 67.

[0210] Item 70. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus, wherein the method includes: determining a target motion vector from a set of candidate motion vectors based on information of neighboring blocks associated with a target block of the video; performing motion estimation of a filtering process based on the target motion vector; and generating a bitstream of the target block according to the motion estimation.

[0211] Item 71. A method for storing a bitstream of a video, comprising: determining a target motion vector from a set of candidate motion vectors based on information of neighboring blocks associated with a target block of the video; performing motion estimation of a filtering process based on the target motion vector; generating a bitstream of the target block according to the motion estimation; and storing the bitstream in a non-transitory computer-readable recording medium.

[0212] Item 72. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: determining an error, the error including neighborhood information of a target block of the video; performing a filtering process based on the error; and generating a bitstream of the target block according to the filtering process.

[0213] Item 73. A method for storing a bitstream of a video, comprising: determining an error, the error including neighborhood information of a target block of the video; performing a filtering process based on the error; generating a bitstream of the target block according to the filtering process; and storing the bitstream in a non-transitory computer-readable recording medium.

[0214] Item 74. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: performing a filtering process on a set of overlapping blocks associated with a target block of the video; and generating a bitstream of the target block based on the filtering process.

[0215] Item 75. A method for storing a bitstream of a video, comprising: performing a filtering process on a set of overlapping blocks associated with a target block of the video; generating a bitstream of the target block based on the filtering process; and storing the bitstream in a non-transitory computer-readable recording medium.

[0216] Item 76. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: determining an encoding method for a frame associated with a target block of the video based on whether a filtering process is applied to the frame; and generating a bitstream of the target block based on the determination.

[0217] Item 77. A method for storing a bitstream of a video, comprising: determining an encoding method for a frame associated with a target block of the video based on whether a filtering process is applied to the frame; generating a bitstream for the target block based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium. Example Device

[0218] Fig.15 A block diagram of a computing device 1500 in which various embodiments of the present disclosure may be implemented is shown. The computing device 1500 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0219] It should be understood that Fig.15 The computing device 1500 shown in FIG. 1 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.

[0220] like Fig.15 As shown, computing device 1500 includes a general computing device 1500. Computing device 1500 may include at least one or more processors or processing units 1510, memory 1520, storage unit 1530, one or more communication units 1540, one or more input devices 1550, and one or more output devices 1560.

[0221] In some embodiments, computing device 1500 can be implemented as any user terminal or server terminal with computing power. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a gaming device, or any combination thereof, including the accessories and peripherals of these devices, or any combination thereof. It is conceivable that computing device 1500 can support any type of interface to the user (such as a "wearable" circuit device, etc.).

[0222] The processing unit 1510 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 1520. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to increase the parallel processing capability of the computing device 1500. The processing unit 1510 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0223] The computing device 1500 typically includes various computer storage media. Such media can be any media accessible by the computing device 1500, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 1520 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (such as a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 1530 can be any removable or non-removable medium, and can include machine-readable media, such as a memory, a flash drive, a disk, or other media that can be used to store information and / or data and can be accessed in the computing device 1500.

[0224] The computing device 1500 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Fig.15 Although not shown in the figure, a disk drive for reading from and / or writing to a removable nonvolatile disk and an optical drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to the bus (not shown) via one or more data medium interfaces.

[0225] The communication unit 1540 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 1500 can be implemented by a single computing cluster or multiple computing machines, which can communicate via a communication connection. Therefore, the computing device 1500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general network nodes.

[0226] Input device 1550 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 1560 may be one or more of various output devices, such as a display, speaker, printer, etc. With the aid of communication unit 1540, computing device 1500 may also communicate with one or more external devices (not shown), such as storage devices and display devices, and may also communicate with one or more devices that enable a user to interact with computing device 1500, or, if necessary, may also communicate with any device (e.g., a network card, a modem, etc.) that enables computing device 1500 to communicate with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).

[0227] In some embodiments, some or all components of the computing device 1500 may also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage services, which will not require the end user to know the physical location or configuration of the system or hardware that provides these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using a suitable protocol. For example, a cloud computing provider provides an application via a wide area network, which can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. The computing resources in a cloud computing environment may be merged or distributed at the location of a remote data center. Cloud computing infrastructure can provide services through a shared data center, although they appear as a single access point to the user. Therefore, the cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server, or may be installed on a client device directly or otherwise.

[0228] In an embodiment of the present disclosure, the computing device 1500 may be used to implement video encoding / decoding. The memory 1520 may include one or more video codec modules 1525 having one or more program instructions. These modules are accessible and executable by the processing unit 1510 to perform the functions of the various embodiments described herein.

[0229] In an example embodiment performing video encoding, input device 1550 may receive video data as input 1570 to be encoded. The video data may be processed, for example, by video codec module 1525 to generate an encoded bitstream. The encoded bitstream may be provided as output 1580 via output device 1560.

[0230] In an example embodiment performing video decoding, input device 1550 may receive an encoded bitstream as input 1570. The encoded bitstream may be processed, for example, by video codec module 1525 to generate decoded video data. The decoded video data may be provided as output 1580 via output device 1560.

[0231] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be appreciated by those skilled in the art that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These modifications are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A video processing method, include: During a conversion between a target block of a video and a bitstream of the target block, determining a target motion vector from a set of candidate motion vectors based on information of neighboring blocks associated with the target block; performing motion estimation of a filtering process based on the target motion vector; as well as The conversion is performed based on the motion estimate. The method according to claim 1 , wherein the information of the neighboring blocks comprises a cost of the neighboring blocks.

3. The method of claim 2, wherein the cost of the neighboring blocks depends on candidate motion vectors in the motion estimation of the target block.

4. The method of claim 2, wherein a final cost of the candidate motion vector to be checked for the target block is determined using a linear function of the costs associated with the target block and the neighboring blocks, or Wherein the final cost of the candidate motion vector to be checked against the target block is determined using a non-linear function of costs associated with the target block and the neighboring blocks. The method of claim 1 , wherein in the motion estimation of the filtering process, motion estimation differences include neighbor information.

6. The method of claim 1 , wherein the difference metric of the candidate motion vectors comprises at least one of the following: a first cost between the target block and a reference block corresponding to the candidate motion vector; or A second cost between a j-th neighboring block of the target block and a j-th neighboring block of the reference block, wherein j is an integer.

7. The method of claim 6, wherein the difference metric is evaluated as: Where W 0 represents the initial value, W j represents the jth initial value, T represents the first cost, K j denotes the second cost, S denotes the total number of neighboring blocks, and j is an integer.

8. The method according to claim 7, wherein W 0 , W 1 ....W s have the same value; or W 0 , W 1 ....W s have different values.

9. The method according to claim 7, wherein W 1 , W 2 ....W s have the same value and the value is the same as W 0 The values ​​of are different.

10. The method of claim 6, wherein at least one of the first cost or the second cost is determined using a distortion metric.

11. The method of claim 10, wherein the distortion metric comprises at least one of: Sum of Absolute Differences (SAD), The sum of squared errors (SSE), or Mean Sum of Squared Error (MSE).

12. The method according to claim 1, wherein the neighboring blocks of the target block include at least one of the following: The top neighboring block of the target block, The bottom neighboring block of the target block, The left neighboring block of the target block, The right neighboring block of the target block, The upper left neighboring block of the target block, The upper right neighboring block of the target block, The lower left neighboring block of the target block, or The lower right neighboring block of the target block.

13. The method according to claim 1, wherein the neighboring blocks of the reference block associated with the target block include at least one of the following: the top neighboring block of the reference block, the bottom neighboring block of the reference block, the left neighboring block of the reference block, the right neighboring block of the reference block, The upper left neighboring block of the reference block, The upper right neighboring block of the reference block, The lower left neighboring block of the reference block, or The lower right neighboring block of the reference block.

14. A method according to claim 12 or 13, wherein different block sizes are used for different neighbouring blocks.

15. The method according to claim 12, wherein the first block size of the neighboring block is the same as the second block size of the target block, or The first block size of the neighboring block is different from the second block size of the target block.

16. The method according to claim 13, wherein the third block size of the neighboring block of the reference block is the same as the fourth block size of the reference block, or The third block size of the neighboring block of the reference block is different from the fourth block size of the reference block. 17 . The method of claim 12 , wherein the neighboring block has a size of W×H, wherein W represents a width of the target block and H represents a height of the target block. 18 . The method of claim 13 , wherein the neighboring block of the reference block has a size of W×H, wherein W represents a width of the target block and H represents a height of the target block.

19. The method according to claim 1, wherein W 0 , W 1 ....W s At least one of is determined based on a block size of one or more neighboring blocks.

20. The method of claim 1, wherein different neighbor information is employed for different layers in a hierarchical motion estimation scheme.

21. The method according to claim 20, wherein W 0 , W 1 ....W s have the same value for different layers in the hierarchical motion estimation scheme, or W 0 , W 1 ....W s There are different values ​​for different layers in the hierarchical motion estimation scheme.

22. The method of claim 20, wherein the total number of neighboring blocks is different for different layers in the hierarchical motion estimation scheme.

23. The method of claim 20, wherein the motion estimation with the neighbor information is applied to L1 layer and L0 layer in the hierarchical motion estimation scheme.

24. The method according to any one of claims 1 to 23, wherein determining the target motion vector based on the information of the neighboring blocks is applied to at least one layer in a hierarchical motion estimation scheme.

25. The method according to any one of claims 1 to 23, wherein determining whether the target motion vector is applied based on the information of the neighboring blocks is based on different sizes of the target blocks in a hierarchical motion estimation scheme in the filtering process.

26. A method of video processing, include: During conversion between a target block of video and a bitstream of the target block, determining an error, the error comprising neighborhood information of the target block; performing a filtering process based on the error; as well as The converting is performed according to the filtering process.

27. The method according to claim 26, wherein the proximity information is expressed as: Where W 0 represents the initial value, W j represents the jth initial value, T represents the first cost between the target block and the reference block corresponding to the candidate motion vector, K j represents a second cost between the j-th neighboring block of the target block and the j-th neighboring block of the reference block, S represents the total number of neighboring blocks, j is an integer and 1≤j≤S.

28. The method of claim 27, wherein W 0 , W 1 ....W s have the same value; or Where W 0 , W 1 ....W s have different values.

29. The method of claim 27, wherein W 1 , W 2 ....W s have the same value and the value is the same as W 0 The values ​​of are different.

30. The method of claim 27, wherein the second cost is determined using a distortion metric.

31. The method of claim 30, wherein the distortion metric comprises at least one of: Sum of Absolute Differences (SAD), The sum of squared errors (SSE), or Mean Sum of Squared Error (MSE).

32. A method of video processing, include: During conversion between a target block of video and a bitstream of the target block, performing a filtering process on a set of overlapping blocks associated with the target block; as well as The converting is performed according to the filtering process.

33. The method of claim 32, wherein a width step and a height step are used, and Wherein the width step and the height step are different from the size of the filter block.

34. A method according to claim 33, wherein at least one of the width step or the height step is smaller than the size of the filter block.

35. The method of claim 33, wherein after a block with position (X, Y) is filtered, the next block to be filtered is at (X+WS, Y), where X represents horizontal position, Y represents vertical position, and WS represents the width step.

36. The method of claim 33, wherein after all blocks with vertical position Y are filtered, the next block to be filtered is at (X, Y + WS), where X represents horizontal position, Y represents vertical position, and WS represents the width step.

37. The method of claim 32, wherein the size of the block to be filtered is one of: B×B, WS×B, B×HS, or WS×HS, Where B represents the size of the filter block, WS represents the width step and HS represents the height step.

38. The method of claim 32, wherein at least one of error or noise for the set of overlapping blocks is determined based on neighboring blocks.

39. The method of claim 38, wherein the error for the set of overlapping blocks is determined by weighting the errors of a portion of the neighboring blocks or the errors of all neighboring blocks, or The error for the set of overlapping blocks is determined by averaging errors of the portion of the neighboring blocks or errors of all the neighboring blocks.

40. The method of claim 38, wherein the noise for the set of overlapping blocks is determined by weighting the noise of a portion of the neighboring blocks or the noise of all neighboring blocks, or The noise for the set of overlapping blocks is determined by averaging noise of the portion of the neighboring blocks or noise of all the neighboring blocks.

41. The method of claim 32, wherein an error of a neighboring block is used as an error for the set of overlapping blocks, or The noise of the neighboring blocks is used as noise for the set of overlapping blocks.

42. A method of video processing, include: During conversion between a target block of video and a bitstream of the target block, determining an encoding manner of the frame associated with the target block based on whether a filtering process is applied to the frame; as well as The converting is performed based on the determination.

43. The method of claim 42, wherein a frame after the filtering process is processed differently than another frame without the filtering process.

44. The method of claim 42, wherein at least one of the following quantization parameters (QP) of a frame with the filtering process applied is in a variation that decreases the P value or increases the P value, Stripe level, Codec Tree Unit (CTU) level, Codec Unit (CU) level, or Block level.

45. The method of claim 44, wherein the decreasing change or the increasing change is applied to a luma QP, or wherein the decreasing change or the increasing change is applied to the chroma QP, or Wherein the decreasing change or the increasing change is applied to both luma QP and chroma QP.

46. ​​The method of claim 42, wherein the intra-frame cost of some / all blocks in a frame with the filtering process applied is reduced by a Q value.

47. The method of claim 42, wherein the skip cost of some / all blocks in a frame with the filtering process applied is increased by a value of V.

48. The method of claim 42, wherein codec information of at least one block is determined differently for frames with the filtering process applied and frames without the filtering process applied.

49. The method according to claim 48, wherein the codec information comprises at least one of the following: Prediction model, Intra prediction mode, Quadtree partition flag, Binary tree partition type, Ternary tree partition type, Motion vector, Merge flag, or Merge index.

50. The method of claim 42, wherein whether and / or how at least one of the following is segmented is different for frames with the filtering process applied and frames without the filtering process applied: piece, Region, or CTU.

51. The method of claim 42, wherein a maximum depth of a CU in a frame utilizing the filtering process applied is increased.

52. The method of claim 42, wherein different motion search methods are utilized for frames with the filtering process applied and frames without the filtering process applied.

53. The method of claim 42, wherein different fast intra mode algorithms are utilized for frames with the filtering process applied and frames without the filtering process applied.

54. The method of claim 42, wherein screen content codec tools are not permitted to be used to encode and decode frames that utilize the filtering process applied.

55. The method according to claim 54, wherein the screen content codec tool comprises at least one of the following: Palette mode, Intra Block Copy (IBC) mode, Block-based delta pulse code modulation (BDPCM), Adaptive Color Transformation (ACT), or Transform skip mode.

56. The method of claim 42, wherein a difference between a block with the filtering process applied and an original block is used as a metric to determine if the block needs to be processed differently in the conversion.

57. The method according to any one of claims 42 to 56, wherein determining the encoding manner of the frame is applied in a condition.

58. The method of claim 57, wherein the condition is that distortion of original pixels and filtered pixels exceeds a first threshold at one of: CTU level, CU level, or block level.

59. The method of claim 57, wherein the condition is that distortion of a filtered current pixel and filtered neighboring pixels exceeds a second threshold at one of: CTU level, CU level, or block level.

60. The method of claim 58 or 59, wherein the distortion comprises one of: SAD, SSE, or MSE.

61. The method of claim 57, wherein the condition is that one of the values ​​in the average motion vector exceeds a third threshold at one of: CTU level, CU level, or block level.

62. A method according to any one of claims 1 to 61, wherein the block size of the target block used in the filtering process is not taken into account.

63. The method of claim 62, wherein at least one of the width or height of the target block is greater than or equal to 4, or wherein at least one of the width or the height of the target block is less than or equal to 64, or Wherein at least one of the width or the height of the target block is equal to 8.

64. A method according to any one of claims 1 to 61, wherein at least one of the width of the target block, the height of the target block, the width step, the height step, the size of the filter block, P, Q, V, X, Y or Z is an integer and depends on: Stripe group type, Chipset type, Image type, Color component, Time domain layer identification, Layer identification in pyramid motion estimation search; Standard grade, the level of the standard, or The standard layer.

65. The method of any one of claims 1 to 61, wherein the filtering process comprises at least one of the following: Motion Compensated Temporal Filter (MCTF), MCTF related variance, Bilateral filter, Low pass filter, High pass filter, or Loop filter.

66. A method according to any one of claims 1 to 65, wherein the conversion comprises encoding the target block into the bitstream.

67. A method according to any one of claims 1 to 65, wherein the conversion comprises decoding the target block from the bitstream.

68. An apparatus for processing video data, comprising a processor and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 67.

69. A non-transitory computer-readable storage medium storing instructions for causing a processor to execute the method according to any one of claims 1 to 67.

70. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method executed by a video processing device, wherein the method include: Determining a target motion vector from a set of candidate motion vectors based on information of neighboring blocks associated with the target block of the video; performing motion estimation of a filtering process based on the target motion vector; as well as A bitstream of the target block is generated based on the motion estimation.

71. A method for storing a bit stream of a video, include: Determining a target motion vector from a set of candidate motion vectors based on information of neighboring blocks associated with the target block of the video; performing motion estimation of a filtering process based on the target motion vector; generating a bitstream of the target block according to the motion estimation; as well as The bit stream is stored in a non-transitory computer-readable recording medium.

72. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method executed by a video processing device, wherein the method include: determining an error, the error comprising neighboring information of a target block of the video; performing a filtering process based on the error; as well as A bitstream of the target block is generated according to the filtering process.

73. A method for storing a bit stream of a video, include: determining an error, the error comprising neighboring information of a target block of the video; performing a filtering process based on the error; generating a bitstream of the target block according to the filtering process; as well as The bit stream is stored in a non-transitory computer-readable recording medium.

74. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method executed by a video processing device, wherein the method include: performing a filtering process on a set of overlapping blocks associated with a target block of the video; as well as A bitstream of the target block is generated according to the filtering process.

75. A method for storing a bit stream of a video, include: performing a filtering process on a set of overlapping blocks associated with a target block of the video; generating a bitstream of the target block according to the filtering process; as well as The bit stream is stored in a non-transitory computer-readable recording medium.

76. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method executed by a video processing device, wherein the method include: determining an encoding manner for a frame associated with a target block of the video based on whether a filtering process is applied to the frame; as well as A bitstream of the target block is generated based on the determination.

77. A method for storing a bit stream of a video, include: determining an encoding manner for a frame associated with a target block of the video based on whether a filtering process is applied to the frame; generating a bitstream of the target block based on the determining; as well as The bit stream is stored in a non-transitory computer-readable recording medium.