Video processing method and device, medium and electronic equipment

Through intra-frame block copy prediction and filtering processing, the reference block and template area fitting filter are used to solve the discontinuity problem between the image block and the surrounding pixels, and improve the image encoding and decoding quality.

CN120676172APending Publication Date: 2025-09-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410334152.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In traditional video coding and decoding schemes, discontinuous boundaries often appear between image blocks and surrounding pixels, resulting in poor image quality.

Method used

Through intra-block copy prediction and filtering, the current block is filtered using the reference block and template region fitting filter to eliminate spatial discontinuity.

Benefits of technology

The image encoding and decoding quality is improved and the spatial discontinuity problem between the current block and the surrounding pixels is eliminated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676172A_ABST
    Figure CN120676172A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of video encoding and decoding, and particularly relates to a video processing method, a video processing device, a computer readable medium, electronic equipment and a computer program product. The method comprises the following steps: performing intra-frame block copy prediction according to a reference block corresponding to a current block to obtain a first prediction value of the current block; determining a first template region corresponding to the reference block and a second template region corresponding to the current block, the first template region comprising one or more reconstructed image regions adjacent to the reference block and the second template region comprising one or more reconstructed image regions adjacent to the current block; and fitting a filter according to the first template region and the second template region, and filtering the first predicted value of the current block according to the filter to obtain a second predicted value of the current block. The coding and decoding quality of the audio and video data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of video coding and decoding technology, and specifically relates to a video processing method, a video processing device, a computer-readable medium, an electronic device, and a computer program product. Background Art

[0002] In order to adapt to the large-scale data transmission of audio and video data, it is usually necessary to encode the original audio and video data at the data sending end to form a compressed data stream. After transmitting the data stream to the data receiving end, the data stream is decoded and restored to obtain the predicted and reconstructed audio and video data.

[0003] Taking video coding as an example, traditional video coding solutions need to divide video images into non-overlapping image blocks. In some coding modes, discontinuous boundaries between image blocks and surrounding pixels are likely to appear, resulting in poor image quality. Summary of the Invention

[0004] The present application provides a video processing method, a video processing device, a computer-readable medium, an electronic device, and a computer program product, the purpose of which is to improve the encoding and decoding quality of audio and video data.

[0005] According to one aspect of an embodiment of the present application, a video processing method is provided, the method comprising:

[0006] performing intra-block copy prediction based on a reference block corresponding to a current block to obtain a first prediction value of the current block, wherein the current block is an image block to be processed in a current video frame, and the reference block is a reconstructed image block in the current video frame;

[0007] determining a first template region corresponding to the reference block and a second template region corresponding to the current block, wherein the first template region includes one or more reconstructed image regions adjacent to the reference block, and the second template region includes one or more reconstructed image regions adjacent to the current block;

[0008] A filter is fitted according to the first template area and the second template area, and a first prediction value of the current block is filtered according to the filter to obtain a second prediction value of the current block.

[0009] According to one aspect of an embodiment of the present application, a video processing device is provided, the device comprising:

[0010] a prediction module configured to perform intra-block copy prediction based on a reference block corresponding to a current block to obtain a first prediction value of the current block, wherein the current block is an image block to be processed in a current video frame, and the reference block is a reconstructed image block in the current video frame;

[0011] a determining module configured to determine a first template region corresponding to the reference block and a second template region corresponding to the current block, wherein the first template region includes one or more reconstructed image regions adjacent to the reference block, and the second template region includes one or more reconstructed image regions adjacent to the current block;

[0012] The filtering module is configured to fit a filter according to the first template area and the second template area, and filter the first prediction value of the current block according to the filter to obtain a second prediction value of the current block.

[0013] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the video processing method in the above technical solution is implemented.

[0014] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement the video processing method in the above technical solution.

[0015] According to one aspect of an embodiment of the present application, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the video processing method in the above technical solution is implemented.

[0016] In the technical solution provided in the embodiments of the present application, a first prediction value of the current block is obtained by performing intra-block copy prediction based on a reference block corresponding to the current block. A filter is then fitted based on a first template region corresponding to the reference block and a second template region corresponding to the current block, and the first prediction value of the current block is filtered using the filter to obtain a second prediction value of the current block. By filtering the first prediction value of the current block using the fitted filter, spatial discontinuity between the current block and surrounding pixels can be eliminated, thereby improving image encoding and decoding quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0018] Figure 1 A schematic diagram of a system architecture to which the technical solutions of the embodiments of the present application can be applied is shown.

[0019] Figure 2 The diagram shows how the video encoding device and the video decoding device are placed in a streaming environment.

[0020] Figure 3 The basic flow chart of the encoding process performed by the video encoder is shown.

[0021] Figure 4 A schematic diagram showing the principle of the inter-frame prediction mode is shown.

[0022] Figure 5 A schematic diagram showing the principle of the intra block copy prediction mode is shown.

[0023] Figure 6 A flowchart of a video processing method in one embodiment of the present application is shown.

[0024] Figure 7 A schematic diagram of the mapping form of the filter in one embodiment of the present application is shown.

[0025] Figure 8 A schematic diagram of the filter shape used in some application scenarios of the embodiments of the present application is shown.

[0026] Figure 9 A schematic diagram showing the selection of sampling positions based on the position coordinates of pixel points in one embodiment of the present application is shown.

[0027] Figure 10 A schematic diagram of selecting sampling positions based on a round-trip scanning method in one embodiment of the present application is shown.

[0028] Figure 11 A schematic diagram of selecting sampling positions based on the ZigZag scanning method in one embodiment of the present application is shown.

[0029] Figure 12 A schematic diagram showing the distribution of template areas corresponding to image blocks in one embodiment of the present application is shown.

[0030] Figure 13 A schematic diagram showing a type of forming different template areas with some sub-regions and image blocks in one embodiment of the present application.

[0031] Figure 14 A schematic diagram of a template area selected for an image block in one embodiment of the present application is shown.

[0032] Figure 15 A schematic diagram showing the distribution of template areas used in an application scenario in an embodiment of the present application is shown.

[0033] Figure 16A schematic diagram shows the effect of expanding the template area in an application scenario according to an embodiment of the present application.

[0034] Figure 17 The structural block diagram of the video processing device provided in an embodiment of the present application is schematically shown.

[0035] Figure 18 The following schematically shows a block diagram of a computer system structure of an electronic device suitable for implementing an embodiment of the present application. DETAILED DESCRIPTION

[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.

[0037] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0038] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0039] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0040] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0041] Video coding generally refers to the processing of a sequence of pictures to form a video or video sequence. In the field of video coding, the terms "picture", "frame" or "image" can be used as synonyms. The video coding used in the embodiments of the present application represents video encoding or video decoding. Video encoding is performed on the source side and generally includes processing (for example, by compression) the original video image to reduce the amount of data required to represent the video image, thereby more efficiently storing and / or transmitting. Video decoding is performed on the destination side and generally includes inverse processing relative to the encoder to reconstruct the video image. The "encoding" of the video frames involved in the embodiments should be understood as "encoding" or "decoding" involving a sequence of video images. The combination of the encoding part and the decoding part is also called codec (encoding and decoding).

[0042] Each image in a video image sequence is typically divided into a set of non-overlapping blocks, which are typically encoded at the block level. In other words, the encoder typically processes, or encodes, the video at the block level (also called image blocks or video blocks), for example, by generating a prediction block through spatial (intra-image) and temporal (inter-image) prediction, subtracting the prediction block from the current block (the block currently being processed or to be processed) to obtain a residual block, transforming the residual block in the transform domain, and quantizing the residual block to reduce the amount of data to be transmitted (compressed). The decoder then applies the inverse of the encoder's processing to the coded or compressed block to reconstruct the current block for representation. Furthermore, the encoder replicates the decoder processing loop, so that the encoder and decoder generate the same predictions (e.g., intra-frame predictions and inter-frame predictions) and / or reconstructions for processing, or encoding, subsequent blocks.

[0043] The term "block" refers to a portion of an image or frame. In the embodiments of the present application, the current block refers to the block currently being processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.

[0044] Figure 1 A schematic diagram of a system architecture to which the technical solutions of the embodiments of the present application can be applied is shown.

[0045] like Figure 1 As shown, the system architecture 100 includes a plurality of terminal devices, which can communicate with each other via, for example, a network 150. For example, the system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via the network 150. Figure 1 In the embodiment of the present invention, the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.

[0046] For example, the first terminal device 110 can encode video data (such as a video picture stream captured by the terminal device 110) for transmission to the second terminal device 120 via the network 150. The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device 120 can receive the encoded video data from the network 150, decode the encoded video data to restore the video data, and display the video picture based on the restored video data.

[0047] In one embodiment of the present application, the system architecture 100 may include a third terminal device 130 and a fourth terminal device 140 for performing bidirectional transmission of encoded video data, such as during a video conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode video data (e.g., a video image stream captured by the terminal device) for transmission to the other of the third terminal device 130 and the fourth terminal device 140 via a network 150. Each of the third terminal device 130 and the fourth terminal device 140 may also receive the encoded video data transmitted by the other of the third terminal device 130 and the fourth terminal device 140, decode the encoded video data to recover the video data, and display the video image on an accessible display device based on the recovered video data.

[0048] exist Figure 1 In the embodiment of the present invention, the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140 may be servers, personal computers, and smartphones, but the principles disclosed in this application are not limited thereto. The embodiments disclosed in this application are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 150 represents any number of networks that transmit encoded video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140, including, for example, wired and / or wireless communication networks. The communication network 150 can exchange data using circuit-switched and / or packet-switched channels. The network may include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this application, unless otherwise explained below, the architecture and topology of network 150 may be irrelevant to the operations disclosed in this application.

[0049] In one embodiment of the present application, Figure 2 The present invention illustrates the placement of a video encoding device and a video decoding device in a streaming environment. The subject matter disclosed herein is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV (television), and storing compressed video on digital media such as CDs, DVDs, and memory sticks.

[0050] The streaming system may include an acquisition subsystem 213, which may include a video source 201, such as a digital camera, that creates an uncompressed video picture stream 202. In one embodiment, the video picture stream 202 includes samples captured by the digital camera. The video picture stream 202 is depicted as a thicker line to emphasize the higher data volume of the video picture stream compared to the encoded video data 204 (or the encoded video stream 204). The video picture stream 202 may be processed by an electronic device 220, which includes a video encoding device 203 coupled to the video source 201. The video encoding device 203 may include hardware, software, or a combination of hardware and software to implement or embody various aspects of the disclosed subject matter, as described in greater detail below. The encoded video data 204 (or the encoded video stream 204) is depicted as a thinner line to emphasize the lower data volume of the encoded video data 204 (or the encoded video stream 204), which may be stored on a streaming server 205 for future use. One or more streaming client subsystems, such as Figure 2 , client subsystem 206 and client subsystem 208 in the streaming server 205 can access the streaming server 205 to retrieve the copies 207 and 209 of the encoded video data 204. The client subsystem 206 can include, for example, a video decoding device 210 in the electronic device 230. The video decoding device 210 decodes the incoming copy 207 of the encoded video data and produces an output video picture stream 211 that can be presented on a display 212 (e.g., a display screen) or another presentation device. In some streaming systems, the encoded video data 204, video data 207, and video data 209 (e.g., video bitstreams) can be encoded according to certain video encoding / compression standards.

[0051] It should be noted that the electronic device 220 and the electronic device 230 may include other components not shown in the figure. For example, the electronic device 220 may include a video decoding device, and the electronic device 230 may also include a video encoding device.

[0052] In one embodiment of the present application, taking the international video coding standards HEVC (High Efficiency Video Coding, H.265), VVC (Versatile Video Coding, H.266), and China's national video coding standard AVS (Audio Video Coding Standard) as examples, after a video frame image is input, the video frame image will be divided into several non-overlapping processing units according to a block size, and each processing unit will perform similar compression operations. This processing unit is called CTU (Coding Tree Unit, coding tree unit), or LCU (Largest Coding Unit). The CTU can continue to be divided more finely to obtain one or more basic coding units CU. CU is the most basic element in a coding link.

[0053] Figure 3 The basic flow chart of the encoding process performed by the video encoder is shown, in which intra-frame prediction is used as an example for explanation.

[0054] Among them, the original image signal s k [x,y] and predicted image signal Perform difference operation to obtain the residual signal u k [x,y], residual signal u k [x,y] is transformed and quantized to obtain quantized coefficients. The quantized coefficients are entropy coded to obtain the encoded bit stream, and the reconstructed residual signal is obtained by inverse quantization and inverse transformation. u'k[x,y] , predict image signal and the reconstructed residual signal u' k [x,y] superposition generates image signal Image signal On the one hand, it is input to the intra-frame mode decision module and the intra-frame prediction module for intra-frame prediction processing, and on the other hand, the reconstructed image signal s' is output through loop filtering. k [x,y], reconstructed image signal s' k [x,y] can be used as the reference image for the next frame for motion estimation and motion compensation prediction. Then based on the result of motion compensation prediction s' r [x+m x ,y+m y ] and intra prediction results Get the predicted image signal of the next frame And continue to repeat the above process until the encoding is completed.

[0055] The encoding operations for each CU involved in the above video encoding process are described in detail as follows.

[0056] Predictive Coding: Predictive coding includes methods such as intra-frame prediction and inter-frame prediction. The original video signal is predicted by a selected reconstructed video signal to obtain a residual video signal. The encoder needs to decide which predictive coding mode to use for the current CU and inform the decoder. Intra-frame prediction refers to the predicted signal coming from an already coded and reconstructed area within the same image; inter-frame prediction refers to the predicted signal coming from a previously coded image different from the current image (called a reference image).

[0057] Transform & Quantization: After the residual video signal undergoes transformations such as the Discrete Fourier Transform (DFT) and Discrete Cosine Transform (DCT), it is converted to the transform domain, where coefficients are stored. The transform coefficients are then subjected to a lossy quantization operation, which removes some information and makes the quantized signal more suitable for compression. Some video coding standards may offer more than one transform scheme, so the encoder must select one for the current CU and inform the decoder. The level of quantization is typically determined by the quantization parameter (QP). A larger QP value means that coefficients with a larger value range will be quantized to the same output, which typically results in greater distortion and a lower bitrate. Conversely, a smaller QP value means that coefficients with a smaller value range will be quantized to the same output, which typically results in less distortion and a higher bitrate.

[0058] Entropy Coding or Statistical Coding: The quantized transform domain signal is statistically compressed and encoded based on the frequency of occurrence of each value, and finally a binary (0 or 1) compressed code stream is output. At the same time, the encoding generates other information, such as the selected coding mode, motion vector data, etc., which also need to be entropy coded to reduce the bit rate. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to express the same signal. Common statistical coding methods include variable length coding (VLC) or context-based binary arithmetic coding (CABAC).

[0059] The context-based adaptive binary arithmetic coding (CABAC) process consists of three main steps: binarization, context modeling, and binary arithmetic coding. After binarization, the input syntax elements can be encoded using both the normal coding mode and the bypass coding mode. In the bypass coding mode, instead of assigning a specific probability model to each binary bit, the input binary bit values ​​are directly encoded using a simple bypass encoder, speeding up both encoding and decoding. Generally, different syntax elements are not completely independent, and even the same syntax elements have some memory. Therefore, based on conditional entropy theory, conditional coding using other coded syntax elements can further improve coding performance compared to independent or memoryless coding. This coded symbol information used as a condition is called context. In the normal coding mode, the binary bits of the syntax elements are sequentially fed into the context modeler. The encoder assigns an appropriate probability model to each input binary bit based on the values ​​of previously coded syntax elements or binary bits. This process is known as context modeling. The context model corresponding to the syntax element can be located using ctxIdxInc (contextindex increment) and ctxIdxStart (context index Start). After the bin value and the assigned probability model are fed into the binary arithmetic encoder for encoding, the context model needs to be updated based on the bin value, which is the adaptive process in encoding.

[0060] Loop Filtering: The changed and quantized signal will be reconstructed through inverse quantization, inverse transformation and prediction compensation operations to obtain a reconstructed image. Compared with the original image, due to the influence of quantization, some information of the reconstructed image is different from the original image, that is, the reconstructed image will produce distortion. Therefore, the reconstructed image can be filtered, such as deblocking filter (DB), SAO (Sample Adaptive Offset) or ALF (Adaptive Loop Filter) and other filters, which can effectively reduce the degree of distortion caused by quantization. Since these filtered reconstructed images will be used as a reference for subsequent encoded images to predict future image signals, the above filtering operation is also called loop filtering, that is, filtering operation within the encoding loop.

[0061] Based on the above encoding process, at the decoding end, after obtaining the compressed code stream (i.e., bitstream), entropy decoding is performed on each CU to obtain various mode information and quantization coefficients. The quantization coefficients are then dequantized and inversely transformed to obtain a residual signal. Furthermore, based on the known coding mode information, the prediction signal corresponding to the CU can be obtained. The residual signal is then added to the prediction signal to obtain a reconstructed signal. The reconstructed signal then undergoes loop filtering and other operations to produce the final output signal.

[0062] Current mainstream video coding standards, such as HEVC, VVC, AVS3, AV1, and AV2, all use a block-based hybrid coding framework. They divide the original video data into a series of coding blocks and combine video coding methods such as prediction, transform, and entropy coding to achieve video data compression.

[0063] Among them, motion compensation is a commonly used prediction method for video coding. Motion compensation is based on the redundant characteristics of video content in the time domain or spatial domain, and derives the predicted value of the current coding block from the coded area. This type of prediction method includes: inter-frame prediction, intra-frame block copy prediction, intra-frame string copy prediction, etc. In specific coding implementations, these prediction methods may be used alone or in combination. For coding blocks using these prediction methods, it is usually necessary to explicitly or implicitly encode one or more two-dimensional displacement vectors in the code stream to indicate the displacement of the current block (or the current block's co-located block) relative to its one or more reference blocks.

[0064] It's important to note that displacement vectors may have different names in different prediction modes and implementations. This article uniformly describes them as follows: 1) The displacement vector in inter prediction is called a motion vector (MV); 2) The displacement vector in intra block copy is called a block vector (BV); 3) The displacement vector in intra string copy is called a string vector (SV). The following describes the relevant technologies in inter prediction and intra block copy prediction.

[0065] Figure 4 FIG. 1 shows a schematic diagram of the principle of the inter-frame prediction mode. Figure 4 As shown in the figure, inter-frame prediction exploits temporal correlations in the video, using pixels from neighboring coded images to predict pixels in the current image. This effectively removes temporal redundancy and saves bits in the coded residual data. Here, P is the current frame, Pr is the reference frame, B is the current block to be coded, and Br is B's reference block. B' and B have the same coordinates in the image: Br's coordinates are (xr, yr) and B''s coordinates are (x, y). The displacement between the current coded block and its reference block is called a motion vector (MV): MV = (xr - x, yr - y).

[0066] Considering the strong correlation between adjacent blocks in the temporal or spatial domain, MV prediction technology can be used to further reduce the bits required to encode MV. In H.265 / HEVC, inter-frame prediction includes two MV prediction technologies: Merge and AMVP.

[0067] Figure 5 FIG. 1 shows a schematic diagram of the principle of the intra-block copy prediction mode. Figure 5 As shown in the figure, Intra Block Copy (IBC) can be considered a special inter-frame prediction mode. The implementation principle of Intra Block Copy is almost the same as the motion compensation scheme in the inter-frame prediction model. The difference is that inter-frame prediction selects the reference block for motion compensation in a reference video frame different from the current video frame, while Intra Block Copy selects the reference block for motion compensation within the current video frame. In Intra Block Copy mode, the block vector represents the relative displacement from the current block position to the reference block position within the current video frame.

[0068] Intra-block copying (IBC) is an intra-frame coding tool adopted in the HEVC Screen Content Coding (SCC) extension, significantly improving the coding efficiency of screen content. IBC is also adopted in AVS3, VVC, and AV1 to improve the performance of screen content coding. IBC exploits the spatial correlation of screen content video and uses the pixels of the previously encoded image in the current image to predict the pixels of the current block to be coded, effectively saving the bits required to encode the pixels.

[0069] Figure 6 The flowchart of the video processing method in one embodiment of the present application is shown. The video processing method can be executed by a terminal device or a server that sends or receives video encoding data. The embodiment of the present application is described by taking the method executed by the terminal device as an example. The terminal device can be, for example, Figure 2 The video decoding device 210 or the video encoding device 220 is shown.

[0070] like Figure 6 As shown, the video processing method in the embodiment of the present application includes the following steps S610 to S630.

[0071] S610: Perform intra-block copy prediction based on a reference block corresponding to a current block to obtain a first prediction value of the current block, where the current block is an image block to be processed in a current video frame, and the reference block is a reconstructed image block in the current video frame.

[0072] S620: Determine a first template area corresponding to the reference block and a second template area corresponding to the current block, the first template area includes one or more reconstructed image areas adjacent to the reference block, and the second template area includes one or more reconstructed image areas adjacent to the current block.

[0073] S630: Fitting a filter according to the first template area and the second template area, and filtering the first prediction value of the current block according to the filter to obtain a second prediction value of the current block.

[0074] The image block in the embodiment of the present application refers to the basic unit for encoding or decoding processing, and may include, for example: a coding unit, a luminance coding unit, a chrominance coding unit, a coding block, a luminance coding block, a chrominance coding block, a prediction unit, a luminance prediction unit, a chrominance prediction unit, a luminance prediction block, a chrominance prediction block, and the like.

[0075] In the embodiments of the present application, the reference block is an image block determined by shifting the current block according to the block vector in intra block copy prediction mode. Based on this, the reconstructed values ​​of pixels within the reference block are equivalent to the predicted values ​​of pixels within the current block. Correspondingly, the reconstructed values ​​of pixels within the first template area are equivalent to the predicted values ​​of pixels within the current block.

[0076] In the video processing method provided in an embodiment of the present application, a first prediction value of the current block can be obtained after intra-block copy prediction is performed based on a reference block corresponding to the current block. Then, a filter is fitted based on a first template area corresponding to the reference block and a second template area corresponding to the current block, and the first prediction value of the current block is filtered according to the filter to obtain a second prediction value of the current block. By filtering the first prediction value of the current block by fitting the filter, the problem of spatial discontinuity between the current block and the surrounding pixels can be eliminated, thereby improving the image encoding and decoding quality. This video processing method can be applied to video codecs or video compression products that use IBC technology.

[0077] In one embodiment of the present application, the reference block corresponding to the current block may be located in a sub-pixel area. Images in natural scenes are generally analog and continuous, and the motion of objects in the image is also continuous, so the motion offset will not be a jump-like motion of integer pixels. In order to improve the accuracy of the prediction, motion estimation with sub-pixel accuracy is introduced into video compression coding technology. The sub-pixel area can only be obtained by a certain interpolation calculation of the integer pixel area. Between the integer pixel motion estimation and the sub-pixel motion estimation, it is necessary to perform interpolation calculation of the sub-pixel points within the search range of the reference frame image. On this basis, the reference block corresponding to the current block may be located in an integer pixel area, or it may be located in a sub-pixel area.

[0078] In one embodiment of the present application, the method of filtering the first prediction value of the current block according to the filter may further include: obtaining a filtering condition of the current block, the filtering condition being used to determine whether to filter the current block; and when the filtering condition is met, filtering the first prediction value of the current block according to the filter.

[0079] In one embodiment of the present application, the filtering condition includes at least one of the following conditions:

[0080] (1) The code stream parameters corresponding to the current block have specified values.

[0081] (2) The image feature of the current block meets the preset first feature condition.

[0082] (3) The image features of the reference block meet the preset second feature condition.

[0083] (4) The image features of the template area meet the preset third feature condition.

[0084] The specific implementation methods of the above five conditions are explained below.

[0085] In one embodiment of the present application, the filtering condition may include condition (1): a code stream parameter corresponding to the current block has a specified value.

[0086] In one embodiment of the present application, there are syntax elements related to video coding and decoding in the video code stream where the current block is located. These syntax elements can be used as code stream parameters to guide related operations of video coding and decoding.

[0087] The bitstream syntax description method is similar to that of the C language. Bitstream syntax elements are represented in boldface. Each syntax element is described by its name (underline-separated lowercase letters), syntax, and semantics. Syntax element values ​​in syntax tables and the main text are represented in regular font. In some cases, syntax tables may use other variable values ​​derived from syntax elements. Such variables are named in syntax tables or the main text using a mix of lowercase and uppercase letters without underscores. Variables beginning with an uppercase letter are used to decode the current and related syntax structures and can also be used to decode subsequent syntax structures. Variables beginning with a lowercase letter are used only within the section in which they appear. The relationship between mnemonics for syntax element values ​​and variable values ​​and their values ​​is explained in the main text. In some cases, the two are used equivalently. Mnemonics are represented by one or more underline-separated letters, each beginning with an uppercase letter and possibly including multiple uppercase letters. Hexadecimal notation may be used when the length of a bit string is an integer multiple of 4. The prefix of hexadecimal is "0x", for example "0x1a" represents the bit string "0001 1010".

[0088] In one embodiment of the present application, the code stream parameters include a filtering flag, which is used to indicate whether to perform filtering on the image block. The filtering flag may include one or more of the following flags:

[0089] Sequence header filter flag, used to indicate whether the image blocks in the video frame sequence are to be filtered;

[0090] Image header filtering flag, used to indicate whether the image block in the video frame is filtered;

[0091] Slice header filter flag, used to indicate whether the image blocks in the image slice are to be filtered;

[0092] Block-level filtering flag, used to indicate whether the current block is filtered.

[0093] In one embodiment of the present application, when the block-level filtering flag exists in the code stream parameters, the block-level filtering flag is decoded when a preset decoding condition is met, and is prohibited from being decoded when the decoding condition is not met.

[0094] In one embodiment of the present application, the decoding conditions include one or more of the following conditions:

[0095] (1) The high-level syntax element in the code stream parameter indicates that filtering processing is performed on the image block. The high-level syntax element includes at least one of a sequence header filter flag, a picture header filter flag, and a slice header filter flag.

[0096] (2) The area of ​​the current block is within a preset value range; for example, the area of ​​the current block is greater than 32.

[0097] (3) The position coordinates of the current block are within a preset coordinate range. For example, the position coordinates (horizontal coordinates and / or vertical coordinates) of the current block are greater than or equal to a specific threshold.

[0098] (4) The area range occupied by the reference area of ​​the current block in the current video frame meets the preset range restriction, and the reference area includes one or more of the current block, the template area of ​​the current block, and the extended area of ​​the template area.

[0099] In an optional embodiment, the current video frame is divided into multiple image regions, for example, M*N image regions. The number of image regions occupied by sample points in the reference region in the current video frame is K. In this case, this number K can be used as the region range occupied by the reference region in the current video frame. The preset range limit can be determined based on the block size of the current block.

[0100] For example, the preset range limit may be that the region range (i.e., the number K) is less than or equal to a number threshold. The number threshold may be expressed as W*H*TH, where W is the width of the current block, H is the height of the current block, and TH is a preset scale factor.

[0101] (5) The reference area of ​​the current block meets the reference range restriction of intra block copy prediction.

[0102] In one embodiment of the present application, the preset coordinate range includes a value greater than or equal to a coordinate threshold; and a method for determining the coordinate threshold includes at least one of the following methods:

[0103] (1) Determine the coordinate threshold according to a preset fixed value.

[0104] (2) Determine the coordinate threshold according to the template size of the template area; when the template area is the image area located above the current block, the template size is the height of the template area; when the template area is the image area located to the left of the current block, the template size is the width of the template area.

[0105] In one embodiment of the present application, when the current block does not include a chrominance component, for example, the current block is a luminance block including only a luminance component, the template size of the template area is used as the coordinate threshold;

[0106] When the current block contains chroma components, for example, a chroma block containing only chroma components or an image block containing both luma and chroma components, a weighting coefficient is determined based on the sampling ratio of the luma and chroma components, and the product of the weighting coefficient and the template size of the template area is used as the coordinate threshold. For example, for a YUV420 image, the length and width of the luma component are twice that of the color components, so the weighting coefficient can be determined as 2.

[0107] In an optional embodiment, the code stream includes a high-level syntax element indicating whether to use the intra block copy adaptive prediction filter (IBC-APF) provided in an embodiment of the present application, and whether the current block is allowed to use the IBC-APF-based filtering method is determined according to the high-level syntax element.

[0108] High-level syntax elements may include, for example, one or more of the following: sequence header (seq_ibc_apf_flag), picture header (pic_ibc_apf_flag), and slice header syntax elements (slice_ibc_apf_flag). The decoding priority is sequence header > picture header > slice header > block level. If a high-priority syntax element indicates that IBC-APF is not allowed, the lower-level syntax elements do not need to be decoded.

[0109] A video sequence is the highest-level syntactic structure of a bitstream. A video sequence begins with the first sequence header. A sequence end code or video editing code indicates the end of a video sequence. The sequence headers between the first sequence header and the first occurrence of the sequence end code or video editing code are repeated sequence headers. Each sequence header is followed by one or more coded pictures, each preceded by a picture header. Coded pictures are arranged in bitstream order within the bitstream, which should be the same as the decoding order. The decoding order may differ from the display order.

[0110] A picture can be a frame or a field. Its coded data begins with a picture start code and ends with a sequence start code, a sequence end code, or the next picture start code. In the bitstream, the coded data for the two fields of an interlaced picture can appear sequentially or interleaved. The decoding and display order of the two fields is specified in the picture header. Picture types include I-pictures, P-pictures, and B-pictures.

[0111] A slice is a rectangular area in an image, which contains the parts of several maximum coding units in the image, and slices should not overlap.

[0112] In another optional implementation, a high-level syntax element exists in the bitstream, indicating whether to use any intra block copy prediction filter (IBC-PF), wherein the IBC-PF filter may include multiple different filtering modes.

[0113] In another optional implementation, a block-level filtering flag cu_ibc_apf_flag exists in the bitstream, which is used to indicate whether the current block is filtered using the IBC-APF mode.

[0114] In one embodiment of the present application, the code stream parameters further include a filtering index, where the filtering index is used to indicate a filtering mode for filtering the image block, and the filtering mode includes one or more different filtering methods.

[0115] For example, the block-level flag cu_ibc_pf_flag in the bitstream indicates whether the current block uses IBC-PF filtering. If IBC-PF can include multiple filtering methods, the block-level index cu_ibc_pf_index in the bitstream indicates the IBC-PF method used by the current block. The decoding of cu_ibc_pf_index depends on cu_ibc_pf_flag.

[0116] Table 1 shows the syntax structure of code stream parameters related to the IBC-PF filter in an application scenario of an embodiment of the present application.

[0117] Table 1

[0118] cu_ibc_pf_flag ae(v) if(CuIbcPfFlag){ cu_ibc_pf_index ae(v) }

[0119] The block-level block copy intra prediction filter flag, cu_ibc_pf_flag, is a binary variable. A value of '1' indicates that the IBCPF mode can be used; a value of '0' indicates that the IBCPF mode should not be used. The value of cuIbcPfFlag is equal to the value of cu_ibc_pf_flag. If cu_ibc_pf_flag does not exist in the bitstream, the value of cuIbcPfFlag is 0.

[0120] The block-level block copy intra prediction filter index cu_ibc_pf_index indicates the IBCPF mode used. The value of CuIbcPfIndex is equal to the value of cu_ibc_pf_index. If cu_ibc_pf_index does not exist in the codestream, the value of CuIbcPfIndex is 0.

[0121] Taking three filtering methods as examples, Table 2 shows the values ​​of the above code stream parameters and their corresponding semantics.

[0122] Table 2

[0123]

[0124]

[0125] When the value of cu_ibc_pf_flag is 0, it indicates that the current block does not use the IBCPF mode for filtering, so there is no need to decode cu_ibc_pf_index.

[0126] When cu_ibc_pf_flag is set to 1, it indicates that the current block is filtered using the IBCPF mode. At this point, you can continue decoding cu_ibc_pf_index to determine which filtering mode is used for the current block. IBCPF mode 3 represents the intra-block copy adaptive prediction filtering mode IBC-APF provided in an embodiment of the present application. IBCPF mode 1 and IBCPF mode 2 are filtering methods other than the filtering mode provided in an embodiment of the present application, such as mean filtering, median filtering, Gaussian filtering, and the like.

[0127] In one embodiment of the present application, in addition to indicating the use of filtering for an image block, the filter flag is also used to indicate the filtering mode used for filtering the image block. The filtering mode includes one or more different filtering methods. Taking block-level code stream parameters as an example, only the filter flag cu_ibc_pf_flag can be used to indicate the IBC mode used for the current block.

[0128] Table 3 shows the value and semantics of the embodiment of the present application when only the filtering flag cu_ibc_pf_flag is used.

[0129] Table 3

[0130] cu_ibc_pf_flag Semantics 0 Do not use IBCPF mode 1 Using IBCPF Mode 1 2 Using IBCPF Mode 2 3 Using IBCPF Mode 3 (the IBC-APF mode)

[0131] In one embodiment of the present application, the code stream parameters further include a filter type field for indicating a filter type.

[0132] When the filter type field takes the first value, the filter flag is used to indicate whether to perform filtering on the image block, and the filtering mode for performing filtering on the image block is the preset first mode;

[0133] When the filter type field takes the second value, the filter flag is used to indicate whether to perform filtering on the image block, and the filtering mode for filtering the image block is the preset second mode, which is a filtering mode different from the first mode.

[0134] For example, there is a type field ibc_type in the code stream indicating the IBC type (which can be a flag of sequence level, picture level, slice level, or block level). When ibc_type has different values, different filtering methods are used.

[0135] Table 4 shows the values ​​and semantics of the filter type field and filter flag in an application scenario of an embodiment of the present application.

[0136] Table 4

[0137] ibc_type cu_ibc_pf_flag Semantics 0 0 Do not use IBCPF mode 0 1 Using IBCPF Mode 1 0 2 Using IBCPF Mode 2 1 0 Do not use IBCPF mode 1 1 Using IBCPF Mode 3 (the IBC-APF mode)

[0138] In one embodiment of the present application, the binarization / debinarization method of the above-mentioned filter flag cu_ibc_pf_flag or filter index cu_ibc_pf_index can use variable-length code or fixed-length code, and the variable-length code includes truncated unary code, truncated binary code, K-order exponential Golomb code, etc.

[0139] In one embodiment of the present application, the value of the filtering flag is determined according to the coding loss, and the coding loss is calculated by a preset cost function;

[0140] If the coding loss after filtering the template region of the image block is less than the coding loss before filtering the template region of the image block, assigning a filtering flag to a first value, the first value being used to indicate that filtering is to be performed on the image block;

[0141] If the coding loss after filtering the template area of ​​the image block is greater than the coding loss before filtering the template area of ​​the image block, the filtering flag is assigned a second value, which is used to indicate that filtering of the image block is prohibited.

[0142] For example, an embodiment of the present application can use the Sum of Absolute Difference (SAD) as a cost function to calculate the coding loss of the current block before and after filtering. If the coding loss of filtering using IBCPF is smaller, it is determined to use the IBCPF mode for the current block; otherwise, the IBCPF mode is prohibited for the current block.

[0143] In one embodiment of the present application, the value of the filter flag is determined according to the number of transform coefficients in the image block;

[0144] If the transform coefficients in the image block meet a preset first quantity condition, assigning a filtering flag to a first value, the first value being used to indicate that filtering processing is to be performed on the image block;

[0145] If the transform coefficients in the image block meet a preset second quantity condition, the filtering flag is assigned a second value, and the second value is used to indicate that filtering processing on the image block is prohibited.

[0146] For example, embodiments of the present application may count the number of coefficients in the current block whose transform coefficients are even, and then determine whether to use the IBCPF based on the parity of this number of coefficients. For example, if the number of coefficients in the previous block whose transform coefficients are even is m, the first number condition may be that m is an odd number, and the second number condition may be that m is an even number.

[0147] In some other optional implementations, other implicit coefficients may also be used to derive parameter values ​​of the filtering flag.

[0148] In one embodiment of the present application, determining the filtering condition for performing filtering processing on the current block may include condition (2): the image feature of the current block satisfies a preset first feature condition.

[0149] In one embodiment of the present application, the first characteristic condition includes one or more of the following conditions:

[0150] (2.1) The current video frame where the current block is located has a specified image type.

[0151] If IBC-APF is not allowed based on the picture type, then there is no need to decode the syntax elements at the picture header and below. For example, IBC-APF mode is only allowed in I pictures, and there is no need to decode IBC-APF-related syntax elements in the block level, slice header, and picture header in non-I pictures.

[0152] (2.2) The current block has a specified distribution position in the current video frame.

[0153] For example, it is only available when the horizontal and vertical coordinates of the current block are larger than the size of the filter template.

[0154] (2.3) The image size of the current block is within a preset size range, where the image size includes one or more of width, height, or area.

[0155] IBC-APF is only allowed to be used when the block size meets the conditions. If the current block size does not meet the conditions, there is no need to decode the block-level IBC-APF related syntax elements. The block size may include one or more of width, height or area. The conditions are greater than, greater than or equal to, less than, or less than or equal to a specific threshold. The conditions related to the image size of the current block may include multiple conditions, each corresponding to a different threshold. For example, IBCPF is only used when the following conditions are met at the same time: the current block area (width * height) is greater than 64, the current block width is less than or equal to 64, and the current block height is less than or equal to 64.

[0156] For another example, filtering is allowed to be performed on the current block only when the width*height of the current block is greater than 32.

[0157] (2.4) The block vector resolution of the current block is within the preset resolution range.

[0158] IBC-APF is only allowed when the Adaptive Block Vector Resolution (ABVR) of the current block meets certain conditions. If the block vector resolution size of the current block does not meet the conditions, the block-level IBC-APF related syntax elements do not need to be decoded. The condition can be that the current block vector resolution is a specific subset of the available ABVR list.

[0159] For example, currently ABVR allows {1-pel, 4pel}, which is only used when the block vector resolution of the current block is 1-pel.

[0160] For example, currently ABVR allows {1-pel,4pel}, which is only used when the block vector resolution of the current block is 4-pel.

[0161] For example, currently ABVR allows {1 / 4pel, 1-pel, 4pel}, which is only used when the block vector resolution of the current block is 1-pel.

[0162] For example, currently ABVR allows {1 / 4pel, 1-pel, 4pel}, which is only used when the block vector resolution of the current block is 1-pel or 4-pel.

[0163] (2.5) The block vector residual of the current block is within a preset residual range, and the block vector residual includes one or more of a horizontal residual and a vertical residual.

[0164] IBC-APF is only allowed if the current block vector residual (BVD) meets the conditions. If the BVD size of the current block does not meet the conditions, the block-level IBC-APF related syntax elements do not need to be decoded. The conditions are that the absolute value of the horizontal BVD or (and) the vertical BVD is less than (or equal to, less than or equal to, greater than or equal to, or greater than) a specific threshold.

[0165] For example, the IBCPF is allowed to be used only when both the horizontal BVD and the vertical BVD of the current block are equal to 0.

[0166] (2.6) The block vector index of the current block is within the preset index range.

[0167] IBC-APF is only allowed when the block vector index of the current block meets the conditions. If the predicted block vector index of the current block does not meet the conditions, the block-level IBCPF-related syntax elements do not need to be decoded. The conditions can be, for example, that the absolute value of the block vector index bvp_idx is less than (or equal to, less than or equal to, greater than or equal to, or greater than) a specific threshold.

[0168] (2.7) The current block is an image block with a specified color component.

[0169] IBC-APF is only allowed to be used when the current block is a specific color component. For example, IBCPF can be used only for the luma component Y, or only for the chroma component U or V, or both.

[0170] In one embodiment of the present application, determining the filtering conditions for performing filtering processing on the current block may include condition (3): the image feature of the reference block satisfies a preset second feature condition.

[0171] In one embodiment of the present application, the second characteristic condition includes

[0172] (3.1) The sample points of the reference block meet the preset sample point availability conditions, which include that the sample points are located within the image independent decoding boundary or the sample points have been decoded and reconstructed.

[0173] IBC-APF is only allowed when the block vector of the current prediction block meets the validity conditions. The availability of the reference block and template area is determined by conditions such as: the sample point is within the image's independent decoding boundary, or the sample point has been decoded and reconstructed. Other restrictions may also be included, such as meeting hardware implementation limitations.

[0174] In one embodiment of the present application, determining the filtering conditions for performing filtering processing on the current block may include condition (4): the image features of the template area meet a preset third feature condition.

[0175] In one embodiment of the present application, the third characteristic condition may include one or more of the following conditions:

[0176] (4.1) The area of ​​the template region is within the preset area range.

[0177] (4.2) The sample points in the template area meet the preset sample point availability conditions, which include that the sample points are within the image independent decoding boundary or the sample points have completed decoding and reconstruction.

[0178] (4.3) The number of sample pairs collected in the template area is greater than the number of model parameters of the filter.

[0179] The above embodiments provide a variety of filtering conditions for whether to filter the current block. The embodiments of the present application may use one or more of the above filtering conditions. Different filtering methods may correspond to different filtering conditions, and the embodiments of the present application do not impose any special restrictions on this.

[0180] In one embodiment of the present application, a method for fitting a filter according to a first template area and a second template area may include: selecting a first sample point from the first template area, and selecting a second sample point corresponding to the first sample point from the second template area; fitting a filter according to a sample pair consisting of the first sample point and the second sample point, wherein the input item of the filter includes the pixel value of the first sample point, the output item of the filter includes the pixel value of the second sample point, and the pixel value includes a predicted value or a reconstructed value.

[0181] In one embodiment of the present application, the filter input further includes pixel values ​​of one or more neighboring pixel points, where the neighboring pixel points are the nearest or next nearest neighbor pixels of the first sample point.

[0182] For example, the filter model can be expressed as: y=p0*x0+p1*x1+p2*x2+…+pn*xn.

[0183] Where {p0, p1, …, pn} are the filter parameters, {x0, x1, …, xn} are the input values, and y is the output value. xi is derived based on the reference sample corresponding to the current sample to be filtered and its adjacent samples.

[0184] Figure 7 FIG. 1 shows a schematic diagram of the mapping form of the filter in one embodiment of the present application. Figure 7As shown, the first sample point C is selected from the first template area where the reference block is located, and the second sample point C' is selected from the second template area where the current block is located.

[0185] The position of the first sample point C is the filter center, which is determined by the position of the second sample point C′ after displacement according to the block vector.

[0186] The neighborhood pixel points of the first sample point C may include multiple nearest neighbor pixel points, such as Figure 7 As shown, multiple pixel points N, S, W and E are located above, below, to the left and to the right of the first sample point C.

[0187] The neighboring pixel points of the first sample point C may include multiple pixel points of the next nearest neighbors, such as Figure 7 As shown, multiple pixel points NW, NE, SW and SE are located at the upper left, upper right, lower left and lower right positions of the first sample point C.

[0188] Each time such a mapping relationship is established between the first sample point C and the second sample point C', a sample pair can be formed, and a mapping relationship equation can be generated. Multiple such mapping relationship equations can be used to solve the weighted parameters of the model.

[0189] In one embodiment of the present application, the shape of the filter can be of various types. Figure 8 A schematic diagram of the filter shape used in some application scenarios of the embodiments of the present application is shown.

[0190] like Figure 8 As shown, shape 1 indicates that the input items of the first type of filter include: a first sample point C and four nearest neighboring pixel points located above, below, to the left, and to the right of the first sample point C.

[0191] Shape 2 indicates that the input items of the second type of filter include: the first sample point C, the four nearest neighboring pixel points located above, below, left, and right of the first sample point C, and the four next nearest neighboring pixel points located at the upper left, upper right, lower left, and lower right positions of the first sample point C.

[0192] Shape three indicates that the input items of the third type of filter include: the first sample point C, and the four next-nearest neighboring pixel points located at the upper left, upper right, lower left and lower right positions of the first sample point C.

[0193] Shape four indicates that the input items of the fourth type of filter include: the first sample point C, the four nearest neighboring pixel points located above, below, left, and right of the first sample point C, the four next nearest neighboring pixel points located at the upper left, upper right, lower left, and lower right positions of the first sample point C, and the four next nearest neighboring pixel points located above, below, left, and right of the first sample point C (that is, the outer pixel points adjacent to the nearest neighbor pixel points in the corresponding directions).

[0194] In one embodiment of the present application, the pixel boundaries of the template area need to be expanded according to the shape of the filter. Figure 8 The second type of filter with shape 2 shown in FIG needs to expand the pixel points outward by one unit at the edge of the template area. Figure 8 The fourth type of filter with shape four shown needs to expand the pixels outward by two units at the edge of the template area.

[0195] In one embodiment of the present application, the filter includes one or more combination items with independent weighting parameters, and the combination item takes the pixel values ​​of one or more pixel points as input items.

[0196] The filter in the embodiment of the present application may be composed of at least one monomial, wherein each monomial may have an independent weighting parameter. When a monomial has at least two pixel values ​​as input items, the monomial is called a combination item.

[0197] In one embodiment of the present application, the filter includes at least two combinations of terms with different orders, where the order is the highest power of the input terms in the combination. The power operation can introduce nonlinear factors to improve the filter's fit to the sample point mapping relationship.

[0198] Assume that Rj represents the input of the filter, and xi can be any combination of Rj, such as Rj, m*Rj, Rj k and at least one of the constant bias term B, or a combination of any of these terms multiplied or added together. Here, m and k are non-zero integers, and k is an integer greater than 1.

[0199] In one embodiment of the present application, the filter may include one or more of the following multiple candidate models.

[0200] (1)y=p0C+p1N+p2S+p3W+p4E+p5C 2 +p6B.

[0201] (2)y=p0C+p1N+p2S+p3W+p4E+p5B.

[0202] (3)y=p0C+p1B.

[0203] (4)

[0204] (5)

[0205] (6)

[0206] (7)

[0207] (8)

[0208] (9)

[0209] (10)

[0210] (11)

[0211] (12)y=p0C+p1S+p2W+p3E+p4SW+p5SE+p6C 2 +p7B.

[0212] In one embodiment of the present application, the sample points used to fit the filter can be all or part of the pixels selected from the first template area and the second template area. When a part of the pixels are selected to fit the filter, pixel sampling can be performed in the corresponding template area according to a preset sampling rule.

[0213] In one embodiment of the present application, the method for selecting sample points in the template area may further include: obtaining a pixel sampling mode of the current block, the pixel sampling mode including full pixel sampling or partial pixel sampling; when the pixel sampling mode of the current block is full pixel sampling, selecting all pixels in the template area as sample points; when the pixel sampling mode of the current block is partial pixel sampling, selecting partial pixels with specified sampling positions in the template area as sample points.

[0214] In one embodiment of the present application, the designated sampling position includes at least one of the following three sampling positions.

[0215] The first sampling position: a designated sampling position where the position coordinates of the pixel point meet the preset coordinate value conditions.

[0216] In one embodiment of the present application, the coordinate value condition includes: at least one of the horizontal position coordinate or the vertical position coordinate of the pixel point is an even number; or, at least one of the horizontal position coordinate or the vertical position coordinate of the pixel point is an odd number.

[0217] Figure 9 A schematic diagram showing the selection of sampling positions based on the position coordinates of pixel points in one embodiment of the present application is shown.

[0218] like Figure 9 As shown, in the template area, a coordinate system is established with the upper left corner as the coordinate origin (0, 0), the horizontal position coordinate x represents the sequential position of the pixel points arranged from left to right in the horizontal direction, and the vertical position coordinate y represents the sequential position of the pixel points arranged from top to bottom in the horizontal direction.

[0219] According to the preset coordinate value conditions, even positions, odd positions or all positions can be selected as designated sampling positions in the horizontal direction, and even positions, odd positions or all positions can be selected as designated sampling positions in the vertical direction.

[0220] For example, in Figure 9 In the embodiment shown, pixel positions that are even positions in the horizontal direction and all positions in the vertical direction are selected as designated sampling positions, that is, pixel positions in the shaded portion in the figure are selected as designated sampling positions.

[0221] The second sampling position: a specified sampling position selected along the preset pixel scanning direction.

[0222] In one embodiment of the present application, pixels can be scanned in the template area using any scanning method, such as a round-trip scan or a ZigZag scan, so that designated sampling positions are selected along a preset pixel scanning direction in accordance with a sequence and a preset sampling rule. The preset sampling rule may be, for example, interval sampling, where a designated sampling position is selected after one or more scanned pixels have passed.

[0223] Figure 10 A schematic diagram of selecting sampling positions based on a round-trip scanning method in one embodiment of the present application is shown.

[0224] like Figure 10 As shown in the figure, a row of pixels is scanned from left to right within the template area. After reaching the boundary, the next row of pixels is scanned from right to left. During the pixel scanning process, a designated sampling position is selected for every scanned pixel along the scanning direction. For example, the arrow in the figure indicates the scanning direction, and the pixel positions in the shaded area are the selected designated sampling positions.

[0225] Figure 11A schematic diagram of selecting sampling positions based on the ZigZag scanning method in one embodiment of the present application is shown.

[0226] like Figure 11 As shown, pixels are scanned from the upper left corner to the lower right corner of the template area using a ZigZag scanning method. During the pixel scanning process, a designated sampling position is selected for every scanned pixel along the scanning direction. For example, the arrow in the figure indicates the scanning direction, and the shaded pixel positions are the selected designated sampling positions.

[0227] The third sampling position: a specified sampling position where the pixel value falls within a preset value range.

[0228] Taking the decoded and reconstructed pixels in the template area as an example, only when the reconstructed value of the pixel is greater than / less than a certain threshold value, it is used for filter fitting.

[0229] For example, pixels in the template region whose reconstructed values ​​are greater than a preset threshold can be fitted with one filter, while pixels in the template region whose reconstructed values ​​are less than or equal to the preset threshold can be fitted with another filter. When filtering is required for the first predicted value of a pixel in the current block, different filters can be selected for filtering based on the numerical relationship between the first predicted value and the preset threshold.

[0230] In one embodiment of the present application, the template area is composed of one or more nearest neighbor areas or second nearest neighbor areas, the nearest neighbor area includes an image area with a specified image size located above or to the left of the image block, and the second nearest neighbor area includes an image area with a specified image size located above, below, or above the right of the image block.

[0231] Figure 12 A schematic diagram showing the distribution of template areas corresponding to image blocks in one embodiment of the present application is shown.

[0232] like Figure 12 As shown, the template region corresponding to the image block 1201 may be composed of multiple sub-regions 1202, wherein each sub-region 1202 may be a nearest neighbor region or a second nearest neighbor region of the image block. The nearest neighbor region may, for example, include an image region B located above the image block 1201 or an image region D located to the left of the image block 1201, and the second nearest neighbor region may, for example, include an image region A located to the upper left of the image block 1201, an image region E located to the lower left of the image block 1201, or an image region C located to the upper right of the image block 1201.

[0233] The template area of ​​the image block 1201 can be formed by combining one or more of the image areas AE.

[0234] In one embodiment of the present application, the image sizes of the sub-regions constituting the template region are specified as follows.

[0235] The nearest neighbor region B above the image block 1201 has the same image size as the image block 1201 in the horizontal direction, and the nearest neighbor region B above the image block 1201 has a specified image size in the vertical direction.

[0236] The nearest neighbor region D on the left side of the image block 1201 has the same image size as the image block 1201 in the vertical direction, and has a specified image size in the horizontal direction.

[0237] The next neighboring region C to the upper right of the image block 1201 has the same image size as the image block 1201 in the horizontal direction, and has a specified image size in the vertical direction.

[0238] The next neighboring region E below the left of the image block 1201 has the same image size as the image block 1201 in the vertical direction, and the next neighboring region below the left of the image block 1201 has a specified image size in the horizontal direction.

[0239] The next neighboring area A at the upper left of the image block 1201 has a specified image size in both the horizontal and vertical directions.

[0240] The designated image size may be a preset value greater than or equal to one, for example, the designated image size may be set to 6. When the designated image size is greater than one, filtering may be performed using multiple layers of neighboring pixels, thereby improving the continuity between pixels in the current block and neighboring pixels.

[0241] In one embodiment of the present application, the sub-regions of the template region constituting the image block may have the same specified image size or different specified image sizes. For example, the vertical size of image region C may be the same as or different from the horizontal size of image region E.

[0242] In one embodiment of the present application, for image blocks with different color components, the template sizes of the corresponding template areas may also be different. For example, the template size of the template area of ​​the luminance block is different from the template size of the template area of ​​the chrominance block.

[0243] In one embodiment of the present application, pixels that have been decoded and reconstructed or are allowed to be available can be selected from the sub-region as sample points for filter fitting of the current block. For example, when some pixels in image region C have been decoded and reconstructed, while other pixels have not been decoded and reconstructed, the pixels that have been decoded and reconstructed can be selected from image region C as sample points for filter fitting of the current block.

[0244] In one embodiment of the present application, when the decoded and reconstructed pixels in a sub-region do not meet the aforementioned size requirements, the sub-neighboring region can be configured as unavailable. For example, when the lower right corner of image region C is not reconstructed or exceeds the image boundary, image region C can be configured as unavailable; for another example, when the lower right corner of image region E is not reconstructed or exceeds the image boundary, image region E can be configured as unavailable.

[0245] In one embodiment of the present application, Figure 12 All of the sub-regions shown are combined to form a template region, or a portion of the sub-regions can be combined to form a template region.

[0246] Figure 13 A schematic diagram showing how some sub-regions and image blocks are combined to form different template regions in one embodiment of the present application is shown. Figure 13 As shown, based on the combination of different sub-regions, ten types of optional template regions can be formed as examples. The embodiment of the present application can specify one or more candidate template regions for the image block.

[0247] Figure 14 FIG. 1 shows a schematic diagram of a template region selected for an image block in one embodiment of the present application. Figure 14 As shown, in the embodiment of the present application, the template area corresponding to the image block includes at least one of a full area combination, a left area combination, or an upper area combination.

[0248] The full region combination includes the nearest neighbor regions located on the left and above the image block and the next nearest neighbor regions located on the upper left, lower left and upper right of the image block.

[0249] The left region combination includes the nearest neighbor region located on the left side of the image block and the next nearest neighbor region located on the lower left side of the image block.

[0250] The upper region combination includes the nearest neighbor region located above the image block and the next nearest neighbor region located to the upper right of the image block.

[0251] In one embodiment of the present application, the region type field can be used in the video stream to identify the region type of the template region used by the image block. For example, when the region type field takes a value of 1, it means that the template region selected by the image block is Figure 10 The full area combination shown in ; when the area type field value is 01, it means that the template area selected by the image block is Figure 10 The left area combination shown in ; when the area type field value is 00, it means that the template area selected by the image block is Figure 10 The upper area combination shown in .

[0252] In one embodiment of the present application, after determining the template area used when performing filter fitting on an image block, availability information of one or more sub-areas constituting the reference area can be obtained, and then the area range of the template area can be adjusted based on the availability information of the one or more sub-areas.

[0253] In one embodiment of the present application, adjusting the area range of the template area based on the availability information of one or more sub-areas may further include: removing sub-areas in an unavailable state from the template area; and configuring the template area to an unavailable state when all sub-areas in the template area are in an unavailable state.

[0254] For example, by parsing the video code stream, the indicator field corresponding to the image block is set to 1, indicating that the template area used for color component prediction of the image block is Figure 10 The full area combination including the five sub-areas of AE is shown in .

[0255] When filtering the image block, availability information of each sub-region in the template region may be obtained, and the region range of the template region may be adjusted according to the availability information.

[0256] For example, sub-regions A, B, and C are image regions located above the image block. If sub-regions A, B, and C have not yet completed encoding and reconstruction, sub-regions A, B, and C are in an unavailable state. At this time, the area range of the template area can be adjusted from A+B+C+D+E to D+E.

[0257] For another example, sub-regions D and E are image regions located to the left of the image block. If encoding and reconstruction of sub-regions D and E have not yet been completed, sub-regions D and E are also in an unusable state. In this case, all five sub-regions A to E are in an unusable state, so the entire template region based on this full region combination can be configured as unusable.

[0258] In one embodiment of the present application, the current block may be filtered using multiple different filters, and then the filtering results of the multiple filters may be weighted to obtain a final prediction value for the current block. In one embodiment of the present application, the final predicted image may be obtained by weighting a first prediction value obtained by intra block copy prediction and a second prediction value obtained by filtering the first prediction value.

[0259] In one embodiment of the present application, after obtaining the second prediction value of the current block, a weighted operation is performed on the first prediction value and the second prediction value according to a preset weight coefficient to obtain a third prediction value of the current block.

[0260] For example, in an embodiment of the present application, the third prediction value pred of the current block may be determined according to the following formula.

[0261] pred=w*pred1+(1-w)*pred2

[0262] Among them, pred1 represents the first prediction value obtained by intra-frame block copy prediction, pred2 represents the second prediction value obtained after predicting the first prediction value based on the adaptive filtering scheme provided by the above embodiment, and w represents the weight coefficient for weighted operation of the two prediction values.

[0263] The following describes in detail the specific implementation of the video processing method in some application scenarios of the embodiment of the present application, taking the decoding process performed by the video decoding end as an example.

[0264] In one application scenario, a video decoder can decode a video stream and, based on the decoding result, determine whether the image block to be processed (using a coding block as an example) uses the IBC filtering method provided in the above embodiments of this application. Determining whether to use the IBC filtering method can be implemented in the following three ways.

[0265] In the first implementation, the codestream includes a high-level syntax element indicating whether to use IBC-APF. Based on this high-level syntax element, it can be determined whether the current block is allowed to use IBC-APF. The syntax structure of the relevant syntax elements is shown in Table 5.

[0266] Table 5

[0267]

[0268] pic_ibc_flag represents the picture-level block copy intra prediction flag. This flag is a binary variable. A value of '1' indicates that IBC mode can be used; a value of '0' indicates that IBC mode should not be used. The value of PicIbcFlag is equal to the value of pic_ibc_flag. If pic_ibc_flag is not present in the bitstream, the value of PicIbcFlag is 0.

[0269] pic_ibc_type indicates the picture-level block copy intra prediction mode type index. A value of '0' indicates that the first type of IBC mode can be used; a value of '1' indicates that the second type of IBC mode can be used. The value of PicIbcType is equal to the value of pic_ibc_type. If pic_ibc_type does not exist in the codestream, the value of PicIbcType is 0.

[0270] pic_ibc_apf_flag represents the picture-level block copy intra prediction adaptive filtering flag. This flag is a binary variable. A value of '1' indicates that the IBC adaptive filtering mode can be used; a value of '0' indicates that the IBC adaptive filtering mode should not be used. The value of PicIbcApfFlag is equal to the value of pic_ibc_apf_flag. If pic_ibc_apf_flag is not present in the bitstream, the value of PicIbcApfFlag is 0.

[0271] In the second implementation, a high-level syntax element in the bitstream indicates whether to use IBC-PF (IBC filtering method, which can include multiple filtering methods). The use of IBC-PF is determined based on this syntax element. The syntax structure of the relevant syntax elements is shown in Table 6.

[0272] Table 6

[0273]

[0274] pic_ibc_flag represents the picture-level block copy intra prediction flag. This flag is a binary variable. A value of '1' indicates that IBC mode can be used; a value of '0' indicates that IBC mode should not be used. The value of PicIbcFlag is equal to the value of pic_ibc_flag. If pic_ibc_flag is not present in the bitstream, the value of PicIbcFlag is 0.

[0275] pic_ibc_pf_index represents the picture-level block copy intra prediction filter index. This index is a binary variable. A value of '0' indicates that the first IBC filter mode can be used; a value of '1' indicates that the second IBC filter mode can be used. The value of PicIbcPfFlag is equal to the value of pic_ibc_pf_flag. If pic_ibc_pf_flag is not present in the bitstream, the value of PicIbcPfFlag is 0.

[0276] In a third implementation, a block-level flag included in the code stream is obtained by decoding the code stream, and the block-level flag is used to indicate whether to perform filtering processing on the current block.

[0277] In some application scenarios, the block-level flag is decoded only when a preset decoding condition is met. The decoding condition may include one or more of the following conditions.

[0278] (1) According to the high-level syntax information of the image header, the current image is allowed to use the IBC-APF mode.

[0279] (2) The area of ​​the current block (width * height) is greater than 32.

[0280] (3) The position (horizontal coordinate and / or vertical coordinate) of the current block is greater than or equal to a specific threshold.

[0281] The specific threshold can be determined according to the template size. Assume that the template size is tpl_size (for example, a value of 4). The template size represents the height of the template above the image block, or the width of the template on the left.

[0282] The threshold is determined based on the partition tree type of the current block (including luma, chroma, and luminance-chroma trees). If the partition tree type is luma only, the threshold is tpl_size. If the current block includes a chroma tree, the threshold is tpl_size * scale_ratio. The scale_ratio is determined by the image color format. For example, for a YUV420 image, the length and width of the luma component are twice that of the chroma components. In this case, the scale_ratio is 2, meaning the threshold is tpl_size * 2.

[0283] In some application scenarios, only the flag cu_ibc_pf_flag may be used to indicate the IBC-APF mode used by the current block. The semantic information corresponding to different values ​​of the flag cu_ibc_pf_flag is shown in Table 7.

[0284] Table 7

[0285] cu_ibc_pf_flag Semantics 0 Do not use IBCPF mode 1 Using IBC-APF mode (using the upper and left templates) 2 Use IBC-APF mode (use the template above) 3 Use IBC-APF mode (use the template on the left)

[0286] In some application scenarios, the bitstream also contains a flag, ibc_type, that indicates the filter type (which can be a sequence-level, picture-level, slice-level, or block-level flag). Different values ​​of ibc_type indicate different filtering methods. Table 8 shows the semantic information corresponding to different values ​​of the ibc_type and cu_ibc_pf_flag flags.

[0287] Table 8

[0288] ibc_type cu_ibc_pf_flag Semantics 0 0 Do not use IBCPF mode 0 1 Using IBCPF Mode 1 0 2 Using IBCPF Mode 2 1 0 Do not use IBCPF mode 1 1 Using IBC-APF mode (using the upper and left templates) 1 2 Use IBC-APF mode (use the template above) 1 3 Use IBC-APF mode (use the template on the left)

[0289] The binarization / debinarization method of the above flags may use truncated unary codes.

[0290] If it is determined that the current block uses IBC-APF, a filter is fitted according to the following method, and the IBC prediction value of the current block is filtered based on the filter.

[0291] The filter model of IBC-APF can be expressed as: y = p0*x0+p1*x1+p2*x2+…+pn*xn.

[0292] Among them, {p0, p1, …, pn} are filter parameters, {x0, x1, …, xn} are input values, and y is the output value.

[0293] xi is derived based on the reference sample corresponding to the current sample to be filtered and its adjacent samples.

[0294] by Figure 7 Taking the sample points shown in the figure as an example, the filter center C is the reference sample point for the current sample point C' to be filtered, determined based on the block vector. N, S, W, E, NW, NE, SW, and SE represent the sample points above, below, left, right, upper left, upper right, lower left, and lower right of the spatial position of sample point C, respectively.

[0295] The filter shape can be extended to other shapes (including but not limited to Figure 8 The filter center is located at the reference sample point determined by the block vector of the current sample point to be filtered.

[0296] Let Rj represent the input of the filter, and xi can be any combination of Rj, such as Rj, m*Rj, Rj k , and at least one of the constant bias term B, or a combination of any of these terms multiplied or added together. Here, m and k are non-zero integers, and k is an integer greater than 1.

[0297] In some application scenarios, the filter may select one or more of the following candidate models.

[0298] a) y = p0C + p1N + p2S + p3W + p4E + p5C 2 +p6B

[0299] b) y = p0C + p1N + p2S + p3W + p4E + p5B

[0300] c)y=p0C+p1(N+S)+p2(W+E)+p3B

[0301] d)y=p0C+p1(N+W)+p2(S+E)+p3B

[0302] e)y=p0C+p1N+p2S+p3B

[0303] f)y=p0C+p1E+p2W+p3B

[0304] g)y=p0C+p1B

[0305] h)

[0306] i)

[0307] j)

[0308] k)

[0309] l)

[0310] m)

[0311] n)

[0312] o)

[0313] p)

[0314] Based on the above filter model, let y be the sample points in the current block template area and the input be the sample points in the corresponding reference block template area. This generates an equation. Multiple such equations are constructed using the sample points in the template area to solve for the model parameters pi. The model parameter solution process can be coupled with the encoding and decoding process, that is, the model parameters are calculated online for each encoded block using adjacent templates.

[0315] Figure 15 The diagram shows the distribution of template regions used in an application scenario according to an embodiment of the present application. The template regions may be all adjacent regions of an image block shown in the diagram.

[0316] For example, the top+left template indicates that the template area includes the adjacent areas A, B, and D above and to the left of the image block.

[0317] The upper template indicates that the template area includes the adjacent area B above the image block.

[0318] The left template indicates that the template area includes the adjacent area D to the left of the image block.

[0319] The size of region B is (blk_w, 2), and the size of region D is (2, blk_h). blk_w represents the width of the image block, and blk_h represents the height of the image block.

[0320] In some application scenarios, multiple templates can exist at the same time, and the specific template to be used is determined based on the code stream analysis results. Figure 14 When the value of the corresponding syntax element in the codestream is 1, it indicates that the full region combination template (LT type) is used. When the value of the corresponding syntax element in the codestream is 01, it indicates that the left region combination template (L type) is used. When the value of the corresponding syntax element in the codestream is 00, it indicates that the top region combination template (T type) is used.

[0321] After the filter is determined, the boundary of the template used as input needs to be expanded according to the shape of the filter.

[0322] For example, the filter selection Figure 8 For the shape shown, the four neighboring pixel points N, S, W and E around the central sample point C are needed. At this time, the template needs to be expanded outward by 1 unit.

[0323] Figure 16 A schematic diagram shows the effect of expanding the template area in an application scenario according to an embodiment of the present application.

[0324] like Figure 16 As shown, the original template region includes adjacent regions A, B, and D located above and to the left of the image block. After the template region is expanded outward, corresponding expanded regions PT, PL, PB, and PR can be determined.

[0325] PT and PB have the same width as the current block, and PL and PR have the same height as the current block.

[0326] The height of PT and PB is pad_size, and the height of PL and PR is pad_size. Pad_size is determined by the filter shape. For example, if the filter needs to use pixels 1 unit away from the center position, the pad_size value is 1.

[0327] For a region (e.g., PT and PL) in {PT, PB, PR, PL}, first determine whether the region is within an available reconstructed area (e.g., within the image boundary). If so, use the value of that region; otherwise, set its value to the value of the adjacent reconstructed sample. For example, PT is set to the value of the sample below, and PL is set to the value to the right.

[0328] For some regions in {PT, PB, PR, PL} (such as PR and PB), their values ​​are set to the values ​​of their adjacent reconstructed samples. For example, PR is set to the value of the sample to the left, and PB is set to the value of the sample above.

[0329] The model parameters of the filter model are calculated using all sample points in the template area.

[0330] Based on the available samples in the template area, multiple equations can be established for the filter model, that is, Ax = b. Solving for x can obtain the values ​​of all model parameters p_i. Where A is a matrix with M rows and N columns, M and N represent the number of equations established based on the template and the number of model parameters, respectively. There are many solutions to this equation, such as the LDL decomposition method or Gaussian elimination method. Here we take the LDL decomposition method as an example:

[0331] (1)A T Ax=AT b.

[0332] (2) To A T A decomposes into LDL T x=A T b.

[0333] (3) Solve LY = A T b gets Y.

[0334] (4) Solving DL T x=Y to get x, which is the model parameter p i .

[0335] After solving the model parameters, the IBC prediction value (or the modified intermediate value of the IBC prediction value) can be filtered according to the filter model. The IBC prediction value of the current block is used as input, and the boundary of the current block is expanded according to the filter shape. Boundary pixels can be directly copied, and boundary samples can be used directly if available. The derived model parameters are then used for filtering.

[0336] Based on the introduction of the above embodiments and application scenarios, it can be seen that the embodiments of the present application can adaptively calculate the filtering parameters based on the reconstructed area adjacent to the current block and the reference block, and filter the IBC prediction block, which can eliminate the spatial discontinuity between the prediction block and the surrounding pixels, and is conducive to improving the coding performance.

[0337] It should be noted that although the steps of the method of the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all steps must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0338] The following introduces an embodiment of the device of the present application, which can be used to execute the video processing method in the above embodiment of the present application. Figure 17 The structure block diagram of the video processing device provided by the embodiment of the present application is schematically shown. Figure 17 As shown, the video processing device 1700 includes:

[0339] A prediction module 1710 is configured to perform intra block copy prediction based on a reference block corresponding to a current block to obtain a first prediction value of the current block, where the current block is an image block to be processed in a current video frame and the reference block is a reconstructed image block in the current video frame;

[0340] a determining module 1720 configured to determine a first template region corresponding to the reference block and a second template region corresponding to the current block, wherein the first template region includes one or more reconstructed image regions adjacent to the reference block, and the second template region includes one or more reconstructed image regions adjacent to the current block;

[0341] The filtering module 1730 is configured to fit a filter according to the first template area and the second template area, and filter the first prediction value of the current block according to the filter to obtain a second prediction value of the current block.

[0342] In one embodiment of the present application, based on the above embodiments, the filtering module 1730 can be further configured to: obtain the filtering condition of the current block, and the filtering condition is used to determine whether to filter the current block; when the filtering condition is met, filter the first prediction value of the current block according to the filter.

[0343] In one embodiment of the present application, based on the above embodiments, the filtering condition includes at least one of the following conditions:

[0344] The code stream parameter corresponding to the current block has a specified value;

[0345] The image feature of the current block satisfies a preset first feature condition;

[0346] The image feature of the reference block satisfies a preset second feature condition;

[0347] The image feature of the template area meets a preset third feature condition, and the template area includes a first template area and a second template area.

[0348] In one embodiment of the present application, based on the above embodiments, the code stream parameters include a filtering flag, which is used to indicate whether to perform filtering processing on the image block. The filtering flag includes one or more of the following flags:

[0349] Sequence header filter flag, used to indicate whether the image blocks in the video frame sequence are to be filtered;

[0350] Image header filtering flag, used to indicate whether the image block in the video frame is filtered;

[0351] Slice header filter flag, used to indicate whether the image blocks in the image slice are to be filtered;

[0352] Block-level filtering flag, used to indicate whether the current block is filtered.

[0353] In one embodiment of the present application, based on the above embodiments, when the block-level filtering flag exists in the code stream parameters, the block-level filtering flag is decoded when a preset decoding condition is met, and decoding of the block-level filtering flag is prohibited when the decoding condition is not met.

[0354] In one embodiment of the present application, based on the above embodiments, the code stream parameters further include a filtering index, which is used to indicate a filtering mode for filtering the image block, and the filtering mode includes one or more different filtering methods.

[0355] In one embodiment of the present application, based on the above embodiments, the filtering flag is further used to indicate a filtering mode for filtering the image block, and the filtering mode includes one or more different filtering methods.

[0356] In one embodiment of the present application, based on the above embodiments, the code stream parameter further includes a filter type field for indicating a filter type;

[0357] When the value of the filter type field is the first value, the filter flag is used to indicate whether to perform filtering processing on the image block, and the filtering mode for performing filtering processing on the image block is the preset first mode;

[0358] When the filtering type field takes a second value, the filtering flag is used to indicate whether the image block is filtered, and the filtering mode for filtering the image block is a preset second mode, which is a filtering mode different from the first mode.

[0359] In one embodiment of the present application, based on the above embodiments, the value of the filtering flag is determined according to the coding loss, and the coding loss is calculated by a preset cost function;

[0360] If the coding loss after filtering the template region of the image block is less than the coding loss before filtering the template region of the image block, assigning the filtering flag to a first value, where the first value is used to indicate that filtering is to be performed on the image block;

[0361] If the coding loss after filtering the template area of ​​the image block is greater than the coding loss before filtering the template area of ​​the image block, the filtering flag is assigned a second value, which is used to indicate that filtering of the image block is prohibited.

[0362] In one embodiment of the present application, based on the above embodiments, the value of the filter flag is determined according to the number of transform coefficients in the image block;

[0363] If the transform coefficients in the image block meet a preset first quantity condition, assigning the filter flag a first value, where the first value is used to indicate that filtering processing is to be performed on the image block;

[0364] If the transformation coefficients in the image block meet a preset second quantity condition, the filtering flag is assigned a second value, and the second value is used to indicate that filtering processing on the image block is prohibited.

[0365] In one embodiment of the present application, based on the above embodiments, the first characteristic condition includes one or more of the following conditions:

[0366] The current video frame where the current block is located has a specified image type;

[0367] The current block has a specified distribution position in the current video frame;

[0368] The image size of the current block is within a preset size range, where the image size includes one or more of width, height, or area;

[0369] The block vector resolution of the current block is within a preset resolution range;

[0370] The block vector residual of the current block is within a preset residual range, and the block vector residual includes one or more of a horizontal residual and a vertical residual;

[0371] The block vector index of the current block is within a preset index range;

[0372] The current block is an image block having a specified color component.

[0373] In one embodiment of the present application, based on the above embodiments, the second characteristic condition includes: the sample points of the reference block meet a preset sample point availability condition, and the sample point availability condition includes that the sample point is located within the image independent decoding boundary or the sample point has completed decoding and reconstruction.

[0374] In one embodiment of the present application, based on the above embodiments, the third characteristic condition includes one or more of the following conditions:

[0375] The area of ​​the template region is within a preset area range;

[0376] The sample points of the template area meet a preset sample point availability condition, where the sample point availability condition includes that the sample point is within an image independent decoding boundary or the sample point has completed decoding and reconstruction;

[0377] The number of sample pairs collected in the template area is greater than the number of model parameters of the filter.

[0378] In one embodiment of the present application, based on the above embodiments, the filtering module 1730 is further configured to: select a first sample point from the first template area, and select a second sample point corresponding to the first sample point from the second template area; fit a filter according to a sample pair consisting of the first sample point and the second sample point, the input item of the filter includes the pixel value of the first sample point, the output item of the filter includes the pixel value of the second sample point, and the pixel value includes a predicted value or a reconstructed value.

[0379] In one embodiment of the present application, based on the above embodiments, the input item of the filter also includes pixel values ​​of one or more neighborhood pixels, and the neighborhood pixels are the pixels that are the nearest neighbors or the next nearest neighbors of the first sample point.

[0380] In one embodiment of the present application, based on the above embodiments, the filter includes one or more combination items with independent weighting parameters, and the combination item uses the pixel values ​​of one or more pixel points as input items.

[0381] In one embodiment of the present application, based on the above embodiments, the filter includes at least two combination terms with different orders, and the order is the highest power of the input terms in the combination terms.

[0382] In one embodiment of the present application, based on the above embodiments, a method for selecting sample points in a template area includes: obtaining a pixel sampling mode of the current block, where the pixel sampling mode includes full pixel sampling or partial pixel sampling; when the pixel sampling mode of the current block is full pixel sampling, selecting all pixels in the template area as sample points; when the pixel sampling mode of the current block is partial pixel sampling, selecting partial pixels with specified sampling positions in the template area as sample points.

[0383] In one embodiment of the present application, based on the above embodiments, the designated sampling position includes at least one of the following sampling positions:

[0384] The pixel point's position coordinates meet the specified sampling position of the preset coordinate value conditions;

[0385] A designated sampling position is selected along a preset pixel scanning direction;

[0386] The pixel value falls within the specified sampling position within the preset value range.

[0387] In one embodiment of the present application, based on the above embodiments, the coordinate value conditions include:

[0388] At least one of the horizontal position coordinate and the vertical position coordinate of the pixel point is an even number;

[0389] Alternatively, at least one of the horizontal position coordinate and the vertical position coordinate of the pixel point is an odd number.

[0390] In one embodiment of the present application, based on the above embodiments, the template area is composed of one or more nearest neighbor areas or second nearest neighbor areas, the nearest neighbor area includes an image area with a specified image size located above or to the left of the image block, and the second nearest neighbor area includes an image area with a specified image size located in the upper left, lower left or upper right of the image block.

[0391] In one embodiment of the present application, based on the above embodiments, the next-nearest neighboring region to the upper right of the image block has the same image size as the image block in the horizontal direction, and the next-nearest neighboring region to the upper right of the image block has a specified image size in the vertical direction, and the specified image size is greater than or equal to one;

[0392] The next nearest neighboring area at the lower left of the image block has the same image size as the image block in the vertical direction, and the next nearest neighboring area at the upper right of the image block has the specified image size in the horizontal direction;

[0393] The next-nearest neighboring area at the upper left of the image block has the specified image size in both the horizontal direction and the vertical direction.

[0394] In one embodiment of the present application, based on the above embodiments, the template area includes at least one of a full area combination, a left area combination, and an upper area combination;

[0395] The full region combination includes the nearest neighbor regions located on the left and above the image block and the next nearest neighbor regions located on the upper left, lower left and upper right of the image block;

[0396] The left region combination includes a nearest neighbor region located on the left side of the image block and a next nearest neighbor region located on the lower left side of the image block;

[0397] The upper region combination includes a nearest neighbor region located above the image block and a next nearest neighbor region located to the upper right of the image block.

[0398] In one embodiment of the present application, based on the above embodiments, the code stream parameter corresponding to the current block includes a region type field, and the region type field is used to indicate the region type of the template region corresponding to the current block.

[0399] In one embodiment of the present application, based on the above embodiments, the fitting module 1730 is further configured to: perform a weighted operation on the first prediction value and the second prediction value according to a preset weight coefficient to obtain a third prediction value of the current block.

[0400] The specific details of the video processing device provided in each embodiment of the present application have been described in detail in the corresponding method embodiments and will not be repeated here.

[0401] Figure 18 The block diagram schematically shows a computer system structure of an electronic device used to implement an embodiment of the present application.

[0402] It should be noted that Figure 18 The computer system 1800 of the electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.

[0403] like Figure 18 As shown, the computer system 1800 includes a central processing unit 1801 (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1802 (ROM) or the program loaded from the storage part 1808 into the random access memory 1803 (RAM). Various programs and data required for system operation are also stored in the random access memory 1803. The central processing unit 1801, the read-only memory 1802 and the random access memory 1803 are connected to each other via a bus 1804. An input / output interface 1805 (i.e., an I / O interface) is also connected to the bus 1804.

[0404] The following components are connected to the input / output interface 1805: an input section 1806 including a keyboard, a mouse, and the like; an output section 1807 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 1808 including a hard disk; and a communication section 1809 including a network interface card such as a local area network card or a modem. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to the input / output interface 1805 as needed. Removable media 1811, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1810 as needed, so that computer programs read therefrom can be installed into the storage section 1808 as needed.

[0405] In particular, according to an embodiment of the present application, the processes described in the various method flow charts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1809 and / or installed from a removable medium 1811. When the computer program is executed by the central processing unit 1801, the various functions defined in the system of the present application are performed.

[0406] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0407] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0408] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0409] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0410] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.

[0411] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A video processing method, characterized in that: include: performing intra-block copy prediction based on a reference block corresponding to a current block to obtain a first prediction value of the current block, wherein the current block is an image block to be processed in a current video frame, and the reference block is a reconstructed image block in the current video frame; determining a first template region corresponding to the reference block and a second template region corresponding to the current block, wherein the first template region includes one or more reconstructed image regions adjacent to the reference block, and the second template region includes one or more reconstructed image regions adjacent to the current block; A filter is fitted according to the first template area and the second template area, and a first prediction value of the current block is filtered according to the filter to obtain a second prediction value of the current block.

2. The method according to claim 1, characterized in that Performing filtering processing on the first prediction value of the current block according to the filter, comprising: Acquire a filtering condition of the current block, where the filtering condition is used to determine whether to perform filtering processing on the current block; When the filtering condition is met, filtering processing is performed on the first prediction value of the current block according to the filter.

3. The method according to claim 2, characterized in that The filtering condition includes at least one of the following conditions: The code stream parameter corresponding to the current block has a specified value; The image feature of the current block satisfies a preset first feature condition; The image feature of the reference block satisfies a preset second feature condition; The image features of the template area meet a preset third feature condition, and the template area includes the first template area and the second template area.

4. The method according to claim 3, characterized in that The code stream parameters include a filtering flag, which is used to indicate whether to perform filtering processing on the image block. The filtering flag includes one or more of the following flags: Sequence header filter flag, used to indicate whether the image blocks in the video frame sequence are to be filtered; Image header filtering flag, used to indicate whether the image block in the video frame is filtered; Slice header filter flag, used to indicate whether the image blocks in the image slice are to be filtered; The block-level filtering flag is used to indicate whether filtering is performed on the current block.

5. The method according to claim 4, characterized in that When the block-level filter flag exists in the code stream parameters, the block-level filter flag is decoded when a preset decoding condition is met, and the decoding of the block-level filter flag is prohibited when the decoding condition is not met.

6. The method according to claim 5, characterized in that The decoding conditions include one or more of the following conditions: A high-level syntax element in the code stream parameter indicates that filtering processing is performed on the image block, and the high-level syntax element includes at least one of a sequence header filter flag, a picture header filter flag, and a slice header filter flag; The area of ​​the current block is within a preset value range; The position coordinates of the current block are within a preset coordinate range; The reference area of ​​the current block in the current video frame satisfies a preset range restriction, and the reference area includes one or more of the current block, a template area of ​​the current block, and an extended area of ​​the template area; The reference area of ​​the current block meets the reference range restriction of intra block copy prediction.

7. The method according to claim 6, characterized in that The preset coordinate range includes a range greater than or equal to a coordinate threshold; and the method for determining the coordinate threshold includes at least one of the following methods: Determining the coordinate threshold according to a preset fixed value; Determining the coordinate threshold according to the template size of the template area; When the template area is an image area located above the current block, the template size is the height of the template area; when the template area is an image area located to the left of the current block, the template size is the width of the template area.

8. The method according to claim 7, characterized in that Determining the coordinate threshold according to the template size of the template area includes: When the current block does not include a chroma component, using a template size of the template area as the coordinate threshold; When the current block includes a chroma component, a weighting coefficient is determined according to a sampling ratio of a luminance component to a chroma component, and a product of the weighting coefficient and a template size of the template area is used as the coordinate threshold.

9. The method according to claim 4, characterized in that The code stream parameters further include a filtering index, where the filtering index is used to indicate a filtering mode for filtering the image block, and the filtering mode includes one or more different filtering methods.

10. The method according to claim 4, characterized in that The filtering flag is also used to indicate a filtering mode for filtering the image block, and the filtering mode includes one or more different filtering methods.

11. The method according to claim 10, characterized in that The code stream parameters further include a filter type field for indicating a filter type; When the value of the filter type field is the first value, the filter flag is used to indicate whether to perform filtering processing on the image block, and the filtering mode for performing filtering processing on the image block is the preset first mode; When the filtering type field takes a second value, the filtering flag is used to indicate whether the image block is filtered, and the filtering mode for filtering the image block is a preset second mode, which is a filtering mode different from the first mode.

12. The method according to claim 4, characterized in that The value of the filtering flag is determined according to the coding loss, and the coding loss is calculated by a preset cost function; If the coding loss after filtering the template region of the image block is less than the coding loss before filtering the template region of the image block, assigning the filtering flag to a first value, where the first value is used to indicate that filtering is to be performed on the image block; If the coding loss after filtering the template area of ​​the image block is greater than the coding loss before filtering the template area of ​​the image block, the filtering flag is assigned a second value, which is used to indicate that filtering of the image block is prohibited.

13. The method according to claim 4, characterized in that The value of the filtering flag is determined according to the number of transform coefficients in the image block; If the transform coefficients in the image block meet a preset first quantity condition, assigning the filter flag a first value, where the first value is used to indicate that filtering processing is to be performed on the image block; If the transformation coefficients in the image block meet a preset second quantity condition, the filtering flag is assigned a second value, and the second value is used to indicate that filtering processing on the image block is prohibited.

14. The method according to claim 3, characterized in that The first characteristic condition includes one or more of the following conditions: The current video frame where the current block is located has a specified image type; The current block has a specified distribution position in the current video frame; The image size of the current block is within a preset size range, where the image size includes one or more of width, height, or area; The block vector resolution of the current block is within a preset resolution range; The block vector residual of the current block is within a preset residual range, and the block vector residual includes one or more of a horizontal residual and a vertical residual; The block vector index of the current block is within a preset index range; The current block is an image block having a specified color component.

15. The method according to claim 3, characterized in that The second characteristic condition includes: the sample points of the reference block meet a preset sample point availability condition, and the sample point availability condition includes that the sample points are located within an image independent decoding boundary or the sample points have completed decoding and reconstruction.

16. The method according to claim 3, characterized in that The third characteristic condition includes one or more of the following conditions: The area of ​​the template region is within a preset area range; The sample points of the template area meet a preset sample point availability condition, where the sample point availability condition includes that the sample point is within an image independent decoding boundary or the sample point has completed decoding and reconstruction; The number of sample pairs collected in the template area is greater than a preset number threshold.

17. The method according to claim 16, characterized in that The number threshold comprises the number of model parameters in the filter.

18. The method according to any one of claims 1 to 17, characterized in that Fitting a filter according to the first template area and the second template area includes: Selecting a first sample point from the first template area, and selecting a second sample point corresponding to the first sample point from the second template area; A filter is fitted according to a sample pair consisting of the first sample point and the second sample point, where an input item of the filter includes a pixel value of the first sample point, an output item of the filter includes a pixel value of the second sample point, and the pixel value includes a predicted value or a reconstructed value.

19. The method according to claim 18, characterized in that The input item of the filter further includes pixel values ​​of one or more neighboring pixel points, where the neighboring pixel points are the nearest neighbor or the second nearest neighbor of the first sample point.

20. The method according to claim 19, characterized in that The filter includes one or more combination items with independent weighting parameters, and the combination item takes the pixel value of one or more pixel points as an input item.

21. The method according to claim 20, characterized in that The filter includes at least two combination terms with different orders, the order being the highest power of input terms in the combination terms.

22. The method according to claim 18, wherein Methods for selecting sample points in the template area include: Obtaining a pixel sampling mode of the current block, where the pixel sampling mode includes full pixel sampling or partial pixel sampling; When the pixel sampling mode of the current block is full pixel sampling, all pixels in the template area are selected as sample points; When the pixel sampling mode of the current block is partial pixel sampling, some pixels with designated sampling positions are selected in the template area as sample points.

23. The method according to claim 22, characterized in that The designated sampling location includes at least one of the following sampling locations: The pixel point's position coordinates meet the specified sampling position of the preset coordinate value conditions; A designated sampling position is selected along a preset pixel scanning direction; The pixel value falls within the specified sampling position within the preset value range.

24. The video decoding method according to claim 23, wherein: The coordinate numerical conditions include: At least one of the horizontal position coordinate and the vertical position coordinate of the pixel point is an even number; Alternatively, at least one of the horizontal position coordinate and the vertical position coordinate of the pixel point is an odd number.

25. The method according to any one of claims 1 to 17, characterized in that The template area is composed of one or more nearest neighbor areas or second nearest neighbor areas, the nearest neighbor area includes an image area with a specified image size located above or to the left of the image block, and the second nearest neighbor area includes an image area with a specified image size located above, below, or above the right of the image block.

26. The method according to claim 25, characterized in that The next-nearest neighboring region to the upper right of the image block has the same image size as the image block in the horizontal direction, and the next-nearest neighboring region to the upper right of the image block has a specified image size in the vertical direction, and the specified image size is greater than or equal to one; The next nearest neighboring area at the lower left of the image block has the same image size as the image block in the vertical direction, and the next nearest neighboring area at the upper right of the image block has the specified image size in the horizontal direction; The next-nearest neighboring area at the upper left of the image block has the specified image size in both the horizontal direction and the vertical direction.

27. The method according to claim 25, characterized in that The template area includes at least one of a full area combination, a left area combination and an upper area combination; The full region combination includes the nearest neighbor regions located on the left and above the image block and the next nearest neighbor regions located on the upper left, lower left and upper right of the image block; The left region combination includes a nearest neighbor region located on the left side of the image block and a next nearest neighbor region located on the lower left side of the image block; The upper region combination includes a nearest neighbor region located above the image block and a next nearest neighbor region located to the upper right of the image block.

28. The method according to claim 25, characterized in that The code stream parameters corresponding to the current block include a region type field, where the region type field is used to indicate a region type of a template region corresponding to the current block.

29. The method according to any one of claims 1 to 17, characterized in that After obtaining the second prediction value of the current block, the method further includes: A weighted operation is performed on the first prediction value and the second prediction value according to a preset weight coefficient to obtain a third prediction value of the current block.

30. A video processing device, characterized in that: include: a prediction module configured to perform intra-block copy prediction based on a reference block corresponding to a current block to obtain a first prediction value of the current block, wherein the current block is an image block to be processed in a current video frame, and the reference block is a reconstructed image block in the current video frame; a determining module configured to determine a first template region corresponding to the reference block and a second template region corresponding to the current block, wherein the first template region includes one or more reconstructed image regions adjacent to the reference block, and the second template region includes one or more reconstructed image regions adjacent to the current block; The filtering module is configured to fit a filter according to the first template area and the second template area, and filter the first prediction value of the current block according to the filter to obtain a second prediction value of the current block.

31. A computer-readable medium, characterized in that The computer-readable medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 29 is implemented.

32. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the executable instructions to implement the method according to any one of claims 1 to 29.

33. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 29 is implemented.