Image processing method, device, equipment, medium and program product

By adopting a new template area design in video encoding and decoding, the prediction and filtering effects of the template area are enhanced, the problem of insufficient adaptability of fixed-shape template areas is solved, and the efficiency and performance of video encoding and decoding are improved.

CN120658874APending Publication Date: 2025-09-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410317462.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The fixed-shape template regions in existing video codecs are difficult to adapt to complex and changing video content, resulting in poor prediction or filtering effects and difficulty in effectively processing different video sequences.

Method used

A new template area design is adopted to ensure that the pixel correlation between the sample points in the template area and the target pixel points corresponding to the pixel points to be processed in the current coding block is greater than the correlation threshold, and the template length in any direction is greater than the length threshold, including horizontal, vertical and diagonal directions, to improve the way the template area is determined during the video encoding and decoding process.

Benefits of technology

It improves the prediction and filtering effects of video encoding and decoding, enhances encoding and decoding efficiency and performance, and improves pixel accuracy and filtering quality by processing a wider range of neighborhood samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120658874A_ABST
    Figure CN120658874A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device, equipment, a medium and a program product. The method comprises the following steps: determining a current coding block in a compressed code stream; determining a template area for the current coding block; and decoding the current coding block based on the template region to obtain a reconstructed image of the current coding block. According to the embodiment of the invention, an effective template area can be provided, so that the video coding and decoding performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio and video technology, in particular to the field of video coding and decoding, and specifically to an image processing method, an image processing apparatus, an image processing device, a computer-readable storage medium, and a computer program product. Background Art

[0002] The template area is a reference area used in video coding and decoding technology to predict or filter the pixel values ​​or motion vectors of the coding blocks.

[0003] Currently, video encoding and decoding is often based on one or more fixed-shape template regions for encoding and decoding; however, existing fixed-shape template regions have the disadvantage of poor prediction or filtering effects, which makes the template regions unable to adapt to complex and changeable video content and difficult to effectively process different video sequences. Summary of the Invention

[0004] The embodiments of the present application provide an image processing method, apparatus, device, medium, and program product, which can provide an effective template area, thereby improving video encoding and decoding performance.

[0005] In one aspect, an embodiment of the present application provides an image processing method, the method comprising:

[0006] Determining a current coded block in a compressed code stream;

[0007] Determine a template region for the current coding block; a pixel correlation between a sample point in the template region and a target pixel point corresponding to a pixel point to be processed in the current coding block is greater than a correlation threshold, where the target pixel point is a pixel point corresponding to the pixel point to be processed in the original image; and a template length of the template region in any direction along a middle sample point of the template region is greater than a length threshold, where the any direction includes at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction.

[0008] The current coding block is decoded based on the template area to obtain a reconstructed image of the current coding block.

[0009] On the other hand, an embodiment of the present application provides an image processing method, the method comprising:

[0010] Determining a current encoding block to be encoded;

[0011] Determine a template region for the current coding block; a pixel correlation between a sample point in the template region and a target pixel point corresponding to a pixel point to be processed in the current coding block is greater than a correlation threshold, where the target pixel point is a pixel point corresponding to the pixel point to be processed in the original image; and a template length of the template region in any direction along a middle sample point of the template region is greater than a length threshold, where the any direction includes at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction.

[0012] The current coding block is encoded based on the template area to generate a compressed code stream.

[0013] In another aspect, an embodiment of the present application provides an image processing device, comprising:

[0014] a determination unit, configured to determine a current coding block in a compressed code stream;

[0015] a processing unit configured to determine a template region for a current coding block; wherein a pixel correlation between a sample point in the template region and a target pixel point corresponding to a pixel point to be processed in the current coding block is greater than a correlation threshold, the target pixel point being a pixel point corresponding to the pixel point to be processed in the original image; and wherein a template length of the template region in any direction along a middle sample point of the template region is greater than a length threshold, the any direction comprising at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction;

[0016] The processing unit is further configured to perform decoding processing on the current coding block based on the template area to obtain a reconstructed image of the current coding block.

[0017] In another aspect, an embodiment of the present application provides an image processing device, comprising:

[0018] a determination unit, configured to determine a current coding block to be encoded;

[0019] a processing unit configured to determine a template region for a current coding block; wherein a pixel correlation between a sample point in the template region and a target pixel point corresponding to a pixel point to be processed in the current coding block is greater than a correlation threshold, the target pixel point being a pixel point corresponding to the pixel point to be processed in the original image; and wherein a template length of the template region in any direction along a middle sample point of the template region is greater than a length threshold, the any direction comprising at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction;

[0020] The processing unit is further configured to perform encoding processing on the current encoding block based on the template area to generate a compressed code stream.

[0021] On the other hand, an embodiment of the present application provides an image processing device, the image processing device comprising:

[0022] a processor for loading and executing computer programs;

[0023] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image processing method is implemented.

[0024] On the other hand, the present application provides a computer-readable storage medium, which stores a computer program. The computer program is suitable for being loaded by a processor and executing the above-mentioned image processing method.

[0025] On the other hand, the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the computer device executes the above-mentioned image processing method.

[0026] In an embodiment of the present application, an effective template area is provided, and the effectiveness of the template area can be reflected in the following aspects: on the one hand, the pixel correlation between the sample points in the template area and the original pixel points (or target pixel points) corresponding to the pixel points to be processed (such as the pixel points to be filtered or the pixel points to be predicted) in the current coding block in the compressed code stream is greater than the pixel correlation of the traditional template area; in this way, when the current coding block is decoded (such as prediction or filtering) using the model area, the pixel points in the current coding block can be processed using the neighboring sample points with higher correlation belonging to the template area, which greatly improves the prediction effect or filtering effect of the pixel points in the current coding block. On the other hand, the template length of the template area in any direction along the middle sample points in the template area is longer than the corresponding template length in the traditional template area; in this way, when the current coding block is decoded using the template area, the pixel points in the current coding block can be processed using the neighboring sample points within a larger reference range within the template area, which significantly improves the prediction accuracy or filtering accuracy of the pixel points in the current coding block. For example, the template area is a filter template (or filter shape) used in various loop filtering tools; under this implementation method, if the sample points within the current pixel (the middle sample point, that is, the sample point in the filter template that spatially overlaps with the pixel point to be filtered in the current coding block) in the filter template have a higher pixel correlation with the target pixel point corresponding to the pixel point to be processed in the current coding block, and the template length of the filter template in any direction along the middle sample point is longer (that is, it contains more sample points), then when the current coding block is filtered based on the filter template, the distortion can be reduced to a greater extent, the filtering quality can be enhanced, and the encoding and decoding efficiency and performance can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0028] Figure 1 It is a schematic diagram of a video encoder framework;

[0029] Figure 2 It is a schematic diagram of a framework of a cross-component sample adaptive compensation technology;

[0030] Figure 3a It is a schematic diagram of the shape of a diamond filter template;

[0031] Figure 3b It is a schematic diagram of the shape of a filter template in an adaptive loop filter;

[0032] Figure 4 It is a schematic diagram of a framework of cross-component adaptive loop filtering;

[0033] Figure 5 It is a schematic diagram of the framework of bilateral filtering;

[0034] Figure 6 is a schematic diagram of the architecture of an image processing system provided by an exemplary embodiment of the present application;

[0035] Figure 7 is a flowchart of an image processing method provided by an exemplary embodiment of the present application;

[0036] Figure 8 is a schematic diagram of a pixel correlation strategy provided by an exemplary embodiment of the present application;

[0037] Figure 9 is a schematic diagram for introducing the direction of a template area provided by an exemplary embodiment of the present application;

[0038] Figure 10a 1 is a schematic diagram of the shape of a filter template provided by an exemplary embodiment of the present application;

[0039] Figure 10b is a schematic diagram of the shape of another filter template provided by an exemplary embodiment of the present application;

[0040] Figure 10c This is a schematic diagram of the shape of another filter template provided by an exemplary embodiment of the present application;

[0041] Figure 10d1 is a schematic diagram of the shape of another filter template provided by an exemplary embodiment of the present application;

[0042] Figure 11a is a schematic diagram of a coded block in the spatial domain provided by an exemplary embodiment of the present application;

[0043] Figure 11b is a schematic diagram of a coded block in the time domain provided by an exemplary embodiment of the present application;

[0044] Figure 12 is a flowchart of another image processing method provided by an exemplary embodiment of the present application;

[0045] Figure 13 is a structural diagram of an image processing device provided by an exemplary embodiment of the present application;

[0046] Figure 14 is a schematic structural diagram of another image processing device provided by an exemplary embodiment of the present application;

[0047] Figure 15 It is a structural diagram of an image processing device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0049] In the embodiment of the present application, an image processing solution is proposed, which mainly involves video coding and decoding technology. The following is a brief introduction to the technical terms and related concepts involved in the image processing solution provided in the embodiment of the present application:

[0050] 1. Video coding technology.

[0051] Video coding technology, also known as video codec technology, refers to the encoding method that converts an original video file into another format through compression. Video decoding is the reverse process of video coding. A video file consists of at least two video frames (or image frames) connected in sequence; in other words, a video frame is the smallest or most basic unit of video. When a video is played, multiple video frames are output continuously in chronological order. When the continuous video frame rate exceeds 24 frames per second, the human eye perceives each frame as smooth and continuous, based on the principle of persistence of vision. Video is represented by a video signal, typically an electrical signal. Transmitting video signals enables video transmission and storage across networks. Based on the acquisition method, video signals can be captured by a camera or generated by a computer. Due to the different statistical characteristics of different video signals, the corresponding compression encoding methods may also vary.

[0052] The following is an introduction to the existing mainstream video coding technologies:

[0053] Modern mainstream video coding technologies, such as the international video coding standards HEVC (High Efficiency Video Coding), such as HEVC / H.265, VVC (Versatile Video Coding), such as VVC / H.266, and AVS (Audio Video Coding Standard), use a hybrid coding framework to perform the following operations and processing on the input raw video signal:

[0054] 1) Block partition structure: According to the size of the input image (such as the video frame that needs to be compressed, encoded or decoded in the video), the input image is divided into several non-overlapping processing units; then similar compression operations can be performed on each processing unit during encoding and decoding, avoiding the difficulties brought about by directly encoding and decoding a frame of image. Among them, the divided processing unit can be called CTU (Coding Tree Unit) or LCU (Largest Coding Unit). The processing unit CTU can also continue to be divided more finely to obtain one or more basic coding units, which are called CU (Coding Unit or Coding Block). Among them, each CU unit is the most basic element in a coding and decoding link. The subsequent embodiments of this application use each CU as an example to explain the relevant encoding and decoding.

[0055] 2) Predictive Coding: Predictive coding is based on the correlation characteristics between discrete signals (such as the spatial correlation between pixels in different parts of the same video frame in the spatial domain, or the temporal correlation between pixels in different video frames in a video sequence (including multiple frames played sequentially) in the temporal domain). It uses one or more signals before the current signal to predict the predicted value of the current signal, thereby encoding the residual (or prediction error) between the actual value of the current signal and the predicted value. This avoids the high computational complexity and waste of compression resources caused by directly compressing all video frames.

[0056] Predictive coding mainly includes intra-frame prediction and inter-frame prediction. Among them: ① Intra-frame prediction: The prediction signal used to predict the current coding unit comes from an already coded and reconstructed area within the same image; ② Inter-frame prediction: The prediction signal used to predict the current coding unit comes from another image (which can be called a reference image) that has been coded and is different from the image to which the current coding unit belongs. During the video encoding and decoding process, when the encoder encodes the unit to be coded (such as the CU mentioned above) in the original video signal (such as a video frame), if any predictive coding method (such as intra-frame prediction or inter-frame prediction) is used, it is necessary to use the reconstructed video signal in the original video signal (for example, if the predictive coding method is intra-frame prediction, the reconstructed video signal belongs to the current image; if the predictive coding method is inter-frame prediction, the reconstructed video signal comes from an image reconstructed before the current image) to predict the unit to be coded, and obtain the residual video signal of the current unit to be coded (such as the residual mentioned above). Then, this residual video signal is compressed and encoded to generate a bitstream, which is then transmitted to the decoder. Correspondingly, the encoder also needs to inform the decoder of any predictive coding method used in the encoding process, so that after receiving the compressed code stream (that is, the code stream mentioned above, or called image code stream, video code stream, encoding code stream, etc.), the decoder uses the same predictive coding method as the encoding process to reconstruct the image during the decoding process of the compressed code stream.

[0057] 3) Transform coding and quantization (Transform & Quantization): The residual video signal undergoes transformation operations such as DFT (Discrete Fourier Transform) and DCT (Discrete Cosine Transform), which can convert the residual video signal into a transform domain, which is called a transform coefficient. In this way, the signal in the transform domain can be further subjected to lossy quantization operations, losing some redundant information, so that the quantized signal is conducive to compression expression. In some video coding standards, there may be one or more change methods; therefore, in the video encoding and decoding process, the encoder needs to select a transform method for the current encoded CU and inform the decoder of the transform method so that the decoder can use the corresponding transform method for inverse transformation during the decoding process. It is worth noting that the degree of quantization fineness of the above-mentioned quantization operations is usually determined by the quantization parameter (QP). A larger value of the quantization parameter QP means that coefficients with a larger value range will be quantized into the same output, which usually results in greater distortion and a lower bit rate (i.e., the number of data bits transmitted per unit time during data transmission). Conversely, a smaller value of the quantization parameter QP means that coefficients with a smaller value range will be quantized into the same output, which usually results in less distortion and a higher bit rate.

[0058] 4) Entropy Coding or Statistical Coding: The quantized transform domain signal will be statistically compressed according to statistical coding (that is, statistically compressed according to the frequency of occurrence of each value), and finally a binary (0 or 1) compressed code stream will be output. At the same time, the encoding generates other information, such as the selected mode (such as prediction mode) and motion vector, which also need to be entropy coded to reduce the bit rate. The statistical coding mentioned above is a lossless coding method that can effectively reduce the bit rate required to express the same signal; among them, statistical coding can include but is not limited to: variable length coding (VLC) or context-based binary arithmetic coding (CABAC).

[0059] 5) Loop Filtering: Based on the image encoded in the aforementioned steps, a series of operations, including inverse quantization, inverse transformation, and prediction compensation (i.e., the reverse operations of 2) to 4) above, are performed to obtain a reconstructed decoded image. Due to the effects of quantization, the reconstructed image (i.e., the decoded image) differs from the original image in some information, resulting in distortion. Therefore, filtering the reconstructed image using filters can effectively reduce the degree of distortion caused by quantization. These filters may include, but are not limited to, deblocking filters (DF), sample-adaptive offset (SAO), or adaptive loop filters (ALF). These filtered reconstructed images can serve as reference information for subsequent encoded images and are used to predict future signals. Therefore, the filtering operation described above is also called loop filtering, or the filtering operation within the encoding loop.

[0060] The following combination Figure 1 The basic process of video encoding given above (i.e. steps 1) to 5)) is introduced with the video encoder shown in FIG. Figure 1 The current coding block to be coded (or called the current coding unit) is the k-th CU in the current image frame (such as Figure 1 The s shown k [x,y]) is used as an example, k is a positive integer, and k is less than or equal to the total number of CUs contained in the current image frame. k [x, y] represents the pixel point (referred to as pixel) with coordinates [x, y] in the k-th CU, where x represents the horizontal coordinate of the pixel and y represents the vertical coordinate of the pixel; s k [x,y] can obtain the prediction signal after motion compensation or intra-frame prediction. The prediction signal and the original signal s k [x,y] performs difference operation to obtain the residual video signal u k [x,y]; then the residual video signal u k After [x,y] is transformed and quantized, the quantized data is obtained. Among them, the data output by the quantization process has two data flows:

[0061] Data Flow 1: The encoder sends the quantized output data to an entropy encoder for entropy encoding, generating an encoded bitstream. This bitstream is then stored in a buffer pending transmission to the decoder. After receiving the bitstream, the decoder first performs entropy decoding on each CU to obtain various mode information and quantized transform coefficients for the current CU. Each transform coefficient is then dequantized and inversely transformed to obtain a residual signal. Furthermore, based on the known mode information from the encoder, the decoder can obtain a prediction signal corresponding to the current CU. The residual signal and the prediction signal are then added together to produce a reconstructed signal. Finally, the reconstructed value (or reconstructed signal) of the decoded image is filtered through a loop filter to produce the final output signal.

[0062] Data flow 2: The encoder can perform inverse quantization and inverse transformation on the quantized output data to obtain the inverse transformed residual video signal u′ k [x, y]; then, the inverse transformed residual video signal u′ k [x,y] and prediction signal Add up to get a new prediction signal And the new prediction signal The new prediction signal is sent to the buffer of the current image. After intra-frame prediction processing, we get And the new prediction signal After loop filtering, the reconstructed signal s′ can be obtained k [x,y], and reconstruct the signal s′ k [x,y] is sent to the decoded image buffer for storage to generate the reconstructed video. Reconstructed signal s′ k [x,y] is obtained through motion compensation prediction in Can represent the reference block, m x and m y Represents the horizontal and vertical components of the motion vector of the reference block, respectively.

[0063] Based on the above Figure 1As shown in the detailed description of the encoding process, at the decoding end, after receiving the compressed bitstream, the decoder first performs entropy decoding on each CU to obtain the various mode information used in the encoding process and the quantized transform coefficients. These mode information and quantized transform coefficients are then dequantized and inversely transformed to produce a residual signal (the aforementioned residual video signal). Furthermore, based on the known coding mode information, a prediction signal corresponding to the CU can be obtained. The prediction signal and the residual signal are then added together to produce the reconstructed signal (or decoded image) for the CU. Furthermore, to account for distortion and other issues, the reconstructed value of the decoded image undergoes filtering to produce the final output signal. During this series of encoding processes, the coding framework primarily relies on the rate-distortion optimization criterion (RDO) to evaluate and select the optimal coding parameters and results. The RDO criterion simultaneously considers both bitrate and distortion in the cost function calculation, ensuring both low distortion and low bitrate, which is more conducive to video stream transmission.

[0064] 2. Template Area

[0065] Template regions are areas used as references during the video encoding and decoding process. These regions can be used to search for the area that best matches the current block to be encoded or decoded, allowing prediction of the current block using the coded blocks within these regions, thereby enabling image reconstruction (e.g., on the decoder side) or compressed bitstream transmission (e.g., on the encoder side). Different template regions are involved in different stages of the video encoding and decoding process. For example, during the intra-frame prediction or inter-frame prediction stages of video encoding and decoding, a template region of a certain shape is involved. This template region can be referred to as a prediction template or prediction shape. During the intra-frame prediction or inter-frame prediction process, the prediction template uses the coded blocks within the prediction template in the image frame, either in the spatial or temporal domain, to weight the pixel prediction values ​​of the current block, thereby implementing the prediction process. Another example is the filtering stage of the video encoding and decoding process, which involves a template region of a certain shape. This template region can be referred to as a filter template or filter shape. Considering that quantization and other factors can cause some information to differ from the original image and result in distortion, a filter template of a certain shape is used to filter the reconstructed image to reduce the degree of distortion caused by quantization.

[0066] For ease of explanation, the image processing solution provided by the embodiment of the present application will be introduced using the template area as a filter template as an example. The following is an introduction to the various filter tools deployed by the filter template; among them:

[0067] (1) Deblocking Filter (DBF).

[0068] Since the transform and quantization encoding process of each block in video coding is independent, the quantization error and distribution characteristics introduced by each block are also independent of each other. Therefore, there will be discontinuity problems at the boundaries of adjacent blocks in the video frame (or image). In addition, in the motion compensation prediction process, the prediction values ​​of adjacent blocks may come from different positions in different images, which will also cause the boundaries between different blocks to be discontinuous. Currently, video coding standards such as HEVC (High Efficiency Video Coding, international video coding standard HEVC / H.265), VVC (Versa tile Video Coding, international video coding standard VVC / H.266) and AVS (Audio Video Coding Standard, video coding standard AVS) use DBF technology to eliminate blocking effects. Among them, DBF technology can specifically reduce or eliminate blocking effects by smoothing block boundaries, making the image look smoother and more natural. Specifically, DBF includes two steps: filtering decision and filtering operation. First, filtering decision is made to obtain the maximum filtering length of the boundary, the filtering strength of the boundary (such as no filtering, short-tap weak filtering, short-tap strong filtering and long-tap filtering, etc.) and its filtering parameters. Then, based on the determined maximum filtering length and boundary filtering strength, the filtering parameters are used to adaptively correct the block boundary.

[0069] (2) Sample Adaptive Offset (SAO).

[0070] The quantization process in video coding causes the loss of high-frequency information, resulting in a ripple effect at the edges of detailed textures in the image, known as ringing. Visually, this effect results in oscillations due to dramatic grayscale changes in the image. To suppress this ringing effect and minimize both subjective and objective quality loss, video coding standards such as HEVC, VVC, and AVS employ SAO (Saturation Allocation) to compensate pixel values. SAO uses the CTU as the basic unit, classifying reconstructed pixels and applying different compensation values ​​to each class. SAO primarily includes two compensation methods: edge offset (EO) and band offset (BO). Edge offset adaptively compensates pixel values ​​in border regions where ringing occurs, reducing the potential for moiré-like distortion during the encoding and decoding process. Band offset divides all pixel values ​​in the image into several sidebands, selecting four consecutive sidebands for compensation within each sideband.

[0071] (3) Cross-Component Sample Adaptive Offset (CC-SAO).

[0072] The cross-component sample adaptive compensation technology is similar to the aforementioned sample adaptive compensation technology. It also adaptively classifies different reconstructed pixels and then compensates the offset values ​​of pixels of different categories. In detail, the sample adaptive compensation technology mainly compensates pixel values ​​within the same color component, while the cross-component sample adaptive compensation technology allows compensation between different color components. For example, the technical framework diagram of the cross-component sample adaptive compensation can be seen in Figure 2 ;like Figure 2 As shown, the input of the cross-component sample adaptive compensation framework comes from different components, namely Y component, U component and V classification, and its output is added to the output result of the sample adaptive compensation for YUV components to remove redundancy between components.

[0073] YUV is a digital representation of color, specifically a pixel format that represents luminance and chrominance components separately. YUV signals are derived from RGB signals in video encoding technology. RGB (Red, Green, Blue) is a color model used in computer systems. A variety of colors can be created by varying and superimposing the three color channels. "Y" in a YUV signal represents brightness (luminance or luma), or grayscale values; "U" and "V" represent chrominance (chroma), describing the image's color and saturation, and are used to specify pixel color. Compared to RGB signals, encoding and transmitting YUV signals requires significantly less bandwidth (RGB requires three separate video signals to be transmitted simultaneously). Luminance is established from the RGB input signals by superimposing specific portions of the RGB signals. "Chroma" defines two aspects of color: hue and saturation, represented by Cr and Cb, respectively. Cr reflects the difference between the red portion of the RGB input signal and the luminance value of the RGB signal, while Cb reflects the difference between the blue portion of the RGB input signal and the luminance value of the RGB signal. The importance of using the YUV color format lies in the separation of its luminance signal (Y) and chrominance signals (U and V). If an image only includes the Y component (luminance) and does not include the U and V components (chrominance), then the image is a black and white grayscale image.

[0074] (4) Adaptive Loop Filter (ALF)

[0075] Adaptive loop filtering aims to construct the Wiener-Hoff equation based on the principle of Wiener filtering through the original image information and the reconstructed image information to be filtered, and solve a series of filter coefficients with minimum mean square error to achieve the purpose of reducing decoding errors and improving PSNR (Peak Signal-to-Noise Ratio); the peak signal ratio is an indicator for evaluating image or video quality; it is calculated by measuring the mean square error between the original image (or video) and the processed image (or video); the higher the value of the calculated peak signal-to-noise ratio, the smaller the image distortion, that is, the better the image quality. Specifically, adaptive loop filtering can be based on a filter template of a certain shape, using a fixed set of filter coefficients or an online trained set of filter coefficients to convolve with the current reconstructed image to be filtered (or the current coding block to be encoded), and perform operations on neighboring pixels to achieve the purpose of enhancing the image. Among them, the shape diagrams of the two diamond filter templates in the international video coding standard VVC / H.266 can be seen. Figure 3a ,like Figure 3a The filter coefficients of the filter template are shown as C0-C2; the shape diagram of the filter template in the adaptive loop filter in the video coding standard AVS can be seen in Figure 3b ,like Figure 3b The filter template shown has filter coefficients C0-C14.

[0076] (5) Cross-Component Adaptive Loop Filter (CC-ALF)

[0077] Considering that the brightness of a video usually contains more detailed information, it is possible to consider using the brightness information for ALF to compensate for the details of the chrominance components, thereby improving the quality of the chrominance components. In practical applications, since the chrominance components are usually encoded and transmitted at a lower resolution, there may be a mismatch or distortion between the chrominance components and the brightness components; for this reason, cross-component adaptive loop filtering can reduce this distortion by applying a filter template to the chrominance components during the decoding process. The technical framework of cross-component adaptive loop filtering can be found in Figure 4 ;like Figure 4 As shown, the adaptive loop filtering based on a filter template of a certain shape across components inputs the luminance pixel Luma to be filtered, outputs the chrominance compensation value, and adds it to the processing result of the adaptive loop filtering ALF chrominance. The result after addition can improve the quality of the chrominance component.

[0078] (6) Bilateral Filter (BIF).

[0079] In order to further eliminate the ringing effect caused by transform quantization, international video coding standards such as VVC / H.266 have proposed bilateral filtering technology, which can enhance the quality of the decoded image after inverse transformation. For example, the filtering process of bilateral filtering can be seen in Figure 5 ;like Figure 5 The filter template shown is a cross-shaped filter, filtering the current pixel using spatially adjacent pixels within an 8×8 TU, thereby enhancing image quality and improving coding efficiency. A TU (Transform Unit) is the basic unit for transformation and quantization during video encoding and decoding.

[0080] It should be noted that the above (1)-(6) are several exemplary filters provided in the embodiments of the present application and do not limit the embodiments of the present application.

[0081] Based on the above-mentioned introduction to basic content such as video coding and decoding technology and template areas, the image processing solution proposed in the embodiment of the present application improves the existing template areas in traditional video coding, proposes a variety of more effective template areas, and proposes a method for determining the template area during the video coding and decoding process to enhance the prediction or filtering effect of the template area, thereby improving video coding efficiency and performance. In short, the embodiment of the present application mainly improves video coding and decoding efficiency and performance from two aspects. On the one hand, it proposes a new and more effective template area, and on the other hand, it provides an indication method for determining the template area during the video coding and decoding process. The following is a brief introduction to these two aspects, including:

[0082] (1) The embodiments of the present application propose a new and more effective template area.

[0083] The new template area proposed in the embodiment of the present application, on the one hand: the pixel correlation between the sample points (such as pixel points) in the new template area and the original pixel points (or called target pixel points, which refers to the pixel points corresponding to the pixel points to be processed in the original image to which the current coding block belongs) to the pixel points to be processed in the current coding block is greater than the pixel correlation between the sample points in the traditional template area and the original pixel points corresponding to the pixel points to be processed in the current coding block; in this way, when using the new template area to predict or filter the pixel points to be processed in the current coding block, it is possible to refer to the neighboring sample points with a greater pixel correlation with the pixel points to be processed in the current coding block, thereby improving the prediction effect or filtering effect of the pixel points to be processed in the current coding block. For example, the template area is a filter template, and the distribution of the samples contained in the filter template is different from the distribution of the samples contained in the traditional filter template. Therefore, the video pixels referenced by the new filter template are different from the video pixels referenced by the traditional filter template, and the pixel correlation between the samples in the new filter template and the target pixels corresponding to the pixels to be processed in the current coding block is large; in this way, a better filter coefficient of the filter template can be obtained, and the current coding block can be filtered based on the better filter coefficient, which can enhance the filtering effect and improve the video encoding and decoding efficiency.

[0084] On the other hand, the new template region proposed in the embodiments of the present application has a longer template length in any direction of the intermediate sample point (e.g., the current pixel) than the traditional template region. This longer template length allows for a larger reference range for the coding block to be processed. Thus, when using the template region to predict the current coding block, a more accurate prediction or filtering effect can be achieved based on pixels within a larger reference range, significantly improving video encoding and decoding performance.

[0085] (2) The embodiments of the present application provide multiple indication methods for determining the template area.

[0086] The embodiments of the present application provide a display indication method and an implicit indication method for template area switching. Among them, the display indication method refers to the determination of the template area through the flag bit transmission method, and the implicit indication method refers to the determination of the template area through the non-flag bit transmission method (such as the default method, the reference frame method, the inheritance method or the derivative method, etc.). The display indication method and the implicit indication method provided by the embodiments of the present application can determine different template areas for different coding blocks during the video encoding and decoding process, or determine the same template area for different coding blocks; by designing a variety of indication methods, the method of determining the template area during the video encoding and decoding process is enriched to meet greater encoding and decoding requirements.

[0087] As described above, a template region can be a template for different stages of video encoding and decoding, such as a filter template or a prediction template. Depending on the encoding and decoding stage used by the template region, the aforementioned template region is different, and the indication method for indicating the template region is also different. For example, if the template region is a filter template, the explicit indication method and the implicit indication method are used to indicate the filter template determined for the coding block; for another example, if the template region is a prediction template, the explicit indication method and the implicit indication method are used to indicate the prediction template determined for the coding block.

[0088] Thus, on the one hand, the embodiment of the present application provides a variety of new template areas, the distribution of the samples contained in the new template area is different from the distribution of the samples contained in the traditional template area, and the new template area has at least the following characteristics: the pixel correlation between the samples in the new template area and the original pixels corresponding to the pixels to be processed in the current coding block is greater, and the new template area has a longer template length in any direction at the intermediate samples (i.e., the pixels in the new template area that coincide with the original pixels of the current pixel to be filtered). Therefore, based on the samples with greater pixel correlation between the original pixels corresponding to the current pixel to be processed in the neighborhood of the intermediate samples in the new template area, when the current coding block is processed (such as predicting or filtering), the prediction effect or filtering effect can be improved. In addition, based on the characteristic that the template length of the new template area in any direction of the intermediate samples is longer, a larger reference range can be determined for the original pixels of the current pixel to be filtered, so that when the original pixels of the current pixel to be filtered are processed based on the larger reference range, a more accurate prediction result or filtering result can be obtained, thereby improving video encoding and decoding efficiency and performance.

[0089] Furthermore, embodiments of the present application provide multiple ways to indicate template regions. These multiple indications allow different coding blocks in a video sequence to switch between multiple template regions, and also allow different coding blocks in a video sequence to use the same template region. By designing multiple indications, the methods for determining template regions during video encoding and decoding are enriched, meeting greater encoding and decoding requirements, significantly improving video encoding and decoding efficiency, and enhancing video encoding and decoding performance.

[0090] The image processing solution provided in the embodiments of the present application can be applied to products with video encoding and decoding capabilities or video compression capabilities, and the products here may include image processing devices or applications.

[0091] An application refers to any application with video editing capabilities. For example, an application is a loop filter tool used in digital signal processing. This loop filter tool can be one of the aforementioned filters, offering advantages such as filtering capabilities, dynamic feedback loop characteristics, improved loop detection performance, and enhanced anti-interference performance. An application can be a computer program designed to perform one or more specific tasks. Based on how applications operate, applications can include: clients installed on terminals, mini-programs (subprograms of clients) that can be used without downloading or installing, web (World Wide Web) applications opened through browsers, and so on. Based on the type of functionality, applications can include, but are not limited to, IM (Instant Messaging) applications and content interaction applications. Instant messaging applications refer to internet-based applications for instant messaging and social interaction. These applications include, but are not limited to, social applications with communication capabilities, map applications with social interaction capabilities, and game applications. Content interaction applications refer to applications that enable content interaction, such as online banking, sharing platforms, personal spaces, and news applications.

[0092] The image processing device may refer to a physical device with video encoding and decoding functions. The physical device may include a video codec, a terminal / server deployed with a video codec, or a terminal / server deployed with a loop filtering tool. Among them, the video codec can be further divided into a video encoder and a video decoder. The video encoder is used to convert the original video data (such as the original video sequence) into a compressed code stream in a compressed format; and the video decoder is used to convert the compressed video data (i.e., the compressed code stream) back to the original video sequence. Among them, the terminal may include but is not limited to: a smartphone (such as a smartphone deploying an Android system, or a smartphone deploying an Internetworking Operating System (IOS)), a tablet computer, a portable personal computer, a mobile Internet device (Mobile Internet Devices, MID), a vehicle-mounted device, a head-mounted device and other terminal devices. The embodiment of the present application does not limit the type of terminal device, which is explained here. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.

[0093] It should be noted that the above is an exemplary introduction to the products to which the image processing scheme is applicable in the embodiments of the present application, and does not limit the type of products used in the image processing scheme provided in the embodiments of the present application. For example, the image processing scheme provided in the embodiments of the present application can also be software provided by a plug-in, or deployed in an application or computer device in the form of a plug-in. For ease of explanation, the image processing scheme will be deployed using an image processing device (such as a video encoder or video decoder), that is, the image processing device will be used to implement video encoding and decoding using the image processing scheme as an example for introduction, which will be specifically described here.

[0094] An exemplary architecture diagram of an image processing system based on video coding technology can be found in Figure 6 ;like Figure 6 The system shown includes terminal 601 and terminal 602. The embodiment of the present application does not limit the number and type of terminals involved in the system. Among them, terminal 601 can be a terminal device held by a user with a video sharing demand, and terminal 602 can be a terminal device held by a user with a video receiving demand. In a specific implementation, when user 1 holding terminal 601 wants to share a video with user 2 holding terminal 602, user 1 can send the video through terminal 601. At this time, terminal 601 (specifically, the video encoder deployed in terminal 601) can encode the video (such as the corresponding processing of steps 1)-5) in the aforementioned video encoding and decoding technology). In the prediction process of the video encoding process, the template area provided by the embodiment of the present application (the template area is used as a prediction template at this time) can be used to perform inter-frame prediction processing or intra-frame prediction processing. After the video encoding process obtains the encoded video, the encoded video can be subjected to inverse quantization, inverse transformation and prediction compensation to obtain a reconstructed decoded image, and then the template area provided by the embodiment of the present application (the template area is used as a filter template at this time) is used to filter the compressed code stream, effectively reducing the degree of distortion caused by quantization. The reconstructed image after the filtering operation can be stored in the decoded image cache as a reference for subsequent encoded images and used to predict future signals.

[0095] Then, terminal 601 transmits the encoded compressed code stream directly to terminal 602. Finally, after receiving the compressed code stream, terminal 602 decodes the received compressed code stream. The decoding process can be regarded as the inverse process of the encoding process, and its purpose is to restore the original video. Specifically, the terminal 602 (specifically, the video decoder deployed in the terminal 602) performs entropy decoding on the compressed code stream, and the entropy decoding obtains various mode information and transform coefficients used for the current coding block (such as the CU unit currently to be decoded) in the compressed code stream during the encoding process (the entropy decoding process can use the explicit indication method and the implicit indication method provided in the embodiment of the present application to determine the template area used in the encoding process). The various mode information and transform coefficients are dequantized and inversely transformed to obtain a residual signal of the current coding block; the prediction signal of the current coding block is predicted based on the encoding information, etc., and the template area provided in the embodiment of the present application (in this case, the template area is a prediction template) can be used in this prediction process to predict the prediction signal; the residual signal and the prediction signal of the current coding block are then added to obtain a reconstructed signal of the current coding block; the reconstructed signal can also be filtered through the template area (in this case, the template area is a filter template) to ultimately restore the original video.

[0096] It should be noted that the above Figure 6 This is only a schematic diagram of the architecture of an exemplary image processing system provided in an embodiment of the present application. In actual applications, the architecture can be adaptively changed. For example, the system also includes a server 603, which is an intermediate device that interacts with the terminal 601 and the terminal 602 to provide technical services and technical support to the terminal 601 and the terminal 602, and can realize data forwarding and caching, etc. The number of servers 603 included in the system can be one or more, such as distributed servers. The terminals and servers in the system can be directly or indirectly connected via wired or wireless communication, and this application does not impose any restrictions on this.

[0097] It should also be noted that the collection and processing of relevant data in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations. The acquisition of personal information must be subject to the knowledge or consent of the individual subject (or the presence of a legal basis for information acquisition), and subsequent data use and processing must be carried out within the scope of authorization of laws and regulations and the subject of personal information. For example, when the embodiments of this application are applied to specific products or technologies, such as when a terminal sends a video, it is necessary to obtain the permission or consent of the uploader or creator of the video, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant region.

[0098] Based on the above introduction to the image processing solution and the applied product or scenario architecture, a more detailed image processing method proposed in the embodiment of the present application is introduced below in conjunction with the accompanying drawings.

[0099] See Figure 7 , Figure 7 is a flowchart of an image processing method provided by an exemplary embodiment of the present application; Figure 7 The flowchart shown may be a flowchart of the decoding side, and may be specifically executed by an image processing device (such as a video decoder) held by the decoding end. The method may include but is not limited to steps S701-S703:

[0100] S701: Determine a current coding block in a compressed code stream.

[0101] The current coding block is the coding area to be decoded in the current image during the decoding process. The coding area can be a coding block at the slice level, CTU level, block level or image level. Among them, a slice can be understood as a data block or data piece in an image, which is obtained by segmenting an image. An image can include one or more slices after segmentation; a slice can include multiple macroblocks; a macroblock is a coding block smaller than a slice obtained by dividing the image. CTU (Coding Tree Unit) is similar to a macroblock. It is a coding block obtained by dividing an image, and the divided CTU can be further divided into smaller coding blocks, such as coding units (CU), prediction units (PU) and transform units (TU). Based on this, the embodiment of the present application does not limit the type of the current coding block. For example, the current coding block may include at least one of the following: image, data block, coding tree unit and macroblock. In other words, the current coding block can refer to an entire image (or video frame) to be decoded in a compressed code stream, or it can be a partial image area to be decoded in an image in a compressed code stream.

[0102] S702: Determine a template area for the current coding block.

[0103] This template region is a new template region provided in the embodiment of the present application. Compared with the traditional template region, this new template region is more effective. That is, when the new template region is used to perform decoding processing (such as prediction processing or filtering operation) on the current coding block, a better processing effect (such as prediction effect or filtering effect) can be obtained. Among them, the new template region provided in the embodiment of the present application has the following design features:

[0104] (1) The pixel correlation between the sample point in the template area and the target pixel point corresponding to the pixel point to be processed in the current coding block is greater than the correlation threshold. Among them, the target pixel point corresponding to the pixel point to be processed in the current coding block refers to the pixel point corresponding to the pixel point to be processed in the original image at the coding end, so the target pixel point can also be called the original pixel point. Specifically, after the template area is designed, the coding cost of the designed template area and the traditional template area can be compared at the coding end (such as a video encoder) through a certain coding cost comparison method (or cost comparison strategy); if the coding cost of the designed template area is less than the coding cost of the traditional template area (such as the traditional template area), it indicates that the designed template area has a more significant prediction effect or filtering effect, so the designed template area can be used as the template area used in encoding and decoding, and the designed template area is more effective.

[0105] The above-mentioned cost comparison strategies include at least one of the following: pixel correlation strategy and distortion cost strategy. Among them: ① The distortion cost strategy is to select the template region with the lowest coding cost, that is, the lowest distortion, by calculating the degree of distortion after processing using the template region. Distortion cost strategies may include but are not limited to: Rate Distortion Optimization (RDO) and Peak Signal-to-Noise Ratio (PSNR) strategies. ② The pixel correlation strategy is to select the template region with the lowest coding cost, that is, the higher the pixel correlation, by calculating the pixel correlation. The higher the pixel correlation between the sample points in the template region and the original pixel points corresponding to the pixel points to be processed in the current coding block, the closer the sample points in the template region are to the original pixel points corresponding to the pixel points to be processed in the current coding block in terms of brightness, color, etc., and thus the template region is more suitable for processing the pixel points to be processed in the current coding block, that is, the template region can achieve better results when used to process the pixel points to be processed in the current coding block. For example, the template area is a filter template. Since the pixel correlation between the sample points in the filter template and the original pixel points corresponding to the pixel points to be processed in the current coding block is high, training can obtain more accurate filtering coefficients for each sample point in the filter template, so that filtering operations based on high-accuracy filter coefficients can obtain better filtering effects.

[0106] The following is combined with Figure 8 Introduce the relevant content of pixel relevance strategy; such as Figure 8As shown, assuming that the template area that the encoding end can originally use is template area 1, the embodiment of the present application designs a new template area 2. Then, at the encoding end, pixel correlation calculation is performed on the sample points in template area 1 and the original pixel points corresponding to the pixel points to be processed in the current coding block (at this time, the coding block is the coding area in the original image in the encoding end), and a correlation result 1 corresponding to template area 1 is obtained; the correlation result 1 is used to indicate the pixel correlation between the sample points in template area 1 and the original pixel points corresponding to the pixel points to be processed in the current coding block. Similarly, at the encoding end, pixel correlation calculation can be performed on the sample points in template area 2 and the original pixel points corresponding to the pixel points to be processed in the current coding block (at this time, the coding block is the coding area in the original image in the encoding end), and a correlation result 2 corresponding to template area 2 is obtained; the correlation result 2 is used to indicate the degree of pixel correlation between the sample points in template area 2 and the original pixel points corresponding to the pixel points to be processed in the current coding block (the higher the pixel correlation degree, the higher the pixel correlation, that is, the closer the sample points in the template area and the original pixel points corresponding to the pixel points to be processed in the current coding block are).

[0107] Furthermore, a correlation threshold is obtained, which is the correlation result corresponding to the traditional template area (such as the correlation result 1 corresponding to the template area 1). In this way, the correlation result 2 corresponding to the newly designed template area 2 of the embodiment of the present application is compared with the correlation threshold. If the correlation result 2 is greater than the correlation threshold, it is determined that the template area 2 has a better processing effect when applied to video encoding and decoding, that is, the newly designed template area 2 of the embodiment of the present application has a better processing effect than the traditional template area (such as the template area 1).

[0108] It is worth noting that ① the above-mentioned pixel correlation calculation process can be roughly described as: respectively calculating the correlation between all the sample points (or other sample points except the middle sample point) in the template area and the pixel points to be processed in the current coding block, and then performing a merging operation (such as averaging, weighted square, etc.) on the correlation results corresponding to each sample point to obtain the pixel correlation result corresponding to the template area. ② The above is to introduce the advantages of the template area provided by the embodiment of the present application on the coding side (that is, the pixel correlation between the sample points in the template area and the original pixel points corresponding to the pixel points to be processed in the current coding block of the coding end is higher than the pixel correlation between the sample points in the traditional template area and the original pixel points corresponding to the pixel points to be processed in the current coding block of the coding end) by giving the process of pixel correlation calculation on the coding end. Accordingly, the template area used in the decoding processing at the decoding end is consistent with the template area used in the encoding processing at the coding side, then the decoding side can determine that the pixel correlation between the sample points in the template area and the target pixel points corresponding to the pixel points to be processed in the current coding block (at this time, the current coding block is the coding block to be decoded by the decoding end) is greater than the correlation threshold.

[0109] (2) The template length of the template area in any direction along the middle sample point of the template area is greater than the length threshold. Here, any direction along the middle sample point of the template area may include at least one of the following: horizontal direction, vertical direction and diagonal direction. The longer the template length of the template area in any direction along the middle sample point of the template area, the more sample points are contained in the any direction, and the more pixel points can be referenced in the any direction when the template area is used to process the current coding block, which can achieve a higher processing effect. The length threshold can be set based on the template length of the traditional template area. For example, if the template length of the traditional template area in the horizontal direction along the middle sample point is 3 (such as 3 pixel units), the length threshold in the horizontal direction can be taken as 3. For example, the schematic diagram of any direction along the middle sample point of the template area can be seen in Figure 9 ;like Figure 9 As shown, the number of sample points of the filter template is 29, and the range of filter coefficients corresponding to the sample points is C0-C 14 , and C0-C 13 Each filter coefficient corresponds to two sample points, C 14 The coefficient corresponds to a sample point. Figure 9 In the filter template shown, the center sample point is C 14 For the sample point where the filter template is located, the template length in any direction along the center sample point in the filter template can be expressed by the number of samples, such as the template length in the horizontal direction is 9, the template length in the vertical direction is 9, and the template length in the diagonal direction is 3.

[0110] Specifically, after the template area is designed, the encoding end (such as a video encoder) can determine whether the designed template area is better than the traditional template area by comparing the template lengths of the designed template area and the traditional template area in the same direction. For example, taking the directions that need to be compared including the horizontal direction, the length direction and the diagonal direction (in some cases, the length comparison can be performed in only two directions or one direction), assuming that the template lengths of the traditional template area along the horizontal direction, the vertical direction and the diagonal direction of the middle sample point are all 5, an optional length threshold value is: the length threshold in the horizontal direction, the length threshold in the vertical direction and the length threshold in the diagonal direction are all 5. If the template lengths of the newly designed template area in at least two directions of the three directions (i.e., the horizontal direction, the vertical direction and the diagonal direction) are greater than the corresponding length thresholds, then it can be determined that the newly designed template area can refer to a larger range when used for video encoding and decoding, and the newly designed template area is determined to be better.

[0111] It should be noted that the embodiment of the present application is based on the above-mentioned (1) and (2) to jointly determine whether the newly designed template area is sufficiently effective. In this way, not only can the prediction effect or filtering effect be improved based on the advantage of greater pixel correlation, but also the prediction result or filtering result can be more accurately predicted based on a larger reference range (i.e., a larger template length), thereby improving video encoding and decoding efficiency and performance.

[0112] It should also be noted that the samples in the new template area have the characteristic of symmetrical distribution. In this way, when the new template area is used to process the pixels to be processed in the current coding block, the domain pixels of the middle sample point can be fully referenced, thereby improving the video encoding and decoding effect. In a specific implementation, it is assumed that the template area includes N sample points, and N is a positive integer; wherein, the area in the template area along the horizontal direction of the middle sample point includes N1 horizontal sample points, the area in the template area along the vertical direction of the middle sample point includes N2 vertical sample points, and the area in the template area along the diagonal direction of the middle sample point includes N3 diagonal sample points; N1, N2, and N3 are positive integers, and N1, N2, and N3 are all less than N. Then, the N1 horizontal sample points are symmetrically distributed with the vertical direction of the middle sample point as the symmetry axis, the N2 vertical sample points are symmetrically distributed with the horizontal direction of the middle sample point as the symmetry axis; and the N3 diagonal sample points are symmetrically distributed with the diagonal direction of the middle sample point as the symmetry axis.

[0113] Taking the template area as the filter template as an example, the shapes of several filter templates provided by the embodiment of the present application that meet the above-described pixel correlation characteristics, template length characteristics and symmetric distribution characteristics are given. Figure 10a The template shape of the filter template shown can be described as follows: the filter template includes the number of samples N=29, the 29 samples include: the middle sample C 14 , along the middle sample point C 14 The number of horizontal sample points in the horizontal direction is N1=8, along the middle sample point C 14 The number of vertical sample points in the vertical direction is N2=8, along the middle sample point C 14 The number of diagonal sample points N3 on any of the two diagonal lines in the diagonal direction of is 2, and two sample points are distributed adjacent to a sample point on any of the two diagonal lines passing through the middle sample point in the template area (for example, sample points C2 and C5 are distributed adjacent to sample point C6 on the diagonal line). Figure 10b 、 Figure 10c and Figure 10d The template shape of the filter template shown can be described as follows: the filter template includes the number of samples N=29, the 29 samples include: the middle sample C 14 , along the middle sample point C 14 N1 horizontal sample points in the horizontal direction, along the middle sample point C 14The vertical direction of the N2 vertical sample points, along the middle sample point C 14 2*N3 diagonal sample points in the two diagonal directions of the middle sample point; wherein, the N3 sample points distributed on the two diagonal lines of the middle sample point can be expressed as N3=(N-1-N1-N2) / 2.

[0114] Depend on Figure 10a 、 Figure 10b 、 Figure 10c and Figure 10d The four filter templates shown and Figure 3b Compared with the traditional template area shown in the figure, the distribution of the video pixels referenced by these four filter templates is different from Figure 3b The distribution of the video pixels referenced by the traditional template area shown is completely different, and the template lengths of the four filter templates in the horizontal and vertical directions are greater than Figure 3b The traditional template area shown in the figure has a template length in the corresponding direction. Therefore, when using the four filter templates for filtering operations, the reference pixel range is larger, which improves the filtering effect and thus improves the video encoding and decoding efficiency. It should be noted that ① each sample point in the template area corresponds to a coefficient, and the coefficients of the symmetrically distributed sample points in the template area are the same; then when the template area is a filter template, such as Figure 10a 、 Figure 10b 、 Figure 10c and Figure 10d Each sample point in the filter template shown corresponds to a filter coefficient (such as C0-C 14 ), and except C 14 The filter coefficients outside the template are also symmetrically distributed. In addition, since the template area provided by the embodiment of the present application has the advantages of high pixel correlation and longer template length, the filter coefficients corresponding to the samples in the trained template area are more accurate. When filtering operations are performed based on more accurate filter coefficients, better filtering effects can be achieved, thereby improving video encoding and decoding efficiency. ② The above Figure 10a 、 Figure 10b 、 Figure 10c and Figure 10d The filter coefficient range is C0-C 14 ) is used as an example. In practical applications, the number of sample points of the filter template can be more, and the range of the filter coefficient can be larger accordingly. There is no limitation on this. Figure 10a 、 Figure 10b 、 Figure 10c and Figure 10d The distribution of filter coefficients in the filter templates shown are all examples, and the position distribution of filter coefficients in different filter templates is not fixed.

[0115] Furthermore, the above mainly introduces the characteristics of the template area provided by the embodiment of the present application. The following describes the process of determining the template area indicated for the current coding block during the video decoding process in the embodiment of the present application. As described above, the method for determining the template area indicated for the current coding block in the embodiment of the present application can be roughly divided into an explicit indication method and an implicit indication method, wherein:

[0116] (1) Display indication mode.

[0117] Display indication is achieved by transmitting a flag (or other information with a marking function) between the video encoder and the video decoder. Specifically, during compression encoding, the video encoder can add the template area used in the current encoding process to the compressed bitstream via a flag. Correspondingly, upon receiving the compressed bitstream, the video decoder can parse the flag of the current encoding block in the compressed bitstream to obtain the template area of ​​the current encoding block.

[0118] In a specific implementation, when the decoder needs to use a template region to process the current coding block in the compressed code stream (for example, when a filter template needs to be used to filter the current coding block (specifically, the reconstruction information corresponding to the current coding block)), the decoder can obtain the flag bit of the current coding block from the compressed code stream; then, by parsing the flag bit of the current coding block, the template region of the current coding block can be obtained. Therefore, using a flag bit to indicate the display indication method of the template region of the current coding block on the decoder side not only does not require consuming more resources on the decoder side, but also ensures that the codec uses the same template region for the current coding block.

[0119] Furthermore, when a video sequence is transmitted between the encoding end and the decoding end (in the scenario where the video encoding end compresses and transmits the video sequence, the coding block in the compressed code stream belongs to the video sequence), the embodiment of the present application supports the subdivision of the display indication method into sequence-level display indication and non-sequence-level display indication according to the relationship between the flag bit and the coding block in the video sequence. Among them, the sequence-level display indication means that the flag bit can be used to indicate the template area used by all or part of the coding blocks in the entire video sequence, while the non-sequence-level display indication means that the flag bit is only used to indicate the template area used by a single coding block in the video sequence. The specific implementation process of the sequence-level display indication and the non-sequence-level display indication is introduced below, wherein:

[0120] 1) Sequence level display indication.

[0121] In the case where the display indication mode is a sequence-level display indication, the flag bit can be called a sequence-level flag bit; the sequence-level flag bit can be expressed as a flag bit at the SPS (Sequence Parameter Set) level, which can be used to indicate: whether the coding block in the compressed code stream uses the target template area under the target coding mode; wherein the target coding mode can be any of the multiple coding modes in the video coding and decoding technology, and the target template area can be any of the multiple template areas in the video coding and decoding set. In other words, the sequence-level flag bit can directly indicate the usage of the template area of ​​all or part of the coding blocks in the compressed code stream, such as indicating whether a certain template area is used under a certain coding mode. In this way, batch parsing of the template area of ​​the coding block can be achieved at the decoding end, which greatly improves the video decoding efficiency and enhances the decoding performance of the video decoding end.

[0122] In a specific implementation, after the video decoding end obtains the sequence-level flag of the current coding block from the compressed code stream, it can obtain the current coding mode of the current coding block, which is the target coding mode mentioned above. Then, the sequence-level flag of the current coding block is parsed to obtain the parsing result of the sequence-level flag. Then, based on the parsing result of the sequence-level flag and the current coding mode of the current coding block, the template area of ​​the current coding block is determined. Among them, the method for obtaining the current coding mode of the current coding block may include but is not limited to: negotiated setting between the video decoding end and the video encoding end, or parsed from the compressed code stream, etc. The embodiment of the present application does not limit the method for determining the current coding mode of the current coding block.

[0123] As described above, the sequence-level flag can be used to indicate whether the coding block in the compressed code stream uses the target template area. Based on this, the sequence-level flag of the current coding block is parsed at the video decoding end. After obtaining the parsing result of the sequence-level flag, it is specifically determined whether the current coding block uses the target template area based on the parsing result of the sequence-level flag.

[0124] Optionally, the parsing result of the sequence-level flag indicates whether the current coding block in the current coding mode uses the target template area. In this implementation, the method of determining the template area of ​​the current coding block according to the parsing result of the sequence-level flag and the current coding mode of the current coding block includes: according to the indication of the parsing result of the sequence-level flag, determining that the current coding block in the current coding mode does not use the target template area (such as the current coding block in the default AI (All intra) coding mode (i.e., the target coding mode) does not enable the square filter template (i.e., the target template area), or the current coding block in the RA (Rand om Access) coding mode does not enable the square filter template). Or, according to the indication of the parsing result of the sequence-level flag, determining that the current coding block in the current coding mode needs to use the target template area (such as the current coding block in the LD (Low Delay) coding mode enables the square filter template).

[0125] Optionally, the parsing result of the sequence-level flag indicates that the current coding block in the current coding mode uses the target template area. In this implementation, the method of determining the template area of ​​the current coding block based on the parsing result of the sequence-level flag and the current coding mode of the current coding block includes: obtaining the default correspondence between the current coding mode (i.e., the aforementioned target template area) and the template area of ​​the current coding block based on the parsing result of the sequence-level flag, and the default correspondence is used to indicate that the current coding block in the current coding mode is allowed to use the target template area. For example, when the default correspondence indicates that the current coding mode is the LD coding mode, the template area allowed to be used by the current coding block is a square filter template. Then, based on the default correspondence, the target template area that has a correspondence with the current coding mode is determined as the template area of ​​the current coding block.

[0126] For example, the parsing result of the sequence-level flag (i.e., the SPS-level flag) parsed by the video decoding end is used to indicate: the template area used by the current entire video sequence is taken as an example. An example of the parsing result obtained by parsing the sequence-level flag can be seen in Table 1:

[0127] Table 1

[0128]

[0129]

[0130] As shown in Table 1, when the result of parsing the sequence-level flag is 0, it indicates that the entire current video sequence can use a square-shaped template area, and when the result of parsing the sequence-level flag is 1, it indicates that the entire current video sequence does not use a square-shaped template area, but uses another template area. It should be understood that the shapes of the various template areas (such as filter templates) shown in Table 1 are examples and do not limit the embodiments of the present application.

[0131] It should be noted that the specific implementation process of the sequence-level display indication described above is combined with the coding mode of the coding block in the compressed code stream; that is, the parsing result of the sequence-level flag bit needs to be combined with the current coding mode of the current coding block to determine the template area for the current coding block. For example, the coding mode of coding block 1 and coding block 2 in the compressed code stream is coding mode 1, while the coding mode of coding block 3 is coding mode 2; then after parsing the sequence-level flag bits corresponding to coding block 1 and coding block 2 respectively, the two parsing results obtained are that the template areas determined by coding block 1 and coding block 2 under coding mode 1 are the same, while after parsing the sequence-level flag bits corresponding to coding block 3, the parsing result obtained is that the template area determined by coding block 3 under coding mode 2 may be different from the template area of ​​coding block 1 (or coding block 2). In actual applications, the sequence-level flag bit of the current coding block can also be used to determine the template area without combining it with the current coding mode of the current coding block. Instead, the sequence-level flag bit can be used to directly determine the same template area for all coding blocks in the entire video sequence; that is, the sequence-level flag bit is used to indicate that all coding blocks in the compressed code stream use the same template area. In this case, the video decoding end only needs to parse the sequence-level flag once, and then parse it again only when the sequence-level flag is updated, which greatly reduces the workload of parsing the flag and improves the efficiency of video decoding.

[0132] 2) Non-sequence level display indication.

[0133] In the case where the display indication mode is a non-sequential display indication, the flag bit can be called a non-sequential flag bit. Unlike the sequential flag bit, a non-sequential flag bit is only used to indicate the template area used by a coding block in the compressed code stream; in this case, the compressed code stream includes a non-sequential flag bit corresponding to each coding block in the multiple coding blocks in the compressed code stream; the non-sequential flag bit of the current coding block is used to indicate the template area used by the current coding block. Therefore, when the video decoding end obtains the template area of ​​a single coding block, it determines the specific template area used by the single coding block by directly parsing the non-sequential flag bit corresponding to the single coding block (such as the slice-level / CTU-level / block-level coding block).

[0134] For example, the video decoding end directly analyzes the non-sequence level flag of the current coding block, and a schematic situation of determining the template area used by the current coding block can be seen in Table 2:

[0135] Table 2

[0136]

[0137] As shown in Table 2, when the non-sequential level flag of the current coding block is 00, the parsing result obtained by parsing the non-sequential level flag is 0, indicating that the current coding block uses a filter template (i.e., template area) with a square shape. Similarly, when the non-sequential level flag of the current coding block is 01, the parsing result obtained by parsing the non-sequential level flag is 1, indicating that the current coding block uses a filter template (i.e., template area) with a diamond shape. It should be understood that the shapes of the various template areas (such as filter templates) shown in Table 2 are examples and do not limit the embodiments of the present application.

[0138] It should be noted that the prediction or filtering process of video encoding and decoding actually involves predicting or filtering the different components of the current coding block. That is, if the current coding block includes a first component, a second component, and a third component, then the prediction or filtering of the current coding block is the prediction or filtering of the first component, the second component, and the third component. The first component, the second component, and the third component included in the current coding block can be the YUV components mentioned above. If an image only includes the Y component (luminance component) but does not include the U and V components (chrominance components), then the image is a black and white grayscale image. Among them, the first component, the second component and the third component included in the current coding block may specifically include any of the following cases: the first component is a Y component, the second component is a U component, and the third component is a V component; or, the first component is a Y component, the second component is a V component, and the third component is a U component; the first component is a U component, the second component is a V component, and the third component is a Y component; or, the first component is a U component, the second component is a Y component, and the third component is a V component; or, the first component is a V component, the second component is a Y component, and the third component is a U component; or, the first component is a V component, the second component is a U component, and the third component is a Y component.

[0139] Based on this, the embodiments of the present application support the template areas used by the first component, the second component and the third component included in the current coding block, respectively. This can be indicated by parsing the same flag bit (i.e., a non-sequential level flag bit); or different components may correspond to their own flag bits (i.e., a non-sequential level flag bit), so that the template area used by each component is determined by separately parsing the flag bits corresponding to different components. Optionally, when the same flag bit is parsed to indicate the template area used by each component in the YUV component of the current coding block, the non-sequential level flag of the current coding block is used to indicate that the first component, the second component and the third component included in the current coding block all use the same template area, such as specifically indicating that the first component, the second component and the third component included in the current coding block all use a template area with a square shape. Optionally, when different flag bits are parsed to indicate the template area used by each component in the YUV components of the current coding block, the non-sequential level flag bits of the current coding block are: one or more of the first sub-flag bit of the first component, the second sub-flag bit of the second component, and the third sub-flag bit of the third component; wherein the first sub-flag bit of the first component is used to indicate the template area used by the first component; the second sub-flag bit of the second component is used to indicate the template area used by the second component; and the third sub-flag bit of the third component is used to indicate the template area used by the third component.

[0140] (2) Implicit indication method.

[0141] In contrast to the aforementioned explicit indication method, the implicit indication method doesn't achieve uniformity in the video codec's pattern area by adding a flag to the compressed bitstream. Instead, the video encoder and decoder agree in advance to use the same method to determine the template area for the same coding block. This agreed-upon method of determining the template area is called the implicit indication method. This eliminates the need for the video encoder to transmit flags, and the video decoder to decode them. This reduces the amount of compressed bitstream data transmitted, thereby improving video codec performance.

[0142] In the embodiment of the present application, three implicit indication methods are provided, namely, the default usage method (or referred to as default usage, default method, etc.), the derived method and the reference frame method (or referred to as reference frame dependency). Among them, the default usage refers to the video encoding end and the video decoding end setting through negotiation, and the default template area is used for processing for the same coding block. The derived method refers to the video encoding end and the video decoding end setting through negotiation, and the same method is used to determine the coded blocks related to the coding block for the same coding block to determine the template area of ​​the coding block. The reference frame method refers to the video encoding end and the video decoding end setting through negotiation, and the template area of ​​the coding block is determined according to the number of available reference frames of the video frame (or image) to which the coding block belongs for the same coding block. The specific implementation processes of the default usage method, the derived method and the reference frame method are introduced below, wherein:

[0143] 1) Default usage.

[0144] For the video decoding end, when it needs to determine the template area for the current coding block in the compressed code stream, it can obtain the default setting information, which is negotiated and set by the video encoding end and the video decoding end, and is used to indicate the information of the default template area. In this way, the video decoding end can use the default template area as the template area of ​​the current coding block based on the default setting information. For example, assuming that the video encoding end and the video decoding end negotiate and set (such as offline negotiation or online negotiation) to use a square template area to process the current coding block during video encoding, and also use a square template area to introduce the current coding block during video decoding; then when the video encoding end encodes the current coding block, it can use the square template area by default to encode the current coding block. Similarly, when the video decoding end receives the compressed code stream and needs to decode the current coding block in the compressed code stream, it can use the square template area by default to decode the current coding block.

[0145] As described above, the current coding block includes a first component, a second component, and a third component. Therefore, the embodiment of the present application supports the video codec to default to using only template areas of the same shape for all components (Y, U, and V) of the current coding block, that is, the three components of the current coding block use template areas of the same shape for processing. Under this implementation, the default setting information of the video decoding end can be specifically used to indicate that different components in the current coding block correspond to the same default template area, so that the video decoding end uses the default template area as the template area of ​​the first component, the second component, and the third component in the current coding block, respectively. Furthermore, when decoding the current coding block, the default template area is specifically used to process the first component, the second component, and the third component in the current coding block, respectively.

[0146] The embodiment of the present application also supports the video codec to default all components (Y, U and V) of the current coding block to use their own default template areas respectively; that is, there is no restriction on the three components of the current coding block to use template areas of the same shape. Under this implementation, the default setting information of the video decoding end can be specifically used to indicate that different components of the current coding block correspond to their own default template areas, so that the video decoder can set a default first template area for the first component in the current coding block, a default second template area for the second component, and a default third template area for the third component based on the default setting information; wherein the first template area, the second template area and the third template area are not the same at the same time, that is, at least two template areas among the first template area, the second template area and the third template area are different. For example, assuming that the default setting information indicates that the default template area of ​​the first component in the current coding block is template area A, the default template area of ​​the second component is template area B, and the default template area of ​​the third component is template area C, then when decoding the current coding block, template area A is used to decode the first component of the current coding block, template area B is used to decode the second component of the current coding block, and template area C is used to decode the third component of the current coding block; wherein, the shapes of at least two template areas among template area A, template area B and template area C are different.

[0147] 2) Derivation method.

[0148] Derivation can be simply understood as the process of implicitly indicating the template area used by the current coded block to be decoded based on the pixel distribution of the decoded coded block in the compressed bitstream (such as the reference block / CTU-level / slice-level coded block). In other words, the video decoder can use the template area used by the decoded coded block as the template area for the current coded block.

[0149] In a specific implementation, the video decoder can obtain the coded block from the compressed code stream. Among them: ① The coded block can be a coded block in the same image (or video frame) as the current coded block in the spatial domain, and the coded block can have a certain distance or be adjacent to the current coded block in the spatial domain. For example, a coded block adjacent to the current coded block in the image, or a coded block that is separated from the current coded block by one or more pixels in the image and is located in a certain direction of the current coded block (such as the upper left / left / above). Figure 11a As shown, the coded blocks of the current coding block 1101 may be coded blocks 1102 and 1103 that have been decoded in the image and are adjacent to the current coding block; of course Figure 11aThe positions of the coded blocks shown in the image are only examples. Alternatively, ② the coded block may also be a coded block whose position in the image in the time domain is the same as the position of the current coded block in the current image, that is, the coded block may come from another image different from the current image to which the current coded block belongs; the other image and the current image may be adjacent or have a certain distance in the time domain. For example, other images refer to: images that are located before the current image and adjacent to the current image in the image order, or images that are located before the current image and are separated from the current image by one or more frames in the image order; the image order may include the image display order (Picture Order Count, POC) or the image encoding order (Encoding Order Count, EOC), the image display order may be the playback order of each video frame in the video sequence, and the image encoding order may be the order in which the encoding end encodes each video frame in the video sequence. Furthermore, the coded blocks in other images refer to the coded blocks in other images that have the same position as the current coded block in the current image, i.e., coded blocks at the same position; from a visual perspective, when other images and the current image completely overlap, the coded blocks in other images and the current coded block in the current image also completely overlap. Figure 11b As shown, assuming that the current image to which the current coding block 1101 belongs is the i-th video frame in the video sequence, i is a positive integer, and the coded block 1102 there may belong to the coding block at the same position in the j-th video frame in the video sequence; j is a positive integer, and j is less than i.

[0150] Then, the pixel distribution between the coded block and the current coded block is calculated to obtain a pixel distribution result. Specifically, the difference between the content distribution of the current coded block and the content distribution of the coded block is calculated using indicators such as pixel correlation and / or gradient distribution. That is, the pixel distribution result can be used to indicate the difference in content distribution between the current coded block and the coded block. Pixel correlation refers to the similarity or association between adjacent pixels in an image. Considering the continuity and smoothness of objects in an image, the pixel values ​​of adjacent pixels usually have a certain degree of similarity. Therefore, if the content distribution of the coded block (i.e., the pixels in the coded block) and the content distribution of the current coded block have a high pixel similarity (or a low difference), then it indicates that the template area used by the coded block can be used as the template area of ​​the current coded block. Gradient distribution refers to the degree and direction of change in pixel values ​​in an image. The gradient here is a measure of the change in pixel values ​​in an image, which reflects the presence of edges, textures, and other high-frequency components in the image. Therefore, analyzing the gradient distribution of the coded block and the gradient distribution of the current coded block can obtain the difference in pixel value changes between the current coded block and the coded block. If the difference is large, it indicates that the two are relatively similar, and the current coded block can use the same template area as the coded block.

[0151] Finally, the video decoder determines the template area for the current coding block based on the pixel distribution result. In detail, if the pixel distribution result indicates that the difference in content distribution between the current coding block and the coded block is less than or equal to the difference threshold, the template area used by the coded block is used as the template area of ​​the current coding block. On the contrary, if the pixel distribution result indicates that the difference in content distribution between the current coding block and the coded block is greater than the difference threshold, a reference method is used to determine the template area for the current coding block; the reference method here can be any other method except the derived method provided in the embodiment of the present application, such as the reference method can include at least one of the following: the default method (see the aforementioned description of the default usage method under the implicit indication method), the flag method (see the aforementioned description of the display indication method) and the reference frame method (see the following 3)).

[0152] 3) Reference frame method.

[0153] A reference frame is a frame used for prediction and reference of the current frame during the video encoding and decoding process. Specifically, in video encoding and decoding, a series of video frames included in a video sequence can be divided into multiple types, including I-frames (Intra-frame, key frame), P-frames (Predictive frame, forward prediction frame) and B-frames (Bo-directional frame, bidirectional prediction frame). When predicting different types of frames in a video sequence, the type and number of reference frames used are different. In detail: an I-frame is an independent frame, and other frames in the video sequence cannot be used for prediction or reference when predicting this frame; P-frames and B-frames are non-independent frames, and P-frames are encoded based on predictions of forward frames in the video sequence (such as forward I-frames or P-frames), and B-frames are encoded based on predictions of previous and next frames in the video sequence (such as forward and / or backward I-frames or B-frames).

[0154] Based on this, the embodiment of the present application supports the video codec to select different model areas for the current coding block based on the number of frames of the available reference frames of the current image (or called the current video frame, current frame) to which the current coding block belongs. Specifically, when it is necessary to determine the template area for the current coding block, the number of frames of the reference frame corresponding to the current coding block can be obtained first, specifically, the current image to which the current coding block belongs and the number of frames of the reference frame corresponding to the current frame are obtained. Then, a template area is determined for the current coding block based on the number of frames; wherein, depending on the number of frames, the template area determined for the current coding block may be different or the same. For example, examples of template areas corresponding to different frame types can be found in Table 3:

[0155] Frame type of the current frame The number of reference frames Template area used I frame 0 Template area 1 Unidirectional prediction frame 1 Template Area 2 Bidirectionally predicted frames 2 Template Area 3

[0156] As shown in Table 3, if the frame type of the current frame to which the current coding block belongs in the video sequence is an I frame, and the frame number of the reference frame of the I frame is 0, then the template area 1 is determined to be the template area of ​​the current coding block. Similarly, if the frame type of the current frame is a unidirectional prediction frame (such as a P frame, a B frame with only a unidirectional reference), and the frame number of the reference frame of the current frame is 1, then the template area 2 is determined to be the template area of ​​the current coding block. Similarly, if the frame type of the current frame is a bidirectional prediction frame (i.e., a B frame), and the frame number of the reference frame of the current frame is at least 2, then the template area 3 is determined to be the template area of ​​the current coding block. It should be understood that the frame numbers and the shapes of the corresponding template areas in Table 3 are examples and do not limit the embodiments of the present application, and are specifically explained here.

[0157] In summary, the embodiments of the present application can determine the template area for the current coding block through the above-described explicit indication method and implicit indication method. By providing multiple indication methods, the method of determining the template area during the video encoding and decoding process is enriched to meet the user's various encoding and decoding needs.

[0158] It is worth noting that the above descriptions are all based on the example of determining a template area for the current coding block by explicit indication or implicit indication. However, in actual applications, the number of template areas determined by the video codec for the current coding block can be at least two; that is, the embodiment of the present application supports the use of one or more template areas provided by the embodiment of the present application for the same coding block. In addition, different coding blocks in a video sequence can use different template areas, that is, it supports switching among multiple template areas (such as filter shapes), and the methods for determining the template areas used by different coding blocks can also be different.

[0159] S703: Decode the current coding block based on the template area to obtain a reconstructed image of the current coding block.

[0160] After determining the template area for the current coding block based on the aforementioned step S702, the video decoder can use the template area to decode the current coding block to reconstruct the current coding block and obtain a reconstructed image of the current coding block; then all coding blocks in the compressed code stream are reconstructed, that is, the video sequence can be restored.

[0161] As mentioned above, the image processing scheme provided by the embodiment of the present application can be applied to the filtering stage of video coding and decoding, and can also be applied to the prediction stage (such as intra-frame prediction or inter-frame prediction) based on a template of a certain shape in video coding and decoding. Therefore, when the image processing method is applied to the filtering stage, the template area is a filter template, then after the filter template is determined for the current coding block, the current coding block can be filtered based on the filter template of the current coding block to obtain a reconstructed image of the current coding block. Similarly, when the image processing method is applied to the prediction stage, the template area is a prediction template, then after the prediction template is determined for the current coding block, the current coding block can be predicted based on the prediction template of the current coding block (such as the prediction processing includes at least one of the following: inter-frame prediction processing and intra-frame prediction processing) to obtain a reconstructed image of the current coding block.

[0162] In summary, on the one hand, the embodiments of the present application provide a variety of effective template areas, and the pixel correlation between the samples in these template areas and the original pixels corresponding to the pixels to be processed in the current coding block in the compressed code stream in the original image is greater than the corresponding pixel correlation in the traditional template area; in this way, when the model area is used to decode the current coding block, the neighboring samples with higher correlation within the template area can be used to process the pixels in the current coding block, which greatly improves the prediction effect or filtering effect of the pixels in the current coding block. Moreover, the template length of these template areas in any direction along the middle samples in the template area is longer than the corresponding template length in the traditional template area; in this way, when the template area is used to decode the current coding block, the pixel points in the current coding block can be processed using the domain samples within a larger reference range within the template area, which significantly improves the prediction accuracy or filtering accuracy of the pixels in the current coding block. On the other hand, the embodiments of the present application provide a display indication method and an implicit indication method to determine the template for the current coding block; by providing a variety of indication methods, the method of determining the template area in the video encoding and decoding process is enriched to meet the user's various encoding and decoding needs.

[0163] See Figure 12 , Figure 12 is a flowchart of another image processing method provided by an exemplary embodiment of the present application; Figure 12 The flowchart shown may be a flowchart of the encoding side, and may be specifically executed by an image processing device (such as a video encoder) held by the encoding end; the method may include but is not limited to steps S1201-S1203:

[0164] S1201: Determine a current coding block to be encoded.

[0165] On the encoder side, the current coding block is the coding area to be encoded in the current image during the encoding process. This coding area can be a coding block at the slice level, CTU level, block level, or image level. When encoding an image (such as a video frame in a video sequence), the image needs to be divided into blocks to improve the image compression rate and reduce storage and transmission costs. Specifically, the image can be divided into blocks according to the aforementioned description of the block division structure to obtain multiple non-overlapping coding blocks; the current coding block refers to the coding block that the encoder currently needs to encode.

[0166] S1202: Determine a template area for the current coding block.

[0167] The embodiments of the present application provide a variety of effective template areas. At the encoding end, these template areas have the following characteristics: ① The pixel correlation between the sample points in the template area and the target pixel points corresponding to the pixel points to be processed in the current coding block (that is, the original pixel points in the original image) is greater than the correlation threshold. The correlation threshold is determined by the encoding end after performing pixel correlation calculations on the new template area and the traditional template area and the coding block respectively; the pixel correlation between the sample points in the new template area and the target pixel points corresponding to the pixel points to be processed in the current coding block is greater than the correlation threshold, indicating that the new template area provided by the embodiment of the present application has a better processing effect, such as a better filtering effect when the template area is a filter template. ② The template length of the template area in any direction along the middle sample point of the template area is greater than the length threshold; wherein, any direction along the middle sample point of the template area can include at least one of the following: horizontal direction, vertical direction and diagonal direction. It should be noted that the longer the template length of the template area in any direction along the middle sample point of the template area, the more sample points are included in any direction. Then, when the template area is used to process the current coding block, the number of pixel points that can be referenced in any direction is greater, thereby achieving a higher processing effect.

[0168] In a specific implementation, if a video encoder needs to use a template region to encode a current coding block in a video sequence, there are multiple candidate template regions to select from (e.g., the shapes of the candidate template regions may include but are not limited to: square, diamond, rectangle, cross, triangle, and combinations thereof). After obtaining the multiple candidate template regions to be selected, the video encoder may use a cost comparison strategy to perform a cost estimation calculation on each of the multiple candidate template regions to obtain a cost estimation result corresponding to each candidate template region; the cost estimation result corresponding to any candidate template region is used to indicate the degree of loss when the corresponding candidate template region is used to encode the current coding block. The video encoder then compares the cost estimation results corresponding to the multiple candidate template regions, aiming to select the candidate template region with the smallest degree of loss indicated by the cost estimation result from the multiple candidate template regions, and select one or more candidate template regions with a degree of loss less than a certain threshold as the template region for the current coding block (e.g., the candidate template region with the smallest degree of loss is used as the template region for the current coding block, that is, the candidate template region with the optimal shape is selected as the template region for the current coding block). It can be seen that since the loss result corresponding to the selected candidate template area is smaller, when the candidate template area is used to encode the current coding block, the distortion degree can be minimized to the greatest extent, thereby improving the coding quality and enhancing the coding performance.

[0169] The aforementioned cost comparison strategies include at least one of the following: a pixel correlation strategy and a distortion cost strategy. The following applies: ① The distortion cost strategy selects the template region with the lowest encoding cost, i.e., the lowest distortion, by calculating the degree of distortion after processing the template region. Distortion cost strategies may include, but are not limited to, rate distortion optimization (RDO) and peak signal-to-noise ratio (PSNR).

[0170] ② The pixel correlation strategy is to select a template area with the lowest coding cost, that is, a higher pixel correlation, by calculating pixel correlation. The higher the pixel correlation between the sample points in the template area and the target pixel points corresponding to the pixel points to be processed in the current coding block, the closer the sample points in the template area are to the target pixel points corresponding to the pixel points to be processed in the current coding block in terms of brightness, color, etc., so that the template area is more suitable for processing the target pixel points corresponding to the pixel points to be processed in the current coding block, that is, the template area can achieve better results when used to process the target pixel points corresponding to the pixel points to be processed in the current coding block. For example, the template area is a filter template. Since the pixel correlation between the sample points in the filter template and the original pixel points corresponding to the pixel points to be filtered in the current coding block is high, training can obtain more accurate filter coefficients for each sample point in the filter template, so that filtering operations based on high-accuracy filter coefficients can achieve better filtering effects. When actually calculating pixel correlation, the pixel correlation between the neighboring pixels in the template area except the current pixel (or intermediate sample, i.e., the pixel that overlaps with the pixel to be processed (i.e., the target pixel) in the current coding block) and the target pixel corresponding to the pixel to be processed in the current coding block (e.g., the original pixel corresponding to the pixel to be filtered in the current coding block) can be calculated. Alternatively, the pixel correlation between all pixels in the template area (including the current pixel) and the target pixel corresponding to the pixel to be processed in the current coding block can be calculated; when calculating pixel correlation, the present embodiment of the application does not limit the sample selected from the template area for pixel correlation calculation.

[0171] It should be noted that the above description uses the example of selecting the optimal template region from multiple candidate template regions for the current coding block using a cost comparison strategy. In the actual selection process, at least two template regions can be selected for the current coding block based on the cost estimation results, depending on the user's coding requirements. This embodiment of the application does not limit the number of template regions selected by the encoder for the current coding block, and the number of template regions selected can also vary for different coding blocks.

[0172] It should also be noted that, in addition to determining the template region for the current coding block by the cost comparison method described above, embodiments of the present application also support the video coding end using other methods to determine the template region for the current coding block. For example, when the current coding block is not the first coding block (such as the first coding block in the first video frame in a video sequence), the template region of the current coding block can be determined using the previous coded block of the current coding block.

[0173] Optionally, the video encoding end adopts a default template area as the template area of ​​the current coding block. Specifically, the video encoding end adopts the default template area set by negotiation with the video decoding end (such as offline negotiation or online negotiation, etc.) as the template area of ​​the current coding block. In more detail, the current coding block includes a first component, a second component, and a third component. Then, the embodiment of the present application supports the video codec to use the same shape of the template area for the three components of the current coding block by default. Under this implementation, the video decoding end can obtain default setting information, which is used to indicate that different components in the current coding block correspond to the same default template area, so that the video encoding end uses the default template area as the template area of ​​the first component, the second component, and the third component in the current coding block. Furthermore, when encoding the current coding block, the default template area is specifically used to process the first component, the second component, and the third component in the current coding block respectively. Alternatively, the embodiment of the present application also supports the video encoding end to default all components (Y, U, and V) of the current coding block to use their respective default template areas. In this case, the default setting information of the video encoding end can be specifically used to indicate that different components of the current coding block correspond to their own default template areas, so that the video encoder can set template areas for the three components in the current coding block based on the default setting information, and the template areas of at least two of the three components are different.

[0174] Optionally, the video encoding end determines the template area for the current coding block by derivation. The derivation method refers to the process of implicitly indicating the template area used by the current coding block to be encoded based on the pixel distribution of the coded block (such as the reference block / CTU-level / slice-level coding block) that has been encoded during the encoding process. In a specific implementation, the video encoder can obtain a coded block, which can be a coding block in the same image (or video frame) as the current coding block in the spatial domain (adjacent or separated by one or more pixels), or the coded block is a coding block with the same position in the image as the current coding block in the current image in the time domain. Then, the pixel distribution between the coded block and the current coding block is calculated to obtain a pixel distribution result. Finally, the video encoder determines the template area for the current coding block based on the pixel distribution result.

[0175] Optionally, the video encoder determines the template area for the current coding block by using a reference frame method. In a specific implementation, the video encoder is supported to select different model areas for the current coding block based on the number of available reference frames of the current image (or current video frame, current frame) to which the current coding block belongs. If the number of available reference frames of the current frame is 0, the current coding block uses a square template area by default; if the number of available reference frames of the current frame is 1, the current editing block uses a diamond template area by default; and so on.

[0176] It should be noted that after the video encoder processes the current coding block in any of the three optional methods mentioned above, the video decoder needs to process the current coding block in the same way. The video encoding end and the video decoding end can negotiate with each other so that the video decoding end can know the template area used by the video encoding end when encoding, so the way the video decoding end determines the template area is called an implicit indication method. In addition to this implicit indication method, the embodiment of the present application also supports transmitting the flag indicating the template area used by the current encoder to the video decoding end through the flag bit transmission method, so as to ensure that the video decoding end can use the same template area as the video encoding end for decoding processing when decoding the corresponding current coding block, thereby ensuring the consistency of the encoding and decoding ends. Among them, the above-mentioned flag bit transmission method is the display indication method provided by the embodiment of the present application; according to the coding block range that can be covered by the flag bit indication, the display indication method can be subdivided into sequence-level display indication and non-sequence-level display indication.

[0177] Among them: ① The flag bit under the sequence-level display indication can be called the sequence-level flag bit; the sequence-level flag bit can be expressed as a flag bit at the SPS (Sequence Paramater Set) level, which can be used to indicate whether the coding block uses the target template area in the target coding mode. Considering that the sequence-level flag bit has an indicative effect on the entire video sequence, the sequence-level flag bit can be transmitted once when the video encoding end starts transmitting the compressed code stream, and can be updated and transmitted when the meaning of the subsequent sequence-level flag bit indication changes, which is beneficial to saving transmission overhead and improving video coding efficiency and performance. ② The flag bit under the non-sequence-level display indication can be called a non-sequence-level flag bit. On the video encoding end, the video encoder needs to compress the template area used by each coding block into the compressed code stream in the form of a non-sequence-level flag bit, so that the video decoder can accurately parse the template area used by the coding block according to the corresponding non-sequence-level flag bit.

[0178] S1203: Encode the current coding block based on the template area to generate a compressed code stream.

[0179] After determining the template region for the current coding block based on the aforementioned step S1202, the video encoder can use the template region to perform encoding processing on the current coding block to implement an encoded video sequence and obtain a compressed bitstream. Optionally, when the image processing method is applied to the filtering stage of video coding, the template region is a filter template, and the encoding processing here can specifically be a filtering operation. Optionally, when the image processing method is applied to the prediction stage of video coding, the template region is a prediction template, and the encoding processing here is a prediction processing.

[0180] It should be noted that some contents of the encoding process of the video encoder shown in steps S1201-S1203 (such as the display indication mode, implicit indication mode, prediction process and filtering operation, etc.) are similar to the above Figure 7 The relevant contents of the decoding process of the video decoder shown are the same. Therefore, only a brief introduction is given in steps 1201-S1203. The detailed implementation process can be found in the aforementioned Figure 7 Related description of the specific implementation process shown in steps S701-S703 in the illustrated embodiment.

[0181] In summary, on the one hand, the embodiments of the present application provide a variety of effective template areas, and the pixel correlation between the sample points in these template areas and the original pixel points corresponding to the pixel points to be processed in the current coding block in the compressed code stream in the original image is greater than the corresponding pixel correlation in the traditional template area; and the template length of the template area in any direction along the middle sample point in the template area is longer than the corresponding template length in the traditional template area, which significantly improves the prediction accuracy or filtering accuracy of the pixel points in the current coding block. On the other hand, the embodiments of the present application provide a display indication method and an implicit indication method to determine the template for the current coding block; by providing a variety of indication methods, the method of determining the template area in the video encoding and decoding process is enriched to meet the user's various encoding and decoding needs.

[0182] The method of the embodiment of the present application is described in detail above. In order to facilitate the above-mentioned scheme of the embodiment of the present application to be better implemented, accordingly, the device of the embodiment of the present application is provided below. In the embodiment of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuit or memory) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the module or unit function.

[0183] Figure 13 A schematic diagram of the structure of an image processing device provided by an exemplary embodiment of the present application is shown; the image processing device can be used to perform Figure 7 Some or all of the steps in the method embodiment shown.

[0184] See Figure 13 , the device includes the following units:

[0185] A determining unit 1301 is configured to determine a current coding block in a compressed code stream;

[0186] A processing unit 1302 is configured to determine a template region for a current coding block; a pixel correlation between samples in the template region and a target pixel corresponding to a pixel to be processed in the current coding block is greater than a correlation threshold, where the target pixel is a pixel corresponding to the pixel to be processed in the original image; and a template length of the template region in any direction along a middle sample of the template region is greater than a length threshold, where the any direction includes at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction.

[0187] The processing unit 1302 is further configured to perform decoding processing on the current coding block based on the template region to obtain a reconstructed image of the current coding block.

[0188] In one implementation, any direction includes at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction.

[0189] In one implementation, the template area includes N sample points, where N is a positive integer; an area in the template area along a horizontal direction of a middle sample point includes N1 horizontal sample points, an area in the template area along a vertical direction of the middle sample point includes N2 vertical sample points, and an area in the template area along a diagonal direction of the middle sample point includes N3 diagonal sample points; N1, N2, and N3 are positive integers, and N1, N2, and N3 are all less than N;

[0190] Among them, N1 horizontal sample points are symmetrically distributed with the vertical direction of the middle sample point as the symmetry axis, N2 vertical sample points are symmetrically distributed with the horizontal direction of the middle sample point as the symmetry axis; N3 diagonal sample points are symmetrically distributed with the diagonal direction of the middle sample point as the symmetry axis;

[0191] Each sample point in the template area corresponds to a coefficient, and the coefficients of the symmetrically distributed sample points are the same.

[0192] In one implementation, N=29, N1=8, N2=8, N3=2, and the template shape of the template area includes:

[0193] There are two sample points distributed adjacent to a sample point on any diagonal line passing through the middle sample point in the template area.

[0194] In one implementation, N=29, and the template shape of the template area includes:

[0195] There are N3 sample points distributed on the two diagonal lines of the middle sample point, where N3 = (N-1-N1-N2) / 2.

[0196] In one implementation, the processing unit 1302, when determining the template region for the current coding block, is specifically configured to:

[0197] Get the flag of the current encoding block from the compressed code stream;

[0198] Parse the flag bit of the current coding block to obtain the template area of ​​the current coding block;

[0199] The current coding block includes at least one of the following: an image, a data block, a coding tree unit, and a macroblock.

[0200] In one implementation, the flag bit is a sequence-level flag bit, and the sequence-level flag bit is used to indicate: whether a coding block in a compressed code stream uses a target template area in a target coding mode; the target coding mode is any one of multiple coding modes, and the target template area is any one of multiple template areas;

[0201] The processing unit 1302 is configured to parse the flag bit of the current coding block to obtain the template area of ​​the current coding block, specifically to:

[0202] Get the current encoding mode of the current encoding block; the current encoding mode is the target encoding mode;

[0203] Parse the sequence-level flag of the current coding block to obtain the parsing result;

[0204] According to the parsing result and the current encoding mode, the template area of ​​the current encoding block is determined.

[0205] In one implementation, the parsing result indicates whether the current coding block in the current coding mode uses the target template area. The processing unit 1302 is configured to determine the template area of ​​the current coding block based on the parsing result and the current coding mode, specifically to:

[0206] According to the indication of the parsing result, it is determined that the current coding block in the current coding mode does not use the target template area; or,

[0207] According to an indication of the parsing result, it is determined that the current coding block in the current coding mode uses a target template area.

[0208] In one implementation, the parsing result indicates that the current coding block in the current coding mode uses the target template area. The processing unit 1302 is configured to determine the template area of ​​the current coding block based on the parsing result and the previous coding mode, specifically to:

[0209] Obtain the default correspondence between the current coding mode and the template area according to the parsing result; the default correspondence is used to indicate that the current coding block in the current coding mode is allowed to use the target template area; the current coding mode is the target coding mode;

[0210] According to the default corresponding relationship, a target template region having a corresponding relationship with the current coding mode is determined as the template region of the current coding block.

[0211] In one implementation, the flag bit is a sequence-level flag bit, and the sequence-level flag bit is used to indicate that: all coding blocks in the compressed code stream use the same template area;

[0212] Among them, all coding blocks in the compressed code stream belong to the video sequence.

[0213] In one implementation, the flag is a non-sequential flag, and the compressed code stream includes a non-sequential flag corresponding to each coding block in the compressed code stream; the non-sequential flag of the current coding block is used to indicate the template area used by the current coding block.

[0214] In one implementation, the current coding block includes a first component, a second component, and a third component; the non-sequential level flag of the current coding block is used to indicate that the first component, the second component, and the third component all use the same template region.

[0215] In one implementation, the current coding block includes a first component, a second component, and a third component; the non-sequence-level flag of the current coding block is: a first sub-flag of the first component, a second sub-flag of the second component, and a third sub-flag of the third component;

[0216] Among them, the first sub-flag of the first component is used to indicate the template area used by the first component; the second sub-flag of the second component is used to indicate the template area used by the second component; and the third sub-flag of the third component is used to indicate the template area used by the third component.

[0217] In one implementation, the processing unit 1302, when determining the template region for the current coding block, is specifically configured to:

[0218] Get default setting information, the default setting information indicates the default template area;

[0219] Based on the default setting information, the default template area is used as the template area of ​​the current coding block; the default setting information is negotiated and set by the encoding end and the decoding end.

[0220] In one implementation, the current coding block includes a first component, a second component, and a third component; the default setting information indicates that different components in the current coding block correspond to the same default template region;

[0221] The processing unit 1302 is configured to use the default template area as the template area of ​​the current coding block based on the default setting information, specifically to:

[0222] The default template area is used as the template area of ​​the first component, the second component and the third component in the current coding block respectively.

[0223] In one implementation, the current coding block includes a first component, a second component, and a third component; the default setting information indicates that different components of the current coding block correspond to respective default template regions;

[0224] The processing unit 1302 is configured to use the default template area as the template area of ​​the current coding block based on the default setting information, specifically to:

[0225] Based on the default setting information, a first template area is set for the first component in the current coding block, a second template area is set for the second component, and a third template area is set for the third component; at least two template areas among the first template area, the second template area and the third template area are different.

[0226] In one implementation, the processing unit 1302, when determining the template region for the current coding block, is specifically configured to:

[0227] Obtain a coded block from the compressed code stream; the coded block is a coded block that is in the same image as the current coded block in the spatial domain, or a coded block that has the same position in the image as the current coded block in the temporal domain;

[0228] Calculate the pixel distribution between the coded block and the current coded block to obtain a pixel distribution result;

[0229] A template area is determined for the current coding block according to the pixel distribution result.

[0230] In one implementation, the processing unit 1302 is configured to determine the template area for the current coding block according to the pixel distribution result, specifically to:

[0231] If the pixel distribution result indicates that the difference in content distribution between the current coding block and the coded block is less than or equal to the difference threshold, the template area used by the coded block is used as the template area of ​​the current coding block.

[0232] In one implementation, the processing unit 1302 is configured to determine the template area for the current coding block according to the pixel distribution result, specifically to:

[0233] If the pixel distribution result indicates that the difference in content distribution between the current coding block and the coded block is greater than the difference threshold, a template region is determined for the current coding block using a reference method;

[0234] The reference mode includes at least one of the following: a default mode, a flag mode, and a reference frame mode.

[0235] In one implementation, the processing unit 1302, when determining the template region for the current coding block, is specifically configured to:

[0236] Get the number of reference frames corresponding to the current coding block;

[0237] A template region is determined for the current coding block based on the number of frames.

[0238] In one implementation, the template region is a filter template; and the processing unit 1302 is configured to decode the current coding block based on the template region to obtain a reconstructed image of the current block to be coded, and is configured to:

[0239] The current coding block is filtered based on the filter template of the current coding block to obtain a reconstructed image of the current coding block.

[0240] In one implementation, the template region is a prediction template; and the processing unit 1302 is configured to decode the current coding block based on the template region to obtain a reconstructed image of the current block to be coded, specifically to:

[0241] Performing prediction processing on the current coding block based on the prediction template of the current coding block to obtain a reconstructed image of the current coding block;

[0242] The prediction process includes at least one of the following: inter-frame prediction process and intra-frame prediction process.

[0243] According to one embodiment of the present application, Figure 13 The various units in the image processing device shown can be individually or fully combined into one or several other units to form a whole, or one (or some) of the units can be further divided into multiple functionally smaller units to form a whole, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the image processing device may also include other units. In actual applications, these functions can also be implemented with the assistance of other units and can be implemented by the collaboration of multiple units. According to another embodiment of the present application, the image processing device can be executed by running on a general-purpose computing device such as a computer that includes processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM). Figure 7 A computer program (including program code) for each step of the corresponding method shown in FIG. Figure 13 The image processing apparatus shown in and the image processing method of the embodiment of the present application are implemented. The computer program can be recorded on a computer-readable recording medium, for example, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.

[0244] In an embodiment of the present application, a valid template region is provided. The validity of the template region can be reflected in: firstly, the pixel correlation between the samples in the template region and the original pixels (or target pixels) corresponding to the pixels to be processed (e.g., pixels to be filtered or pixels to be predicted) in the current coding block in the compressed code stream. Secondly, the template length of the template region in any direction along the intermediate samples in the template region is longer than the corresponding template length in a traditional template region.

[0245] Figure 14 A schematic diagram of the structure of another image processing device provided by an exemplary embodiment of the present application is shown; the image processing device can be used to perform Figure 12 Some or all of the steps in the method embodiment shown. Figure 14 , the device includes the following units:

[0246] A determining unit 1401 is configured to determine a current coding block to be encoded;

[0247] A processing unit 1402 is configured to determine a template region for a current coding block; a pixel correlation between samples in the template region and a target pixel corresponding to a pixel to be processed in the current coding block is greater than a correlation threshold, where the target pixel is a pixel corresponding to the pixel to be processed in the original image; and a template length of the template region in any direction along a middle sample of the template region is greater than a length threshold, where the any direction includes at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction.

[0248] The processing unit 1402 is further configured to perform encoding processing on the current coding block based on the template region to generate a compressed code stream.

[0249] In one implementation, the processing unit 1402, when determining the template region for the current coding block, is specifically configured to:

[0250] Obtain multiple candidate template regions;

[0251] A cost comparison strategy is used to perform a cost estimation calculation on each candidate template region among a plurality of candidate template regions, and obtain a cost estimation result corresponding to each candidate template region; the cost estimation result corresponding to the candidate template region is used to indicate the degree of loss when the current coding block is encoded using the corresponding candidate template region; wherein the cost comparison strategy includes at least one of the following: a pixel correlation strategy and a distortion cost strategy;

[0252] A candidate template region with the smallest degree of loss indicated by a cost estimation result is selected from the multiple candidate template regions as the template region of the current coding block.

[0253] According to one embodiment of the present application, Figure 14The various units in the image processing device shown can be individually or fully combined into one or several other units to form a whole, or one (or some) of the units can be further divided into multiple functionally smaller units to form a whole, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the image processing device may also include other units. In actual applications, these functions can also be implemented with the assistance of other units and can be implemented by the collaboration of multiple units. According to another embodiment of the present application, the image processing device can be executed by running on a general-purpose computing device such as a computer that includes processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM). Figure 12 A computer program (including program code) for each step of the corresponding method shown in FIG. Figure 14 The image processing apparatus shown in and the image processing method of the embodiment of the present application are implemented. The computer program can be recorded on a computer-readable recording medium, for example, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.

[0254] In an embodiment of the present application, a valid template region is provided. The validity of the template region can be reflected in: firstly, the pixel correlation between the samples in the template region and the original pixels (or target pixels) corresponding to the pixels to be processed (e.g., pixels to be filtered or pixels to be predicted) in the current coding block in the compressed code stream. Secondly, the template length of the template region in any direction along the intermediate samples in the template region is longer than the corresponding template length in a traditional template region.

[0255] Figure 15 FIG2 shows a schematic diagram of the structure of an image processing device provided by an exemplary embodiment of the present application. Figure 15, the image processing device includes a processor 1501, a communication interface 1502 and a computer-readable storage medium 1503. The processor 1501, the communication interface 1502 and the computer-readable storage medium 1503 can be connected via a bus or other means. The communication interface 1502 is used to receive and send data. The computer-readable storage medium 1503 can be stored in the memory of the image processing device. The computer-readable storage medium 1503 is used to store computer programs. The computer programs include program instructions. The processor 1501 is used to execute the program instructions stored in the computer-readable storage medium 1503. The processor 1501 (or CPU (Central Processing Unit)) is the computing core and control core of the image processing device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to realize the corresponding method flow or corresponding function.

[0256] The embodiment of the present application also provides a computer-readable storage medium (Memory), which is a memory device in the image processing device for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the image processing device and, of course, the extended storage medium supported by the image processing device. The computer-readable storage medium provides a storage space that stores the processing system of the image processing device. In addition, one or more instructions suitable for being loaded and executed by the processor 1501 are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer-readable storage medium located away from the aforementioned processor.

[0257] In one embodiment, the computer-readable storage medium stores one or more instructions; the processor 1501 loads and executes the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above-mentioned image processing method embodiment; in a specific implementation, the processor 1501 loads the one or more instructions in the computer-readable storage medium and executes the following steps:

[0258] Determining a current coded block in a compressed code stream;

[0259] Determine a template region for the current coding block; a pixel correlation between a sample point in the template region and a target pixel corresponding to a pixel to be processed in the current coding block is greater than a correlation threshold, where the target pixel is a pixel corresponding to the pixel to be processed in the original image; and a template length of the template region in any direction along a middle sample point of the template region is greater than a length threshold, where the any direction includes at least one of the following: horizontal, vertical, and diagonal directions.

[0260] The current coding block is decoded based on the template area to obtain a reconstructed image of the current coding block.

[0261] In one implementation, any direction includes at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction.

[0262] In one implementation, the template area includes N sample points, where N is a positive integer; an area in the template area along a horizontal direction of a middle sample point includes N1 horizontal sample points, an area in the template area along a vertical direction of the middle sample point includes N2 vertical sample points, and an area in the template area along a diagonal direction of the middle sample point includes N3 diagonal sample points; N1, N2, and N3 are positive integers, and N1, N2, and N3 are all less than N;

[0263] Among them, N1 horizontal sample points are symmetrically distributed with the vertical direction of the middle sample point as the symmetry axis, N2 vertical sample points are symmetrically distributed with the horizontal direction of the middle sample point as the symmetry axis; N3 diagonal sample points are symmetrically distributed with the diagonal direction of the middle sample point as the symmetry axis;

[0264] Each sample point in the template area corresponds to a coefficient, and the coefficients of the symmetrically distributed sample points are the same.

[0265] In one implementation, N=29, N1=8, N2=8, N3=2, and the template shape of the template area includes:

[0266] There are two sample points distributed adjacent to a sample point on any diagonal line passing through the middle sample point in the template area.

[0267] In one implementation, N=29, and the template shape of the template area includes:

[0268] There are N3 sample points distributed on the two diagonal lines of the middle sample point, where N3 = (N-1-N1-N2) / 2.

[0269] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when executing to determine a template region for a current coding block, specifically perform the following steps:

[0270] Get the flag of the current encoding block from the compressed code stream;

[0271] Parse the flag bit of the current coding block to obtain the template area of ​​the current coding block;

[0272] The current coding block includes at least one of the following: an image, a data block, a coding tree unit, and a macroblock.

[0273] In one implementation, the flag bit is a sequence-level flag bit, and the sequence-level flag bit is used to indicate: whether a coding block in a compressed code stream uses a target template area in a target coding mode; the target coding mode is any one of multiple coding modes, and the target template area is any one of multiple template areas;

[0274] One or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when parsing the flag bit of the current coding block to obtain the template area of ​​the current coding block, specifically perform the following steps:

[0275] Get the current encoding mode of the current encoding block; the current encoding mode is the target encoding mode;

[0276] Parse the sequence-level flag of the current coding block to obtain the parsing result;

[0277] According to the parsing result and the current encoding mode, the template area of ​​the current encoding block is determined.

[0278] In one implementation, the parsing result indicates whether the current coding block in the current coding mode uses the target template area. The one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when executing, determine the template area of ​​the current coding block based on the parsing result and the current coding mode, specifically perform the following steps:

[0279] According to the indication of the parsing result, it is determined that the current coding block in the current coding mode does not use the target template area; or,

[0280] According to an indication of the parsing result, it is determined that the current coding block in the current coding mode uses a target template area.

[0281] In one implementation, the parsing result indicates that the current coding block in the current coding mode uses the target template region. The one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when executing the instructions, determine the template region of the current coding block based on the parsing result and the previous coding mode, specifically perform the following steps:

[0282] Obtain the default correspondence between the current coding mode and the template area according to the parsing result; the default correspondence is used to indicate that the current coding block in the current coding mode is allowed to use the target template area; the current coding mode is the target coding mode;

[0283] According to the default corresponding relationship, a target template region having a corresponding relationship with the current coding mode is determined as the template region of the current coding block.

[0284] In one implementation, the flag bit is a sequence-level flag bit, and the sequence-level flag bit is used to indicate that: all coding blocks in the compressed code stream use the same template area;

[0285] Among them, all coding blocks in the compressed code stream belong to the video sequence.

[0286] In one implementation, the flag is a non-sequential flag, and the compressed code stream includes a non-sequential flag corresponding to each coding block in the compressed code stream; the non-sequential flag of the current coding block is used to indicate the template area used by the current coding block.

[0287] In one implementation, the current coding block includes a first component, a second component, and a third component; the non-sequential level flag of the current coding block is used to indicate that the first component, the second component, and the third component all use the same template region.

[0288] In one implementation, the current coding block includes a first component, a second component, and a third component; the non-sequence-level flag of the current coding block is: a first sub-flag of the first component, a second sub-flag of the second component, and a third sub-flag of the third component;

[0289] Among them, the first sub-flag of the first component is used to indicate the template area used by the first component; the second sub-flag of the second component is used to indicate the template area used by the second component; and the third sub-flag of the third component is used to indicate the template area used by the third component.

[0290] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when executing to determine a template region for a current coding block, specifically perform the following steps:

[0291] Get default setting information, the default setting information indicates the default template area;

[0292] Based on the default setting information, the default template area is used as the template area of ​​the current coding block; the default setting information is negotiated and set by the encoding end and the decoding end.

[0293] In one implementation, the current coding block includes a first component, a second component, and a third component; the default setting information indicates that different components in the current coding block correspond to the same default template region;

[0294] The one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when executing the process of using the default template region as the template region of the current coding block based on the default setting information, specifically perform the following steps:

[0295] The default template area is used as the template area of ​​the first component, the second component and the third component in the current coding block respectively.

[0296] In one implementation, the current coding block includes a first component, a second component, and a third component; the default setting information indicates that different components of the current coding block correspond to respective default template regions;

[0297] The one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when executing the process of using the default template region as the template region of the current coding block based on the default setting information, specifically perform the following steps:

[0298] Based on the default setting information, a first template area is set for the first component in the current coding block, a second template area is set for the second component, and a third template area is set for the third component; at least two template areas among the first template area, the second template area and the third template area are different.

[0299] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when executing to determine a template region for a current coding block, specifically perform the following steps:

[0300] Obtain a coded block from the compressed code stream; the coded block is a coded block that is in the same image as the current coded block in the spatial domain, or a coded block that has the same position in the image as the current coded block in the temporal domain;

[0301] Calculate the pixel distribution between the coded block and the current coded block to obtain a pixel distribution result;

[0302] A template area is determined for the current coding block according to the pixel distribution result.

[0303] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when determining a template area for a current coding block based on a pixel distribution result, specifically perform the following steps:

[0304] If the pixel distribution result indicates that the difference in content distribution between the current coding block and the coded block is less than or equal to the difference threshold, the template area used by the coded block is used as the template area of ​​the current coding block.

[0305] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when determining a template area for a current coding block based on a pixel distribution result, specifically perform the following steps:

[0306] If the pixel distribution result indicates that the difference in content distribution between the current coding block and the coded block is greater than the difference threshold, a template region is determined for the current coding block using a reference method;

[0307] The reference mode includes at least one of the following: a default mode, a flag mode, and a reference frame mode.

[0308] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when executing to determine a template region for a current coding block, specifically perform the following steps:

[0309] Get the number of reference frames corresponding to the current coding block;

[0310] A template region is determined for the current coding block based on the number of frames.

[0311] In one implementation, the template region is a filter template; one or more instructions in a computer-readable storage medium are loaded by the processor 1501 and, when performing decoding processing on a current coding block based on the template region to obtain a reconstructed image of the current block to be coded, perform the following steps:

[0312] The current coding block is filtered based on the filter template of the current coding block to obtain a reconstructed image of the current coding block.

[0313] In one implementation, the template region is a prediction template; one or more instructions in a computer-readable storage medium are loaded by the processor 1501 and, when performing decoding processing on a current coding block based on the template region to obtain a reconstructed image of the current block to be coded, specifically perform the following steps:

[0314] Performing prediction processing on the current coding block based on the prediction template of the current coding block to obtain a reconstructed image of the current coding block;

[0315] The prediction process includes at least one of the following: inter-frame prediction process and intra-frame prediction process.

[0316] In another embodiment, the computer-readable storage medium stores one or more instructions; the processor 1501 loads and executes the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above-mentioned image processing method embodiment; in a specific implementation, the processor 1501 loads the one or more instructions in the computer-readable storage medium and executes the following steps:

[0317] Determining a current encoding block to be encoded;

[0318] Determine a template region for the current coding block; a pixel correlation between a sample point in the template region and a target pixel corresponding to a pixel to be processed in the current coding block is greater than a correlation threshold, where the target pixel is a pixel corresponding to the pixel to be processed in the original image; and a template length of the template region in any direction along a middle sample point of the template region is greater than a length threshold, where the any direction includes at least one of the following: horizontal, vertical, and diagonal directions.

[0319] The current coding block is encoded based on the template area to generate a compressed code stream.

[0320] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1501 and, when executing to determine a template region for a current coding block, specifically perform the following steps:

[0321] Obtain multiple candidate template regions;

[0322] A cost comparison strategy is used to perform a cost estimation calculation on each candidate template region among a plurality of candidate template regions, and obtain a cost estimation result corresponding to each candidate template region; the cost estimation result corresponding to the candidate template region is used to indicate the degree of loss when the current coding block is encoded using the corresponding candidate template region; wherein the cost comparison strategy includes at least one of the following: a pixel correlation strategy and a distortion cost strategy;

[0323] A candidate template region with the smallest degree of loss indicated by a cost estimation result is selected from the multiple candidate template regions as the template region of the current coding block.

[0324] Based on the same inventive concept, the principles and beneficial effects of solving the problems of the image processing device provided in the embodiment of the present application are similar to the principles and beneficial effects of solving the problems of the image processing method provided in the method embodiment of the present application. Please refer to the principles and beneficial effects of the implementation of the method. For the sake of concise description, they will not be repeated here.

[0325] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a blockchain node device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described image processing method.

[0326] Those skilled in the art will appreciate that the units and algorithmic steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0327] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD) or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0328] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any technical object of a person skilled in the art that is within the technical scope disclosed in the present application and that can be easily conceived of by a person skilled in the art is within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection of the claims.

Claims

1. An image processing method, characterized in that: include: Determining a current coded block in a compressed code stream; Determining a template region for the current coding block; A pixel correlation between a sample point in the template area and a target pixel point corresponding to a pixel point to be processed in the current coding block is greater than a correlation threshold, and the target pixel point is a pixel point corresponding to the pixel point to be processed in the original image; a template length of the template area in any direction along a middle sample point of the template area is greater than a length threshold, and the any direction includes at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction; The current coding block is decoded based on the template area to obtain a reconstructed image of the current coding block.

2. The method according to claim 1, wherein The template area includes N sample points, where N is a positive integer; an area in the template area along the horizontal direction of the middle sample point includes N1 horizontal sample points, an area in the template area along the vertical direction of the middle sample point includes N2 vertical sample points, and an area in the template area along the diagonal direction of the middle sample point includes N3 diagonal sample points; N1, N2, and N3 are positive integers, and N1, N2, and N3 are all less than N; The N1 horizontal sample points are symmetrically distributed with the vertical direction of the middle sample point as the symmetry axis, the N2 vertical sample points are symmetrically distributed with the horizontal direction of the middle sample point as the symmetry axis; and the N3 diagonal sample points are symmetrically distributed with the diagonal direction of the middle sample point as the symmetry axis. Each sample point in the template area corresponds to a coefficient, and the coefficients of the sample points that are symmetrically distributed are the same.

3. The method according to claim 2, wherein N=29, N1=8, N2=8, N3=2, the template shape of the template area includes: Two sample points are distributed adjacent to a sample point on any diagonal line passing through the middle sample point in the template region.

4. The method according to claim 2, wherein N=29, the template shapes of the template area include: There are N3 sample points distributed on the two diagonal lines of the middle sample point, where N3=(N-1-N1-N2) / 2.

5. The method according to claim 1, wherein The determining of a template region for the current coding block includes: Obtaining a flag bit of the current coding block from the compressed code stream; Parsing the flag bit of the current coding block to obtain a template area of ​​the current coding block; The current coding block includes at least one of the following: an image, a data block, a coding tree unit, and a macroblock.

6. The method according to claim 5, wherein The flag bit is a sequence-level flag bit, and the sequence-level flag bit is used to indicate whether the coding block in the compressed code stream uses the target template area under the target coding mode; the target coding mode is any one of multiple coding modes, and the target template area is any one of multiple template areas; The parsing of the flag bit of the current coding block to obtain the template area of ​​the current coding block includes: Obtaining a current coding mode of the current coding block; the current coding mode is the target coding mode; Parsing the sequence-level flag of the current coding block to obtain a parsing result; Determine a template area of ​​the current coding block according to the parsing result and the current coding mode.

7. The method according to claim 6, wherein The parsing result indicates whether the current coding block in the current coding mode uses a target template area, and determining the template area of ​​the current coding block according to the parsing result and the current coding mode includes: According to an indication of the parsing result, determining that the current coding block in the current coding mode does not use the target template area; or, According to an indication of the parsing result, it is determined that the current coding block in the current coding mode uses the target template area.

8. The method according to claim 6, wherein The parsing result indicates that the current coding block in the current coding mode uses a target template area, and determining the template area of ​​the current coding block according to the parsing result and the current coding mode includes: Obtaining a default correspondence between the current coding mode and the template region according to the parsing result; the default correspondence is used to indicate that the current coding block in the current coding mode is allowed to use the target template region; the current coding mode is the target coding mode; According to the default corresponding relationship, a target template region having a corresponding relationship with the current coding mode is determined as the template region of the current coding block.

9. The method according to claim 5, wherein The flag bit is a sequence-level flag bit, and the sequence-level flag bit is used to indicate that all coding blocks in the compressed code stream use the same template area; Wherein, all coding blocks in the compressed code stream belong to a video sequence.

10. The method according to claim 5, wherein The flag bit is a non-sequential flag bit, and the compressed code stream includes a non-sequential flag bit corresponding to each coding block in the compressed code stream; the non-sequential flag bit of the current coding block is used to indicate the template area used by the current coding block.

11. The method according to claim 10, wherein The current coding block includes a first component, a second component and a third component; the non-sequential level flag of the current coding block is used to indicate that the first component, the second component and the third component all use the same template area.

12. The method according to claim 10, wherein The current coding block includes a first component, a second component, and a third component; the non-sequence-level flag bits of the current coding block are: a first sub-flag bit of the first component, a second sub-flag bit of the second component, and a third sub-flag bit of the third component; Among them, the first sub-flag of the first component is used to indicate the template area used by the first component; the second sub-flag of the second component is used to indicate the template area used by the second component; and the third sub-flag of the third component is used to indicate the template area used by the third component.

13. The method according to claim 1, wherein The determining of a template region for the current coding block includes: Acquire default setting information, where the default setting information indicates a default template area; Based on the default setting information, the default template area is used as the template area of ​​the current coding block; the default setting information is negotiated and set by the encoding end and the decoding end.

14. The method according to claim 13, wherein The current coding block includes a first component, a second component, and a third component; the default setting information indicates that different components in the current coding block correspond to the same default template area; The step of using the default template area as the template area of ​​the current coding block based on the default setting information includes: The default template area is used as the template area of ​​the first component, the second component and the third component in the current coding block respectively.

15. The method according to claim 13, wherein The current coding block includes a first component, a second component, and a third component; the default setting information indicates that different components of the current coding block correspond to respective default template regions; The step of using the default template area as the template area of ​​the current coding block based on the default setting information includes: Based on the default setting information, respectively set a first template region for the first component in the current coding block, set a second template region for the second component, and set a third template region for the third component; At least two template regions among the first template region, the second template region and the third template region are different.

16. The method according to claim 1, wherein The determining of a template region for the current coding block includes: Acquire a coded block from the compressed code stream; the coded block is a coded block that is in the same image as the current coded block in the spatial domain, or a coded block that has the same position in the image as the current coded block in the temporal domain; Calculating pixel distribution between the coded block and the current coded block to obtain a pixel distribution result; A template area is determined for the current coding block according to the pixel distribution result.

17. The method according to claim 16, wherein The determining a template area for the current coding block according to the pixel distribution result includes: If the pixel distribution result indicates that the difference in content distribution between the current coding block and the coded block is less than or equal to a difference threshold, the template area used by the coded block is used as the template area of ​​the current coding block.

18. The method according to claim 16, wherein The determining a template area for the current coding block according to the pixel distribution result includes: If the pixel distribution result indicates that the difference in content distribution between the current coding block and the coded block is greater than a difference threshold, determining a template region for the current coding block in a reference manner; The reference mode includes at least one of the following: a default mode, a flag mode, and a reference frame mode.

19. The method according to claim 1, wherein The determining of a template region for the current coding block includes: Obtaining the number of reference frames corresponding to the current coding block; A template area is determined for the current coding block according to the number of frames.

20. The method according to any one of claims 1 to 19, wherein The template area is a filter template; The decoding process of the current coding block based on the template area to obtain a reconstructed image of the current block to be coded includes: The current coding block is filtered based on the filter template of the current coding block to obtain a reconstructed image of the current coding block.

21. The method according to any one of claims 1 to 19, wherein: The template area is a prediction template; and decoding the current coding block based on the template area to obtain a reconstructed image of the current block to be coded includes: Performing prediction processing on the current coding block based on a prediction template of the current coding block to obtain a reconstructed image of the current coding block; The prediction process includes at least one of the following: inter-frame prediction process and intra-frame prediction process.

22. An image processing method, characterized in that: include: Determining a current encoding block to be encoded; Determine a template region for the current coding block; a pixel correlation between a sample point in the template region and a target pixel point corresponding to a pixel point to be processed in the current coding block is greater than a correlation threshold, the target pixel point being a pixel point corresponding to the pixel point to be processed in the original image; a template length of the template region in any direction along a middle sample point of the template region is greater than a length threshold, the any direction including at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction; The current coding block is coded based on the template region to generate a compressed code stream.

23. The method according to claim 22, wherein The determining of a template region for the current coding block includes: Obtain multiple candidate template regions; Performing a cost estimation calculation on each candidate template region among the multiple candidate template regions using a cost comparison strategy to obtain a cost estimation result corresponding to each candidate template region; the cost estimation result corresponding to the candidate template region is used to indicate a degree of loss when encoding the current coding block using the corresponding candidate template region; wherein the cost comparison strategy includes at least one of the following: a pixel correlation strategy and a distortion cost strategy; A candidate template region with the smallest degree of loss indicated by a cost estimation result is selected from the multiple candidate template regions as the template region of the current coding block.

24. An image processing device, characterized in that: include: a determination unit, configured to determine a current coding block in a compressed code stream; a processing unit, configured to determine a template region for the current coding block; A pixel correlation between a sample point in the template area and a target pixel point corresponding to a pixel point to be processed in the current coding block is greater than a correlation threshold, and the target pixel point is a pixel point corresponding to the pixel point to be processed in the original image; a template length of the template area in any direction along a middle sample point of the template area is greater than a length threshold, and the any direction includes at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction; The processing unit is further configured to perform decoding processing on the current coding block based on the template area to obtain a reconstructed image of the current coding block.

25. An image processing device, characterized in that: include: A determination unit, configured to determine a current coding block to be coded; a processing unit, configured to determine a template region for the current coding block; wherein a pixel correlation between a sample point in the template region and a target pixel point corresponding to a pixel point to be processed in the current coding block is greater than a correlation threshold, the target pixel point being a pixel point corresponding to the pixel point to be processed in the original image; and a template length of the template region in any direction along a middle sample point of the template region is greater than a length threshold, the any direction including at least one of the following: a horizontal direction, a vertical direction, and a diagonal direction; The processing unit is further configured to perform encoding processing on the current coding block based on the template area to generate a compressed code stream.

26. An image processing device, characterized in that a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the image processing method according to any one of claims 1 to 21 is implemented, or the image processing method according to any one of claims 22 to 23 is implemented.

27. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is suitable for being loaded by a processor and executing the image processing method according to any one of claims 1 to 21, or implementing the image processing method according to any one of claims 22 to 23.

28. A computer program product, characterized in that The computer program product includes computer instructions, and when the computer instructions are executed by a processor, the image processing method according to any one of claims 1 to 21 is implemented, or the image processing method according to any one of claims 22 to 23 is implemented.