Intra-frame prediction method and device, storage medium and electronic equipment

By fusing more pixel information of reconstructed blocks within the template matching framework, and adopting an adaptively weighted template matching prediction mode, the problem of low accuracy of template matching prediction mode is solved, improving video encoding performance and saving code rate.

CN120568076APending Publication Date: 2025-08-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410217268.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The accuracy of template matching prediction mode is not high, resulting in limited intra prediction performance.

Method used

Fuse more pixel information of reconstructed blocks within the template matching framework, and adopt adaptively weighted template matching prediction mode to improve prediction accuracy.

Benefits of technology

It improves the accuracy of template matching prediction mode, improves the performance of video encoding, and saves the code rate required for video encoding by about 0.3%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120568076A_ABST
    Figure CN120568076A_ABST
Patent Text Reader

Abstract

The invention discloses an intra-frame prediction method and device, a storage medium and electronic equipment, and belongs to the technical field of audio and video coding and decoding. The method comprises the following steps: determining a first coding block and a first template, wherein the first template and the first coding block have a first position relationship; at least two second templates are determined in a coded pixel area in the frame, the second templates and the first templates are the same in shape, and template matching distortion between the second templates and the first templates meets a preset requirement; a second coding block corresponding to each second template is determined, the second template and the second coding block have a second position relation, and the first position relation and the second position relation are the same position relation; and fusing the pixel values of the second coding blocks to obtain a predicted value corresponding to the first coding block. According to the method, the fusion of pixel information of more reconstructed blocks is realized, the intra-frame coding performance under the intra-frame prediction mode of template matching is effectively improved, and the video coding performance is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio and video coding and decoding technology, and in particular to an intra-frame prediction method, device, storage medium, and electronic device. Background Art

[0002] Intra-frame prediction is a key technology in video coding. It primarily leverages spatial correlations in the video domain, using previously encoded pixels in the current image to predict the value of the current pixel, thereby removing spatial redundancy. This prediction is typically performed within the same frame, hence the name intra-frame prediction. In intra-frame prediction, the encoder selects the most appropriate prediction mode, which predicts the current block based on the relationship between the pixel values ​​of the current block and its neighboring pixels. Template matching, a related technique, is a method for intra-frame prediction. However, template matching suffers from low accuracy, creating a bottleneck that restricts coding performance. Summary of the Invention

[0003] The embodiments of the present application provide an intra-frame prediction method, apparatus, storage medium, and electronic device to solve the problem that the accuracy of the template matching prediction mode in the related art is not high, thereby creating a bottleneck that restricts encoding performance.

[0004] According to one aspect of an embodiment of the present application, a method for intra-frame prediction is provided, the method comprising:

[0005] Determine a first coding block and a first template, where the first template and the first coding block have a first positional relationship;

[0006] Determining at least two second templates in the encoded pixel region within the frame, wherein the second templates have the same shape as the first template, and template matching distortion between the second templates and the first template meets a preset requirement;

[0007] Determine a second coding block corresponding to each second template, the second template and the second coding block have a second positional relationship, and the first positional relationship and the second positional relationship are the same positional relationship;

[0008] The pixel values ​​of each of the second coding blocks are fused to obtain a prediction value corresponding to the first coding block.

[0009] According to one aspect of an embodiment of the present application, an intra-frame prediction apparatus is provided, the apparatus comprising:

[0010] A coding block determining module, configured to determine a first coding block and a first template, wherein the first template and the first coding block have a first positional relationship;

[0011] A search module is configured to determine at least two second templates in an encoded pixel region within a frame, wherein the second templates have the same shape as the first template and the template matching distortion between the second templates and the first template meets a preset requirement;

[0012] A prediction module is used to determine the second coding block corresponding to each second template, the second template and the second coding block have a second positional relationship, and the first positional relationship and the second positional relationship are the same positional relationship; and to fuse the pixel values ​​of each second coding block to obtain a prediction value corresponding to the first coding block.

[0013] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the above-mentioned intra-frame prediction method.

[0014] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the above-mentioned intra-frame prediction method.

[0015] According to one aspect of an embodiment of the present application, a computer program product is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to implement the above-described intra-frame prediction method.

[0016] The technical solutions provided in the embodiments of the present application can bring the following beneficial effects:

[0017] The present invention provides an intra-frame prediction method that specifically improves the template matching intra-frame prediction mode. This improved intra-frame prediction method integrates pixel information from more reconstructed blocks (encoded blocks) within the template matching framework, thereby breaking the current block prediction model based on a single reconstructed block in related technologies, raising the upper limit of the prediction accuracy of the template matching prediction mode, and thus improving video coding performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] FIG1( a ) is a schematic diagram of a newly added prediction direction of a VCC prediction mode provided in one embodiment of the present application;

[0020] FIG1( b ) is a schematic diagram of a DC prediction mode provided by an embodiment of the present application;

[0021] FIG1( c ) is a schematic diagram of a Planar prediction mode provided by an embodiment of the present application;

[0022] FIG1( d ) is a schematic diagram of a template matching prediction mode provided by an embodiment of the present application;

[0023] Figure 2 is a schematic diagram of an application program operating environment provided by an embodiment of the present application;

[0024] Figure 3 This is a flowchart of an intra-frame prediction method provided by an embodiment of the present application;

[0025] Figure 4 is a schematic diagram of an exemplary shape and position of a first template provided in one embodiment of the present application;

[0026] Figure 5 This is a flow chart of a second template determination method provided by an embodiment of the present application;

[0027] Figure 6 is a schematic diagram of a search area provided by an embodiment of the present application;

[0028] Figure 7 This is a schematic diagram of searching for a second coding block provided by an embodiment of the present application;

[0029] Figure 8 This is a flow chart of a method for determining a prediction value corresponding to a first coding block provided by an embodiment of the present application;

[0030] Figure 9 This is a block diagram of an intra-frame prediction device provided by one embodiment of the present application;

[0031] Figure 10 This is a structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0032] Before introducing the method embodiments provided in the present application, a brief introduction is first given to the relevant terms or nouns that may be involved in the method embodiments of the present application to facilitate understanding by those skilled in the art in the field of the present application.

[0033] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to enable data computing, storage, processing, and sharing. Cloud technology is a general term for network, information technology, integration technology, management platform technology, and application technology, all based on the cloud computing business model. It can form a resource pool that can be used on demand with flexibility and convenience. Cloud computing technology will become a crucial support. Backend services for technical network systems, such as video websites, image websites, and more portals, require extensive computing and storage resources. With the rapid development and application of the internet industry, every item will likely have its own unique identification mark and will need to be transmitted to backend systems for logical processing. Data of varying levels will be processed separately, and data from all industries will require a strong system backend, which can only be achieved through cloud computing.

[0034] Video encoding and decoding: Video encoding and decoding are common processing technologies in digital media. Video encoding, also known as video compression, primarily uses specific compression techniques to convert raw video data into another format, typically smaller in size, to reduce storage space and transmission bandwidth. The core of video encoding is to reduce the video data rate while preserving visual quality. Common compression techniques used in video encoding can be categorized as lossy or lossless. Lossy compression alters the image, reducing information content and resulting in a decrease in image quality, but it offers higher compression ratios. Lossless compression does not cause any loss in the video image, allowing for a complete restoration of the file, but at a relatively low compression ratio. Video decoding is the reverse process of video encoding. Its primary task is to decompress the encoded (compressed) video data and restore it to a playable video stream. Video decoding must ensure that the quality and detail of the original video are restored as closely as possible. In general, video encoding and decoding are interdependent and complementary technologies that together enable the compression, storage, transmission, and playback of video data. With increasing demand for high-definition video, video encoding and decoding technologies are constantly evolving to improve compression efficiency, reduce distortion, and enhance video quality.

[0035] Before describing the embodiments of the present application in detail, the relevant technical background related to the embodiments of the present application is introduced to facilitate understanding by those skilled in the art in the field of the present application.

[0036] Intra-frame prediction is an important technology in video coding. It mainly uses the correlation in the spatial domain of the video and uses the encoded pixels of the current image to predict the value of the current pixel to achieve the purpose of removing the spatial redundancy of the video. This prediction is usually performed within the same frame, so it is called intra-frame prediction. In intra-frame prediction, the encoder selects the most appropriate prediction mode, which predicts the current block based on the relationship between the pixel value of the current block and its adjacent pixel values. For example, if an area is smooth in the image, the prediction may be based on the average or median of the adjacent pixels. If an area contains obvious edges or textures, the prediction may be based on the direction and intensity of these edges or textures. The embodiment of the present application briefly introduces the intra-frame prediction mode in the related art:

[0037] VVC Intra-frame Prediction Modes: VVC (Versatile Video Coding, also known as H.266) is a next-generation video coding standard designed to provide higher coding efficiency and a wider range of applications than HEVC (H.265). In VVC, intra-frame prediction technology has been further optimized and expanded. First, the number of intra-frame prediction modes in VVC has increased from 33 in HEVC to 65, primarily to accommodate prediction requirements for a wider range of directions and textures. Furthermore, VVC introduces new intra-frame prediction technologies, such as wide-angle intra prediction and block-based quadtree partitioning. To accommodate a wider range of prediction directions, the number of intra-frame angular prediction modes in VVC has increased to 65. Together with the DC mode and planar mode, the total number of intra-frame prediction modes in VVC is 67. Please refer to Figure 1(a), which shows a schematic diagram of the newly added prediction directions in the VVC prediction mode. The dashed lines in Figure 1(a) indicate the prediction directions added to VVC compared to HEVC.

[0038] DC prediction mode: DC prediction mode is an intra-frame prediction mode, which is mainly suitable for predicting large flat areas, that is, scenes where the pixel values ​​in the area are basically unchanged. In DC prediction mode, the pixel value of the current block is obtained by the average value of the reference pixels above and to the left of it. Please refer to Figure 1(b), which shows a schematic diagram of the DC prediction mode. According to the shape of the current block, for a square current block, the average value of the reference pixels on the top and left is calculated, and for a non-square current block, the average value of the long side is calculated, and then the current block is filled, as follows: when the width is equal to the height, the average value of the left reference pixels and the upper reference pixels is used as the prediction value to fill the entire current block; when the width is greater than the height, the average value of the upper reference pixels is used as the prediction value to fill the entire current block; when the width is less than the height, the average value of the left reference pixels is used as the prediction value to fill the entire current block.

[0039] Planar prediction mode: Planar prediction mode is an intra-frame prediction mode suitable for pixel gradients, that is, areas where pixel values ​​change slowly. See Figure 1(c), which shows a schematic diagram of the Planar prediction mode. The Planar prediction mode is calculated by taking a weighted average of the pixel values ​​corresponding to four preset parameters: a, b, c, and d, with the weights being distance-dependent.

[0040] MIP (Matrix Weighted Intra Prediction) prediction mode: This intra prediction technique uses linear affine transformations. First, a current block with a width of W and a height of H is predicted. The prediction references the reconstructed pixels W above and H to the left. The reconstructed pixels are obtained in the same way as traditional intra prediction. These (W + H) pixels are then averaged, affine transformed, and upsampled to obtain the final prediction value.

[0041] MRL (Multiple Reference Line) prediction mode: In HEVC, only the left column and the top row of pixels are used as reference pixels, while VVC allows the use of multiple reference lines (MRL). By using multiple reference lines, the MRL prediction mode can better capture local changes and texture information in the image, thereby providing more accurate predictions. In VVC, the MRL prediction mode can use multiple different reference row indexes, such as reference rows 0, 1, 3, etc. The encoder selects the appropriate reference row index based on the image content and prediction accuracy requirements. For each reference row index, the encoder generates a prediction block and uses rate-distortion optimization (Rate-Distortion Optimization) technology to select the best prediction block as the final prediction result.

[0042] IBC (Intra block copy) prediction mode: IBC is a block-level coding mode. IBC coding is regarded as the third prediction mode in addition to the intra-frame or inter-frame prediction mode. Similar to the inter-frame technology, the encoding end performs motion search (Block Matching, Block Maching, BM) to find the best block vector (Block Vector, also called Motion Vector) for each current block. The block vector is used to indicate the displacement from the current block to the reference block. The difference from the inter-frame technology is that the best block vector of IBC is searched in the reconstructed area of ​​the frame where the current block is located, while the inter-frame motion vector is obtained by searching in the adjacent reference frames.

[0043] Template Matching prediction mode: Template matching prediction technology is used for intra-frame prediction to improve intra-frame prediction performance. Template matching prediction can use the reconstructed pixels on the left and top sides of the current block as templates, such as Г-shaped reconstructed pixels as templates, and search in the encoded and reconstructed area of ​​the current frame to obtain the target template that best matches the current Г-shaped template. The reconstructed block corresponding to the target template is directly used as the prediction value of the current block. Please refer to Figure 1(d), which shows a schematic diagram of the template matching prediction mode. In this diagram, the blocks participating in intra-frame prediction can be represented by CU, that is, coding unit. CU is a core concept in video coding. It represents an image area block that is processed independently. By flexibly selecting the division method and size of CU, the encoder can better adapt to different image content and achieve efficient video compression. In one embodiment, the current coding block (CU) is searched within a given search range (R1 to R4 in Figure 1(d)) based on a G-shaped template. The best matching target template is obtained, as shown in the upper left corner of Figure 1(d). The blocks adjacent to the searched target template are then used as matching prediction blocks for the current CU's template and are directly used as the prediction result for the current CU. For template matching prediction, the encoder only needs to indicate whether template matching prediction is used, without identifying other information, as this process can also be completed on the decoder.

[0044] The template matching prediction mode in related technologies directly uses the reconstructed block corresponding to the matched target template as the prediction result for the current block. However, the actual difference between the selected reconstructed block and the current block may still be significant, resulting in low template matching accuracy. Directly using the reconstructed block corresponding to the matched target template as the prediction result for the current block cannot further incorporate pixel information from other related reconstructed blocks. Therefore, the loss of important pixel information may limit the upper limit of template matching prediction accuracy.

[0045] In view of this, the embodiments of the present application make specific improvements to the template matching prediction mode and propose an improved intra-frame prediction method. This method integrates pixel information from more reconstructed blocks within the template matching framework, thereby breaking the current block prediction model based on a single reconstructed block in related technologies and raising the upper limit of the prediction accuracy of the template matching prediction mode. Specifically, this method effectively improves the intra-frame coding performance under the template matching prediction mode by proposing a template matching prediction mode based on adaptive weighting, thereby improving the performance of video coding.

[0046] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0047] Please refer to Figure 2 , which shows a schematic diagram of an application program running environment provided by an embodiment of the present application. The application program running environment may include: a terminal 10 and a server 20.

[0048] The terminal 10 includes, but is not limited to, electronic devices such as mobile phones, computers, intelligent voice interaction devices, smart home appliances, car terminals, game consoles, e-book readers, multimedia playback devices, wearable devices, etc. The terminal 10 may be installed with a client of an application.

[0049] In an embodiment of the present application, the above-mentioned application may be any application that can provide or use intra-frame prediction services. Typically, the application may be an audio and video application. Of course, in addition to audio and video application applications, other types of applications may also provide services that rely on or use intra-frame prediction services. For example, news applications, social applications, interactive entertainment applications, browser applications, shopping applications, content sharing applications, virtual reality (VR) applications, augmented reality (AR) applications, etc., which are not limited in this embodiment of the present application. Optionally, a client of the above-mentioned application is running in the terminal 10.

[0050] The server 20 is used to provide background services for the client of the application in the terminal 10. For example, the server 20 can run an audio and video application system or an audio and video encoding and decoding service system. For example, the server 20 can be the background server of the above-mentioned application. The server 20 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the server 20 provides background services for applications in multiple terminals 10 at the same time.

[0051] Optionally, the terminal 10 and the server 20 may communicate with each other via a network 30. The terminal 10 and the server 20 may be directly or indirectly connected via wired or wireless communication, which is not limited in this application.

[0052] Please refer to Figure 3 , which shows a flow chart of an intra-frame prediction method provided by an embodiment of the present application. The method can be run in a computer device, which refers to an electronic device with data calculation and processing capabilities. For example, the execution subject of each step can be Figure 2The encoder or codec in the server 20 or terminal 10 in the application execution environment is shown.

[0053] S301. Determine a first coding block and a first template, wherein the first template and the first coding block have a first positional relationship.

[0054] The embodiments of the present application can be implemented in a scenario where an encoder performs intra-frame prediction encoding on a video, encoding each coding unit within the frame to obtain an encoding result. The coding unit can be understood as a coding block, and the first coding block is the coding block currently undergoing intra-frame prediction. This coding block has not yet been encoded, and the encoding method of the embodiments of the present application can complete the encoding of this first coding block. The first template is a pixel interval formed by pixels having a first positional relationship with the above-mentioned first coding block, and is a pixel interval that has already been encoded.

[0055] The template matching prediction mode is a prediction mode that searches for a related reconstructed block for predicting the first coding block in an already coded pixel region based on a template. The first template is a search basis used in the search process. The embodiments of the present application do not limit the specific shape and position of the first template. The first template can be understood as a pixel region including a number of pixels, and the pixel region has a unique positional relationship with the first coding block, namely the first positional relationship. Typically, the first template is composed of pixels adjacent to the first coding block.

[0056] Please refer to Figure 4 , which shows a schematic diagram of an exemplary shape and position of the first template in an embodiment of the present application. Figure 4 The black part and the gray part correspond to the first template and the first coding block respectively. That is, the pixel area formed by the adjacent pixels on the left and the adjacent pixels on the top of the first coding block is the first template. Figure 4 CU represents the first coding block.

[0057] S302. Determine at least two second templates in the encoded pixel area within the frame, wherein the second templates have the same shape as the first template, and the template matching distortion between the second templates and the first template meets a preset requirement.

[0058] In the embodiment of the present application, the second template is selected by template matching distortion. That is, a pixel region that meets the following requirements is searched for in the encoded pixel region within the frame: (1) the pixel region has the same shape as the first template; (2) the template matching distortion between the pixel region and the first template is small, small enough to meet the preset requirements. Pixel regions that meet the above two requirements can be used as the second template. Of course, the embodiment of the present application does not limit the specific second template selection method.

[0059] In one embodiment, please refer to Figure 5 , which shows a flow chart of a second template determination method according to an embodiment of the present application. The above-mentioned determination of at least two second templates in the pixel region encoded within the frame includes:

[0060] S501. Determine a search area of ​​a preset shape in the encoded pixel area within the frame;

[0061] In one embodiment, all pixel regions that have been coded in the frame can be used as search regions. In some other embodiments, a region of a preset shape that is close to the first coding block in the pixel regions that have been coded in the frame can be used as the search region to increase the search speed. The embodiment of the present application does not limit the shape of the search region. Please refer to Figure 6 , which shows a schematic diagram of the search area of ​​an embodiment of the present application. Figure 6 The middle grey shaded area is the search area, which includes the coded pixel areas on the left, top and upper left sides of the first coding block (CU).

[0062] S502. Determine, in the search area, pixel distributions corresponding to pixel intervals of the same shape as the first template; calculate the sum of absolute pixel differences between the pixel distribution corresponding to each pixel interval and the pixel distribution of the first template, where the sum of absolute pixel differences is used to quantify template matching distortion;

[0063] Each pixel region and the first template include a number of pixels, forming a pixel distribution. Since the pixel region and the first template have the same shape, the template matching distortion can be quantified based on the difference between the pixels corresponding to the corresponding positions. In one embodiment, the template matching distortion can be quantified by the pixel absolute error. The pixel absolute error sum refers to the sum of the absolute errors between the pixels corresponding to the corresponding positions. The sum of absolute differences (SAD) is an error measurement method commonly used in video coding. In the context of video coding, SAD is generally used to evaluate the accuracy of prediction. For example, in intra-frame prediction, the encoder selects a prediction mode that predicts the value of the current pixel based on the encoded pixels. In order to evaluate the accuracy of the prediction, the encoder calculates the SAD. A smaller SAD value generally means a more accurate prediction, so the encoder can select this prediction mode to encode the current block.

[0064] In the embodiment of the present application, To calculate SAD, where L0 represents the first template, Li represents a certain (i-th) pixel region, M and N represent the length and width of the template shape, respectively, and Loss(i) represents the pixel absolute error of the pixel region. For example, if the pixels of the first template L0 are [100, 100, 101, 102] and the search pixel region Li is [99, 100, 100, 101], then the template matching distortion or the pixel absolute error Loss(i) is 1+0+1+1=3.

[0065] S503. Determine the pixel absolute error and the pixel interval that meets the preset requirement as the second template.

[0066] The embodiments of the present application do not limit the preset requirements. For example, the second template may be limited to the absolute pixel error and the minimum preset number of pixel intervals, or the second template may be limited to the absolute pixel error and the pixel intervals less than the preset error and threshold. The preset number and the preset error and threshold can be set according to actual conditions and do not constitute an implementation obstacle to the embodiments of the present application.

[0067] In one embodiment, determining the pixel absolute errors and the pixel intervals that meet the preset requirements as the second template includes: sorting the pixel absolute errors in ascending order, and determining the first preset number of pixel absolute errors and their corresponding pixel intervals in the sorted results as the second template, where the preset number is greater than or equal to 2. For example, the preset number may be 3, i.e., three second templates are searched.

[0068] In a specific embodiment, a specific embodiment of a quick search for a second template is disclosed:

[0069] Step 1: Downsample the G-shaped template (first template) and region R (search area) around the first coding block by 1 / 2 in both horizontal and vertical directions to obtain template L0' and region R'. Based on the same operation, downsample template L0' and region R' again by 1 / 2 in both horizontal and vertical directions to obtain template L0" and region R".

[0070] Step 2: Use the L0” template to search in area R” to obtain several best-matching templates. Sort them in ascending order of template matching distortion to obtain n matching templates, which are represented by L1”, L2”…, Ln”.

[0071] Step 3: Upsample L1”, L2”…, Ln” to obtain the corresponding search reference template in the R’ search area.

[0072] Step 4: In the region R', multiple templates that best match the template L0' are searched within the range S'xS' around each search reference template, and finally a total of m best matching templates are screened.

[0073] Step 5: Upsample the m templates to obtain Li' corresponding to the third search area, where i is greater than 0.

[0074] Step 6: Search within the range S x S around each Li' to ultimately obtain the w best-matching second templates. In this embodiment, the search ranges S' x S' and S x S (S' > S), and n > m > w, where w is an integer greater than 1. For example, S' = 9, S = 3, n = 9, m = 6, and w = 3 can be set.

[0075] S303. Determine a second coding block corresponding to each second template, wherein the second template and the second coding block have a second positional relationship, and the first positional relationship and the second positional relationship are the same positional relationship.

[0076] Please refer to Figure 7 , which shows a schematic diagram of searching for the second coding block in an embodiment of the present application. During the search process, according to the first template in the black part, several search areas can be searched, such as L1, L2, L3, and L4. Taking L1, L2, L3, and L4 as the second template as an example, the second coding blocks P1, P2, P3, and P4 can be obtained, where the positional correspondence between the second coding block P1 and the second template L1 is the second positional relationship. Similarly, the positional correspondence between the second coding block P2 and the second template L2, the positional correspondence between the second coding block P3 and the second template L3, and the positional correspondence between the second coding block P4 and the second template L4 are all second positional relationships, while the positional correspondence between the first coding block CU and the first template L0 is the first positional relationship. Obviously, the first positional relationship and the second positional relationship are the same, both indicating that the left side and the upper part of the coding block are surrounded by templates with corresponding relationships.

[0077] S304. Fuse the pixel values ​​of the second coding blocks to obtain a prediction value corresponding to the first coding block.

[0078] Unlike the solution in the related art that directly obtains the template matching prediction result, the embodiment of the present application screens multiple second coding blocks and fuses the pixel information to improve the accuracy of the prediction value. The above-mentioned fusion of the pixel values ​​of each of the above-mentioned second coding blocks to obtain the prediction value corresponding to the above-mentioned first coding block includes: determining the weight of each of the above-mentioned second coding blocks, the sum of the weights of each of the above-mentioned second coding blocks being a preset value; fusing the pixel values ​​of each of the above-mentioned second coding blocks based on the above-mentioned weights to obtain the prediction value corresponding to the above-mentioned first coding block. The embodiment of the present application does not set the preset value, for example, the preset value can be 1.

[0079] In some implementations, the fusion effects of different weighting schemes may be selected to obtain the optimal prediction value corresponding to the first coding block. In one embodiment, determining the weights of each of the second coding blocks includes determining at least two sets of weights, each set of weights including the weights of each of the second coding blocks.

[0080] In one embodiment, the template matching distortion of the second template corresponding to each of the above-mentioned second coding blocks can be calculated; the weights corresponding to each of the above-mentioned second coding blocks are determined according to the proportion of the template matching distortion corresponding to each of the above-mentioned second templates in the total template matching distortion, and the above-mentioned total template matching distortion is the sum of the template matching distortions corresponding to each of the above-mentioned second templates.

[0081] Taking the existence of three second templates as an example, the template matching distortions corresponding to the three second templates are Loss1, Loss2, and Loss3 respectively, then the total template matching distortion loss_All is loss_All = Loss(1) + Loss(2) + Loss(3), and the weights corresponding to each of the above second coding blocks are: wi = ((loss_All–Loss(i)) / (2*loss_All), i = 1, 2, 3.

[0082] For example, assuming Loss1 = 1, Loss2 = 2, and Loss3 = 3, the calculated weights are w1 = 5 / 12, w2 = 4 / 12, and w3 = 3 / 12. Assume that the pixel value at position (0,0) of the first coding block is 100, while the pixels at positions (0,0) of the three second coding blocks are 101, 102, and 103, respectively. The current weighted prediction result based on this weight is calculated as 101*5 / 12+102*4 / 12+103*3 / 12=101.833333, with the trailing decimals deleted to obtain a predicted value of 101. Of course, if there are other positions in the first coding block, predictions are also performed based on the same inventive concept to obtain corresponding predicted values, thereby obtaining the predicted value of the entire first coding block.

[0083] In one embodiment, each of the second coding blocks can be assigned the same weight. For example, if there are three second templates, the weights are all 1 / 3. Continuing with the previous example, w1 = w2 = w3 = 1 / 3, based on this weight, the prediction result for the position (0, 0) of the first coding block is 102.

[0084] In one embodiment, the weight of any of the second coding blocks can be determined as the preset value, and the weights of the other second coding blocks can be determined as 0. Continuing with the above example, w1=1, w2=w3=0, it means that only the second coding block with the smallest template matching distortion is used to predict the first coding block. Continuing with the above example, the prediction result of the position (0,0) of the first coding block is 102.

[0085] Please refer to Figure 8 , which shows a flow chart of a method for determining a prediction value corresponding to a first coding block. The above-mentioned fusion of the pixel values ​​of each of the above-mentioned second coding blocks based on the above-mentioned weights to obtain the prediction value corresponding to the above-mentioned first coding block includes:

[0086] S801. Perform weighted pixel fusion on each of the above second coding blocks based on each set of weights to obtain a corresponding fusion result;

[0087] The process of weighted pixel fusion has been described in the previous article. For example, the process of predicting the position (0,0) of the first coding block in the previous article belongs to the process of weighted pixel fusion, so that the prediction results of each position of the first coding block can be obtained, that is, the fusion result corresponding to the first coding block can be obtained.

[0088] S802. Calculate the pixel differences between each of the fusion results and the first coding block; and determine the prediction value corresponding to the first coding block based on the fusion result with the smallest pixel difference.

[0089] Based on the above different groups of weights, different fusion results are obtained for the first coding block. Taking three groups of weights as an example, assuming that three different fusion results are obtained, namely P1, P2 and P3, the best template matching prediction result is selected by comparing the pixel differences between these three fusion results and the first coding block, and then compared with other intra-frame prediction modes to select the best intra-frame prediction mode. If the pixel differences between the prediction results of other intra-frame prediction modes and the first coding block are greater than the pixel differences between the best template matching prediction result and the first coding block, the best template matching prediction result is used as the prediction value corresponding to the first coding block. Otherwise, the prediction result of the other intra-frame prediction modes with the smallest pixel difference with the first coding block is used as the prediction value corresponding to the first coding block.

[0090] There can be many intra-frame prediction modes, for example, they can include intra-frame prediction modes based on directional interpolation, such as the 65 angle prediction modes and DC mode and Planar mode in the VVC standard, matrix-based prediction mode (MIP), intra-frame prediction mode based on multiple reference strips (MRL), and intra-frame block copy prediction mode (IBC).

[0091] This embodiment of the present application does not limit the quantization method of pixel differences. For example, the square sum of the errors between the prediction result (including the fusion result obtained in the embodiment of the present application) and the original pixels of the first coding block can be used to quantize the pixel differences, that is, dist represents the sum of squared errors, orig(x,y) represents the original pixels of the first coded block, and pred(x,y) is the predicted pixel value. Assume that the dist values ​​obtained based on P1, P2, P3, and other intra-frame predictions are 100, 110, 120, and 130, respectively. Then, P1 is used as the predicted value for the first coded block.

[0092] In one embodiment, after the pixel values ​​of the above-mentioned second coding blocks are fused to obtain the prediction value corresponding to the above-mentioned first coding block, the above-mentioned method also includes: in the case where the prediction value corresponding to the above-mentioned first coding block is a prediction value obtained by fusing the pixel values ​​of the above-mentioned second coding blocks, the first prediction identifier is determined as a template prediction identifier, the above-mentioned first prediction identifier is used to indicate the prediction mode to the decoding end, and the above-mentioned template prediction identifier indicates the template matching prediction mode; according to the weight determination method of the weight used when fusing the pixel values ​​of the above-mentioned second coding blocks, the second prediction identifier is set, and the above-mentioned second prediction identifier is used to indicate the weight determination method; the above-mentioned first prediction identifier and the above-mentioned second prediction identifier are encoded to obtain an encoding result.

[0093] The encoding end needs to use the first prediction flag to indicate whether the first coding block selects the template matching prediction in the embodiment of the present application or other traditional intra-frame prediction modes. If the template matching prediction method in the embodiment of the present application is selected, a second prediction flag is further required to identify which weighted weight is selected. As shown in the following table, weighted index 1 represents the first group of weights calculated based on template matching distortion, index 2 represents a fixed (1 / 3, 1 / 3, 1 / 3) weight, and index 3 represents a fixed (1, 0, 0) weight.

[0094] Table 1

[0095] Weighted weight index 1 2 3 Binary representation 0 10 11

[0096] At the decoding end, the prediction mode of the first coding block can be obtained by parsing the code stream. If the template prediction identifier is obtained by parsing the code stream, the second prediction identifier is further parsed to obtain the weight acquisition method, so that the template matching intra-frame prediction can be completed using the method of the embodiment of the present application.

[0097] Of course, after encoding the first prediction identifier and the second prediction identifier and obtaining the encoding result, the method further includes: adding the pixel area of ​​the first coding block to the pixel area already encoded in the frame; determining a third coding block and a third template, the third template and the third coding block have the first positional relationship, and the third coding block is the next coding block of the first coding block; determining at least two fourth templates in the pixel area already encoded in the frame, the fourth template has the same shape as the third template, and the template matching distortion between the third template and the fourth template meets the preset requirements; determining the fourth coding block corresponding to each of the fourth templates, the fourth template and the fourth coding block have the second positional relationship; fusing the pixel values ​​of each of the fourth coding blocks to obtain the prediction value corresponding to the third coding block. This process of obtaining the prediction value corresponding to the third coding block is based on the same inventive concept as steps S301 to S304 above, and will not be repeated here.

[0098] The embodiment of the present application proposes an intra-frame prediction method, which is a specific improvement to the intra-frame prediction mode of template matching, and proposes an improved intra-frame prediction method. This method realizes the fusion of pixel information of more reconstructed blocks within the template matching framework, thereby breaking the mode of predicting the current block based on a single reconstructed block in the related art, improving the upper limit of the prediction accuracy of the prediction mode of template matching, and thus improving the performance of video encoding. Specifically, the way in which the embodiment of the present application realizes the fusion of pixel information of more reconstructed blocks is a template matching prediction method based on adaptive weighting. Under the same compression quality, experiments have confirmed that the embodiment of the present application can save about 0.3% of the bit rate required for video encoding. The embodiment of the present application can be applied to short video compression and video call scenarios to save bandwidth, reduce freezes, and improve user experience.

[0099] The embodiments of the present application can be applied to various applications or application scenarios related to video coding, such as video calls, short videos, video websites, remote conferences, etc. For example, it can be applied to the background compression scenario of the video service of the instant messaging program to improve the compression performance of the video service and save bandwidth.

[0100] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0101] Please refer to Figure 9 , which shows a block diagram of an intra-frame prediction device provided by an embodiment of the present application. The device can be a computer device or can be set in a computer device. The device can include:

[0102] A coding block determining module 901 is configured to determine a first coding block and a first template, wherein the first template and the first coding block have a first positional relationship;

[0103] A search module 902 is configured to determine at least two second templates in the encoded pixel region within the frame, wherein the second templates have the same shape as the first template and the template matching distortion between the second templates and the first template meets a preset requirement;

[0104] The prediction module 903 is used to determine the second coding block corresponding to each of the above-mentioned second templates, the above-mentioned second template and the above-mentioned second coding block have a second positional relationship, and the above-mentioned first positional relationship and the above-mentioned second positional relationship are the same positional relationship; and to fuse the pixel values ​​of each of the above-mentioned second coding blocks to obtain the predicted value corresponding to the above-mentioned first coding block.

[0105] In one embodiment, the search module 902 is configured to perform the following operations:

[0106] Determine a search area of ​​a preset shape in the encoded pixel area in the frame;

[0107] Determine, in the search area, pixel distributions corresponding to respective pixel intervals having the same shape as the first template;

[0108] Calculating the pixel absolute error sum between the pixel distribution corresponding to each pixel interval and the pixel distribution of the first template, wherein the pixel absolute error sum is used to quantify the template matching distortion;

[0109] The pixel absolute error and the pixel interval that meets the preset requirements are determined as the second template.

[0110] In one embodiment, the search module 902 is configured to perform the following operations:

[0111] The pixel absolute errors are sorted in ascending order, and the first preset number of pixel absolute errors and their corresponding pixel intervals in the sorting result are determined as the second template, where the preset number is greater than or equal to 2.

[0112] In one embodiment, the prediction module 903 is configured to perform the following operations:

[0113] Determine a weight of each of the second coding blocks, wherein the sum of the weights of the second coding blocks is a preset value;

[0114] The pixel values ​​of each of the second coding blocks are fused based on the weights to obtain a prediction value corresponding to the first coding block.

[0115] In one embodiment, the prediction module 903 is configured to perform the following operations:

[0116] Determining at least two groups of weights, each group of weights including weights of each of the second coding blocks;

[0117] Performing weighted pixel fusion on each of the second coding blocks based on each set of weights to obtain a corresponding fusion result;

[0118] Calculating pixel differences between each of the fusion results and the first coding block;

[0119] The prediction value corresponding to the first coding block is determined according to the fusion result with the smallest pixel difference.

[0120] In one embodiment, the prediction module 903 is configured to perform the following operations:

[0121] Including at least one of the following methods:

[0122] Calculating the template matching distortion of the second template corresponding to each of the second coding blocks; determining the weight corresponding to each of the second coding blocks according to the proportion of the template matching distortion corresponding to each of the second templates in the total template matching distortion, where the total template matching distortion is the sum of the template matching distortions corresponding to each of the second templates;

[0123] Determining each of the second coding blocks to have the same weight;

[0124] The weight of any of the second coding blocks is determined to be the preset value, and the weights of the other second coding blocks are determined to be 0.

[0125] In one embodiment, the prediction module 903 is configured to perform the following operations:

[0126] In a case where the prediction value corresponding to the first coding block is a prediction value obtained by fusing pixel values ​​of each of the second coding blocks, determining the first prediction flag as a template prediction flag, the first prediction flag being used to indicate a prediction mode to a decoding end, and the template prediction flag indicating a template matching prediction mode;

[0127] Setting a second prediction flag according to a weight determination method for weights used when fusing pixel values ​​of the second coding blocks, wherein the second prediction flag is used to indicate the weight determination method;

[0128] The first prediction identifier and the second prediction identifier are encoded to obtain an encoding result.

[0129] In one embodiment, the prediction module 903 is configured to perform the following operations:

[0130] Adding the pixel area of ​​the first coding block to the pixel area already coded in the frame;

[0131] Determine a third coding block and a third template, wherein the third template and the third coding block have the first positional relationship, and the third coding block is a next coding block of the first coding block;

[0132] Determining at least two fourth templates in the encoded pixel region within the frame, wherein the fourth templates have the same shape as the third template, and the template matching distortion between the third template and the fourth template meets the preset requirement;

[0133] Determine a fourth coding block corresponding to each of the fourth templates, wherein the fourth template and the fourth coding block have the second positional relationship;

[0134] The pixel values ​​of the fourth coding blocks are fused to obtain a prediction value corresponding to the third coding block.

[0135] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0136] Please refer to Figure 10 , which shows a block diagram of a computer device provided by an embodiment of the present application. The computer device may be a server for executing the above intra-frame prediction method. Specifically:

[0137] Computer device 1000 includes a central processing unit (CPU) 1001, a system memory 1004 including a random access memory (RAM) 1002 and a read-only memory (ROM) 1003, and a system bus 1005 connecting system memory 1004 and CPU 1001. Computer device 1000 also includes a basic input / output system (I / O system) 1006 for facilitating information transfer between various components within the computer, and a mass storage device 1007 for storing an operating system 1013, application programs 1014, and other program modules 1015.

[0138] The basic input / output system 1006 includes a display 1008 for displaying information and an input device 1009, such as a mouse and keyboard, for user input. Both the display 1008 and the input device 1009 are connected to the central processing unit 1001 via an input / output controller 1010 connected to the system bus 1005. The basic input / output system 1006 may also include an input / output controller 1010 for receiving and processing input from a variety of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1010 also provides output to a display screen, printer, or other types of output devices.

[0139] The mass storage device 1007 is connected to the central processing unit 1001 via a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable media provide non-volatile storage for the computer device 1000. In other words, the mass storage device 1007 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.

[0140] Without loss of generality, computer-readable media may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, tape cassettes, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that computer storage media are not limited to the above-mentioned ones. The above-mentioned system memory 1004 and mass storage device 1007 can be collectively referred to as memory.

[0141] According to various embodiments of the present application, the computer device 1000 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 1000 may be connected to the network 1012 via the network interface unit 1011 connected to the system bus 1005, or the network interface unit 1011 may be used to connect to other types of networks or remote computer systems (not shown).

[0142] The memory further includes a computer program, which is stored in the memory and configured to be executed by one or more processors to implement the intra-frame prediction method.

[0143] In an exemplary embodiment, a computer-readable storage medium is also provided, in which at least one instruction, at least one program, a code set or an instruction set is stored. When the at least one instruction, the at least one program, the code set or the instruction set is executed by a processor, the intra-frame prediction method is implemented.

[0144] Specifically, the intra-frame prediction method includes:

[0145] Determine a first coding block and a first template, wherein the first template and the first coding block have a first positional relationship;

[0146] Determining at least two second templates in the encoded pixel region within the frame, wherein the second templates have the same shape as the first template, and the template matching distortion between the second templates and the first template meets a preset requirement;

[0147] Determine a second coding block corresponding to each of the second templates, wherein the second template and the second coding block have a second positional relationship, and the first positional relationship and the second positional relationship are the same positional relationship;

[0148] The pixel values ​​of the second coding blocks are fused to obtain the prediction value corresponding to the first coding block.

[0149] In one embodiment, the determining of at least two second templates in the pixel region encoded within the frame includes:

[0150] Determine a search area of ​​a preset shape in the encoded pixel area in the frame;

[0151] Determine, in the search area, pixel distributions corresponding to respective pixel intervals having the same shape as the first template;

[0152] Calculating the pixel absolute error sum between the pixel distribution corresponding to each pixel interval and the pixel distribution of the first template, wherein the pixel absolute error sum is used to quantify the template matching distortion;

[0153] The pixel absolute error and the pixel interval that meets the preset requirements are determined as the second template.

[0154] In one embodiment, determining the pixel absolute error and the pixel interval meeting the preset requirement as the second template includes:

[0155] The pixel absolute errors are sorted in ascending order, and the first preset number of pixel absolute errors and their corresponding pixel intervals in the sorting result are determined as the second template, where the preset number is greater than or equal to 2.

[0156] In one embodiment, fusing the pixel values ​​of the second coding blocks to obtain the prediction value corresponding to the first coding block includes:

[0157] Determine a weight of each of the second coding blocks, wherein the sum of the weights of the second coding blocks is a preset value;

[0158] The pixel values ​​of each of the second coding blocks are fused based on the weights to obtain a prediction value corresponding to the first coding block.

[0159] In one embodiment, determining the weights of each of the second coding blocks includes: determining at least two groups of weights, each group of weights including the weights of each of the second coding blocks;

[0160] The above-mentioned fusing of the pixel values ​​of each of the second coding blocks based on the above-mentioned weights to obtain the prediction value corresponding to the above-mentioned first coding block includes:

[0161] Performing weighted pixel fusion on each of the second coding blocks based on each set of weights to obtain a corresponding fusion result;

[0162] Calculating pixel differences between each of the fusion results and the first coding block;

[0163] The prediction value corresponding to the first coding block is determined according to the fusion result with the smallest pixel difference.

[0164] In one embodiment, determining the weight of each of the second coding blocks includes at least one of the following methods:

[0165] Calculating the template matching distortion of the second template corresponding to each of the second coding blocks; determining the weight corresponding to each of the second coding blocks according to the proportion of the template matching distortion corresponding to each of the second templates in the total template matching distortion, where the total template matching distortion is the sum of the template matching distortions corresponding to each of the second templates;

[0166] Determining each of the second coding blocks to have the same weight;

[0167] The weight of any of the second coding blocks is determined to be the preset value, and the weights of the other second coding blocks are determined to be 0.

[0168] In one embodiment, after fusing the pixel values ​​of the second coding blocks to obtain the predicted value corresponding to the first coding block, the method further includes:

[0169] In a case where the prediction value corresponding to the first coding block is a prediction value obtained by fusing pixel values ​​of each of the second coding blocks, determining the first prediction flag as a template prediction flag, the first prediction flag being used to indicate a prediction mode to a decoding end, and the template prediction flag indicating a template matching prediction mode;

[0170] Setting a second prediction flag according to a weight determination method for weights used when fusing pixel values ​​of the second coding blocks, wherein the second prediction flag is used to indicate the weight determination method;

[0171] The first prediction identifier and the second prediction identifier are encoded to obtain an encoding result.

[0172] In one embodiment, after encoding the first predicted identifier and the second predicted identifier to obtain an encoding result, the method further includes:

[0173] Adding the pixel area of ​​the first coding block to the pixel area already coded in the frame;

[0174] Determine a third coding block and a third template, wherein the third template and the third coding block have the first positional relationship, and the third coding block is a next coding block of the first coding block;

[0175] Determining at least two fourth templates in the encoded pixel region within the frame, wherein the fourth templates have the same shape as the third template, and the template matching distortion between the third template and the fourth template meets the preset requirement;

[0176] Determine a fourth coding block corresponding to each of the fourth templates, wherein the fourth template and the fourth coding block have the second positional relationship;

[0177] The pixel values ​​of the fourth coding blocks are fused to obtain a prediction value corresponding to the third coding block.

[0178] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or an optical disk, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0179] In an exemplary embodiment, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described intra-frame prediction method.

[0180] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.

[0181] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0182] In addition, in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0183] The above are merely exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. An intra-frame prediction method, characterized in that: The method comprises: Determine a first coding block and a first template, where the first template and the first coding block have a first positional relationship; Determining at least two second templates in the encoded pixel region within the frame, wherein the second templates have the same shape as the first template, and template matching distortion between the second templates and the first template meets a preset requirement; Determine a second coding block corresponding to each second template, the second template and the second coding block have a second positional relationship, and the first positional relationship and the second positional relationship are the same positional relationship; The pixel values ​​of each of the second coding blocks are fused to obtain a prediction value corresponding to the first coding block.

2. The method according to claim 1, characterized in that The determining of at least two second templates in the pixel region encoded within the frame includes: Determining a search area of ​​a preset shape in the coded pixel area in the frame; Determining pixel distribution corresponding to each pixel interval having the same shape as the first template in the search area; Calculating a sum of pixel absolute errors between a pixel distribution corresponding to each pixel interval and a pixel distribution of the first template, wherein the sum of pixel absolute errors is used to quantify the template matching distortion; The pixel absolute error and the pixel interval that meets the preset requirements are determined as the second template.

3. The method according to claim 2, characterized in that The step of determining the pixel absolute error and the pixel interval meeting the preset requirement as the second template includes: The pixel absolute errors are sorted in ascending order, and the first preset number of pixel absolute errors and their corresponding pixel intervals in the sorting result are determined as the second template, where the preset number is greater than or equal to 2.

4. The method according to any one of claims 1 to 3, characterized in that The fusing pixel values ​​of the second coding blocks to obtain a prediction value corresponding to the first coding block includes: Determine a weight of each second coding block, where the sum of the weights of each second coding block is a preset value; The pixel values ​​of each of the second coding blocks are fused based on the weights to obtain a prediction value corresponding to the first coding block.

5. The method according to claim 4, characterized in that Determining the weight of each second coding block includes: determining at least two groups of weights, each group of weights including the weight of each second coding block; The fusing the pixel values ​​of each second coding block based on the weight to obtain a prediction value corresponding to the first coding block includes: Performing weighted pixel fusion on each of the second coding blocks based on each group of weights to obtain a corresponding fusion result; Calculating pixel differences between each of the fusion results and the first coding block; Determine a prediction value corresponding to the first coding block according to a fusion result with the smallest pixel difference.

6. The method according to claim 4, characterized in that Determining the weight of each second coding block includes at least one of the following methods: Calculating the template matching distortion of the second template corresponding to each second coding block; determining the weight corresponding to each second coding block according to the proportion of the template matching distortion corresponding to each second template in the total template matching distortion, where the total template matching distortion is the sum of the template matching distortions corresponding to each second template; Determining each of the second coding blocks to have the same weight; The weight of any one of the second coding blocks is determined to be the preset value, and the weights of other second coding blocks are determined to be 0.

7. The method according to claim 5, characterized in that After fusing the pixel values ​​of the second coding blocks to obtain the predicted value corresponding to the first coding block, the method further includes: When the prediction value corresponding to the first coding block is a prediction value obtained by fusing pixel values ​​of each of the second coding blocks, determining the first prediction flag as a template prediction flag, the first prediction flag being used to indicate a prediction mode to a decoding end, and the template prediction flag indicating a template matching prediction mode; Setting a second prediction flag according to a weight determination method for weights used when fusing pixel values ​​of the second coding blocks, where the second prediction flag is used to indicate the weight determination method; The first prediction identifier and the second prediction identifier are encoded to obtain an encoding result.

8. The method according to claim 7, characterized in that After encoding the first prediction flag and the second prediction flag to obtain an encoding result, the method further includes: adding a pixel region of the first coding block to an already coded pixel region within the frame; Determine a third coding block and a third template, where the third template and the third coding block have the first positional relationship, and the third coding block is a next coding block of the first coding block; Determining at least two fourth templates in the encoded pixel region within the frame, wherein the fourth templates have the same shape as the third template, and template matching distortion between the third template and the fourth templates meets the preset requirement; Determining a fourth coding block corresponding to each of the fourth templates, wherein the fourth template and the fourth coding block have the second positional relationship; The pixel values ​​of the fourth coding blocks are fused to obtain a prediction value corresponding to the third coding block.

9. An intra-frame prediction device, characterized in that: The device comprises: A coding block determining module, configured to determine a first coding block and a first template, wherein the first template and the first coding block have a first positional relationship; A search module is configured to determine at least two second templates in an encoded pixel region within a frame, wherein the second templates have the same shape as the first template and the template matching distortion between the second templates and the first template meets a preset requirement; A prediction module is used to determine the second coding block corresponding to each second template, the second template and the second coding block have a second positional relationship, and the first positional relationship and the second positional relationship are the same positional relationship; and to fuse the pixel values ​​of each second coding block to obtain a prediction value corresponding to the first coding block.

10. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the intra-frame prediction method according to any one of claims 1 to 8.

11. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes to implement the intra-frame prediction method according to any one of claims 1 to 8.

12. A computer storage medium, characterized in that The storage medium stores at least one instruction, and the at least one instruction and at least one program are loaded by a processor to execute the intra-frame prediction method according to any one of claims 1 to 8.