Encoding method, encoder, and computer-readable storage medium
Patent Information
- Application Number
- CN202211463044.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-11-21
AI Technical Summary
[0004]本申请的发明人发现,上述的预测过程存在一定的局限性,预测过程有待进一步优化
[0011]有益效果是:本申请在构建预测模型时,不再仅仅根据考虑第一参考像素点,而是同时结合第一参考像素点周围像素点的重建像素值、第一参考像素点的重建像素值中的多个值,因此可以提高所构建第一目标预测模型的准确率,能够提高最终对待编码像素点进行预测的准确率,保证最终图像传输的效果。
Smart Images

Figure CN115988212B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of video coding, and in particular relates to a coding method, an encoder, and a computer-readable storage medium. Background Technology
[0002] Because video image data is relatively large, it usually needs to be encoded and compressed before transmission or storage. The encoded data is called video stream.
[0003] Currently, linear prediction is commonly used when encoding video image data. Linear prediction involves constructing a linear model between a reference block and the current coded block, and then using the reconstructed pixels of the reference block to predict the pixel values of the current coded block. The parameters of the linear model are calculated using the reconstructed pixel values of the current coded block and its neighboring reference blocks.
[0004] The inventors of this application have discovered that the above prediction process has certain limitations and needs further optimization. Summary of the Invention
[0005] This application provides an encoding method, an encoder, and a computer-readable storage medium that can optimize the visual effect of images.
[0006] The first aspect of this application provides an encoding method, the method comprising: obtaining a current template and a reference template of a current block, wherein the current template includes a plurality of reconstructed first pixels in the current frame, the reference template includes a plurality of reconstructed second pixels in a reference frame, the current frame includes the current block, and the reference frame corresponds to the current frame; constructing a first target prediction model based on a first reconstructed pixel value of the first pixel and a plurality of corresponding second reconstructed pixel values, wherein the plurality of second reconstructed pixel values include a plurality of values among the reconstructed pixel values of pixels surrounding the first reference pixel and the reconstructed pixel values of the first reference pixel, the first reference pixel being a pixel in the reference template corresponding to the first pixel; and using the first target prediction model to predict the pixel to be encoded in the current block to obtain a first predicted value of the pixel to be encoded.
[0007] A second aspect of this application provides a decoding method, the method comprising: receiving encoded data sent by an encoder; and obtaining a predicted value of a current pixel in a current decoding block by decoding the encoded data; wherein the predicted value of the current pixel in the current decoding block is obtained by processing using any of the encoding methods described above.
[0008] A third aspect of this application provides an encoder, which includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit, respectively. The memory stores program data, and the processor executes the program data in the memory to implement the steps in the above-described encoding method.
[0009] A fourth aspect of this application provides a decoder, which includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit, respectively. The memory stores program data, and the processor executes the program data in the memory to implement the steps in the above-described decoding method.
[0010] A fifth aspect of this application provides a computer-readable storage medium storing a computer program that can be executed by a processor to implement the steps in the above-described method.
[0011] The beneficial effects are: when constructing the prediction model, this application no longer only considers the first reference pixel, but also combines the reconstructed pixel values of the pixels surrounding the first reference pixel and multiple values among the reconstructed pixel values of the first reference pixel. Therefore, it can improve the accuracy of the constructed first target prediction model, improve the accuracy of the final prediction of the pixel to be encoded, and ensure the effect of the final image transmission. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein: Figure 1 This is a simplified structural diagram of linear prediction in existing technologies; Figure 2 This is a flowchart illustrating one embodiment of the coding method of this application; Figure 3 This is a diagram illustrating the current block and the current template; Figure 4 This is a schematic diagram of the reference block and the reference template; Figure 5 This is a schematic diagram of the second pixel and its surrounding pixels; Figure 6 This is a partial flowchart of the coding method of this application in the first application scenario; Figure 7 This is a flowchart illustrating step S120 of the coding method in the second application scenario of this application; Figure 8 This is a flowchart illustrating step S130 of the coding method of this application in the second application scenario; Figure 9 This is a flowchart illustrating step S110 of the coding method of this application in the third application scenario; Figure 10 This is a schematic diagram of the current block and the target region outside the current block; Figure 11 yes Figure 9 A flowchart illustrating step S112; Figure 12 It corresponds Figure 11 A diagram showing the current template and the reference template; Figure 13 This is a partial flowchart illustrating the coding method of this application in the fourth application scenario; Figure 14 This is a partial flowchart illustrating the coding method of this application in the fifth application scenario; Figure 15 This is a flowchart illustrating another embodiment of the coding method of this application; Figure 16 It is a schematic diagram of the encoding blocks of the three color components, their respective reference blocks, the current template, and the reference template; Figure 17 This is a flowchart illustrating another embodiment of the coding method of this application; Figure 18 It is an application scenario Figure 17 The corresponding process diagram; Figure 19 This is a flowchart illustrating another embodiment of the coding method of this application; Figure 20 It is an application scenario Figure 19 The corresponding process diagram; Figure 21 This is a flowchart illustrating one embodiment of the decoding method of this application; Figure 22 This is a schematic diagram of one embodiment of the encoder of this application; Figure 23 This is a schematic diagram of another embodiment of the encoder of this application; Figure 24 This is a schematic diagram of the structure of one embodiment of the decoder of this application; Figure 25 This is a schematic diagram of another embodiment of the decoder in this application; Figure 26 This is a schematic diagram of one embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0014] It should be noted that the terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0015] It should be noted that the terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0016] It should be noted that the terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0017] To better understand the scheme of this application, a brief introduction to the background technology of encoding will be given first: In video encoding and decoding, to improve compression ratio and reduce the number of codewords to be transmitted, the encoder does not directly encode and transmit pixel values. Instead, it uses intra-frame or inter-frame prediction modes, which predict the pixel values of the current block using reconstructed pixels from the coded blocks of the current frame or a reference frame. The pixel value predicted using a certain prediction mode is called the predicted pixel value, and the difference between the predicted pixel value and the original pixel value is called the residual. The encoder only needs to encode a certain prediction mode and the residual generated when using that mode, and the decoder can decode the corresponding pixel value based on this bitstream information, thus greatly reducing the number of codewords required for encoding.
[0018] When encoding an image frame, the coded blocks within the frame are typically encoded in a specific order. For example, they can be encoded from left to right and from top to bottom, or from right to left and from bottom to top. Understandably, when encoding from left to right and from top to bottom, the reconstructed pixels (predicted pixels) adjacent to any given block are located above and to the left of that block. Conversely, when encoding from right to left and from bottom to top, the reconstructed pixels adjacent to any given block are located below and to the right of that block. For clarity, the following explanation will use the left-to-right, top-to-bottom encoding method.
[0019] For ease of explanation, the current block will be defined as the block being encoded at the current moment. When using inter-frame prediction for encoding, the closest reference block to the current block needs to be found in the reference frame (which can be any already encoded image frame) through methods such as motion search, and the motion information between the current block and the reference block, such as the motion vector MV, is recorded.
[0020] See Figure 1 ,exist Figure 1 The current block is a diagonal line-filled part, denoted by label 1. The current template of the current block is a dot-filled part, wherein the current template includes at least one reconstructed pixel around the current block. The reference block of the current block is a horizontal line-filled part, denoted by label 2. The reference template of the reference block is a grid-filled part, wherein the reference template includes at least one reconstructed pixel around the reference block. The current block becomes a prediction block after prediction, denoted by label 3.
[0021] When using LIC (local illumination compensation) mode for encoding, the encoding process is as follows: A linear model is constructed using the illumination variations between all pixels in the current template of the current block and all pixels in the reference template of the reference block. This linear model characterizes the local illumination relationship between the current block and its reference block. The parameters of this linear model include a scaling factor α and an offset β. Then, this linear model is used to compensate for the difference in brightness between the current block and the reference block, i.e., the predicted pixel value of each pixel in the current block is determined using the following formula: qi = α × pi + β, where pi is the initial predicted pixel value of pixel i in the current block, and qi is the predicted pixel value of pixel i. pi is obtained by motion compensation based on the motion vector MV between the current block and the reference block. The determination process of pi will not be described in detail here.
[0022] As can be seen from the above process, in the existing encoding process using LIC mode, only the reference point corresponding to the current pixel in the reference block is considered. However, in actual images, the current pixel may be affected not only by its corresponding reference point but also by other pixels. Therefore, the existing encoding process does not match the actual situation. To overcome the shortcomings of the existing technology, this application proposes the following solution: See Figure 2 , Figure 2 This is a flowchart illustrating one embodiment of the encoding method of this application, which includes: S110: Obtain the current template and reference template of the current block, wherein the current template includes multiple first pixel points that have been reconstructed in the current frame, the reference template includes multiple second pixel points that have been reconstructed in the reference frame, the current frame includes the current block, and the reference frame corresponds to the current frame.
[0023] Specifically, the current block refers to the coding block to be encoded in the current frame, also known as the current coding block. The current block can be a prediction block or a sub-block within a prediction block. Furthermore, the current block can be either a luma block or a chroma block; there are no restrictions on this.
[0024] After acquiring the current block, a matching coded block is searched in the reference frame, and this matching coded block is defined as the reference block. The process of determining the reference block in the reference frame can employ any method available in the prior art, such as using motion search to identify the closest coded block to the current block within the reference frame.
[0025] The current template and reference template can be determined using existing technologies, as described below.
[0026] S120: Construct a first target prediction model based on the first reconstructed pixel value of the first pixel and the corresponding multiple second reconstructed pixel values. The multiple second reconstructed pixel values include multiple values among the reconstructed pixel values of the pixels surrounding the first reference pixel and the reconstructed pixel values of the first reference pixel. The first reference pixel is the pixel in the reference template that corresponds to the first pixel.
[0027] Specifically, the first pixel in the current template corresponds one-to-one with the second pixel in the reference template, and the position of the first pixel in the current template is the same as the position of the corresponding second pixel in the reference template. To better understand this, an example will be provided below: Combination Figure 3 and Figure 4 ,exist Figure 3 The current block is represented by 10, and the pixels filled with dots around the current block are the first pixels in the current template. Figure 4 The reference block is represented by 20. The pixels used to fill the lines around the reference block are the second pixel in the reference template. Figure 3 The second pixel corresponding to the first pixel with index 11 is Figure 4 The pixel with the center number 13, Figure 3 The second pixel corresponding to the first pixel with index 12 is Figure 4 The pixel with the center number 14.
[0028] Meanwhile, the multiple second reconstructed pixel values corresponding to the first pixel point include the reconstructed pixel value of the first reference pixel point and the reconstructed pixel value of at least one pixel point around the first reference pixel point, or the multiple second reconstructed pixel values include the reconstructed pixel values of multiple pixels around the first reference pixel point, or the multiple second reconstructed pixel values include both the reconstructed pixel value of the first reference pixel point and the reconstructed pixel values of pixels around the first reference pixel point.
[0029] To better understand, here we combine Figure 5 Explanation: exist Figure 5 In this application scenario, all the surrounding pixels of the second pixel form a "nine-square grid" with the second pixel. In this application scenario, pixel A is the first reference pixel corresponding to the first pixel. At least one pixel around the first reference pixel includes at least one of the following pixels: pixel B, pixel C, pixel D, pixel E, pixel F, pixel G, pixel H, and pixel I. That is to say, the multiple reconstructed pixel values corresponding to the first pixel at this time include the reconstructed pixel values of multiple pixels among pixel A, pixel B, pixel C, pixel D, pixel E, pixel F, pixel G, pixel H, and pixel I.
[0030] Meanwhile, this implementation method can construct a first target prediction model based on the first reconstructed pixel values of all first pixels and the multiple second reconstructed pixel values corresponding to each first pixel, or it can construct a first target prediction model based on the first reconstructed pixel values of some first pixels and the multiple second reconstructed pixel values corresponding to each of those first pixels.
[0031] The specific reconstructed pixel values corresponding to the first pixel can be predetermined by the designer, or determined based on parameters such as the texture and size of the current block. For example, when the texture of the current block is determined to be a horizontal texture, in Figure 5 In application scenarios, multiple second reconstructed pixel values corresponding to the first pixel can be set to include the reconstructed pixel values of pixels E, A, and C.
[0032] In existing technologies, prediction models are constructed solely based on the first reconstructed pixel value of the first pixel and the reconstructed pixel value of the first reference pixel. This construction process only considers the relationship between the first pixel and the first reference pixel, which has limitations and ultimately leads to inaccurate prediction results, affecting the visual effect of the transmitted image.
[0033] In constructing the prediction model, this application no longer only considers the first reference pixel, but also combines the reconstructed pixel values of the pixels surrounding the first reference pixel and multiple values from the reconstructed pixel values of the first reference pixel. Therefore, it can improve the accuracy of the constructed first target prediction model, improve the accuracy of the final prediction of the pixel to be encoded, and ensure the effect of the final image transmission.
[0034] S130: Use the first target prediction model to predict the pixel to be encoded in the current block and obtain the first predicted value of the pixel to be encoded.
[0035] Specifically, after obtaining the first target prediction model, the model is used to predict the pixels to be encoded in the current block to obtain the corresponding first prediction value.
[0036] In this method, the first target prediction model can be used to predict all the pixels to be encoded in the current block to obtain the corresponding first prediction value, or the matching pixels to be encoded can be predicted to obtain the corresponding first prediction value, as described below.
[0037] As can be seen from the above, this application constructs a first target prediction model based on the first reconstructed pixel value of the first pixel and the corresponding multiple second reconstructed pixel values. Compared with the traditional LIC method, this method can improve the accuracy of the final prediction of the pixel to be encoded, and can ensure the effect of the final image transmission. For ease of explanation, the method used in this application is referred to as the improved LIC method.
[0038] In this embodiment, step S120 specifically includes: S121: Determine multiple model parameters based on the first reconstructed pixel value of the first pixel and the corresponding multiple second reconstructed pixel values.
[0039] Meanwhile, step S130 specifically includes: S131: Determine the first predicted value of the pixel to be encoded based on multiple target values and multiple model parameters, wherein the multiple target values correspond one-to-one with the multiple model parameters.
[0040] Among them, the multiple target values include multiple first target reconstructed pixel values corresponding to the pixel to be encoded and multiple values in the calculated values. The calculated values are obtained by performing calculations on the first target reconstructed pixel values. The multiple first target reconstructed pixel values include the reconstructed pixel values of the first target pixel and the reconstructed pixel values of the pixels surrounding the first target pixel. The first target pixel is the pixel in the reference block that corresponds to the pixel to be encoded.
[0041] Specifically, multiple target values correspond one-to-one with multiple model parameters. In other words, which target value parameters are used for calculation is determined by multiple model parameters, and the model parameters are determined by the process of constructing the first target prediction model. That is to say, step S1301 is related to the process of constructing the first target prediction model.
[0042] Wherein, the position of the first target pixel in the reference block is the same as the position of the pixel to be encoded in the current block, and the multiple reconstructed pixel values of the first target include the reconstructed pixel value of the first target pixel and the reconstructed pixel values of the pixels surrounding the first target pixel, for example in Figure 5 In the application scenario, if pixel A is the first target pixel, then the reconstructed pixel values of the pixels surrounding the first target pixel include the reconstructed pixel values of pixels B, C, D, E, F, G, H, and I.
[0043] Meanwhile, the calculated value is obtained by performing calculations on the reconstructed pixel value of the first target. Specifically, the reconstructed pixel value of the first target can be substituted into a preset function to obtain the calculated value. The preset function can be an identity function, a square function, an absolute value function, etc., and there are no restrictions here.
[0044] In one application scenario, the first predicted value y of the pixel to be encoded can be determined using the following model formula: y = p1X1 + p2X2 + ... + p i X i +……p n X n ; Among them, p1, p2, ... p i ...p n Each of the following represents a model parameter: X1, X2, ... X i、 ...X n Each represents a different objective value, and the model parameter p represents a different objective value. i With target value X i One-to-one correspondence.
[0045] In other words, multiple target values are multiplied by their respective model parameters to obtain multiple first products; then the sum of all first products is determined as the first predicted value of the pixel to be encoded.
[0046] Of course, this application is not limited to this. In other embodiments, after obtaining the sum of all the first products, a series of calculations can be performed on the sum to obtain the first predicted value of the pixel to be encoded.
[0047] Alternatively, in other implementations, other model formulas can be used to calculate the first predicted value of the pixel to be encoded.
[0048] exist Figure 5 In application scenarios, assuming pixel A is the first target pixel, the above model formula can be set as follows: Model Formula 1: a = p1A + p2B + p3C + p4D + p5E + p6c; Model Formula 2: a = p1A + p2B 2 + p3c; Model Formula 3: a = p1A + p2B + p3C + p4D + p5E + p6c + p7A 2 ; Model Formula 4: a = p1A + p2B + p3C + p4D + p5E + p6F + p7G + p8H + p9I + p 10 c); Where a is the first predicted value of the pixel to be encoded, A, B, C, D, E, F, G, H, and I are the reconstructed pixel values of pixels A, B, C, D, E, F, G, H, and I respectively, and c is a preset constant value.
[0049] In other words, Figure 5 In application scenarios, any of the above model formulas can be used to determine the first predicted value of the pixel to be encoded. Of course, other model formulas can also be used to determine the first predicted value of the pixel to be encoded, which will not be listed here.
[0050] As can be seen from the above, the essence of using the first objective prediction model for prediction is to use the model formula for calculation, and the purpose of constructing the first objective prediction model is to determine the multiple model parameters in the model formula.
[0051] For clarity, the following examples will be used to illustrate the point: When determining the first predicted value of the pixel to be encoded, if the first predicted value of the pixel to be encoded is determined using the above model formula, then the process of constructing the first target prediction model is to determine the six model parameters p1 to p6 in y=p1X1+ p2X2+ p3X3+ p4X4+p5X5+ p6c. Specifically, assuming... Figure 5If pixel A is the first reference pixel corresponding to the first pixel, then for multiple first pixels, the first reconstructed pixel value of the first pixel is substituted into y in the above formula. X1 is substituted with the reconstructed pixel value of pixel A, X2 is substituted with the reconstructed pixel value of pixel B, X3 is substituted with the reconstructed pixel value of pixel C, X4 is substituted with the reconstructed pixel value of pixel D, and X5 is substituted with the reconstructed pixel value of pixel E. Thus, multiple equations can be obtained. Then, after the fitting process, the values of p1 to p6 are obtained. Finally, during prediction, model formula one can be used to predict the pixel to be encoded.
[0052] In the first application scenario, see Figure 6 The method of this embodiment further includes: S210: After constructing the first target prediction model into multiple different candidate prediction models in sequence, determine the first prediction value of all pixels to be encoded under each candidate prediction model.
[0053] Steps S120-S130 are executed multiple times. Each time step S120 is executed, the first target prediction model is constructed into a different model, and the first prediction value of each pixel to be encoded under different models is obtained. In other words, for any pixel to be encoded in the current block, its first prediction value under different models can be obtained.
[0054] For example, during the first execution of steps S120-S130, multiple model parameters in the above model formula one are fitted, and then the model formula one with known model parameters is used to predict each pixel to be encoded, obtaining the first predicted value of each pixel to be encoded. During the second execution of steps S120-S130, multiple model parameters in the above model formula two are fitted, and then the model formula three is used to predict each pixel to be encoded. During the third execution of steps S120-S130, multiple model parameters in the above model formula three are fitted, and then the model formula three is used to predict each pixel to be encoded. During the fourth execution of steps S120-S130, multiple model parameters in the above model formula four are fitted, and then the model formula four is used to predict each pixel to be encoded.
[0055] S220: Based on the first predicted value of all pixels to be encoded under each candidate prediction model, determine the cost value corresponding to each candidate prediction model.
[0056] Specifically, for each candidate prediction model, the cost value corresponding to that candidate prediction model is determined based on the first predicted values of all pixels to be encoded in the current block under that candidate prediction model. The cost value is specifically the rate-distortion cost value. In determining the cost value, in addition to using the first predicted values of all pixels to be encoded, other parameters such as quantization parameters can also be combined. However, the process of determining the cost value is existing technology and will not be detailed here.
[0057] S230: The candidate prediction model with the lowest substitution value is selected as the final prediction model.
[0058] S240: Determine the first predicted value of each pixel to be encoded under the final prediction model as the final predicted value of each pixel to be encoded.
[0059] Specifically, step S220 yields the cost value for each candidate prediction model, and then the candidate prediction model with the lowest cost value can be found.
[0060] Among them, the candidate prediction model has the lowest cost, indicating that the prediction result under this model is more accurate. Therefore, this model is determined as the final prediction model, and the first prediction value of the pixel to be encoded under this model is determined as the final prediction value.
[0061] In other words, in this application scenario, multiple candidate prediction models are used to compete, and one is ultimately selected.
[0062] In this application scenario, in order to indicate the final prediction model to the decoding end, a first syntactic element is generated. In response to the instruction of the first syntactic element, steps S210-S240 are executed to generate a second syntactic element, which is used to indicate the final prediction model.
[0063] Specifically, in the LIC method, the syntactic element indicating the use of the LIC method is `lic_enable_flag`. Setting `lic_enable_flag=1` indicates using the LIC method, and `lic_enable_flag=0` indicates not using the LIC method. When `lic_enable_flag=1`, the first syntactic element `lic_mode_flag` is generated. Setting `lic_mode_flag=1` indicates using the improved LIC method, which uses multiple candidate prediction models for competition. Setting `lic_mode_flag=0` indicates using the traditional LIC method, and setting `lic_mode_flag=1` also indicates... The second syntactic element `lic_mode_type_flag` will be generated. Assuming that four candidate prediction models are used for competition, these four candidate prediction models will be represented by 0, 1, 2, and 3 respectively. Specifically, when the candidate prediction model with the number 0 is determined to be the final prediction model, `lic_mode_type_flag=0` is set; when the candidate prediction model with the number 1 is determined to be the final prediction model, `lic_mode_type_flag=1` is set; when the candidate prediction model with the number 2 is determined to be the final prediction model, `lic_mode_type_flag=2` is set; and when the candidate prediction model with the number 3 is determined to be the final prediction model, `lic_mode_type_flag=3` is set.
[0064] When multiple candidate prediction models are not used for competition, there is no need to generate the second syntactic element lic_mode_type_flag.
[0065] In the second application scenario, see Figure 7 Step S120 specifically includes: S1201: Based on the reconstructed pixel value of the second pixel, classify the second pixel to obtain multiple second pixel classes.
[0066] Specifically, after classification, two second pixels with similar reconstructed pixel values are usually assigned to the same second pixel class, while two second pixels with significantly different reconstructed pixel values are usually assigned to different second pixel classes.
[0067] The second pixel can be classified based on a pixel threshold. This process includes: generating classification intervals based on the pixel threshold. For example, if there are three pixel thresholds, A1, A2, and A3, then four classification intervals can be constructed: (-∞, A1), [A1, A2), [A2, A3), and [A3, +∞). If only one pixel threshold, B1, is included, then two classification intervals can be constructed: (-∞, B1) and [B1, +∞). After constructing the classification intervals, second pixels whose reconstructed pixel values fall within the same interval are classified into the same pixel class, thus achieving the classification of the second pixel.
[0068] Alternatively, clustering algorithms can be used to classify the second pixel. The clustering algorithm can be any existing clustering algorithm, such as k-means clustering, DBSCAN clustering, or CLARANS clustering. This application does not limit the type of clustering algorithm.
[0069] For ease of explanation, the second pixel is classified according to a pixel threshold, which is the average value of the reconstructed pixel values of the second pixel. Based on the average value, the second pixel is classified into two second pixel classes. One second pixel class includes second pixel points whose reconstructed pixel values do not exceed the average value, and the other second pixel class includes second pixel points whose reconstructed pixel values exceed the average value.
[0070] S1202: In the current template, determine the first pixel class corresponding to each second pixel class, wherein the first pixel in the first pixel class corresponds one-to-one with the second pixel in the corresponding second pixel class, and the first pixel corresponding to the second pixel is the pixel in the current template that is at the same position as the second pixel.
[0071] Specifically, for each second pixel class, a corresponding first pixel class is determined. This determination involves: for each second pixel in the second pixel class, finding a first pixel at the same position in the current template, and adding that first pixel to the corresponding first pixel class. In other words, there is a one-to-one correspondence between the first pixel in the first pixel class and the second pixel in the corresponding second pixel class, and the first pixel corresponding to the second pixel is the pixel at the same position as the second pixel in the current template.
[0072] S1203: Construct a first target prediction model for each first pixel class based on the first reconstructed pixel value and the corresponding multiple second reconstructed pixel values of the first pixel point in each first pixel class.
[0073] Specifically, for each first pixel class, a corresponding first target prediction model is constructed. This process includes: constructing the first target prediction model based on the first reconstructed pixel value of the first pixel point in the first pixel class and at least one corresponding second reconstructed pixel value. The specific process of constructing the first target prediction model has been described above and will not be repeated here.
[0074] See now Figure 8 Step S130 includes: S1301: Based on the reconstructed pixel value of the third pixel in the reference block, classify the third pixel to obtain multiple third pixel classes. The rule for classifying the third pixel is the same as the rule for classifying the second pixel.
[0075] Specifically, after step S1301, two third pixels with similar reconstructed pixel values are usually assigned to the same third pixel class, while two third pixels with significantly different reconstructed pixel values are usually assigned to different third pixel classes.
[0076] The rule for classifying the second pixel is the same as the rule for classifying the third pixel. This means that the criteria for classifying two second pixels into the same second pixel class is the same as the criteria for classifying two third pixels into the same third pixel class.
[0077] For example, when classifying the second pixel by the average value of the reconstructed pixel value, the third pixel is also classified by the same average value, resulting in two third pixel classes. One third pixel class includes third pixels whose reconstructed pixel values do not exceed the average value, and the other third pixel class includes third pixels whose reconstructed pixel values exceed the average value.
[0078] S1302: Use the first target prediction model corresponding to the pixel to be encoded to predict the pixel to be encoded and obtain the first predicted value of the pixel to be encoded.
[0079] Among them, the first target prediction model corresponding to the pixel to be encoded is the first target prediction model corresponding to the first target pixel class among multiple first pixel classes. The first target pixel class corresponds to the second target pixel class among multiple second pixel classes. The second target pixel class is the same as the third target pixel class among multiple third pixel classes. The third target pixel class includes the pixel in the reference block that corresponds to the pixel to be encoded.
[0080] Specifically, the process of determining the first target prediction model corresponding to the pixel to be encoded includes: first, determining the third pixel corresponding to the pixel to be encoded in the reference block; then, determining the third pixel class to which the third pixel belongs; then, finding the second pixel class that is the same as the third pixel class; then, finding the first pixel class corresponding to the second pixel class; and finally, the first target prediction model corresponding to the first pixel class is the first target prediction model corresponding to the pixel to be encoded.
[0081] After determining the first target prediction model corresponding to the pixel to be encoded, the model is used to predict the pixel to be encoded, and the first predicted value of the pixel to be encoded is obtained.
[0082] It can be understood that if two pixels to be encoded each correspond to a third pixel that belongs to the same third pixel class, then the first target prediction model corresponding to these two pixels to be encoded is the same model.
[0083] The above scheme can use multiple first-target prediction models for prediction, which can adapt well to changes in pixel content.
[0084] It should be noted that in other application scenarios, the second pixel may not be classified. In this case, a first target prediction model is constructed using the first reconstructed pixel value of each first pixel and at least one corresponding second reconstructed pixel value. Then, the first predicted value of each pixel to be encoded is predicted using this model. Alternatively, a first target prediction model can be constructed using the first reconstructed pixel values of some first pixels and at least one corresponding second reconstructed pixel value. Then, the first predicted value of each pixel to be encoded is predicted using this model.
[0085] In the third application scenario, see [reference] Figure 9 Step S110, obtaining the current template and reference template of the current block, includes: S111: Determine the reference block in the reference frame.
[0086] S112: Determine the current template and the reference template based on the reconstructed pixels in the same target region outside the current block and outside the reference block, respectively.
[0087] The target region outside the current block includes at least one of the first sub-region, the second sub-region, the third sub-region, the fourth sub-region, and the fifth sub-region outside the current block. The first sub-region is located on the first side of the current block and its two ends are flush with the two ends of the current block, the second sub-region is located on the second side of the current block and its two ends are flush with the two ends of the current block, the third sub-region connects the first sub-region and the second sub-region, the fourth sub-region is located on the side of the first sub-region away from the third sub-region, and the fifth sub-region is located on the side of the second sub-region away from the third sub-region.
[0088] Specifically, the current template and the reference template are determined based on the reconstructed pixels in the same region outside the current block and the reference block, respectively.
[0089] Combination Figure 10 Taking the current block as an example, the target area outside the current block includes at least one of the first sub-region, the second sub-region, the third sub-region, the fourth sub-region, and the fifth sub-region. The target area can be preset by the designer or determined according to the texture, size, and other parameters of the current block.
[0090] The widths of the first and fourth sub-regions can be equal to the width of the current block, and the heights of the second and fifth sub-regions can be equal to the height of the current block. The heights of the first and fourth sub-regions are equal, and the widths of the second and fifth sub-regions are equal. The heights of the first and fourth sub-regions, as well as the widths of the second and fifth sub-regions, can be set according to actual needs. For example, the first and fourth sub-regions can be set to include 6 rows of pixels, and the second and fifth sub-regions can be set to include 6 columns of pixels.
[0091] After determining the target area, the reconstructed pixels in the target area outside the current block are determined as pixels in the current template, and the reconstructed pixels in the target area outside the reference block are determined as pixels in the reference template.
[0092] Combination Figure 1 In the prior art, the target area is set to include a first sub-region and a second sub-region, with the first sub-region including only one row of pixels and the second sub-region including only one column of pixels, which cannot utilize more spatial information. However, this application sets the target area in the above manner, which can make full use of spatial information and ensure the accuracy of the final prediction result.
[0093] In this application scenario, it's possible that not all reconstructed pixels in the target regions outside the current block and reference block can be acquired. For example, if the target regions outside the current block and reference block exceed the boundaries, only a portion of the reconstructed pixels in these regions can be acquired. This results in a one-to-one correspondence between the reconstructed pixels in the target regions outside the current block and those in the target regions outside the reference block. To avoid this, refer to [the relevant documentation / reference]. Figure 11 Step S112 specifically includes: S1121: Determine the first effective region outside the current block, wherein the first effective region is located in the target region outside the current block, and the first effective region includes the reconstructed pixels that can be acquired outside the current block.
[0094] S1122: Determine a second effective region outside the reference block, wherein the second effective region is located in the target region outside the reference block, and the second effective region includes reconstructed pixels that can be acquired outside the reference block.
[0095] S1123: Determine the common area of the first valid area and the second valid area.
[0096] S1124: Determine the current template and the reference template based on the reconstructed pixels in the common area outside the current block and the reference block, respectively.
[0097] Specifically, the first effective region is located in the target region outside the current block, and the first effective region includes all reconstructed pixels that can be acquired. The second effective region is located in the target region outside the reference block, and the second effective region also includes all reconstructed pixels that can be acquired.
[0098] To better understand, here we combine Figure 12 To illustrate with examples: In the diagram, the area filled with diagonal lines outside the current block is the first effective area outside the current block, and the area filled with dots outside the reference block is the second effective area outside the reference block. For example, in the diagram, assuming that the current block can obtain a maximum of 6 rows of reconstructed pixels within a region equal to 1 block width above it, and a maximum of 6 columns of reconstructed pixels within a region equal to 2 blocks width and height on the left, and a maximum of 4 rows of reconstructed pixels equal to 2 blocks width above the reference block and 0 rows of reconstructed pixels on the left, the common area of the first and second effective areas is obtained: a region of 4 rows equal to 1 block width above it. Finally, the reconstructed pixels in the common area outside the current block and the reference block are determined as the pixels in the current template and the pixels in the reference template, respectively. Finally, these reconstructed pixels are used to construct the model.
[0099] In the fourth application scenario, see Figure 13 The method of this application also includes: S310: Sequentially using the target region as multiple different candidate regions, determine the first predicted value of each pixel to be encoded in each candidate region.
[0100] Specifically, steps S110-S130 are executed multiple times, and each time step S110 is executed, the target region is set as a different candidate region, and the first prediction value of each pixel to be encoded is obtained under different candidate regions.
[0101] For example, multiple candidate regions may include different sub-regions. For instance, one candidate region may include the first, third, and fourth sub-regions; another candidate region may include the second, third, and fifth sub-regions; and yet another candidate region may include the first, second, third, fourth, and fifth sub-regions.
[0102] S320: Determine the cost value corresponding to each candidate region based on the first predicted value of each pixel to be encoded in each candidate region.
[0103] Specifically, for each candidate region, the cost value corresponding to the candidate region is determined based on the first predicted values of all pixels to be encoded within that candidate region. This cost value is specifically the rate-distortion cost value; the process of determining the cost value is prior art and will not be detailed here.
[0104] S330: The candidate region with the lowest substitution value is determined as the final region.
[0105] S340: Determine the first predicted value of each pixel to be encoded in the final region as the final predicted value for each pixel to be encoded.
[0106] Specifically, after obtaining the cost value corresponding to each candidate region, the candidate region with the smallest cost value can be found.
[0107] Among them, the candidate region has the lowest cost, indicating that the final prediction result is more accurate after the target region is determined as the candidate region. Therefore, the candidate region is determined as the final region, and the first prediction value of the pixel to be encoded in the final region is determined as the final prediction value.
[0108] In other words, in this application scenario, multiple candidate regions are used to compete, and one is ultimately selected.
[0109] In this application scenario, in order to indicate the final region to the decoding end, a first syntax element is generated. In response to the instruction of the first syntax element, steps S310-S340 are executed to generate a second syntax element, which is used to indicate the final region.
[0110] Specifically, in the LIC method, the syntax element indicating the use of the LIC method is `lic_enable_flag`. Setting `lic_enable_flag=1` indicates using the LIC method, and `lic_enable_flag=0` indicates not using the LIC method. When `lic_enable_flag=1`, the first syntax element `lic_mode_flag` is generated. Setting `lic_mode_flag=1` indicates using an improved LIC method, which uses multiple candidate regions for competition. When _mode_flag=0, it indicates that the traditional LIC method is used. When lic_mode_flag=1, a second syntax element lic_TL_flag will also be generated. Assuming that three candidate regions are used for competition, these three candidate regions are represented by 0, 1, and 2 respectively. Specifically, when the candidate region with the label 0 is determined to be the final region, lic_TL_flag=0 is set; when the candidate region with the label 1 is determined to be the final region, lic_TL_flag=1 is set; and when the candidate region with the label 2 is determined to be the final region, lic_TL_flag=2 is set.
[0111] When multiple candidate regions are not used for competition, there is no need to generate a second syntactic element.
[0112] In the fifth application scenario, see [reference] Figure 14 The method of this embodiment further includes: S410: Construct a second target prediction model based on the first reconstructed pixel value of the first pixel and the reconstructed pixel value of the first reference pixel.
[0113] S420: Using the second target prediction model, predict each pixel to be encoded in the current block to obtain the second predicted value of each pixel to be encoded.
[0114] Specifically, steps S410 to S420 involve using the traditional LIC method to predict the pixels to be encoded, thereby obtaining a second predicted value for each pixel to be encoded.
[0115] S430: Determine the first generation value based on the first predicted value of all pixels to be encoded in the current block.
[0116] S440: Determine the second generation value based on the second predicted values of all pixels to be encoded in the current block.
[0117] S450: In response to the first generation value being less than the second generation value, the first predicted value of each pixel to be encoded is determined as the final predicted value of each pixel to be encoded; otherwise, the second predicted value of each pixel to be encoded is determined as the final predicted value of each pixel to be encoded.
[0118] Specifically, if the value of the first generation is less than the value of the second generation, it means that the accuracy of prediction using the improved LIC method of this application is higher than the accuracy of prediction using the traditional LIC method. Therefore, the result of prediction using the improved LIC method of this application is taken as the final result; otherwise, the result of prediction using the traditional LIC method is taken as the final result.
[0119] In other words, the results predicted using the existing method and the results predicted using the improved LIC method of this application will compete, and ultimately only one of the two will be chosen.
[0120] When the prediction result using the existing scheme competes with the prediction result using steps S110-S130 of this application, in order to inform the decoder of the final competition result, a first syntactic element is generated. In response to the first syntactic element, steps S410-S450 are executed to generate a second syntactic element. When the value of the first generation is less than the value of the second generation, it indicates that the prediction using the improved LIC method is better, and the value of the first generation is set to a first value, such as 1. Otherwise, it indicates that the prediction using the traditional LIC method is better, and the second syntactic element is set to a second value, such as 0.
[0121] Specifically, in the LIC method, the syntax element that indicates the use of the LIC method is lic_enable_flag. When lic_enable_flag=1, it means the LIC method is used; when lic_enable_flag=0, it means the LIC method is not used. When lic_enable_flag=1, the first syntax element lic_mode_flag is generated. When lic_mode_flag=0, it means the traditional LIC method is used; when lic_mode_flag=1, it means the improved LIC method is used.
[0122] To better understand the settings of syntactic elements, a more detailed explanation follows: In the first instance, when multiple candidate prediction models and multiple candidate regions are not used for competition, if the improved LIC of this application is used directly, there is no need to add new syntactic elements. The original syntactic element lic_enable_flag in the LIC technology is used directly. When lic_enable_flag=1, it means that the improved LIC method is used. If the phenomenon of no solution is found during the prediction process, the traditional LIC method is used for prediction. When lic_enable_flag=0, it means that neither the improved LIC method nor the traditional LIC method is used, and no illumination compensation is performed.
[0123] In the second embodiment, when multiple candidate prediction models and multiple candidate regions are not used for competition, but instead an improved LIC and a traditional LIC are used for competition, the original syntax element lic_enable_flag in the LIC technology is used. When lic_enable_flag=1, it means that the LIC method is used, and when lic_enable_flag=0, it means that the LIC method is not used, that is, no illumination compensation is performed.
[0124] When lic_enable_flag=1 is set, a new syntax element lic_mode_flag is added. When lic_mode_flag=0 is set, the traditional LIC method is used. When lic_mode_flag=1 is set, the improved LIC method is used.
[0125] In the third embodiment, multiple candidate prediction models are used for the first target prediction model or multiple candidate regions are used to compete.
[0126] First, the primitive syntax element lic_enable_flag in the LIC technology is used. When lic_enable_flag=1, it means that the LIC method is used. When lic_enable_flag=0, it means that the LIC method is not used, that is, no lighting compensation is performed.
[0127] Additionally, when lic_enable_flag=1 is set, a new syntax element lic_mode_flag is added. When lic_mode_flag=0 is set, the traditional LIC method is used, and when lic_mode_flag=1 is set, the improved LIC method is used.
[0128] When lic_mode_flag=1, a new syntax element lic_mode_type_flag is added to indicate the final prediction model, or a new syntax element lic_TL_flag is added to indicate the final region. For details, please refer to the relevant introduction above.
[0129] As can be seen from the above, the first application scenario uses multiple candidate prediction models to compete, the second application scenario constructs multiple first target prediction models, the third application scenario takes the common area of the first effective region outside the current block and the second effective region outside the reference block, and the fourth application scenario uses multiple candidate regions to compete.
[0130] In other application scenarios, the solutions of the first, second, third, and fourth application scenarios can be combined arbitrarily. For example, multiple nonlinear predictions of multiple targets can be used for competition, as can multiple candidate regions. Alternatively, multiple first target prediction models can be constructed, and multiple candidate prediction models can be used for competition, etc.
[0131] To facilitate understanding, the following examples will be used to illustrate the point: In the first example: There are three candidate regions: one candidate region includes the first sub-region, the third sub-region, and the fourth sub-region; another candidate region includes the second sub-region, the third sub-region, and the fifth sub-region; and a second target region includes the first sub-region, the second sub-region, the third sub-region, the fourth sub-region, and the fifth sub-region.
[0132] There are four candidate prediction models, corresponding to formulas one through four above.
[0133] By combining any candidate region from the three candidate regions with any candidate prediction model from the four candidate prediction models, 12 combinations can be obtained.
[0134] In this embodiment, the second pixel is not classified. Instead, the first predicted value of the pixel to be encoded under each combination is determined. Based on the first predicted values of all pixels to be encoded under each combination, the cost value corresponding to that combination is determined, thus obtaining the cost values for each of the 12 combinations. Finally, the combination with the smallest cost value is determined as the final combination, and the first predicted value of the pixel to be encoded under that final combination is determined as the final predicted value of the pixel.
[0135] In the second example, a first target prediction model is constructed for each first pixel class, and multiple first target prediction models are used to compete. At the same time, multiple candidate regions are used to compete. It is assumed that there are two first pixel classes, that is, two first target prediction models are constructed, and four candidate prediction models are used to compete, as well as three candidate regions (candidate region 1, candidate region 2 and candidate region 3) are used to compete.
[0136] Firstly, it's understandable that because there are two first pixel classes, meaning two first target prediction models are constructed, the pixels to be encoded in the current block can be divided into two fourth pixel classes during prediction. Points within the same fourth pixel class share the same first target prediction model, while points belonging to different fourth pixel classes have different first target prediction models. For ease of explanation, the two constructed first target prediction models will be referred to as Model 1 and Model 2, and fourth pixel class 1 will correspond to Model 1, and fourth pixel class 2 will correspond to Model 2.
[0137] During encoding, the target region is first set as candidate region 1. Under this candidate region, model 1 is constructed into 4 candidate prediction models in sequence, and model 2 is constructed into 4 candidate prediction models in sequence. Since model 1 and model 2 can be constructed into different models, the candidate prediction models corresponding to model 1 and model 2 can be combined to obtain 4×4=16 combinations. Based on the first prediction value of all pixels to be encoded under each combination, the cost value corresponding to the combination is determined, so that the cost value corresponding to each of the 16 combinations under candidate region 1 can be obtained.
[0138] Then, the target region is set as candidate region 2, and the above steps are repeated to obtain the cost value of each of the above 16 combinations under candidate region 2.
[0139] Next, the target region is set as candidate region 3, and the above steps are repeated to obtain the corresponding cost value of each of the above 16 combinations under candidate region 3.
[0140] Then, find the minimum value among all the obtained values, determine the combination and candidate region corresponding to the minimum value, and finally determine the first predicted value of the current block under the combination and candidate region as the final predicted value.
[0141] See Figure 15 In another embodiment of this application, the decoding method further includes: S510: Obtain the target pixel value of each pixel in the first coding block, wherein the target pixel value is the target prediction value of the pixel or the reconstructed pixel value obtained based on the target prediction value, and take the first coding block or at least one sub-block in the first coding block as the current block, obtain the first prediction value of each pixel in the current block, and take the first prediction value as the target prediction value of the pixel.
[0142] Specifically, the target pixel value of a pixel in the first coding block can be the target predicted value or the reconstructed pixel value obtained based on the target predicted value. However, for ease of explanation, the following will use the target predicted value as an example.
[0143] The target prediction value of at least one sub-block in the first coding block is obtained through steps S110-S130. For details, please refer to the relevant content above, which will not be repeated here.
[0144] Alternatively, the first coded block can be used as the current block described above to obtain the first predicted value for each pixel in the first coded block.
[0145] S520: Construct a first prediction model based on the reconstructed pixel values of the pixels in the current template of the first coding block and the reconstructed pixel values of the pixels in the current template of the second coding block, wherein the first coding block corresponds to the second coding block.
[0146] S530: Based on the first prediction model and the first coding block, obtain the predicted value of each pixel in the second coding block.
[0147] The first coding block is a luma block, and the second coding block is a chroma block corresponding to the first coding block; or, the first coding block is a chroma block, and the second coding block is either a luma block or a chroma block corresponding to the first coding block.
[0148] Specifically, see Figure 16 The coded block for the luma block is represented by Y, and the two coded blocks for the chroma block are represented by C1 and C2 respectively. The reference block, reference template, and current template corresponding to each coded block are respectively represented by... Figure 15 The symbols in the text are used to represent it.
[0149] In this embodiment, if Y is used as the first coding block, the process of rY predicting rC1 is used to build a model, and then the model is applied to Y to obtain the predicted value of each pixel in C1. At the same time, the process of rY predicting rC2 is used to build a model, and then the model is applied to Y to obtain the predicted value of each pixel in C2.
[0150] Alternatively, if C1 is used as the first coding block, the process of rC1 predicting rC2 is used to build a model, and then the model is applied to C1 to predict C2, thus obtaining the predicted value of each pixel in C2; or if C2 is used as the first coding block, the process of rC2 predicting rC1 is used to build a model, and then the model is applied to C2 to predict C1, thus obtaining the predicted value of each pixel in C1.
[0151] Alternatively, if C1 is used as the first coding block, the process of rC1 predicting rY is used to build a model, and then the model is applied to C1 to predict Y, so as to obtain the predicted value of each pixel in Y.
[0152] Alternatively, if C2 is used as the first coding block, the process of rC2 predicting rY is used to build a model, and then the model is applied to C2 to predict Y, so as to obtain the predicted value of each pixel in Y.
[0153] In this process, if Y is used as the second coding block, in order to improve the prediction accuracy, C1 is first used as the first coding block to obtain the predicted value of each pixel in Y, and then C2 is used as the first coding block to obtain the predicted value of each pixel in Y. Finally, the two prediction results are combined to obtain the final prediction result.
[0154] The process of building the model can refer to the process of building the first target prediction model described above. For example, when building the model, combine... Figure 5 The model can be constructed according to the following formula: a=p1A+ p2B+ p3C+ p4D+ p5E+ p6A 2 +p7F.
[0155] See Figure 17 In another embodiment of this application, the encoding method further includes: S610: Obtain the target pixel value of each pixel in the first coding block, wherein the target pixel value is the target prediction value of the pixel or the reconstructed pixel value obtained based on the target prediction value, and take the first coding block or at least one sub-block in the first coding block as the current block, obtain the first prediction value of each pixel in the current block, and take the first prediction value as the target prediction value of the pixel.
[0156] Specifically, this step is the same as step S510, and you can refer to the relevant content above for details, which will not be repeated here.
[0157] S620: Construct a second prediction model based on the reconstructed pixel values of pixels in the first target reference block and the target pixel values of pixels in the first coding block, wherein the first target reference block corresponds to the first coding block.
[0158] S630: Based on the second prediction model and the second target reference block, obtain the predicted value of each pixel in the second coding block, wherein the second target reference block corresponds to the second coding block.
[0159] The first coding block is a luma block, and the second coding block is a chroma block corresponding to the first coding block; or, the first coding block is a chroma block, and the second coding block is either a luma block or a chroma block corresponding to the first coding block.
[0160] Combination Figure 16 In this embodiment, if Y is used as the first coding block, the process of Y' predicting Y is used to build a model, which is then applied to C1' to predict C1, and simultaneously applied to C2' to predict C2. At this time, the process can be used... Figure 18 To express.
[0161] If C1 is used as the first coding block, the process of C1' predicting C1 is used to build a model, and then the model is applied to C2' to predict C2.
[0162] If C2 is used as the first coding block, the process of C2' predicting C2 is used to build a model, and then the model is applied to C1' to predict C1.
[0163] If C1 is used as the first coding block, the process of C1' predicting C1 is used to build a model, and then the model is applied to Y' to predict Y.
[0164] If C2 is used as the first coding block, the process of C2' predicting C2 is used to build a model, and then the model is applied to Y' to predict Y.
[0165] In this process, if Y is used as the second coding block, in order to improve the prediction accuracy, C1 is first used as the first coding block to obtain the predicted value of each pixel in Y, and then C2 is used as the first coding block to obtain the predicted value of each pixel in Y. Finally, the two prediction results are combined to obtain the final prediction result.
[0166] The process of building the model can refer to the process of building the first target prediction model described above, and will not be repeated here. For example, when building the model, combine... Figure 5 The model can be constructed according to the following formula: a=p1A+ p2B+ p3C+ p4D+ p5E+ p6A 2 +p7F.
[0167] See Figure 19 In another embodiment of this application, the encoding method further includes: S710: Obtain the target pixel value of each pixel in the first coding block, wherein the target pixel value is the target prediction value of the pixel or the reconstructed pixel value obtained based on the target prediction value, and take the first coding block or at least one sub-block in the first coding block as the current block, obtain the first prediction value of each pixel in the current block, and take the first prediction value as the target prediction value of the pixel.
[0168] Specifically, this step is the same as step S510, and you can refer to the relevant content above for details, which will not be repeated here.
[0169] S720: Construct a third prediction model based on the reconstructed pixel values of the pixels in the first target reference block and the reconstructed pixel values of the pixels in the second target reference block, wherein the first target reference block corresponds to the first coding block and the second target reference block corresponds to the second coding block.
[0170] S730: Based on the third prediction model and the first coding block, obtain the predicted value of each pixel in the second coding block.
[0171] The first coding block is a luma block, and the second coding block is a chroma block corresponding to the first coding block; or, the first coding block is a chroma block, and the second coding block is either a luma block or a chroma block corresponding to the first coding block.
[0172] Specifically, in combination Figure 16 In this embodiment, if Y is used as the first coding block, and the process of Y'predicting C1' is used to build the model, then the model is applied to Y prediction C1; simultaneously, if the process of Y'predicting C2' is used to build the model, then the model is applied to Y prediction C2. Specifically, this process can be found in [reference needed]. Figure 20 .
[0173] If C1 is used as the first coding block, the process of C1' predicting C2' is used to build a model, and then the model is applied to C1 to predict C2.
[0174] Alternatively, if C2 is the first coding block, the process of C2' predicting C1' is used to build a model, and then that model is applied to C2 to predict C1.
[0175] Alternatively, if C1 is used as the first coding block, the process of C1' predicting Y' is used to build a model, and then the model is applied to C1 to predict Y.
[0176] Alternatively, if C2 is used as the first coding block, the process of C2' predicting Y' is used to build a model, and then that model is applied to C2 to predict Y'.
[0177] In this process, if Y is used as the second coding block, in order to improve the prediction accuracy, C1 is first used as the first coding block to obtain the predicted value of each pixel in Y, and then C2 is used as the first coding block to obtain the predicted value of each pixel in Y. Finally, the two prediction results are combined to obtain the final prediction result.
[0178] The process of building the model can refer to the process of building the first target prediction model described above, and will not be repeated here. For example, when building the model, combine... Figure 5 The model can be constructed according to the following formula: a=p1A+ p2B+ p3C+ p4D+ p5E+ p6F +p7G+p8H+p9I.
[0179] See Figure 21 , Figure 21 This is a flowchart illustrating one embodiment of the decoding method of this application, which includes: S810: Receives encoded data sent by the encoder.
[0180] S820: By decoding the encoded data, the predicted value of the current pixel in the current decoding block is obtained.
[0181] The final predicted value of the current pixel in the current decoding block is obtained by the encoding method in any of the above embodiments. For detailed steps, please refer to the relevant content, which will not be repeated here.
[0182] See Figure 22 , Figure 22 This is a schematic diagram of one embodiment of the encoder of this application. The encoder 300 includes a processor 310, a memory 320, and a communication circuit 330. The processor 310 is coupled to the memory 320 and the communication circuit 330 respectively. The memory 320 stores program data. The processor 310 executes the program data in the memory 320 to implement the steps of the encoding method in any of the above embodiments. The detailed steps can be found in the above embodiments and will not be repeated here.
[0183] The encoder 300 can be any device with algorithm processing capabilities, such as a computer or mobile phone, and there are no restrictions on it.
[0184] See Figure 23 , Figure 23 This is a schematic diagram of another embodiment of the encoder of this application. The encoder 400 includes an acquisition module 410, a construction module 420, and a prediction module 430 connected in sequence.
[0185] The acquisition module 410 is used to acquire the current template and reference template of the current block, wherein the current template includes multiple first pixel points reconstructed in the current frame, the reference template includes multiple second pixel points reconstructed in the reference frame, the current frame includes the current block, and the reference frame corresponds to the current frame.
[0186] The construction module 420 is used to construct a first target prediction model based on the first reconstructed pixel value of the first pixel and the corresponding multiple second reconstructed pixel values. The multiple second reconstructed pixel values include multiple values among the reconstructed pixel values of the pixels surrounding the first reference pixel and the reconstructed pixel values of the first reference pixel. The first reference pixel is the pixel in the reference template that corresponds to the first pixel.
[0187] The prediction module 430 is used to predict the pixel to be encoded in the current block using the first target prediction model to obtain the first predicted value of the pixel to be encoded.
[0188] When the encoder 400 is working, it executes the steps of the encoding method in any of the above embodiments. For details of the steps, please refer to the relevant content above, which will not be repeated here.
[0189] The encoder 400 can be any device with algorithm processing capabilities, such as a computer or mobile phone, and there are no restrictions on it.
[0190] See Figure 24 , Figure 24 This is a schematic diagram of one embodiment of the decoder of this application. The decoder 500 includes a processor 510, a memory 520, and a communication circuit 530. The processor 510 is coupled to the memory 520 and the communication circuit 530 respectively. The memory 520 stores program data. The processor 510 executes the program data in the memory 520 to implement the steps of the decoding method in any of the above embodiments. The detailed steps can be found in the above embodiments and will not be repeated here.
[0191] The decoder 500 can be any device with algorithm processing capabilities, such as a computer or mobile phone, without any restrictions.
[0192] See Figure 25 , Figure 25 This is a schematic diagram of another embodiment of the decoder in this application. The decoder 600 includes an acquisition module 610 and a decoding module 620.
[0193] The acquisition module 610 is used to receive the encoded data sent by the encoder.
[0194] The decoding module 620 is connected to the acquisition module 610 and is used to obtain the predicted value of the current pixel in the current decoding block by decoding the encoded data.
[0195] The predicted value of the current pixel in the current decoding block is obtained by processing the encoding method in any of the above embodiments. For details, please refer to the above content, which will not be repeated here.
[0196] When the decoder 600 is in operation, it executes the steps of the decoding method in any of the above embodiments. For details of the steps, please refer to the relevant content above, which will not be repeated here.
[0197] The decoder 600 can be any device with algorithm processing capabilities, such as a computer or a mobile phone, without any restrictions.
[0198] See Figure 26 , Figure 26 This is a schematic diagram of one embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 700 stores a computer program 710, which can be executed by a processor to implement the steps in any of the above methods.
[0199] Specifically, the computer-readable storage medium 700 can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or a device that can store the computer program 710. Alternatively, it can be a server that stores the computer program 710, which can send the stored computer program 710 to other devices for execution, or it can run the stored computer program 710 itself.
[0200] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An encoding method, characterized in that, The method includes: Obtain the current template and reference template of the current block, wherein the current template includes multiple first pixels that have been reconstructed in the current frame, the reference template includes multiple second pixels that have been reconstructed in the reference frame, the current frame includes the current block, and the reference frame corresponds to the current frame; A first target prediction model is constructed based on the first reconstructed pixel value of the first pixel and the corresponding plurality of second reconstructed pixel values. The plurality of second reconstructed pixel values include multiple values among the reconstructed pixel values of pixels surrounding the first reference pixel and the reconstructed pixel values of the first reference pixel. The first reference pixel is the pixel in the reference template that corresponds to the first pixel. The first target prediction model is used to predict the pixel to be encoded in the current block to obtain the first predicted value of the pixel to be encoded.
2. The method according to claim 1, characterized in that, The step of constructing a first target prediction model based on the first reconstructed pixel value of the first pixel and the corresponding plurality of second reconstructed pixel values includes: Based on the first reconstructed pixel value of the first pixel and the corresponding plurality of second reconstructed pixel values, a plurality of model parameters are determined; The step of using the first target prediction model to predict the pixel to be encoded in the current block and obtaining the first predicted value of the pixel to be encoded includes: The first predicted value of the pixel to be encoded is determined based on multiple target values and multiple model parameters, wherein the multiple target values correspond one-to-one with the multiple model parameters; The plurality of target values include a plurality of first target reconstructed pixel values and a plurality of calculated values corresponding to the pixel to be encoded. The calculated values are obtained by performing calculations on the first target reconstructed pixel values. The plurality of first target reconstructed pixel values include the reconstructed pixel values of the first target pixel and the reconstructed pixel values of the pixels surrounding the first target pixel. The first target pixel is the pixel in the reference block that corresponds to the pixel to be encoded.
3. The method according to claim 2, characterized in that, The step of determining the first predicted value of the pixel to be encoded based on multiple target values and multiple model parameters includes: Each of the target values is multiplied by its corresponding model parameter to obtain a plurality of first products; The first predicted value of the pixel to be encoded is obtained by summing all the first products.
4. The method according to claim 1, characterized in that, The method further includes: After constructing the first target prediction model into multiple different candidate prediction models in sequence, the first prediction value of all the pixels to be encoded under each candidate prediction model is determined respectively. Based on the first predicted value of all the pixels to be encoded under each candidate prediction model, determine the cost value corresponding to each candidate prediction model respectively; The candidate prediction model with the lowest cost value is determined as the final prediction model; The first predicted value of each pixel to be encoded under the final prediction model is determined as the final predicted value of each pixel to be encoded.
5. The method according to claim 4, characterized in that, After determining the first predicted value of each pixel to be encoded under the final prediction model as the final predicted value of each pixel to be encoded, the method further includes: Generate the first syntactic element; In response to the first syntactic element instructing the execution of the step of determining the first predicted value of all the pixels to be encoded under each of the candidate prediction models after sequentially constructing the first target prediction model into multiple different candidate prediction models, up to the step of determining the first predicted value of each pixel to be encoded under the final prediction model as the final predicted value of each pixel to be encoded, a second syntactic element is generated, the second syntactic element being used to indicate the final prediction model.
6. The method according to claim 1, characterized in that, The step of constructing a first target prediction model based on the first reconstructed pixel value of the first pixel and the corresponding plurality of second reconstructed pixel values includes: Based on the reconstructed pixel value of the second pixel, the second pixel is classified to obtain multiple second pixel classes; In the current template, a first pixel class corresponding to each second pixel class is determined, wherein the first pixel in the first pixel class corresponds one-to-one with the second pixel in the corresponding second pixel class, and the first pixel corresponding to the second pixel is the pixel in the current template that has the same position as the second pixel. The first target prediction model corresponding to each first pixel class is constructed based on the first reconstructed pixel value of the first pixel point in each first pixel class and the corresponding plurality of second reconstructed pixel values. The step of using the first target prediction model to predict the pixel to be encoded in the current block and obtaining the first predicted value of the pixel to be encoded includes: Based on the reconstructed pixel value of the third pixel in the reference block, the third pixel is classified to obtain multiple third pixel classes. The rules for classifying the third pixel are the same as those for classifying the second pixel. The first target prediction model corresponding to the pixel to be encoded is used to predict the pixel to be encoded, and the first predicted value of the pixel to be encoded is obtained. Wherein, the first target prediction model corresponding to the pixel to be encoded is the first target prediction model corresponding to the first target pixel class among a plurality of first pixel classes, the first target pixel class corresponds to the second target pixel class among a plurality of second pixel classes, the second target pixel class is the same as the third target pixel class among a plurality of third pixel classes, and the third target pixel class includes the pixel in the reference block corresponding to the pixel to be encoded.
7. The method according to claim 6, characterized in that, The step of classifying the second pixel based on the reconstructed pixel value of the second pixel to obtain multiple second pixel classes includes: Determine the average value of the reconstructed pixel value of the second pixel; Based on the average value, the second pixel is classified into two second pixel classes, wherein one second pixel class includes second pixel points whose reconstructed pixel values do not exceed the average value, and the other second pixel class includes second pixel points whose reconstructed pixel values exceed the average value. The step of classifying the third pixel based on the reconstructed pixel value of the third pixel in the reference block to obtain multiple third pixel classes includes: Based on the average value, the third pixel is classified into two third pixel classes. One third pixel class includes third pixel points whose reconstructed pixel values do not exceed the average value, and the other third pixel class includes third pixel points whose reconstructed pixel values exceed the average value.
8. The method according to claim 4, characterized in that, The steps of obtaining the current template and reference template of the current block include: A reference block is determined in the reference frame; The current template and the reference template are determined based on the reconstructed pixels in the same target region outside the current block and outside the reference block, respectively. The target region outside the current block includes at least one of a first sub-region, a second sub-region, a third sub-region, a fourth sub-region, and a fifth sub-region outside the current block. The first sub-region is located on the first side of the current block and its two ends are flush with the two ends of the current block, the second sub-region is located on the second side of the current block and its two ends are flush with the two ends of the current block, the third sub-region connects the first sub-region and the second sub-region, the fourth sub-region is located on the side of the first sub-region away from the third sub-region, and the fifth sub-region is located on the side of the second sub-region away from the third sub-region.
9. The method according to claim 8, characterized in that, The step of determining the current template and the reference template based on the reconstructed pixels in the same target region outside the current block and outside the reference block, respectively, includes: Determine a first effective region outside the current block, wherein the first effective region is located in the target region outside the current block, and the first effective region includes reconstructed pixels that can be acquired outside the current block; A second effective region is determined outside the reference block, wherein the second effective region is located in the target region outside the reference block, and the second effective region includes reconstructed pixels that can be acquired outside the reference block; Determine the common area of the first valid area and the second valid area; The current template and the reference template are determined based on the reconstructed pixels in the common area outside the current block and outside the reference block, respectively.
10. The method according to claim 8, characterized in that, The method further includes: The target region is sequentially used as multiple different candidate regions to determine the first predicted value of each pixel to be encoded in each of the candidate regions. The cost value corresponding to each candidate region is determined based on the first predicted value of each pixel to be encoded in each candidate region. The candidate region with the lowest cost value is determined as the final region; The first predicted value of each pixel to be encoded in the final region is determined as the final predicted value of each pixel to be encoded.
11. The method according to claim 10, characterized in that, After determining the first predicted value of each of the pixels to be encoded in the final region as the final predicted value for each of the pixels to be encoded, the method further includes: Generate the first syntactic element; In response to the first syntactic element instructing the execution of the step of determining the first predicted value of all the pixels to be encoded under each of the candidate prediction models after sequentially constructing the first target prediction model into multiple different candidate prediction models, up to the step of determining the first predicted value of each pixel to be encoded under the final prediction model as the final predicted value of each pixel to be encoded, a second syntactic element is generated, the second syntactic element being used to indicate the final region.
12. The method according to claim 1, characterized in that, The method further includes: A second target prediction model is constructed based on the first reconstructed pixel value of the first pixel and the reconstructed pixel value of the first reference pixel. Using the second target prediction model, each pixel to be encoded in the current block is predicted to obtain a second predicted value for each pixel to be encoded. The first generation value is determined based on the first predicted value of all the pixels to be encoded in the current block; The second generation value is determined based on the second predicted value of all the pixels to be encoded in the current block; In response to the first generation value being less than the second generation value, the first predicted value of each of the pixels to be encoded is determined as the final predicted value of the pixel to be encoded; otherwise, the second predicted value of each of the pixels to be encoded is determined as the final predicted value of the pixel to be encoded.
13. The method according to claim 12, characterized in that, After determining the first predicted value of each pixel to be encoded as its final predicted value in response to the first generation value being less than the second generation value, and otherwise determining the second predicted value of each pixel to be encoded as its final predicted value, the method further includes: Generate the first syntactic element; In response to the instruction of the first syntactic element to perform the step of constructing a second target prediction model based on the first reconstructed pixel value of the first pixel and the reconstructed pixel value of the first reference pixel, up to the step of determining the first predicted value of each pixel to be encoded as the final predicted value of the pixel to be encoded respectively, in response to the first generation value being less than the second generation value, otherwise determining the second predicted value of each pixel to be encoded as the final predicted value of the pixel to be encoded, a second syntactic element is generated, wherein, in response to the first generation value being less than the second generation value, the second syntactic element is a first value, otherwise the second syntactic element is a second value.
14. The method according to claim 1, characterized in that, The method further includes: Obtain the target pixel value of each pixel in the first coding block, wherein the target pixel value is the target predicted value of the pixel or the reconstructed pixel value obtained based on the target predicted value, and take the first coding block or at least one sub-block in the first coding block as the current block, obtain the first predicted value of each pixel in the current block, and take the first predicted value as the target predicted value of the pixel; A first prediction model is constructed based on the reconstructed pixel values of the pixels in the current template of the first coding block and the reconstructed pixel values of the pixels in the current template of the second coding block. Based on the first prediction model and the first coding block, the predicted value of each pixel in the second coding block is obtained; Wherein, the first coding block is a luma block, and the second coding block is a chroma block corresponding to the first coding block; or, the first coding block is a chroma block, and the second coding block is a luma block or a chroma block corresponding to the first coding block.
15. The method according to claim 1, characterized in that, The method further includes: Obtain the target pixel value of each pixel in the first coding block, wherein the target pixel value is the target predicted value of the pixel or the reconstructed pixel value obtained based on the target predicted value, and take the first coding block or at least one sub-block in the first coding block as the current block, obtain the first predicted value of each pixel in the current block, and take the first predicted value as the target predicted value of the pixel; A second prediction model is constructed based on the reconstructed pixel values of pixels in the first target reference block and the target pixel values of pixels in the first coding block, wherein the first target reference block corresponds to the first coding block; Based on the second prediction model and the second target reference block, the predicted value of each pixel in the second coding block is obtained, wherein the second target reference block corresponds to the second coding block; Wherein, the first coding block is a luma block, and the second coding block is a chroma block corresponding to the first coding block; or, the first coding block is a chroma block, and the second coding block is a luma block or a chroma block corresponding to the first coding block.
16. The method according to claim 1, characterized in that, The method further includes: Obtain the target pixel value of each pixel in the first coding block, wherein the target pixel value is the target predicted value of the pixel or the reconstructed pixel value obtained based on the target predicted value, and take the first coding block or at least one sub-block in the first coding block as the current block, obtain the first predicted value of each pixel in the current block, and take the first predicted value as the target predicted value of the pixel; A third prediction model is constructed based on the reconstructed pixel values of pixels in the first target reference block and the reconstructed pixel values of pixels in the second target reference block, wherein the first target reference block corresponds to the first coding block and the second target reference block corresponds to the second coding block; Based on the third prediction model and the first coding block, the predicted value of each pixel in the second coding block is obtained; Wherein, the first coding block is a luma block, and the second coding block is a chroma block corresponding to the first coding block; or, the first coding block is a chroma block, and the second coding block is a luma block or a chroma block corresponding to the first coding block.
17. A decoding method, characterized in that, The method includes: Receive encoded data sent by the encoder; By decoding the encoded data, the predicted value of the current pixel in the current decoding block is obtained; Wherein, the predicted value of the current pixel in the current decoding block is obtained by processing using the encoding method described in any one of claims 1 to 16.
18. An encoder, characterized in that, The encoder includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit. The memory stores program data. The processor executes the program data in the memory to implement the steps of the method as described in any one of claims 1-16.
19. A decoder, characterized in that, The decoder includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit. The memory stores program data. The processor executes the program data in the memory to implement the steps of the method as described in claim 17.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the steps of the method as described in any one of claims 1-17.
Citation Information
Patent Citations
Image component prediction method, encoder, decoder and storage medium
CN113676732A
Image component prediction method, encoder, decoder, and storage medium
CN113992916A