Image Component Prediction Method, Apparatus, and Computer Storage Medium
By optimizing the selection and processing of adjacent reference pixel points in the image component prediction model, the accuracy and robustness of the prediction model are improved, the problem of inaccurate prediction models in the prior art is solved, and the encoding and codec performance of video images is improved.
Patent Information
- Application Number
- CN201980084457.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2039-07-01
AI Technical Summary
In the existing video encoding technology, in the parameter derivation of the image component prediction model, the selection of adjacent reference pixel points is unreasonable, resulting in inaccurate prediction model and reducing the encoding and codec performance of video images.
By obtaining N adjacent reference pixel points corresponding to the image components to be predicted in the encoded block in the video image, calculating their mean, grouping, determining two fitting points, deriving model parameters, and building a more accurate prediction model.
It improves the robustness of prediction, makes the built prediction model more accurate, and improves the encoding and decoding prediction performance of video images.
Smart Images

Figure CN113196762B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of video coding and decoding, and in particular, to an image component prediction method, apparatus, and computer storage medium. Background Art
[0002] With the improvement of people's requirements for video display quality, new video application forms such as high-definition and ultra-high-definition videos have emerged as the times require. H.265 / High Efficiency Video Coding (HEVC) can no longer meet the rapidly developing needs of video applications. The Joint Video Exploration Team (JVET) has proposed the next-generation video coding standard H.266 / Versatile Video Coding (VVC), and its corresponding test model is the VVC Test Model (VTM) of the reference software test platform for VVC.
[0003] In VTM, an image component prediction method based on a prediction model has been integrated. Through this prediction model, the chrominance component can be predicted from the luminance component of the current coding block (CB). However, when constructing the prediction model, due to the unreasonable selection of adjacent reference pixel points for model parameter derivation, the prediction model is inaccurate, reducing the coding and decoding prediction performance of video images. Summary of the Invention
[0004] The embodiments of the present application provide an image component prediction method, apparatus, and computer storage medium. By optimizing the fitting points used for model parameter derivation, the robustness of prediction can be improved, making the constructed prediction model more accurate, and also capable of enhancing the coding and decoding prediction performance of video images.
[0005] The technical solution of the embodiments of the present application can be implemented as follows:
[0006] In a first aspect, the embodiments of the present application provide an image component prediction method, and the method includes:
[0007] Obtain N adjacent reference pixel points corresponding to the image component to be predicted of the coding block in the video image; where the N adjacent reference pixel points are reference pixel points adjacent to the coding block, and N is a preset integer value;
[0008] Calculate the mean value of the N adjacent reference pixel points to obtain a first mean point;
[0009] Group the second reference pixel set through the first mean point to obtain a first reference pixel subset and a second reference pixel subset;
[0010] Determine two fitting points based on the first reference pixel subset and the second reference pixel subset;
[0011] Determine model parameters based on the two fitting points, and obtain a prediction model corresponding to the image component to be predicted according to the model parameters; wherein, the prediction model is used to implement prediction processing on the image component to be predicted to obtain a predicted value corresponding to the image component to be predicted.
[0012] In a second aspect, an embodiment of the present application provides an image component prediction device, which includes: an acquisition unit, a calculation unit, a grouping unit, a determination unit, and a prediction unit, wherein,
[0013] The acquisition unit is configured to acquire N adjacent reference pixel points corresponding to an image component to be predicted of an encoded block in a video image; wherein, the N adjacent reference pixel points are reference pixel points adjacent to the encoded block, and N is a preset integer value;
[0014] The calculation unit is configured to calculate the mean value of the N adjacent reference pixel points to obtain a first mean point;
[0015] The grouping unit is configured to group the second reference pixel set through the first mean point to obtain a first reference pixel subset and a second reference pixel subset;
[0016] The determination unit is configured to determine two fitting points based on the first reference pixel subset and the second reference pixel subset;
[0017] The prediction unit is configured to determine model parameters based on the two fitting points, and obtain a prediction model corresponding to the image component to be predicted according to the model parameters; wherein, the prediction model is used to implement prediction processing on the image component to be predicted to obtain a predicted value corresponding to the image component to be predicted.
[0018] In a third aspect, an embodiment of the present application provides an image component prediction device, which includes: a memory and a processor;
[0019] The memory is used to store a computer program that can run on the processor;
[0020] The processor is configured to execute the method described in the first aspect when running the computer program.
[0021] In a fourth aspect, an embodiment of the present application provides a computer storage medium, which stores an image component prediction program, and when the image component prediction program is executed by at least one processor, the method described in the first aspect is implemented.
[0022] An embodiment of the present application provides an image component prediction method, apparatus, and computer storage medium. First, N adjacent reference pixel points corresponding to the image component to be predicted of an encoded block in a video image are obtained. The N adjacent reference pixel points are reference pixel points adjacent to the encoded block, and N is a preset integer value. The mean value of the N adjacent reference pixel points is calculated to obtain a first mean point. Then, the second reference pixel set is grouped by the first mean point to obtain a first reference pixel subset and a second reference pixel subset, and two fitting points are determined based on the first reference pixel subset and the second reference pixel subset. Then, model parameters are determined based on the two fitting points, and a prediction model corresponding to the image component to be predicted is obtained according to the model parameters to obtain a predicted value corresponding to the image component to be predicted. In this way, since the first mean point is obtained by taking the mean value of a preset number of adjacent reference pixel points in the second reference pixel set, the preset number of adjacent reference pixel points are grouped and divided according to the first mean point, and then two fitting points for deriving the model parameters are determined, which can improve the robustness of the prediction. That is to say, by optimizing the fitting points used for deriving the model parameters, the constructed prediction model is made more accurate, and the encoding and decoding prediction performance of the video image is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A distribution diagram of an effective adjacent region provided by a related technical solution;
[0024] Figure 2 A distribution diagram of a selection region in three modes provided by a related technical solution;
[0025] Figure 3 A flowchart of a traditional solution for deriving model parameters provided by a related technical solution;
[0026] Figure 4 A schematic diagram of a prediction model under a traditional solution provided by a related technical solution;
[0027] Figure 5 Another schematic diagram of a prediction model under a traditional solution provided by a related technical solution;
[0028] Figure 6 A schematic block diagram of a video encoding system provided by an embodiment of the present application;
[0029] Figure 7 A schematic block diagram of a video decoding system provided by an embodiment of the present application;
[0030] Figure 8 A flowchart of an image component prediction method provided by an embodiment of the present application;
[0031] Figure 9ASchematic diagram of selecting adjacent reference pixel points in the INTRA_LT_CCLM mode provided by an embodiment of the present application;
[0032] Figure 9B Schematic diagram of selecting adjacent reference pixel points in the INTRA_L_CCLM mode provided by an embodiment of the present application;
[0033] Figure 9C Schematic diagram of selecting adjacent reference pixel points in the INTRA_T_CCLM mode provided by an embodiment of the present application;
[0034] Figure 10 Flow chart of a model parameter derivation scheme provided by an embodiment of the present application;
[0035] Figure 11 Schematic diagram of comparison of prediction models between the solution of the present application and traditional solutions provided by an embodiment of the present application;
[0036] Figure 12 Another schematic diagram of comparison of prediction models between the solution of the present application and traditional solutions provided by an embodiment of the present application;
[0037] Figure 13 Another schematic diagram of comparison of prediction models between the solution of the present application and traditional solutions provided by an embodiment of the present application;
[0038] Figure 14 Schematic diagram of the composition structure of an image component prediction device provided by an embodiment of the present application;
[0039] Figure 15 Schematic diagram of the specific hardware structure of an image component prediction device provided by an embodiment of the present application;
[0040] Figure 16 Schematic diagram of the composition structure of an encoder provided by an embodiment of the present application;
[0041] Figure 17 Schematic diagram of the composition structure of a decoder provided by an embodiment of the present application. Detailed implementation manners
[0042] In order to understand the features and technical content of the embodiments of the present application in more detail, the implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The attached drawings are only for reference and explanation, and are not used to limit the embodiments of the present application.
[0043] In a video image, generally, a coding block is characterized by a first image component, a second image component, and a third image component; among them, these three image components are respectively a luminance component, a blue chrominance component, and a red chrominance component. Specifically, the luminance component is usually represented by the symbol Y, the blue chrominance component is usually represented by the symbol Cb or U, and the red chrominance component is usually represented by the symbol Cr or V; thus, the video image can be represented in the YCbCr format or the YUV format.
[0044] In the embodiments of the present application, the first image component may be a luminance component, the second image component may be a blue chrominance component, and the third image component may be a red chrominance component, but the embodiments of the present application do not make specific limitations.
[0045] In the current video image or video encoding / decoding process, for the cross-component prediction technology, it mainly includes the Cross-component Linear Model Prediction (CCLM) mode and the Multi-Directional Linear Model Prediction (MDLM) mode. Whether it is the model parameters derived according to the CCLM mode or the model parameters derived according to the MDLM mode, the corresponding prediction models can all achieve the prediction between image components such as from the first image component to the second image component, from the second image component to the first image component, from the first image component to the third image component, from the third image component to the first image component, from the second image component to the third image component, or from the third image component to the second image component.
[0046] Taking the prediction from the first image component to the second image component as an example, in order to reduce the redundancy between the first image component and the second image component, the CCLM mode is used in the VTM. At this time, the first image component and the second image component are of the same coding block, that is, the prediction value of the second image component is constructed based on the reconstructed value of the first image component of the same coding block, as shown in Equation (1).
[0047] Pred C [i,j] = α · Rec L [i,j] + β(1)
[0048] Among them, i and j represent the position coordinates of the pixel points in the coding block, i represents the horizontal direction, and j represents the vertical direction. Pred C [i,j] represents the prediction value of the second image component corresponding to the pixel point with the position coordinates [i,j] in the coding block, and Pred L [i,j] represents the reconstructed value of the first image component corresponding to the pixel point with the position coordinates [i,j] (after downsampling) in the same coding block. α and β represent the model parameters.
[0049] For a coding block, its adjacent regions may include a left adjacent region, an upper adjacent region, a lower-left adjacent region, and an upper-right adjacent region. In VVC, three cross-component linear model prediction modes may be included, namely: an intra-frame CCLM mode adjacent to the left and upper sides (which can be represented by the INTRA_LT_CCLM mode), an intra-frame CCLM mode adjacent to the left and lower-left sides (which can be represented by the INTRA_L_CCLM mode), and an intra-frame CCLM mode adjacent to the upper and upper-right sides (which can be represented by the INTRA_T_CCLM mode). In these three modes, a preset number (such as 4) of adjacent reference pixel points can be selected for the derivation of the model parameters α and β, and the biggest difference among these three modes lies in that the selection regions corresponding to the adjacent reference pixel points for the derivation of the model parameters α and β are different.
[0050] Specifically, for the coding block corresponding to the second image component with a size of W×H, assuming that the upper selection region corresponding to the adjacent reference pixel points is W', and the left selection region corresponding to the adjacent reference pixel points is H'; thus,
[0051] For the INTRA_LT_CCLM mode, the adjacent reference pixel points can be selected from the upper adjacent region and the left adjacent region, that is, W' = W, H' = H;
[0052] For the INTRA_L_CCLM mode, the adjacent reference pixel points can be selected from the left adjacent region and the lower-left adjacent region, that is, H' = W + H, and W' = 0 is set;
[0053] For the INTRA_T_CCLM mode, the adjacent reference pixel points can be selected from the upper adjacent region and the upper-right adjacent region, that is, W' = W + H, and H' = 0 is set.
[0054] It should be noted that in the latest reference software VTM5.0 of VVC, at most only W-range pixel points are stored in the upper-right adjacent region, and at most only H-range pixel points are stored in the lower-left adjacent region; therefore, although the range of the selection regions of the INTRA_L_CCLM mode and the INTRA_T_CCLM mode is defined as W + H, in actual applications, the selection region of the INTRA_L_CCLM mode will be restricted within H + H, and the selection region of the INTRA_T_CCLM mode will be restricted within W + W; thus,
[0055] For the INTRA_L_CCLM mode, the adjacent reference pixel points can be selected from the left adjacent region and the lower-left adjacent region, H' = min{W + H, H + H};
[0056] For the INTRA_T_CCLM mode, adjacent reference pixel points can be selected from the upper adjacent area and the upper-right adjacent area, and W' = min{W + H, W + W}.
[0057] See Figure 1 , which shows a schematic diagram of the distribution of an effective adjacent area provided by a related technical solution. In Figure 1 , the left adjacent area, the lower-left adjacent area, the upper adjacent area, and the upper-right adjacent area are all effective. On the basis of Figure 1 , the selection areas for three modes are as shown in Figure 2 . Among them, in Figure 2 , (a) shows the selection area of the INTRA_LT_CCLM mode, including the left adjacent area and the upper adjacent area; (b) shows the selection area of the INTRA_L_CCLM mode, including the left adjacent area and the lower-left adjacent area; (c) shows the selection area of the INTRA_T_CCLM mode, including the upper adjacent area and the upper-right adjacent area. In this way, after determining the selection areas of the three modes, reference points for model parameter derivation can be selected within the selection areas. The reference points selected in this way can be called adjacent reference pixel points. Generally, the number of adjacent reference pixel points is at most 4; and for a coded block with a determined size of W×H, the positions of its adjacent reference pixel points are generally determined.
[0058] After obtaining a preset number of adjacent reference pixel points, currently, the model parameters are calculated according to the flow schematic diagram of the traditional model parameter derivation scheme shown in Figure 3 . According to the flow shown in Figure 3 , assuming that the preset number is 4, this flow may include:
[0059] S301: Obtain 4 adjacent reference pixel points;
[0060] S302: Obtain two adjacent reference pixel points with larger luminance components and two adjacent reference pixel points with smaller luminance components through 4 comparisons;
[0061] S303: Calculate the mean points corresponding to the two points with larger luminance components and the mean points corresponding to the two points with smaller luminance components;
[0062] S304: Use the two mean points as two fitting points to derive model parameters;
[0063] S305: Perform prediction processing on the chrominance components according to the constructed prediction model.
[0064] It should be noted that a prediction model is constructed using the principle of "two points determine a straight line"; these two points can be called fitting points. In the current traditional solutions, first, after obtaining 4 adjacent reference pixel points, two adjacent reference pixel points with larger luminance components and two adjacent reference pixel points with smaller luminance components are obtained through 4 comparisons; then, based on the two adjacent reference pixel points with larger luminance components, a mean point (which can be represented by mean max is obtained), and based on the two adjacent reference pixel points with smaller luminance components, another mean point (which can be represented by mean min is obtained), resulting in two mean points mean max and mean min ; then mean max and mean min are used as two fitting points to derive model parameters (which can be represented by α and β), and finally, a prediction model is constructed based on the model parameters, and chrominance component prediction processing is performed according to this prediction model.
[0065] When deriving the model parameters, taking 4 adjacent reference pixel points as an example, currently, the mean values obtained from 2 larger adjacent reference pixel points and the mean values obtained from 2 smaller adjacent reference pixel points are used as two fitting points to derive the model parameters. When these 4 adjacent reference pixel points are relatively evenly distributed, as shown in the schematic diagram of the prediction model under the traditional solution in Figure 4 ; where the abscissa of the coordinate axis represents the luminance value (which can be represented by Luma), the ordinate of the coordinate axis represents the chrominance value (which can be represented by Chroma), the 4 black dots are 4 adjacent reference pixel points, the 2 gray dots are respectively the mean points corresponding to 2 larger adjacent reference pixel points and the mean points corresponding to 2 smaller adjacent reference pixel points among these 4 adjacent reference pixel points (i.e., the two fitting points), and the gray diagonal line represents the prediction model constructed based on these two fitting points; the gray dotted line is the prediction model fitted by these 4 adjacent reference pixel points using the Least Mean Square (LMS) algorithm; it can be seen from Figure 4 that the gray diagonal line is relatively close to the gray dotted line, that is, the prediction model constructed by the traditional solution can relatively accurately fit the distribution of these 4 adjacent reference pixel points.
[0066] However, when these 4 adjacent reference pixel points are unevenly distributed, the prediction model constructed by the traditional solution cannot accurately fit the distribution of these 4 adjacent reference pixel points, as shown in the schematic diagram of the prediction model under another traditional solution in Figure 5 ; it can be seen from Figure 5It can be seen that there is a certain deviation between the gray diagonal line and the gray dotted line, that is, the prediction model constructed by the traditional solution cannot fit well with the distribution of these 4 adjacent reference pixel points. That is to say, the current traditional solution lacks robustness in the process of deriving model parameters and cannot fit well when the 4 adjacent reference pixel points are unevenly distributed.
[0067] To improve the robustness of CCLM prediction and at the same time improve the encoding and decoding performance, the embodiments of the present application provide an image component prediction method, which obtains N adjacent reference pixel points corresponding to the image component to be predicted of an encoded block in a video image. The N adjacent reference pixel points are reference pixel points adjacent to the encoded block, and N is a preset integer value; calculate the mean value of the N adjacent reference pixel points to obtain a first mean point; group the second reference pixel set through the first mean point to obtain a first reference pixel subset and a second reference pixel subset, and determine two fitting points based on the first reference pixel subset and the second reference pixel subset; then determine the model parameters based on the two fitting points, and obtain the prediction model corresponding to the image component to be predicted according to the model parameters to obtain the predicted value corresponding to the image component to be predicted; in this way, since the first mean point is obtained by calculating the mean value of a preset number of adjacent reference pixel points in the second reference pixel set, the preset number of adjacent reference pixel points are grouped and divided according to the first mean point, and then two fitting points for deriving the model parameters are determined, which can improve the robustness of the prediction; that is to say, by optimizing the fitting points used in the derivation of the model parameters, the constructed prediction model is made more accurate, and the encoding and decoding prediction performance of the video image is improved.
[0068] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0069] See Figure 6 , which shows an example of the block diagram of a video coding system provided by the embodiments of the present application; as Figure 6As shown, the video encoding system 600 includes a transform and quantization unit 601, an intra prediction estimation unit 602, an intra prediction unit 603, a motion compensation unit 604, a motion estimation unit 605, an inverse transform and inverse quantization unit 606, a filter control analysis unit 607, a filtering unit 608, an encoding unit 609, a decoded image buffer unit 610, etc. Among them, the filtering unit 608 can implement deblocking filtering and sample adaptive offset (SAO) filtering, and the encoding unit 609 can implement header information encoding and context-based adaptive binary arithmetic coding (CABAC). For the input original video signal, a video coding block can be obtained through the division of a coding tree unit (CTU). Then, the residual pixel information obtained after intra-frame or inter-frame prediction is transformed by the transform and quantization unit 601 for the video coding block, including transforming the residual information from the pixel domain to the transform domain and quantizing the obtained transform coefficients to further reduce the bit rate; the intra prediction estimation unit 602 and the intra prediction unit 603 are used for intra-frame prediction of the video coding block; specifically, the intra prediction estimation unit 602 and the intra prediction unit 603 are used to determine the intra-frame prediction mode to be used for encoding the video coding block; the motion compensation unit 604 and the motion estimation unit 605 are used to perform inter-frame predictive coding of the received video coding block relative to one or more blocks in one or more reference frames to provide temporal prediction information; the motion estimation performed by the motion estimation unit 605 is a process of generating a motion vector, and the motion vector can estimate the motion of the video coding block, and then the motion compensation unit 604 performs motion compensation based on the motion vector determined by the motion estimation unit 605; after determining the intra-frame prediction mode, the intra prediction unit 603 is further used to provide the selected intra-frame prediction data to the encoding unit 609, and the motion estimation unit 605 also sends the calculated and determined motion vector data to the encoding unit 609; in addition, the inverse transform and inverse quantization unit 606 is used for the reconstruction of the video coding block, reconstructing the residual block in the pixel domain, and the reconstructed residual block removes block effect artifacts through the filter control analysis unit 607 and the filtering unit 608, and then adds the reconstructed residual block to a predictive block in the frame of the decoded image buffer unit 610 to generate a reconstructed video coding block; the encoding unit 609 is used to encode various coding parameters and the quantized transform coefficients. In the CABAC-based encoding algorithm, the context can be based on adjacent coding blocks and can be used to encode information indicating the determined intra-frame prediction mode, and output the bitstream of the video signal; while the decoded image buffer unit 610 is used to store the reconstructed video coding block for prediction reference.As the video image encoding progresses, new reconstructed video coding blocks are continuously generated, and these reconstructed video coding blocks are all stored in the decoded picture buffer unit 610.
[0070] See Figure 7 , which shows an example of the block diagram of a video decoding system provided by an embodiment of the present application; as Figure 7 shown, the video decoding system 700 includes a decoding unit 701, an inverse transform and inverse quantization unit 702, an intra prediction unit 703, a motion compensation unit 704, a filtering unit 705, a decoded picture buffer unit 706, etc. Among them, the decoding unit 701 can implement header information decoding and CABAC decoding, and the filtering unit 705 can implement deblocking filtering and SAO filtering. After the input video signal undergoes Figure 6 encoding processing, the bitstream of the video signal is output; this bitstream is input into the video decoding system 700, first passing through the decoding unit 701 to obtain the decoded transform coefficients; the inverse transform and inverse quantization unit 702 processes the transform coefficients to generate a residual block in the pixel domain; the intra prediction unit 703 can be used to generate prediction data for the current video decoding block based on the determined intra prediction mode and data from previously decoded blocks in the current frame or picture; the motion compensation unit 704 determines the prediction information for the video decoding block by analyzing the motion vector and other associated syntax elements, and uses this prediction information to generate a predictive block for the video decoding block being decoded; by summing the residual block from the inverse transform and inverse quantization unit 702 and the corresponding predictive block generated by the intra prediction unit 703 or the motion compensation unit 704, a decoded video block is formed; the decoded video signal passes through the filtering unit 705 to remove block effect artifacts, which can improve the video quality; then the decoded video block is stored in the decoded picture buffer unit 706, and the decoded picture buffer unit 706 stores the reference image for subsequent intra prediction or motion compensation, and is also used for the output of the video signal, that is, the original video signal is recovered.
[0071] The image component prediction method in the embodiments of the present application is mainly applied to the intra prediction unit 603 part as Figure 6 shown and the part as Figure 7The intra prediction unit 703 part shown is specifically applied to the CCLM prediction part in intra prediction. That is to say, the image component prediction method in the embodiments of the present application can be applied to a video coding system, a video decoding system, or even both a video coding system and a video decoding system at the same time, but the embodiments of the present application do not make specific limitations. When this method is applied to the intra prediction unit 603 part, the "encoded block in the video image" specifically refers to the current encoded block in intra prediction; when this method is applied to the intra prediction unit 703 part, the "encoded block in the video image" specifically refers to the current decoded block in intra prediction.
[0072] Based on the above Figure 6 or Figure 7 application scenario example, see Figure 8 , which shows a schematic flowchart of an image component prediction method provided by the embodiments of the present application. As Figure 8 shown, the method may include:
[0073] S801: Obtain N adjacent reference pixel points corresponding to the image component to be predicted of the encoded block in the video image;
[0074] It should be noted that these N adjacent reference pixel points are reference pixel points adjacent to the encoded block, and N is a preset integer value, which can also be referred to as the preset quantity. Among them, the video image can be divided into multiple encoded blocks, and each encoded block can include a first image component, a second image component, and a third image component. The encoded block in the embodiments of the present application is the current block to be encoded in the video image. When the first image component needs to be predicted through a prediction model, the image component to be predicted is the first image component; when the second image component needs to be predicted through a prediction model, the image component to be predicted is the second image component; when the third image component needs to be predicted through a prediction model, the image component to be predicted is the third image component.
[0075] In this way, a preset number of adjacent reference pixel points can be obtained from the reference pixel points adjacent to the encoded block to determine the fitting points for model parameter derivation. In addition, the value of N can generally be 4, but the embodiments of the present application do not make specific limitations.
[0076] In some embodiments, for S801, the obtaining of N adjacent reference pixel points corresponding to the image component to be predicted of the encoded block in the video image may include:
[0077] S801-1: Obtain a first reference pixel set corresponding to the image component to be predicted of the encoded block in the video image;
[0078] It should be noted that when the left adjacent region, the lower left adjacent region, the upper adjacent region, and the upper right adjacent region are all valid regions, for the INTRA_LT_CCLM mode, the first reference pixel set is composed of adjacent reference pixel points in the left adjacent region and the upper adjacent region of the coding block, as shown in Figure 2 (a) of Figure 2 ; for the INTRA_L_CCLM mode, the first reference pixel set is composed of adjacent reference pixel points in the left adjacent region and the lower left adjacent region of the coding block, as shown in Figure 2 (b) of
[0079] ; for the INTRA_T_CCLM mode, the first reference pixel set is composed of adjacent reference pixel points in the upper adjacent region and the upper right adjacent region of the coding block, as shown in Figure 2 (c) of
[0079] In some embodiments, optionally, for S801-1, the obtaining of the first reference pixel set corresponding to the image component to be predicted of the coding block in the video image may include:
[0080] Obtaining reference pixel points adjacent to at least one side of the coding block; wherein, the at least one side includes the left side of the coding block and / or the upper side of the coding block;
[0081] Based on the reference pixel points, forming the first reference pixel set corresponding to the image component to be predicted.
[0082] It should be noted that at least one side of the coding block may refer to the upper side of the coding block, or may refer to the left side of the coding block, or even may refer to the upper side and the left side of the coding block, which is not specifically limited in the embodiments of the present application.
[0083] In this way, for the INTRA_LT_CCLM mode, when the left adjacent region and the upper adjacent region are all valid regions, the first reference pixel set may be composed of reference pixel points adjacent to the left side of the coding block and reference pixel points adjacent to the upper side of the coding block at this time. When the left adjacent region is a valid region and the upper adjacent region is an invalid region, the first reference pixel set may be composed of reference pixel points adjacent to the left side of the coding block at this time; when the left adjacent region is an invalid region and the upper adjacent region is a valid region, the first reference pixel set may be composed of reference pixel points adjacent to the upper side of the coding block at this time.
[0084] In some embodiments, optionally, for S801-1, the obtaining of the first reference pixel set corresponding to the image component to be predicted of the coding block in the video image may include:
[0085] Obtain reference pixel points in a reference row or a reference column adjacent to the coding block; wherein, the reference row is composed of rows adjacent to the upper side and the upper right side of the coding block, and the reference column is composed of columns adjacent to the left side and the lower left side of the coding block;
[0086] Based on the reference pixel points, form a first reference pixel set corresponding to the image component to be predicted.
[0087] It should be noted that the reference row or reference column adjacent to the coding block may refer to the reference row adjacent to the upper side of the coding block, or may refer to the reference column adjacent to the left side of the coding block, or may even refer to the reference row or reference column adjacent to other sides of the coding block. The embodiments of the present application do not make specific limitations. For the convenience of description, in the embodiments of the present application, the reference row adjacent to the coding block will be described by taking the reference row adjacent to the upper side as an example, and the reference column adjacent to the coding block will be described by taking the reference column adjacent to the left side as an example.
[0088] Among them, the reference pixel points in the reference row adjacent to the coding block may include reference pixel points adjacent to the upper side and the upper right side (also referred to as adjacent reference pixel points corresponding to the upper side and the upper right side), where the upper side represents the upper side of the coding block, and the upper right side represents the side length that extends horizontally to the right from the upper side of the coding block and is the same as the height of the current coding block; the reference pixel points in the reference column adjacent to the coding block may also include reference pixel points adjacent to the left side and the lower left side (also referred to as adjacent reference pixel points corresponding to the left side and the lower left side), where the left side represents the left side of the coding block, and the lower left side represents the side length that extends vertically downward from the left side of the coding block and is the same as the width of the current decoding block; however, the embodiments of the present application do not make specific limitations either.
[0089] In this way, for the INTRA_L_CCLM mode, when the left adjacent region and the lower left adjacent region are valid regions, at this time, the first reference pixel set may be composed of reference pixel points in the reference column adjacent to the coding block; for the INTRA_T_CCLM mode, when the upper adjacent region and the upper right adjacent region are valid regions, at this time, the first reference pixel set may be composed of reference pixel points in the reference row adjacent to the coding block.
[0090] S801-2: Perform screening processing on the first reference pixel set to obtain a second reference pixel set; wherein, the second reference pixel set includes N adjacent reference pixel points;
[0091] It should be noted that in the first parameter pixel set, there may be some unimportant reference pixel points (for example, the correlation of these reference pixel points is poor) or some abnormal reference pixel points. To ensure the accuracy of the prediction model, these reference pixel points need to be removed, thus obtaining the second reference pixel set; among them, the number of valid pixel points included in the second reference pixel set is usually selected as 4 in practical applications.
[0092] In some embodiments, for S801-2, the screening process of the first reference pixel set to obtain the second reference pixel set may include:
[0093] Based on the pixel positions and / or image component intensities corresponding to each adjacent reference pixel point in the first reference pixel set, determine the positions of the pixels to be selected;
[0094] According to the determined positions of the pixels to be selected, select the adjacent reference pixel points corresponding to the positions of the pixels to be selected from the first reference pixel set, and form the selected adjacent reference pixel points into the second reference pixel set; among them, the second reference pixel set includes a preset number of adjacent reference pixel points.
[0095] It should be noted that the image component intensity can be represented by an image component value, such as a brightness value, a chroma value, etc.; here, the larger the image component value, the higher the image component intensity. In this way, the screening of the first reference pixel set can be carried out according to the positions of the reference pixels to be selected, or according to the image component intensity (such as brightness value, chroma value, etc.), so as to form the second reference pixel set with the screened reference pixels to be selected. The following will be described by taking the positions of the reference pixels to be selected as an example.
[0096] Exemplarily, assume that the positions of the reference pixel points in the upper selection area W' are S[0, -1], …, S[W' - 1, -1], and the positions of the reference pixel points in the left selection area H' are S[-1, 0], …, S[-1, H' - 1]; the screening method of selecting at most 4 adjacent reference pixel points is as follows:
[0097] For the INTRA_LT_CCLM mode, when both the upper adjacent area and the left adjacent area are valid, at this time, 2 adjacent reference pixel points to be selected can be screened out in the upper selection area W', and their corresponding positions are S[W' / 4, -1] and S[3W' / 4, -1] respectively; 2 adjacent reference pixel points to be selected can be screened out in the left selection area H', and their corresponding positions are S[-1, H' / 4] and S[-1, 3H' / 4] respectively; form these 4 adjacent reference pixel points to be selected into the second reference pixel set, as Figure 9A shown. In Figure 9AAmong them, the left adjacent region and the upper adjacent region of the coding block are both valid. Moreover, in order to keep the luminance component and the chrominance component having the same resolution, it is also necessary to perform downsampling processing on the luminance component so that the downsampled luminance component has the same resolution as the chrominance component.
[0098] For the INTRA_L_CCLM mode, when only the left adjacent region and the lower left adjacent region are valid, at this time, 4 candidate adjacent reference pixel points can be selected in the left selection region H', and their corresponding positions are S[-1, H' / 8], S[-1, 3H' / 8], S[-1, 5H' / 8] and S[-1, 7H' / 8] respectively; these 4 candidate adjacent reference pixel points are combined to form the second reference pixel set, as Figure 9B shown. In Figure 9B Among them, the left adjacent region and the lower left adjacent region of the coding block are both valid. Moreover, in order to keep the luminance component and the chrominance component having the same resolution, it is still necessary to perform downsampling processing on the luminance component so that the downsampled luminance component has the same resolution as the chrominance component.
[0099] For the INTRA_T_CCLM mode, when only the upper adjacent region and the upper right adjacent region are valid, at this time, 4 candidate adjacent reference pixel points can be selected in the upper selection region W', and their corresponding positions are S[W' / 8, -1], S[3W' / 8, -1], S[5W' / 8, -1] and S[7W' / 8, -1] respectively; these 4 candidate adjacent reference pixel points are combined to form the second reference pixel set, as Figure 9C shown. In Figure 9C Among them, the upper adjacent region and the upper right adjacent region of the coding block are both valid. Moreover, in order to keep the luminance component and the chrominance component having the same resolution, it is still necessary to perform downsampling processing on the luminance component so that the downsampled luminance component has the same resolution as the chrominance component.
[0100] In this way, by screening the first reference pixel set, the second reference pixel set can be obtained, and generally 4 adjacent reference pixel points are included in the second reference pixel set; after obtaining the second reference pixel set, the second reference pixel set can be grouped and divided so that the two fitting points used in the model parameter derivation are more accurate, and the robustness of the prediction is improved.
[0101] S802: Calculate the mean value of the N adjacent reference pixel points to obtain the first mean point;
[0102] It should be noted that for the first mean point, it can be obtained by calculating the mean value of a preset number of adjacent reference pixel points in the second reference pixel set. Therefore, in some embodiments, for S802, this step may include:
[0103] Based on the first image component values and the second image component values corresponding to each adjacent reference pixel point in the second reference pixel set, obtain the first mean value corresponding to the first image component and the second mean value corresponding to the second image component, and obtain the first mean value point.
[0104] That is to say, perform a mean calculation on the first image component values corresponding to each adjacent reference pixel point in the second reference pixel set to obtain the first image component mean values corresponding to multiple first image components, which can be called the first mean value of the first image component (which can be represented by mean L ); perform a mean calculation on the second image component values corresponding to each adjacent reference pixel point in the second reference pixel set to obtain the second image component mean values corresponding to multiple second image components, which can be called the first mean value of the second image component (which can be represented by mean C ); here, the first mean value point can be represented by (mean L , mean C ); that is to say, the first image component of the first mean value point is mean L , and the second image component of the first mean value point is mean C .
[0105] S803: Group the second reference pixel set through the first mean value point to obtain a first reference pixel subset and a second reference pixel subset;
[0106] It should be noted that in order to improve the robustness of prediction, at this time, the second reference pixel set can be grouped through the first mean value point. For example, it can be divided into a left category (represented by the Left category) and a right category (represented by the Right category), and then the mean value points of the Left category and the Right category are used as fitting points to deduce the model parameters; among them, the Left category is the first reference pixel subset, and the Right category is the second reference pixel subset.
[0107] Furthermore, in some embodiments, for S803, the grouping of the second reference pixel set through the first mean value point to obtain a first reference pixel subset and a second reference pixel subset may include:
[0108] S803-1: Compare the first image component values corresponding to each adjacent reference pixel point in the second reference pixel set with the first mean value corresponding to the first image component;
[0109] S803-2: Based on the comparison result, if the first image component value corresponding to the adjacent reference pixel point is less than or equal to the first mean value corresponding to the first image component, then put the adjacent reference pixel point into the first reference pixel subset to obtain the first reference pixel subset;
[0110] S803-3: If the first image component value corresponding to an adjacent reference pixel is greater than the first mean value corresponding to the first image component, then put the adjacent reference pixel into the second reference pixel subset to obtain the second reference pixel subset.
[0111] That is to say, after obtaining the first mean value mean corresponding to the first image component L it is possible to compare the first image component value corresponding to each adjacent reference pixel in the second reference pixel set with mean L When the first image component value corresponding to an adjacent reference pixel is less than or equal to mean L then the adjacent reference pixel can be put into the first reference pixel subset to obtain the first reference pixel subset; when the first image component value corresponding to an adjacent reference pixel is greater than mean L then the adjacent reference pixel can be put into the second reference pixel subset to obtain the second reference pixel subset.
[0112] S804: Determine two fitting points based on the first reference pixel subset and the second reference pixel subset;
[0113] It should be noted that the two fitting points include a first fitting point and a second fitting point. Here, according to the first reference pixel subset, the first fitting point can be obtained; according to the second reference pixel subset, the second fitting point can be obtained. Among them, the first fitting point can be obtained by taking the average value of all adjacent reference pixels in the first reference pixel subset, or by taking the average value of some adjacent reference pixels in the first reference pixel subset, or by selecting the middle adjacent reference pixel in the first reference pixel subset as the first fitting point, or even by arbitrarily selecting an adjacent reference pixel in the first reference pixel subset as the first fitting point. The embodiments of the present application do not make specific limitations; similarly, the second fitting point can be obtained by taking the average value of all adjacent reference pixels in the second reference pixel subset, or by selecting the middle adjacent reference pixel in the second reference pixel subset as the first fitting point, or by arbitrarily selecting an adjacent reference pixel in the second reference pixel subset as the first fitting point. The embodiments of the present application also do not make specific limitations. Unless otherwise specified, in the embodiments of the present application, the first fitting point is obtained by taking the average value of all adjacent reference pixels in the first reference pixel subset, and the second fitting point is obtained by taking the average value of all adjacent reference pixels in the second reference pixel subset.
[0114] In some embodiments, optionally, for S804, the determining two fitting points based on the first reference pixel subset and the second reference pixel subset may include:
[0115] S804a-1: Select some adjacent reference pixel points from the first reference pixel subset, calculate the mean value of the some adjacent reference pixel points, and use the calculated mean point as the first fitting point;
[0116] S804a-2: Select some adjacent reference pixel points from the second reference pixel subset, calculate the mean value of the some adjacent reference pixel points, and use the calculated mean point as the second fitting point.
[0117] In some embodiments, optionally, for S804, the determining two fitting points based on the first reference pixel subset and the second reference pixel subset includes:
[0118] S804b-1: Select one of the adjacent reference pixel points from the first reference pixel subset as the first fitting point;
[0119] S804b-2: Select one of the adjacent reference pixel points from the first reference pixel subset as the second fitting point.
[0120] In some embodiments, optionally, for S804, the determining two fitting points based on the first reference pixel subset and the second reference pixel subset may include:
[0121] S804c-1: Based on the first image component value and the second image component value corresponding to each adjacent reference pixel point in the first reference pixel subset, obtain the second mean value corresponding to the first image component and the second mean value corresponding to the second image component, obtain the second mean point, and use the second mean point as the first fitting point;
[0122] S804c-2: Based on the first image component value and the second image component value corresponding to each adjacent reference pixel point in the second reference pixel subset, obtain the third mean value corresponding to the first image component and the third mean value corresponding to the second image component, obtain the third mean point, and use the third mean point as the second fitting point.
[0123] That is to say, calculate the mean value of the first image component values corresponding to each adjacent reference pixel point in the first reference pixel subset, and obtain the first image component mean values corresponding to multiple first image components, which can be called the second mean value of the first image component (which can be represented by mean LeftL ); calculate the mean value of the second image component values corresponding to each adjacent reference pixel point in the first reference pixel subset, and obtain the second image component mean values corresponding to multiple second image components, which can be called the second mean value of the second image component (which can be represented by mean LeftC ), here, the second mean point can be represented by (mean LeftL , meanLeftC ) indicates that the first fitting point can be represented by (mean LeftL , mean LeftC ); that is to say, the first image component of the first fitting point is mean LeftL , and the second image component of the first fitting point is mean LeftC .
[0124] Calculate the mean value of the first image component values corresponding to each adjacent reference pixel point in the second reference pixel subset to obtain the first image component mean corresponding to multiple first image components, which can be called the third mean of the first image component (which can be represented by mean RightL ); calculate the mean value of the second image component values corresponding to each adjacent reference pixel point in the second reference pixel subset to obtain the second image component mean corresponding to multiple second image components, which can be called the third mean of the second image component (which can be represented by mean RightC ). Here, the third mean point can be represented by (mean RightL , mean RightC ), that is, the second fitting point can be represented by (mean RightL , mean RightC ); that is to say, the first image component of the second fitting point is mean RightL , and the second image component of the second fitting point is mean RightC .
[0125] S805: Based on the two fitting points, determine the model parameters, and obtain the prediction model corresponding to the image component to be predicted according to the model parameters; wherein, the prediction model is used to implement the prediction processing of the image component to be predicted to obtain the predicted value corresponding to the image component to be predicted;
[0126] It should be noted that after obtaining the first fitting point and the second fitting point, the model parameters can be determined according to the first fitting point and the second fitting point; here, the model parameters include the first model parameter (which can be represented by α) and the second model parameter (which can be represented by β). Assuming that the image component to be predicted is the chrominance component, according to the model parameters α and β, the prediction model corresponding to the chrominance component as shown in Equation (1) can be obtained.
[0127] In some embodiments, for S805, the model parameters include the first model parameter and the second model parameter, and the determining the model parameters based on the two fitting points may include:
[0128] S805-1: Based on the first fitting point and the second fitting point, calculate the model through the first preset factor to obtain the first model parameter;
[0129] S805-2: Obtain the second model parameter through a second preset factor calculation model based on the first model parameter and the first mean point;
[0130] Further, in some embodiments, after S805-1, the method may further include:
[0131] S805-3: Obtain the second model parameter through a second preset factor calculation model based on the first model parameter and the first fitting point.
[0132] It should be noted that after obtaining the first fitting point (mean LeftL , mean LeftC ) and the second fitting point (mean RightL , mean RightC ), the first model parameter α can be calculated according to the first preset factor calculation model, as shown in Equation (2).
[0133]
[0134] After obtaining the first model parameter α, the second model parameter β can be calculated according to the combination of the first mean point (mean L , mean C ) and the second preset factor calculation model, as shown in Equation (3).
[0135] β = mean C - α × mean L (3)
[0136] In addition, after obtaining the first model parameter α, the second model parameter β can also be calculated according to the combination of the first fitting point (mean LeftL , mean LeftC ) and the second preset factor calculation model, as shown in Equation (4).
[0137] β = mean LeftC - α × mean LeftL (4)
[0138] In this way, after obtaining the first model parameter α and the first model parameter β, a preset model can be constructed. Assuming that the component of the image to be predicted is the chrominance component, the prediction model corresponding to the chrominance component can be obtained according to the model parameters (α and β), as shown in Equation (1); then the chrominance component is predicted using this prediction model to obtain the predicted value corresponding to the chrominance component.
[0139] Further, in some embodiments, after S805, the method may further include:
[0140] Performing prediction processing on the image component to be predicted for each pixel point in the coding block based on the prediction model to obtain the predicted value corresponding to the image component to be predicted for each pixel point.
[0141] It should be noted that after grouping the second reference pixel set by the first mean point, a first reference pixel subset and a second reference pixel subset can be obtained; then, a fitting point is determined according to the first reference pixel subset, and another fitting point is determined according to the second reference pixel subset; in this way, according to the principle of "two points determine a straight line", the slope of the straight line (i.e., the first model parameter) and the intercept of the straight line (i.e., the second model parameter) can be determined, so that the prediction model corresponding to the image component to be predicted can be obtained according to these two model parameters, and the predicted value corresponding to the image component to be predicted for each pixel point in the coding block can be obtained. For example, assuming that the image component to be predicted is the chrominance component, according to the first model parameter α and the second model parameter β, the prediction model corresponding to the chrominance component as shown in Equation (1) can be obtained; then, the prediction model shown in Equation (1) is used to perform prediction processing on the chrominance component of each pixel point in the coding block, so that the predicted value corresponding to the chrominance component of each pixel point can be obtained.
[0142] This embodiment provides an image component prediction method. By obtaining N adjacent reference pixel points corresponding to the image component to be predicted in the coding block of the video image, the N adjacent reference pixel points are reference pixel points adjacent to the coding block, and N is a preset integer value; calculating the mean value of the N adjacent reference pixel points to obtain a first mean point; then grouping the second reference pixel set by the first mean point to obtain a first reference pixel subset and a second reference pixel subset, where the first mean point is obtained by calculating the mean value of a preset number of adjacent reference pixel points in the second reference pixel set; determining two fitting points based on the first reference pixel subset and the second reference pixel subset; determining model parameters based on these two fitting points, and obtaining a prediction model corresponding to the image component to be predicted according to the model parameters, where the prediction model is used to perform prediction processing on the image component to be predicted to obtain the predicted value corresponding to the image component to be predicted; in this way, since the first mean point is obtained by calculating the mean value of a preset number of adjacent reference pixel points in the second reference pixel set, grouping and dividing the preset number of adjacent reference pixel points according to the first mean point, and then determining two fitting points for deriving the model parameters can improve the robustness of the prediction; that is to say, by optimizing the fitting points used for deriving the model parameters, the constructed prediction model is more accurate, and the coding and decoding prediction performance of the video image is improved.
[0143] In another embodiment of the present application, in practical applications, since the adjacent reference pixel points used for deriving the model parameters are generally 4; that is to say, the preset number can be 4. The following will be described in detail with the preset number equal to 4 as an example.
[0144] As Figure 3 shown in the process, after obtaining 4 adjacent reference pixel points, two fitting points can be determined by 4 comparisons and the method of finding the mean point, so as to deduce the model parameters by using the principle of "two points determine a straight line"; thus, according to the model parameters, a prediction model corresponding to the image component to be predicted can be constructed to obtain the predicted value corresponding to the image component to be predicted.
[0145] Specifically, assume that the image component to be predicted is the chrominance component, and the chrominance component is predicted by the luminance component. Suppose the numbers of the 4 selected adjacent reference pixel points are 0, 1, 2, and 3 respectively. By comparing these 4 selected adjacent reference pixel points, based on four comparisons, 2 adjacent reference pixel points with larger luminance values (which may include the pixel point with the largest luminance value and the pixel point with the second largest luminance value) and 2 adjacent reference pixel points with smaller luminance values (which may include the pixel point with the smallest luminance value and the pixel point with the second smallest luminance value) can be further selected. Further, two arrays minIdx[2] and maxIdx[2] can be set to store the two sets of adjacent reference pixel points respectively. Initially, the adjacent reference pixel points numbered 0 and 2 are put into minIdx[2], and the adjacent reference pixel points numbered 1 and 3 are put into maxIdx[2], as shown below,
[0146] Init: minIdx[2] = {0, 2}, maxIdx[2] = {1, 3}
[0147] After that, through four comparisons, the 2 adjacent reference pixel points with smaller luminance values can be stored in minIdx[2], and the 2 adjacent reference pixel points with larger luminance values can be stored in maxIdx[2], as shown below,
[0148] Step1: if (L[minIdx[0]] > L[minIdx[1]], swap(minIdx[0], minIdx[1])
[0149] Step2: if (L[maxIdx[0]] > L[maxIdx[1]], swap(maxIdx[0], maxIdx[1])
[0150] Step3: if (L[minIdx[0]] > L[maxIdx[1]], swap(minIdx, maxIdx)
[0151] Step4: if (L[minIdx[1]] > L[maxIdx[0]], swap(minIdx[1], maxIdx[0])
[0152] In this way, two adjacent reference pixel points with smaller luminance values can be obtained, and the corresponding luminance values are represented by luma 0 min and luma 1 min respectively. The corresponding chrominance values are represented by chroma 0 min and chroma 1 min respectively. At the same time, two adjacent reference pixel points with larger luminance values can also be obtained, and the corresponding luminance values are represented by luma 0 max and luma 1 max respectively. The corresponding chrominance values are represented by chroma 0 max and chroma 1 max respectively. Further, by calculating the average value of the two smaller adjacent reference pixel points, an average value point represented by mean min can be obtained. The luminance value corresponding to this average value point mean min is luma min , and the chrominance value is chroma min . By calculating the average value of the two larger adjacent reference pixel points, a second average value point represented by mean max can be obtained. The luminance value corresponding to this average value point mean max is luma max , and the chrominance value is chroma max . Specifically, it is as follows:
[0153] luma min =(luma 0 min +luma 1 min +1)>>1
[0154] luma max =(luma 0 max +luma 1 max +1)>>1
[0155] chroma min =(chroma 0 min +chroma 1 min +1)>>1
[0156] chroma max=(chroma 0 max +chroma 1 max +1)>>1
[0157] That is to say, after obtaining two mean points mean min (luma min , chroma min ) and mean max (luma max , chroma max ), at this time, these two mean points can be used as two fitting points, and the model parameters can be obtained by these two fitting points through the calculation method of "two points determine a straight line". Specifically, the model parameters α and β can be calculated by Equation (5),
[0158]
[0159] wherein, the model parameter α is the slope in the prediction model, and the model parameter β is the intercept in the prediction model. In this way, after deriving the model parameters, the prediction model corresponding to the chrominance component can be obtained according to the model parameters, as shown in Equation (1); then, the chrominance component is predicted by using the prediction model to obtain the predicted value corresponding to the chrominance component.
[0160] In this way, since the current traditional solution uniformly uses the mean points of the two larger adjacent reference pixel points and the mean points of the two smaller adjacent reference pixel points among 4 adjacent reference pixel points as two fitting points to construct a prediction model, this prediction model lacks robustness. Therefore, when the 4 adjacent reference pixel points are unevenly distributed, the constructed prediction model will not be able to accurately fit the distribution of these 4 adjacent reference pixel points. For example, Figure 5 as shown in the comparison schematic diagram of the prediction model.
[0161] In the embodiment of the present application, for 4 adjacent reference pixel points, after obtaining the luminance mean value of the 4 adjacent reference pixel points, these 4 adjacent reference pixel points can be divided into two categories or two reference pixel subsets by the luminance mean value point; for example, they are divided into a left category (represented by the Left category) and a right category (represented by the Right category), and then the mean points of the Left category and the Right category are used as fitting points to construct a prediction model; wherein, the Left category is the first reference pixel subset, and the Right category is the second reference pixel subset.
[0162] See Figure 10 , which shows a schematic flow chart of a model parameter derivation scheme provided by the embodiment of the present application. As Figure 10 shown, this process may include:
[0163] S1001: Obtain 4 adjacent reference pixel points;
[0164] S1002: Calculate the luminance mean value meanL and the chrominance mean value meanC corresponding to the 4 adjacent reference pixel points;
[0165] S1003: Divide the 4 adjacent reference pixel points into a first reference pixel subset and a second reference pixel subset according to meanL;
[0166] S1004: Calculate the Left mean point corresponding to the first reference pixel subset and the Right mean point corresponding to the second reference pixel subset;
[0167] S1005: Use the Left mean point and the Right mean point as two fitting points to deduce the first model parameter;
[0168] S1006: Deduce the second model parameter according to meanL and meanC;
[0169] S1007: Construct a prediction model according to the two model parameters, and perform prediction processing on the chrominance component according to this prediction model.
[0170] It should be noted that the "two points determine a straight line" principle is used to construct the prediction model; these two points can be called fitting points. First, after obtaining 4 adjacent reference pixel points, by taking the mean value of these 4 adjacent reference pixel points (which can be represented by mean), the luminance mean value meanL and the chrominance mean value meanC corresponding to this mean value are obtained; then, the 4 adjacent reference pixel points are grouped using the luminance mean value meanL, for example, divided into a first reference pixel subset (which can be called the Left class) and a second reference pixel subset (which can be called the Right class); find the Left mean point corresponding to the Left class (which can be represented by meanLeft), the luminance mean value corresponding to this Left mean point is represented by meanLeftLuma, and the chrominance mean value is represented by meanLeftChroma; find the Right mean point corresponding to the Right class (which can be represented by meanRight), the luminance mean value corresponding to this Right mean point is represented by meanRightLuma, and the chrominance mean value is represented by meanRightChroma; use these two mean points (the Left mean point and the Right mean point) as two fitting points, and the model parameters (which can be represented by α and β) can be deduced. Finally, a prediction model as shown in Equation (1) is constructed according to the model parameters, and prediction processing on the chrominance component is performed according to this prediction model.
[0171] Specifically, assume that the image component to be predicted is the chrominance component, and the chrominance component is predicted from the luminance component. Assume that the numbers of the 4 adjacent reference pixel points selected by screening are 0, 1, 2, and 3 respectively. First, calculate the mean point mean(meanL, meanC) corresponding to these 4 adjacent reference pixel points, where meanL represents the luminance mean of the 4 adjacent reference pixel points, and meanC represents the chrominance mean of the 4 adjacent reference pixel points. Assume that these 4 adjacent reference pixel points are selectPix[i](selectLumaPix[i], selectChromaPix[i]) (0 <= i <= 3), the luminance value corresponding to the i-th adjacent reference pixel point is selectLumaPix[i], and the corresponding chrominance value is selectChromaPix[i]. Specifically, meanL and meanC are as follows:
[0172]
[0173]
[0174] After that, by comparing the luminance values selectLumaPix[i] (0 <= i <= 3) of these 4 adjacent reference pixel points with the luminance mean meanL respectively, the pixel points with luminance values less than (or less than or equal to) meanL are put into Left[cntL](LeftLuma[cntL], LeftChroma[cntL]), otherwise, the pixel points are put into Right[cntR](RightLuma[cntR], RightChroma[cntR]) (0 <= cntL <= 3, 0 <= cntR <= 3), specifically as follows:
[0175] int cntL = 0, cntR = 0;
[0176] for(int i = 0; i < 4; i++)
[0177] {
[0178] if(selectLumaPix[i] <= meanL)
[0179] {
[0180] LeftLuma[cntL] = selectLumaPix[i];
[0181] LeftChroma[cntL] = selectChromaPix[i];
[0182] cntL++;
[0183] }
[0184] else
[0185] {
[0186] RightLuma[cntR] = selectLumaPix[i];
[0187] RightChroma[cntR] = selectChromaPix[i];
[0188] cntR++;
[0189] }
[0190] }
[0191] It should be noted that for the judgment condition selectLumaPix[i] <= meanL, it can also be modified to selectLumaPix[i] < meanL.
[0192] Furthermore, calculate the mean points meanLeft (meanLeftLuma, meanLeftChroma) and meanRight (meanRightLuma, meanRightChroma) of the Left class and the Right class respectively, as shown below.
[0193] if (cntL!= 0 && cntR!= 0)
[0194] {
[0195]
[0196]
[0197]
[0198]
[0199] }
[0200] It should also be noted that round represents a rounding integer function. When cntL or cntR is 2, the division by 2 operation can be implemented by a shift operation in computer language; when cntL or cntR is 3, the division by 3 operation can be implemented by a Look-Up Table (LUT) method in computer language; this can achieve the purpose of reducing the computational complexity.
[0201] Further, after obtaining two mean points meanLeft (meanLeftLuma, meanLeftChroma) and meanRight (meanRightLuma, meanRightChroma), the first model parameter α can be derived using these two mean points as fitting points, and then the second model parameter β can be derived using the mean point mean (meanL, meanC), as shown below:
[0202] if (cntL!= 0 && cntR!= 0)
[0203] {
[0204]
[0205] β = meanC - α × meanL
[0206] }
[0207] else
[0208] {
[0209] α = 0
[0210] β = meanC
[0211] }
[0212] In this way, after deriving the two model parameters (α and β), the prediction model corresponding to the chrominance component can be obtained according to the model parameters, as shown in Equation (1); then, the prediction model is used to perform prediction processing on the chrominance component to obtain the predicted value corresponding to the chrominance component.
[0213] Before deriving the two fitting points, the current traditional solution requires 4 comparisons and 4 mean value operations; in the embodiment of the present application, this process requires 4 comparisons and 6 mean value operations. It should be noted that the mean value operations in the current traditional solution are all based on taking the mean value of two adjacent reference pixel points and can be implemented by simple addition and shift methods; in the embodiment of the present application, the mean value operations involve taking the mean value of 1 to 4 adjacent reference pixel points, among which, the mean value operation for 3 adjacent reference pixel points cannot be implemented by the shift method and can be implemented by the look-up table method in the embodiment of the present application.
[0214] In another embodiment of the present application, the preset quantity is 4, that is, the second reference pixel set includes 4 adjacent reference pixel points; in this way, according to the first mean point, the second reference pixel set can be grouped in 3 ways. For example, the first reference pixel subset includes 1 adjacent reference pixel point, and the second reference pixel subset includes 3 adjacent reference pixel points; the first reference pixel subset includes 2 adjacent reference pixel points, or the second reference pixel subset includes 2 adjacent reference pixel points; or the first reference pixel subset includes 3 adjacent reference pixel points, and the second reference pixel subset includes 1 adjacent reference pixel point, etc. Therefore, in some embodiments, before S803, the method may further include:
[0215] Obtain the first image component values corresponding to the 4 adjacent reference pixel points in the second reference pixel set respectively, and obtain the first mean value corresponding to the first image component through mean calculation;
[0216] Correspondingly, for S803, the grouping of the second reference pixel set by the first mean point to obtain a first reference pixel subset and a second reference pixel subset may include:
[0217] Compare the first image component value corresponding to each adjacent reference pixel point in the second reference pixel set with the first mean value corresponding to the first image component;
[0218] If the first image component value corresponding to 1 adjacent reference pixel point in the second reference pixel set is less than or equal to the first mean value corresponding to the first image component, then the first reference pixel subset includes 1 adjacent reference pixel point, and the second reference pixel subset includes 3 adjacent reference pixel points;
[0219] If the first image component values corresponding to 2 adjacent reference pixel points in the second reference pixel set are less than or equal to the first mean value corresponding to the first image component, then the first reference pixel subset includes 2 adjacent reference pixel points, and the second reference pixel subset includes 2 adjacent reference pixel points;
[0220] If the first image component values corresponding to 3 adjacent reference pixel points in the second reference pixel set are less than or equal to the first mean value corresponding to the first image component, then the first reference pixel subset includes 3 adjacent reference pixel points, and the second reference pixel subset includes 1 adjacent reference pixel point.
[0221] It should be noted that assuming that the first image component is the luminance component and the second image component is the chrominance component, after obtaining the first mean value corresponding to the luminance component, the luminance values of each group of the 4 adjacent reference pixel points can be compared with the first mean value to implement the grouping of the 4 adjacent reference pixel points.
[0222] Since the 4 adjacent reference pixel points can be evenly distributed or unevenly distributed, when the 4 adjacent reference pixel points are evenly distributed, after grouping, at this time, the first reference pixel subset includes 2 adjacent reference pixel points, and the second reference pixel subset includes 2 adjacent reference pixel points; when the 4 adjacent reference pixel points are unevenly distributed, after grouping, at this time, the first reference pixel subset includes 1 adjacent reference pixel point, and the second reference pixel subset includes 3 adjacent reference pixel points; or the first reference pixel subset includes 3 adjacent reference pixel points, and the second reference pixel subset includes 1 adjacent reference pixel point; then a first fitting point is determined from the first reference pixel subset, and a second fitting point is determined from the second reference pixel subset, thereby making full use of the characteristics of the 4 adjacent reference pixel points and improving the robustness of the prediction. The following will specifically describe these three cases.
[0223] Optionally, in some embodiments, when the first reference pixel subset includes 2 adjacent reference pixel points and the second reference pixel subset includes 2 adjacent reference pixel points, for S804, determining two fitting points based on the first reference pixel subset and the second reference pixel subset may include:
[0224] For the 2 adjacent reference pixel points in the first reference pixel subset, the second mean of the first image component and the second mean of the second image component are determined by a shifting method to obtain a second mean point as the first fitting point;
[0225] For the 2 adjacent reference pixel points in the second reference pixel subset, the third mean of the first image component and the third mean of the second image component are determined by a shifting method to obtain a third mean point as the second fitting point.
[0226] It should be noted that when the 4 adjacent reference pixel points are evenly distributed, there are two adjacent reference pixel points in each of the first reference pixel subset (such as the Left class) and the second reference pixel subset (such as the Right class). Refer to Figure 11 , which shows a comparison schematic diagram of a prediction model under a solution of the present application and a traditional solution provided by an embodiment of the present application. As Figure 11As shown, the four black dots are four adjacent reference pixel points. The two grey dots are respectively the mean points corresponding to the two larger adjacent reference pixel points and the two smaller adjacent reference pixel points among the four adjacent reference pixel points (i.e., the two fitting points obtained by the traditional solution). Thus, the grey diagonal line is the prediction model constructed according to the traditional solution; the bold black dashed line is the brightness mean of the four adjacent reference pixel points. The left side of the bold black dashed line is the Left class, and the right side of the bold black dashed line is the Right class. The two bold black circles are the mean points of the Left class and the Right class. Thus, the bold black diagonal line is the prediction model constructed according to the solution of the embodiment of the present application; the grey dotted line is the prediction model fitted by using the LMS algorithm for the four adjacent reference pixel points; from Figure 11 it can be seen that the grey diagonal line coincides with the bold black diagonal line and is relatively close to the grey dotted line, that is, the solution of the embodiment of the present application is the same as the prediction model constructed by the traditional solution, and both can accurately fit the distribution of the four adjacent reference pixel points.
[0227] Optionally, in some embodiments, when the first reference pixel subset includes 1 adjacent reference pixel point and the second reference pixel subset includes 3 adjacent reference pixel points, for S804, determining two fitting points based on the first reference pixel subset and the second reference pixel subset may include:
[0228] Regarding the 1 adjacent reference pixel point in the first reference pixel subset as the second mean point to be used as the first fitting point;
[0229] For the 3 adjacent reference pixel points in the second reference pixel subset, determining the third mean of the first image component and the third mean of the second image component through a look-up table method to obtain a third mean point to be used as the second fitting point.
[0230] It should be noted that when the four adjacent reference pixel points are unevenly distributed, there is 1 adjacent reference pixel point in the first reference pixel subset (such as the Left class), and there are 3 adjacent reference pixel points in the second reference pixel subset (such as the Right class). Refer to Figure 12 which shows another comparison schematic diagram of the prediction models under the solution of the present application and the traditional solution provided by the embodiment of the present application. As Figure 12As shown, the four black dots are four adjacent reference pixel points. The two grey dots are respectively the mean points corresponding to two larger adjacent reference pixel points and the mean points corresponding to two smaller adjacent reference pixel points among the four adjacent reference pixel points (i.e., the two fitting points obtained by the traditional solution). In this way, the grey diagonal line is the prediction model constructed according to the traditional solution; the bold black dotted line is the brightness mean of the four adjacent reference pixel points. To the left of the bold black dotted line is the Left class, and to the right of the bold black dotted line is the Right class. The two bold black circles are the mean points of the Left class and the Right class. In this way, the bold black diagonal line is the prediction model constructed according to the solution of the embodiment of the present application; the grey dotted line is the prediction model fitted by using the LMS algorithm for the four adjacent reference pixel points; from Figure 12 it can be seen that the bold black diagonal line and the grey dotted line are closer, that is, the solution of the embodiment of the present application is more in line with the distribution of the four adjacent reference pixel points than the prediction model constructed by the traditional solution.
[0231] Optionally, in some embodiments, when the first reference pixel subset includes three adjacent reference pixel points and the second reference pixel subset includes one adjacent reference pixel point, for S804, determining two fitting points based on the first reference pixel subset and the second reference pixel subset may include:
[0232] For the three adjacent reference pixel points in the first reference pixel subset, determine the second mean of the first image component and the second mean of the second image component in a look-up table manner to obtain a second mean point as the first fitting point;
[0233] Use the one adjacent reference pixel point in the second reference pixel subset as a third mean point as the second fitting point.
[0234] It should be noted that when the four adjacent reference pixel points are unevenly distributed, there are three adjacent reference pixel points in the first reference pixel subset (such as the Left class) and one adjacent reference pixel point in the second reference pixel subset (such as the Right class). Refer to Figure 13 which shows another comparison schematic diagram of the prediction models under the solution of the present application and the traditional solution provided by the embodiment of the present application. As Figure 13As shown, the four black dots are four adjacent reference pixel points. The two gray dots are respectively the mean points corresponding to the two larger adjacent reference pixel points and the two smaller adjacent reference pixel points among the four adjacent reference pixel points (i.e., the two fitting points obtained by the traditional scheme). Thus, the gray slanted line is the prediction model constructed according to the traditional scheme; the bold black dashed line is the brightness mean of the four adjacent reference pixel points. The left side of the bold black dashed line is the Left class, and the right side of the bold black dashed line is the Right class. The two bold black circles are the mean points of the Left class and the Right class. Thus, the bold black slanted line is the prediction model constructed according to the solution of the embodiment of the present application; the gray dotted line is the prediction model fitted by using the LMS algorithm for the four adjacent reference pixel points; from Figure 13 it can be seen that the bold black slanted line and the gray dotted line are closer, that is, the solution of the embodiment of the present application is more in line with the distribution of the four adjacent reference pixel points than the prediction model constructed by the traditional scheme.
[0235] From Figure 12 or Figure 13 it can be seen that when the four adjacent reference pixel points are unevenly distributed, the linear model constructed by the current traditional scheme lacks robustness and cannot well fit the four adjacent reference pixel points, resulting in inaccurate prediction performance. In the embodiment of the present application, by optimizing the fitting points used in the derivation of the model parameters, the CCLM prediction can be made more robust; among them, on the premise of a very small increase in computational complexity, the four adjacent reference pixel points are classified based on the brightness mean, and the mean point of each class is used as the fitting point to construct a prediction model, thereby improving the robustness of the CCLM prediction. For example, based on the latest reference software VTM5.0 of VVC, under the All intra condition, for the test sequences required by JVET with a temporal interval of 24, the average changes of BD-rate on the Y component, Cb component, and Cr component are -0.04%, -0.25%, and -0.20% respectively, which also shows that the solution of the embodiment of the present application brings a certain improvement in prediction performance on the premise of a very small increase in complexity.
[0236] This embodiment provides an image component prediction method. Through the technical solution of this embodiment, the first mean point is used to group and divide a preset number of adjacent reference pixel points, and then the mean point corresponding to each group is used as the fitting point. The model parameters are derived according to the two obtained fitting points. According to the model parameters, a prediction model corresponding to the image component to be predicted can be obtained. The prediction model is used to perform prediction processing on the image component to be predicted to obtain a predicted value corresponding to the image component to be predicted; in this way, the robustness of the CCLM prediction can be improved, so that the constructed prediction model is more accurate, and the prediction performance of video images is also improved.
[0237] Based on the same inventive concept as the foregoing embodiments, refer to Figure 14 , which shows a schematic structural diagram of a composition of an image component prediction device 140 provided in an embodiment of the present application. The image component prediction device 140 may include: an acquisition unit 1401, a calculation unit 1402, a grouping unit 1403, a determination unit 1404, and a prediction unit 1405. Among them,
[0238] The acquisition unit 1401 is configured to acquire N adjacent reference pixel points corresponding to a to-be-predicted image component of an encoded block in a video image; wherein, the N adjacent reference pixel points are reference pixel points adjacent to the encoded block, and N is a preset integer value;
[0239] The calculation unit 1402 is configured to calculate the mean value of the N adjacent reference pixel points to obtain a first mean point;
[0240] The grouping unit 1403 is configured to group the second reference pixel set by the first mean point to obtain a first reference pixel subset and a second reference pixel subset;
[0241] The determination unit 1404 is configured to determine two fitting points based on the first reference pixel subset and the second reference pixel subset;
[0242] The prediction unit 1405 is configured to determine model parameters based on the two fitting points, and obtain a prediction model corresponding to the to-be-predicted image component according to the model parameters; wherein, the prediction model is used to implement prediction processing on the to-be-predicted image component to obtain a predicted value corresponding to the to-be-predicted image component.
[0243] In the above solution, refer to Figure 14 , the image component prediction device 140 may further include a screening unit 1406,
[0244] The acquisition unit 1401 is further configured to acquire a first reference pixel set corresponding to a to-be-predicted image component of an encoded block in a video image;
[0245] The screening unit 1406 is configured to perform screening processing on the first reference pixel set to obtain a second reference pixel set; wherein, the second reference pixel set includes N adjacent reference pixel points.
[0246] In the above solution, the acquisition unit 1401 is specifically configured to acquire reference pixel points adjacent to at least one side of the encoded block; wherein, the at least one side includes the left side and / or the upper side of the encoded block; and based on the reference pixel points, form the first reference pixel set corresponding to the to-be-predicted image component.
[0247] In the above solution, the obtaining unit 1401 is specifically configured to obtain reference pixel points in a reference row or a reference column adjacent to the coding block; wherein, the reference row is composed of rows adjacent to the upper side and the upper right side of the coding block, and the reference column is composed of columns adjacent to the left side and the lower left side of the coding block; and based on the reference pixel points, a first reference pixel set corresponding to the image component to be predicted is formed.
[0248] In the above solution, the determining unit 1404 is further configured to determine the position of the pixel point to be selected based on the pixel positions and / or the intensities of the image components corresponding to each adjacent reference pixel point in the first reference pixel set;
[0249] The screening unit 1406 is specifically configured to, according to the determined position of the pixel point to be selected, select adjacent reference pixel points corresponding to the position of the pixel point to be selected from the first reference pixel set, and form a second reference pixel set with the selected adjacent reference pixel points; wherein, the second reference pixel set includes N adjacent reference pixel points.
[0250] In the above solution, the obtaining unit 1401 is further configured to obtain a first image component mean value corresponding to multiple first image components and a second image component mean value corresponding to multiple second image components in the second reference pixel set based on the first image component value and the second image component value corresponding to each adjacent reference pixel point in the second reference pixel set, so as to obtain the first mean point.
[0251] In the above solution, referring to Figure 14 , the image component prediction device 140 may further include a comparison unit 1407, configured to compare the first image component value corresponding to each adjacent reference pixel point in the second reference pixel set with the first mean value corresponding to the first image component; and based on the comparison result, if the first image component value corresponding to the adjacent reference pixel point is less than or equal to the first mean value corresponding to the first image component, then put the adjacent reference pixel point into the first reference pixel subset to obtain the first reference pixel subset; if the first image component value corresponding to the adjacent reference pixel point is greater than the first mean value corresponding to the first image component, then put the adjacent reference pixel point into the second reference pixel subset to obtain the second reference pixel subset.
[0252] In the above solution, the value of N is 4.
[0253] In the above solution, the obtaining unit 1401 is further configured to obtain the first image component value corresponding to each of the 4 adjacent reference pixel points in the second reference pixel set, and obtain the first image component mean value corresponding to the 4 first image components through mean value calculation;
[0254] Correspondingly, the comparison unit 1407 is specifically configured to compare the first image component values corresponding to each adjacent reference pixel point in the second reference pixel set with the first image component mean value; and if there is 1 adjacent reference pixel point in the second reference pixel set whose corresponding first image component value is less than or equal to the first image component mean value, the first reference pixel subset includes 1 adjacent reference pixel point and the second reference pixel subset includes 3 adjacent reference pixel points; if there are 2 adjacent reference pixel points in the second reference pixel set whose corresponding first image component values are less than or equal to the first image component mean value, the first reference pixel subset includes 2 adjacent reference pixel points and the second reference pixel subset includes 2 adjacent reference pixel points; if there are 3 adjacent reference pixel points in the second reference pixel set whose corresponding first image component values are less than or equal to the first image component mean value, the first reference pixel subset includes 3 adjacent reference pixel points and the second reference pixel subset includes 1 adjacent reference pixel point.
[0255] In the above solution, the obtaining unit 1401 is further configured to obtain the first image component mean value corresponding to multiple first image components and the second image component mean value corresponding to multiple second image components in the first reference pixel subset based on the first image component values and the second image component values corresponding to each adjacent reference pixel point in the first reference pixel subset, obtain a second mean point, and use the second mean point as the first fitting point; and obtain the first image component mean value corresponding to multiple first image components and the second image component mean value corresponding to multiple second image components in the second reference pixel subset based on the first image component values and the second image component values corresponding to each adjacent reference pixel point in the second reference pixel subset, obtain a third mean point, and use the third mean point as the second fitting point.
[0256] In the above solution, the calculation unit 1402 is configured to obtain the first model parameter through a first preset factor calculation model based on the first fitting point and the second fitting point; and obtain the second model parameter through a second preset factor calculation model based on the first model parameter and the first mean point.
[0257] In the above solution, the calculation unit 1402 is further configured to obtain the second model parameter through a second preset factor calculation model based on the first model parameter and the first fitting point.
[0258] In the above solution, the prediction unit 1405 is specifically configured to perform prediction processing on the image component to be predicted of each pixel point in the coding block based on the prediction model to obtain the prediction value corresponding to the image component to be predicted of each pixel point.
[0259] Understandably, in this embodiment, a "unit" may be a part of a circuit, a part of a processor, a part of a program or software, etc. Of course, it may also be a module or non-modular. Moreover, the components in this embodiment may be integrated in one processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional module.
[0260] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0261] Therefore, this embodiment provides a computer storage medium that stores an image component prediction program. When the image component prediction program is executed by at least one processor, it implements the method described in any one of the foregoing embodiments.
[0262] Based on the composition of the above image component prediction device 140 and the computer storage medium, refer to Figure 15 , which shows the specific hardware structure of the image component prediction device 140 provided by the embodiment of the present application. It may include: a network interface 1501, a memory 1502, and a processor 1503; each component is coupled together through a bus system 1504. It can be understood that the bus system 1504 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1504 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 15 all kinds of buses are labeled as the bus system 1504. Among them, the network interface 1501 is used for receiving and sending signals during the process of receiving and sending information to and from other external network elements;
[0263] The memory 1502 is used to store a computer program that can run on the processor 1503;
[0264] A processor 1503, which is configured to execute the following when running the computer program:
[0265] Obtain N adjacent reference pixel points corresponding to the image component to be predicted of an encoded block in a video image; wherein, the N adjacent reference pixel points are reference pixel points adjacent to the encoded block, and N is a preset integer value;
[0266] Calculate the mean value of the N adjacent reference pixel points to obtain a first mean point;
[0267] Group the second reference pixel set by the first mean point to obtain a first reference pixel subset and a second reference pixel subset;
[0268] Determine two fitting points based on the first reference pixel subset and the second reference pixel subset;
[0269] Determine model parameters based on the two fitting points, and obtain a prediction model corresponding to the image component to be predicted according to the model parameters; wherein, the prediction model is used to implement prediction processing on the image component to be predicted to obtain a predicted value corresponding to the image component to be predicted.
[0270] It can be understood that the memory 1502 in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The memory 1502 of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.
[0271] The processor 1503 may be an integrated circuit chip with the ability to process signals. In the implementation process, the steps of the above method can be completed by the integrated logic circuit of the hardware in the processor 1503 or the instructions in the form of software. The above-mentioned processor 1503 may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory 1502, and the processor 1503 reads the information in the memory 1502 and combines its hardware to complete the steps of the above method.
[0272] It can be understood that these embodiments described herein can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For a hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or a combination thereof.
[0273] For a software implementation, the techniques described herein can be implemented by modules (e.g., procedures, functions, etc.) that execute the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented within the processor or external to the processor.
[0274] Optionally, as another embodiment, the processor 1503 is further configured to execute the method described in any one of the foregoing embodiments when running the computer program.
[0275] See Figure 16 , which shows a schematic structural diagram of a composition of an encoder provided by an embodiment of the present application. As Figure 16 shown, the encoder 160 may at least include the image component prediction device 140 described in any one of the foregoing embodiments.
[0276] See Figure 17 , which shows a schematic structural diagram of a composition of a decoder provided by an embodiment of the present application. As Figure 17 shown, the decoder 170 may at least include the image component prediction device 140 described in any one of the foregoing embodiments.
[0277] It should be noted that in the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0278] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0279] The methods disclosed in several method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.
[0280] The features disclosed in several product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.
[0281] The features disclosed in several method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0282] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
[0283] Industrial applicability
[0284] In an embodiment of the present application, first, N adjacent reference pixel points corresponding to a to-be-predicted image component of an encoded block in a video image are obtained. The N adjacent reference pixel points are reference pixel points adjacent to the encoded block, and N is a preset integer value. The mean value of the N adjacent reference pixel points is calculated to obtain a first mean point. Then, the second reference pixel set is grouped by the first mean point to obtain a first reference pixel subset and a second reference pixel subset, and two fitting points are determined based on the first reference pixel subset and the second reference pixel subset. Finally, model parameters are determined based on the two fitting points, and a prediction model corresponding to the to-be-predicted image component is obtained according to the model parameters to obtain a predicted value corresponding to the to-be-predicted image component. In this way, since the first mean point is obtained by calculating the mean value of a preset number of adjacent reference pixel points in the second reference pixel set, the preset number of adjacent reference pixel points are grouped and divided according to the first mean point, and then two fitting points for deriving the model parameters are determined, which can improve the robustness of the prediction. That is to say, by optimizing the fitting points used for deriving the model parameters, the constructed prediction model is made more accurate, and the encoding and decoding prediction performance of the video image is improved.
Claims
1. An image component prediction method, the method comprising: Obtaining a first reference pixel set corresponding to a to-be-predicted image component of an encoded block in a video image; Performing a screening process on the first reference pixel set to obtain a second reference pixel set; wherein, the second reference pixel set includes N adjacent reference pixel points, and the N adjacent reference pixel points are reference pixel points adjacent to the encoded block, and N is a preset integer value; Calculating the mean value of the N adjacent reference pixel points to obtain a first mean point; Grouping the N adjacent reference pixel points by the first mean point to obtain a first reference pixel subset and a second reference pixel subset; Determining two fitting points based on the first reference pixel subset and the second reference pixel subset; Determining model parameters based on the two fitting points, and obtaining a prediction model corresponding to the to-be-predicted image component according to the model parameters; wherein, the prediction model is used to implement a prediction process on the to-be-predicted image component to obtain a predicted value corresponding to the to-be-predicted image component; Wherein, the performing a screening process on the first reference pixel set to obtain a second reference pixel set includes: Determining a to-be-selected pixel point position based on the image component intensity corresponding to each adjacent reference pixel point in the first reference pixel set, or based on the pixel position and the image component intensity corresponding to each adjacent reference pixel point in the first reference pixel set; According to the determined to-be-selected pixel point position, selecting adjacent reference pixel points corresponding to the to-be-selected pixel point position from the first reference pixel set, and forming the selected adjacent reference pixel points into a second reference pixel set.
2. The method according to claim 1, wherein, The obtaining a first reference pixel set corresponding to a to-be-predicted image component of an encoded block in a video image includes: Obtaining reference pixel points adjacent to at least one side of the encoded block; wherein, the at least one side includes the left side and / or the upper side of the encoded block; Based on the reference pixel points, forming a first reference pixel set corresponding to the to-be-predicted image component.
3. The method according to claim 1, wherein, The obtaining a first reference pixel set corresponding to a to-be-predicted image component of an encoded block in a video image includes: Obtaining reference pixel points in a reference row or a reference column adjacent to the encoded block; wherein, the reference row is composed of rows adjacent to the upper side and the upper right side of the encoded block, and the reference column is composed of columns adjacent to the left side and the lower left side of the encoded block; Based on the reference pixel points, forming a first reference pixel set corresponding to the to-be-predicted image component.
4. The method according to claim 1, wherein Before the grouping the N adjacent reference pixel points by the first mean point to obtain a first reference pixel subset and a second reference pixel subset, the method further includes: Obtaining a first image component mean value corresponding to multiple first image components and a second image component mean value corresponding to multiple second image components in the second reference pixel set based on a first image component value and a second image component value corresponding to each adjacent reference pixel point in the N adjacent reference pixel points, to obtain the first mean point.
5. The method according to claim 4, wherein Grouping the N adjacent reference pixel points by the first mean point to obtain a first reference pixel subset and a second reference pixel subset includes: Comparing the first image component value corresponding to each adjacent reference pixel point among the N adjacent reference pixel points with the first mean value corresponding to the first image component; Based on the comparison result, if the first image component value corresponding to an adjacent reference pixel point is less than or equal to the first mean value corresponding to the first image component, putting the adjacent reference pixel point into the first reference pixel subset to obtain the first reference pixel subset; If the first image component value corresponding to an adjacent reference pixel point is greater than the first mean value corresponding to the first image component, putting the adjacent reference pixel point into the second reference pixel subset to obtain the second reference pixel subset.
6. The method according to claim 1, wherein, The value of N is 4.
7. The method according to claim 6, wherein Before grouping the N adjacent reference pixel points by the first mean point to obtain a first reference pixel subset and a second reference pixel subset, the method further includes: Obtaining the first image component values corresponding to 4 adjacent reference pixel points in the second reference pixel set respectively, and obtaining the first image component mean value corresponding to the 4 first image components through mean calculation; Correspondingly, grouping the N adjacent reference pixel points by the first mean point to obtain a first reference pixel subset and a second reference pixel subset includes: Comparing the first image component value corresponding to each adjacent reference pixel point in the second reference pixel set with the first image component mean value; If there is 1 adjacent reference pixel point in the second reference pixel set whose first image component value is less than or equal to the first image component mean value, the first reference pixel subset includes 1 adjacent reference pixel point, and the second reference pixel subset includes 3 adjacent reference pixel points; If there are 2 adjacent reference pixel points in the second reference pixel set whose first image component values are less than or equal to the first image component mean value, the first reference pixel subset includes 2 adjacent reference pixel points, and the second reference pixel subset includes 2 adjacent reference pixel points; If there are 3 adjacent reference pixel points in the second reference pixel set whose first image component values are less than or equal to the first image component mean value, the first reference pixel subset includes 3 adjacent reference pixel points, and the second reference pixel subset includes 1 adjacent reference pixel point.
8. The method according to claim 1, wherein Determining two fitting points based on the first reference pixel subset and the second reference pixel subset includes: Based on the first image component value and the second image component value corresponding to each adjacent reference pixel point in the first reference pixel subset, obtaining the first image component mean value corresponding to multiple first image components and the second image component mean value corresponding to multiple second image components in the first reference pixel subset, obtaining a second mean point, and taking the second mean point as the first fitting point; Based on the first image component values and the second image component values corresponding to each adjacent reference pixel point in the second reference pixel subset, obtain the first image component mean value corresponding to multiple first image components and the second image component mean value corresponding to multiple second image components in the second reference pixel subset, obtain a third mean point, and use the third mean point as the second fitting point.
9. The method according to claim 1, wherein The determining two fitting points based on the first reference pixel subset and the second reference pixel subset includes: Select some adjacent reference pixel points from the first reference pixel subset, calculate the mean value of the some adjacent reference pixel points, and use the calculated mean point as the first fitting point; Select some adjacent reference pixel points from the second reference pixel subset, calculate the mean value of the some adjacent reference pixel points, and use the calculated mean point as the second fitting point.
10. The method according to claim 1, wherein The determining two fitting points based on the first reference pixel subset and the second reference pixel subset includes: Select one of the adjacent reference pixel points from the first reference pixel subset as the first fitting point; Select one of the adjacent reference pixel points from the first reference pixel subset as the second fitting point.
11. The method according to any one of claims 8 to 10, wherein, The model parameters include a first model parameter and a second model parameter. The determining the model parameters based on the two fitting points includes: Based on the first fitting point and the second fitting point, obtain the first model parameter through a first preset factor calculation model; Based on the first model parameter and the first mean point, obtain the second model parameter through a second preset factor calculation model.
12. The method according to claim 11, wherein, After obtaining the first model parameter through the first preset factor calculation model, the method further includes: Based on the first model parameter and the first fitting point, obtain the second model parameter through a second preset factor calculation model.
13. The method according to claim 1, wherein, After obtaining the prediction model corresponding to the image component to be predicted according to the model parameters, the method further includes: Based on the prediction model, perform prediction processing on the image component to be predicted of each pixel point in the coding block to obtain the prediction value corresponding to the image component to be predicted of each pixel point.
14. An image component prediction device, the image component prediction device comprising: An acquisition unit, a screening unit, a calculation unit, a grouping unit, a determination unit, and a prediction unit, where The acquisition unit is configured to acquire a first reference pixel set corresponding to the image component to be predicted of a coding block in a video image; The screening unit is configured to perform screening processing on the first reference pixel set to obtain a second reference pixel set; where the second reference pixel set includes N adjacent reference pixel points, and the N adjacent reference pixel points are reference pixel points adjacent to the coding block, and N is a preset integer value; The calculation unit is configured to calculate the mean value of the N adjacent reference pixel points to obtain a first mean point; The grouping unit is configured to group the N adjacent reference pixel points through the first mean point to obtain a first reference pixel subset and a second reference pixel subset; The determination unit is configured to determine two fitting points based on the first reference pixel subset and the second reference pixel subset; The prediction unit is configured to determine model parameters based on the two fitting points, and obtain a prediction model corresponding to the image component to be predicted according to the model parameters; wherein, the prediction model is used to perform prediction processing on the image component to be predicted to obtain a predicted value corresponding to the image component to be predicted; Wherein, the determination unit is further configured to determine the position of the pixel to be selected based on the intensity of the image component corresponding to each adjacent reference pixel point in the first reference pixel set, or based on the pixel position and the intensity of the image component corresponding to each adjacent reference pixel point in the first reference pixel set; The screening unit is further configured to select, from the first reference pixel set, adjacent reference pixel points corresponding to the position of the pixel to be selected according to the determined position of the pixel to be selected, and form a second reference pixel set with the selected adjacent reference pixel points.
15. An image component prediction device, wherein, The image component prediction device includes: a memory and a processor; The memory is used to store a computer program that can run on the processor; The processor is used to execute the method according to any one of claims 1 to 13 when running the computer program.
16. A computer storage medium, wherein, The computer storage medium stores an image component prediction program, and when the image component prediction program is executed by at least one processor, the method according to any one of claims 1 to 13 is implemented.