Encoding method, decoding method and apparatus, and apparatus
By incorporating a non-linear activation function into the target model, the method addresses the limitations of linear assumptions in VVC, improving chrominance prediction accuracy and compression efficiency for video coding.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-07
- Publication Date
- 2026-04-14
AI Technical Summary
Existing video coding technologies, such as Versatile Video Coding (VVC), assume a linear relationship between luminance and chrominance, leading to limited applicability and adverse effects on compression efficiency for chrominance prediction.
Introduce a non-linear activation function into the construction of the target model to enhance chrominance prediction accuracy, allowing for both linear and non-linear correlations between luminance and chrominance, thereby improving the applicability and compression effect of video coding.
The proposed method achieves high chrominance prediction accuracy for coding units with varying correlations, enhancing the applicability and compression efficiency of video encoding.
Smart Images

Figure 2026512131000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This disclosure claims the priority of a Chinese patent application filed with the China National Intellectual Property Administration on April 12, 2023, with the application number 202310390826.2 and the title of the invention being "Coding Method, Decoding Method and Apparatus, and Equipment", and all its contents are incorporated herein by reference.
[0002] This application relates to the field of video coding and decoding technologies, and particularly to coding methods, decoding methods and apparatuses, and equipment.
Background Art
[0003] In video coding and decoding, since there is a certain correlation between luminance and chrominance, in order to remove the redundancy between different components of luminance and chrominance, in Versatile Video Coding (VVC), an intra - prediction mode based on the Cross - Component Linear Model (CCLM) has been proposed.
[0004] However, CCLM assumes that the chrominance pixel values and their corresponding luminance pixel values in the same coding unit (CU) are in a linear relationship, and uses one linear model to generate the predicted value of the corresponding chrominance pixel from the reconstructed value of the luminance pixel. Therefore, the chrominance prediction method based on CCLM has a low applicable range and has an adverse effect on the compression effect of video coding.
Summary of the Invention
Problems to be Solved by the Invention
[0005] The embodiments of this application provide a coding method, a decoding method and an apparatus, and equipment that are advantageous for improving the applicability of the chrominance prediction method and the compression effect of video coding.
Means for Solving the Problems
[0006] The first aspect provides a decoding method, which is: The method involves obtaining coding information corresponding to a target coding block, wherein the coding information includes at least the index information of the target model, the values of each parameter, the luminance downsampling reconstruction value of the target coding block, and the target difference value. The method involves determining a target model based on the index information of the target model, wherein the target model is constructed based on activation parameters and a first model, and the first model is intended to represent the mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values. Based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block, the chromaticity prediction value of the target coding block is determined. This includes determining the chromaticity reconstruction value of the target coding block based on the chromaticity prediction value of the target coding block and the target difference value.
[0007] A second aspect provides an encoding method, which is: The method involves constructing a target model based on an activation function and a first model, wherein the first model represents a mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values. Based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, the values of each parameter in the target model are determined, Based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block, the chromaticity prediction value of the target coding block is determined. Based on the predicted chromaticity value of the target coded block and the true chromaticity value of the target coded block, the target difference value is determined. This includes generating coding information corresponding to the target coding block based on the index information of the target model, the values of each of the variable parameters, the luminance downsampling reconstruction value of the target coding block, and the target difference value.
[0008] A third aspect provides a decoding device, the device is An information acquisition module for obtaining encoding information corresponding to a target encoding block and a luminance downsampling reconstruction value of the target encoding block, wherein the encoding information includes at least the index information of the target model, the values of each parameter, the luminance downsampling reconstruction value of the target encoding block, and the target difference value. A model determination module for determining a target model based on the index information of the target model, wherein the target model is constructed based on activation parameters and a first model, and the first model represents a mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values. A chromaticity prediction module for determining the chromaticity prediction value of the target coding block based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block, The system includes a chromaticity reconstruction module for determining the chromaticity reconstruction value of the target coding block based on the chromaticity prediction value of the target coding block and the target difference value.
[0009] A fourth aspect provides an encoding device, the device is A first construction module for constructing a target model based on an activation function and a first model, wherein the first model represents a mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values. A first processing module for determining the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, A second processing module for determining the chromaticity prediction value of the target coding block based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block, A third processing module for determining a target difference value based on the chromaticity prediction value of the target coding block and the chromaticity true value of the target coding block, The system includes a first generation module for generating coding information corresponding to the target coding block based on the index information of the target model, the values of each of the variable parameters, the luminance downsampling reconstruction value of the target coding block, and the target difference value.
[0010] The fifth aspect provides a computing device, said computing device, Memory in which computer-readable code is stored, Includes one or more processors, When the computer-readable code is executed by the one or more processors, the computing device executes the decoding method described in the first embodiment or the encoding method described in the second embodiment.
[0011] The sixth aspect provides a computer program including a computer-readable code, wherein when the computer-readable code is executed on a computing device, the computing device causes the computing device to execute the decoding method described in the first aspect or the encoding method described in the second aspect.
[0012] The seventh aspect provides a computer-readable medium on which the computer program described in the seventh aspect is stored.
[0013] The eighth aspect provides a decoding device including a processor and a memory, where the memory stores a program or instructions executable on the processor, and when the program or instructions are executed by the processor, steps of the decoding method according to the first aspect are realized.
[0014] The ninth aspect provides an encoding device including a processor and a memory, where the memory stores a program or instructions executable on the processor, and when the program or instructions are executed by the processor, steps of the encoding method according to the second aspect are realized.
[0015] The tenth aspect provides an encoding / decoding system including an encoding device and a decoding device, where the decoding device is for executing steps of the decoding method according to the first aspect above, and the encoding device is for executing steps of the encoding method according to the second aspect above.
[0016] The eleventh aspect provides a readable storage medium storing a program or instructions, and when the program or instructions are executed by a processor, steps of the decoding method according to the first aspect are executed, or steps of the encoding method according to the second aspect are executed.
[0017] The twelfth aspect provides a chip including a processor and a communication interface, where the communication interface is coupled to the processor, and the processor executes a program or instructions to realize the encoding method according to the second aspect or the decoding method according to the third aspect.
[0018] The thirteenth aspect provides a computer program / product stored in a storage medium and executed by at least one processor to execute steps of the decoding method according to the first aspect or steps of the encoding method according to the second aspect.
Advantages of the Invention
[0019] In the embodiments of the present application, by introducing a non-linear activation function into the construction of the target model, the target model can have high chromaticity prediction accuracy for both target encoding blocks with good or poor linear correlation between luminance and chromaticity. As a result, the applicability of the chromaticity prediction method is improved, and the compression effect of video encoding is also enhanced.
[0020] The above description is only a summary of the technical solutions of the present disclosure. In order to more clearly understand the technical means of the present disclosure and be able to implement them in accordance with the content of the specification, and to more clearly and easily understand the above and other objects, features, and advantages of the present disclosure, specific embodiments of the present disclosure are given below.
Brief Description of the Drawings
[0021] To more clearly explain the technical solutions in the embodiments of the present disclosure or related technologies, the drawings necessary for the description of the embodiments or the prior art are briefly described below. However, the drawings in the following description are only some embodiments of the present disclosure, and it is obvious that those skilled in the art can obtain other drawings based on these drawings without creative labor.
[0022] [Figure 1] It is a mode diagram of a luminance sample block in an embodiment of the present application. " [Figure 2] It is a mode diagram of a chromaticity sample block in an embodiment of the present application. [Figure 3] It is a mode diagram of an example of the luminance-chromaticity value distribution in an embodiment of the present application. [Figure 4] It is a mode diagram of another example of the luminance-chromaticity value distribution in an embodiment of the present application. [Figure 5] It is a mode diagram of each segment region in a multi-mode linear model in an embodiment of the present application. [Figure 6] It is a mode diagram of pixel point grouping in an embodiment of the present application. " [Figure 7] It is a mode diagram of an example of a filter in an embodiment of the present application. [Figure 8] This is a mode diagram of another example of the filter in the embodiment of the present application. [Figure 9] This is a mode diagram of another example of the filter in the embodiment of the present application. [Figure 10] This is a mode diagram of another example of the filter in the embodiment of the present application. [Figure 11] This is a flowchart of an example of a decoding method in the embodiment of the present invention. [Figure 12] This is a flowchart of another example of the encoding method in the embodiment of the present invention. [Figure 13] This is a structural block diagram of an example of a decoding device in an embodiment of the present invention. [Figure 14] This is a structural block diagram of an example of an encoding device in an embodiment of the present invention. [Figure 15] This is a block diagram that modally shows the computing equipment for implementing the method relating to this disclosure. [Figure 16] This figure shows a modal representation of a storage unit for holding or transporting program code that implements the method described herein. [Modes for carrying out the invention]
[0023] To further clarify the objectives, technical solutions, and advantages of the embodiments of this disclosure, the technical solutions of the embodiments of this disclosure will be described clearly and completely below with reference to the drawings of the embodiments of this disclosure, although it is clear that the embodiments described are only a part of the embodiments of this disclosure, not all of them. All other embodiments obtained by those skilled in the art without creative effort based on the embodiments of this disclosure are within the scope of this disclosure.
[0024] The terms “first,” “second,” etc., in the specification and claims of this application are for distinguishing similar objects and are not intended to describe a particular order or priority. It should be understood that the terms used in this manner may be interchangeable where appropriate so that the embodiments of this application may be carried out in an order other than that illustrated or described herein, and the objects distinguished by “first,” “second,” etc., are usually of one type and do not limit the number of objects; for example, the first object may be one or at least two. Furthermore, “and / or” in the specification and claims refers to at least one of the connected objects, and the symbol “ / ” generally indicates that the preceding and following related objects are in an “or” relationship.
[0025] VVC (H.266) adds CCLM (Continuous Computation and Decoding) to the previous generation video encoding and decoding standard, HEVC (H.265). CCLM technology can be subdivided into the following seven technologies.
[0026] 1. CCLM (narrow sense) Since this technique assumes a linear relationship between the chromaticity pixel value and its corresponding luminance pixel value within the same coding block (i.e., coding unit), CCLM uses a single linear (including curved) model to generate a predicted value for the corresponding chromaticity pixel based on the reconstructed value of the luminance pixel.
[0027] The following formula reconstructs the luminance sample using a 6-tap filter, thereby obtaining the reconstructed luminance sample.
number
number
number
[0028] Furthermore, predict the chromaticity sample from the reconstructed luminance sample.
number
number
number
number
number
number
[0029] The linear model defined in CCLM can achieve high coverage only for CUs corresponding to luminance chromaticity value distributions with strong linear correlation (e.g., shown in Figure 3 or Figure 4), thus enabling highly accurate prediction from luminance to chromaticity.
[0030] 2. Slope-adjustable CCLM (CCLM) The Enhanced Compression Model (ECM) optimizes the conventional CCLM by introducing a parameter "u" to adjust the slope, compared to the conventional CCLM:chromaVal=a*lumaVal+b. This results in ECM:chromaVal=a'*lumaVal+b', where a'=a+u and b'=bu*yr. Here, chromaVal is the chromaticity prediction value, lumaVal is the luminance downsampling reconstruction value, a and b are CCLM parameters, yr is the average luminance value, and u is any integer between -4 and 4. It can be seen that the ECM achieves rotation correction around the yr point for the linear model.
[0031] 3. Multi-mode linear model (MM-CCLM), abbreviated as MMLM To better improve the low model coverage in some luminance chromaticity value distribution scenarios, MM-CCLM employs a segmented linear regression method, dividing the luminance chromaticity value distribution into n regions, with different linear model parameters used for each region.
number
[0032] Since segment calculations increase computational complexity, fewer segments are advantageous in terms of computational complexity. Due to relatively low computational complexity, MM-CCLM is generally divided into two segments, and a breakpoint (i.e., threshold) is selected to divide x. As shown in Figure 5, this segment is based on the mean of downsampled luminance samples, with a mean of 108. After segment linear regression, the mean square error (MSE) decreased from 88.44 to 45.69.
[0033] As shown in Figure 6, region P consists of two rows adjacent to the left and above the current chromaticity block (or luminance block after downsampling). The resulting threshold (i.e., the average value of the luminance downsampling reconstruction in region P) is 17. If the luminance downsampling reconstruction value of a pixel point in region P is greater than the threshold, it is classified into group 1 and marked with a white circle. If the luminance downsampling reconstruction value of a pixel point in region P is less than the threshold, it is classified into group 0 and marked with a black circle. Subsequently, within each group, the corresponding linear model parameters are calculated for each region.
[0034] In VVC, there is one parameter (i.e., the LM flag) that indicates whether to apply the cross-component linear prediction model technique. If the LM flag is equal to 1, it indicates that the loss-component linear prediction model technique will be applied. Furthermore, by checking whether the MMLM flag parameter indicates that the multimode cross-component linear prediction model technique will be applied, the encoding and decoding sides can accurately select the model to perform chromaticity prediction.
[0035] 4. Multi-sampling filter linear model (MF-CCLM) MF-CCLM optimizes the reconstructed luminance values, and this reconstruction optimization is achieved based on the default 4:2:0 chromaticity downsampling position. In next-generation encoding / decoding standards (AVC / HEVC), there is support for complementary VUI high-level grammars, so the chromaticity downsampling position can be independently defined based on color display effects.
[0036] Depending on the location, a special downsampling filter can be intentionally applied to luminance downsampling. Downsampling filter 1 (Filter-North):
number
number
number
number
number
number
number
[0037] In MF-CCLM mode, four new modes are added: MFLM-North, MFLM-East, MFLM-South, and MFLM-Center. For example, when MFLM-North is selected, a downsampling filter called Filter-North is used when reconstructing the chromaticity sample.
[0038] 5. Angular-coupled linear models (LAP, LM-Angular Prediction) The background for proposing LAP is that, conventionally, there were only two options for a single block unit: using CCLM or not using CCLM (using conventional angle intraprediction), and the goal was to choose one of these two options. Based on the multi-hypothesis motion compensation prediction concept, LAP combines angle intraprediction and CCLM, resulting in the following:
number
number
number
number
[0039] 6. Convolutional Cross-Component Model (CCCM) This is a technique used in ECM, employing a single 7-tap filter, which consists of 5 taps that are spatial element data with a plus sign shape, plus one nonlinear element P and one fundamental element B. predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B Here, P=(C*C+midVal)>>bitDepth, where bitDepth is the bit depth, midVal is the midpoint of the bit depth, C represents the luminance sample value at the corresponding position of the current chromaticity sample, N, S, E, and W are the neighboring sample values of the current luminance sample, predChromaVal is the predicted value of the current chromaticity sample, B is the scalar offset between input and output, >> is the bit shift operation, and c0, c1, c2, c3, c4, c5, and c6 are the filter coefficients.
[0040] 7. Gradient Linear Model (GLM) This is a technique used in ECM, which uses a luminance gradient G instead of a downsampled luminance sample L:
number
[0041] Compared to CCLM, GLM derives a linear model using the gradient of luminance samples rather than downsampling luminance values. That is, in the case of parameter derivation, the linear model is derived using the gradient G of the luminance samples, not the luminance samples after downsampling, but other design aspects of CCLM (e.g., parameter derivation, linear transformation of predicted samples) remain unchanged. Here, gradient calculation can be implemented using any of the filters shown in Figures 7 to 10.
[0042] As can be seen from the CCLM technology described above, CCLM assumes that there is a linear relationship between the chromaticity pixel value and the corresponding luminance pixel value for the same CU. Therefore, CCLM uses a single linear model to generate a predicted value for the corresponding chromaticity pixel from the reconstructed value of the luminance pixel. However, for CUs where luminance and chromaticity have a nonlinear correlation or where the linear correlation is poor, the residual of the CCLM prediction is large and the compression effect is low.
[0043] To address problems existing in related technologies and to expand the range of applications of video encoding and achieve better compression ratios, this invention provides an encoding / decoding method based on a more broadly applicable Cross-Component General Model (CCGM) prediction.
[0044] In the first embodiment, as shown in Figure 11, is a flowchart of an example of a decoding method in an embodiment of the present application, which may include the following steps S101 to S104. Step S101: Obtain coding information corresponding to the target coding block, the coding information including at least the index information of the target model, the values of each parameter, the luminance downsampling reconstruction value of the target coding block, and the target difference value.
[0045] In practice, the encoding side can determine the values of each parameter when the prediction error of the target model meets a set condition (e.g., the prediction error is smaller than the error threshold) based on the true chromaticity value, the luminance downsampling reconstruction value, the prediction error of the target model, and the mapping relationship between each parameter of the target model. This ensures the accuracy of the target model's chromaticity prediction for the target encoded block.
[0046] In one possible embodiment, to reduce computational complexity, an equal number of reference pixel points associated with the target coding block (e.g., within the target coding block or adjacent target coding blocks to the target coding block) can be selected depending on the number of parameters that need to be numerically determined, and the value of each parameter can be determined based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel points.
[0047] Step S102: Determine the target model based on the index information of the target model.
[0048] Here, the target model is constructed based on the activation parameter and the first model, and the first model is intended to represent the mapping relationship between the luminance downsampling reconstruction value and the chromaticity prediction value. The first model may be a chromaticity prediction model such as a curve model or a CCLM technology-related linear model. Specifically, the CCLM technology-related linear model may be any of CCLM (narrow sense), slope-adjusted CCLM, MM-CCLM, MF-CCLM, LAP, CCCM, or GLM.
[0049] In practice, the encoding side determines a first model to be used to predict the chromaticity of each pixel point within the current CU block (i.e., the target encoding block), and then corrects this first model with an activation function, so that the target model can use the nonlinear attribute of the activation function to handle the general correlation (i.e., linear vs. nonlinear) problem between luminance and chromaticity of CU blocks in encoding within a video frame.
[0050] For example, an activation function term can be introduced to correct the chromaticity prediction values determined by the first model, and if the introduced activation function term includes a higher-order term associated with the luminance downsampling reconstruction value, the target model can also be flexibly switched between different curved and / or linear models (i.e., the activation function zeros out all higher-order terms associated with the luminance downsampling reconstruction value) by the activation function, enabling the target model to achieve high chromaticity prediction accuracy for both CU blocks with a linear correlation between luminance and chromaticity and CU blocks with a nonlinear correlation.
[0051] Step S103: Determine the chromaticity prediction value of the target coding block based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block.
[0052] In practice, after determining the values of each parameter in the target model, the luminance downsampling reconstruction value of each pixel point within the target coding block is input to the target model with the determined parameter values, and the chromaticity prediction value of each pixel point output from the target model can be obtained.
[0053] To make it clear, the target model relating to this application is both a cross-component general-purpose model and a cross-component nonlinear model (CCnLM). By introducing an activation function, this target model can more effectively solve the problem of predicting CU chromaticity when there is a nonlinear correlation between luminance and chromaticity, while guaranteeing the accuracy of CU chromaticity prediction when there is a linear correlation between luminance and chromaticity.
[0054] Step S104: Based on the chromaticity prediction value of the target coding block and the target difference value, the chromaticity reconstruction value of the target coding block is determined.
[0055] In specific implementation, the encoding side determines the target model corresponding to the target encoding block and the values of each parameter in the target model. Based on the target model, the values of each variable parameter, and the luminance downsampling reconstruction value of the target encoding block, it determines the chromaticity prediction value of the target encoding block. Furthermore, based on the chromaticity prediction value of the target encoding block and the true chromaticity value of the target encoding block, it determines the target difference value. Finally, based on the index information of the target model, the values of each parameter, and the target difference value, it can generate encoding information corresponding to the target encoding block. Here, the luminance downsampling reconstruction value can be obtained by downsampling the luminance pixels of the video encoding (for example, luminance pixels whose sampling format is YUV420).
[0056] After obtaining the encoding information corresponding to the target encoding block and the luminance downsampling reconstruction value of the target encoding block, the decoding end can determine the target model constructed by the encoding side for the target encoding block based on the activation function and the first model, based on the index information of the target model. Immediately afterward, the encoding side recovers the chromaticity prediction value of the target encoding block determined by the encoding side based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target encoding block. Furthermore, based on the chromaticity prediction value of the target encoding block and the target difference value, the decoding of the chromaticity quantity information of the target encoding block is completed.
[0057] To make it easier to understand, the target model has good applicability to both CU blocks with linear and nonlinear correlations between luminance and chromaticity. Therefore, the encoding side can reduce the target difference values (e.g., residual values) that the target encoding block generates using the target model, and consequently the data size of the encoded information generated by the encoding side also decreases, thereby improving the compression ratio of video encoding.
[0058] As can be seen from the steps above, by introducing an activation function into the construction of the target model, the target model can have high chromaticity prediction accuracy for both target coding blocks with good and bad linear correlations between luminance and chromaticity. This improves the applicability of the chromaticity prediction method and also improves the compression effect of video coding.
[0059] Embodiment 1 This embodiment describes an example of target model construction.
[0060] In this embodiment, the differences between the first models used to construct the target model can be divided into the following two cases.
[0061] Case 1: The first model is intended to represent a linear mapping (linear correlation) relationship between the luminance downsampling reconstruction value and the chromaticity prediction value.
[0062] In specific implementation, the first model may be a linear model related to CCLM technology.
[0063] To make it easier to understand, CCLM technology reduces crossover redundancy by using a single linear model to generate predicted values for corresponding chromaticity pixels from reconstructed luminance pixel values, and the predicted values for chromaticity pixels are calculated as follows:
number
number
number
[0064] CCLM assumes that luminance and chromaticity have a linear correlation in local textures, but linear models may not accurately represent the correlation between luminance and chromaticity, limiting the applicability of CCLM.
[0065] As can be seen from the correspondence between luminance, chromaticity, and RGB shown below, the linear correlation between luminance and chromaticity in this case is relatively poor, and it is difficult to accurately represent the correlation between luminance and chromaticity using a linear model. Y = 0.299R + 0.587G + 0.114B U = -0.1687R - 0.3313G + 0.5B + 128 V = 0.5R - 0.4187G - 0.0813B + 128 Here, Y represents luminance (i.e., gradation value), U and V represent chromaticity and are used to describe the color and saturation of the image, and R, G, and B represent the three channels of color: red, green, and blue.
[0066] In case 1, the present invention optimizes the CCLM technology-related model by introducing an activation function, for example, by adding an activation function term to the CCLM technology-related model (i.e., the third equation representing the first model) that takes a luminance downsampling reconstruction value as an argument, and / or by using the activation function term to correct the parameter item in the third equation where at least one luminance downsampling reconstruction value is located, so that the optimized CCLM technology-related model (i.e., the target model) can also fit well to discrete points outside the line fitted to the linear model, thereby enabling highly accurate chromaticity prediction for CUs with relatively poor linear correlation between luminance and chromaticity, expanding the application range of CCLM, and increasing the compression ratio of video encoding.
[0067] Case 2: The first model is intended to represent the mapping (non-linear correlation) relationship between the luminance downsampling reconstruction value and the chromaticity prediction value.
[0068] In specific implementation, the first model can represent a nonlinear mapping relationship between the luminance downsampling reconstruction value and the chromaticity prediction value in a curve model related to chromaticity prediction.
[0069] To make it easier to understand, since the curve model can better illustrate the correlation between luminance and chromaticity, the activation function and the curve model can be combined to construct a target model. For example, the activation function can correct at least one parameter term in the third equation representing the first model, and it can also fit well to discrete points outside the curve that are fitted to the curve model, thereby further improving the chromaticity prediction accuracy of CUs where the linear correlation between luminance and chromaticity by the target model is relatively poor.
[0070] For example, in case 2, the target model can represent the following two nonlinear mapping relationships: Nonlinear mapping relationship 1:
number
number
number
number
[0071] Nonlinear mapping relationship 2:
number
number
number
number
[0072] To make it easier to understand, in nonlinear mapping relationship 2, the quadratic term associated with the luminance downsampling reconstruction value is corrected using an activation function. Therefore, when the activation function is zero, nonlinear mapping relationship 2 degenerates into the corresponding linear mapping (linear correlation) relationship, thereby enabling flexible switching of the target model between the linear and nonlinear models. Here, the intermediate bit depth is mainly used to ensure that the obtained chromaticity prediction value does not overflow, and the intermediate bit depth in the activation function is also used to perform rounding operations on the arguments of the activation function.
[0073] As one possible embodiment,
number
number
[0074] In Embodiment 1 described above, the activation function is not limited to the ReLU function; in specific implementations, the activation function may be any of the following: ReLU function, Sigmoid function, Tanh function, Leaky ReLU function, ELU function, PReLU function, Softmax function, Swish function, Maxout function, or Softplus function.
[0075] Embodiment 2 This embodiment describes an example of parameters in a target model.
[0076] In this embodiment, based on the true chromaticity value and luminance downsampling reconstruction value of the target coded block, the values of each parameter in the target model, and the chromaticity prediction error of the target coded block by the target model, the values of each variable parameter in the target model can be determined when the chromaticity prediction error of the target coded block by the target model is minimized.
[0077] In the above embodiment 2, using target models 1 and 2 as examples representing nonlinear mapping relationship 1 and nonlinear mapping relationship 2, respectively, the values of each variable parameter in target models 1 and 2 can be determined by the first or second equation based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block (for example, an adjacent pixel point of the target coding block).
[0078] The first equation above is as follows:
number
number
number
number
[0079] The second equation above is as follows:
number
number
number
number
[0080] To make it clear, the first and second equations above actually estimate the chromaticity prediction error of the target coded block by the target model by adding the chromaticity prediction error for each reference pixel point to the target model based on the Linear Minimum Mean Square Error (LMMSE). Here, the estimation of the chromaticity prediction error for the target coded block can also be determined by other forms of reference functions (i.e., loss functions), and the present invention is not limited to these.
[0081] Since the target model has four variable parameters, determining the values of these four variable parameters d requires at least four reference pixel points with true chromaticity values and luminance downsampling reconstruction values. On the other hand, the minimum size of a VVC video chromaticity block (i.e., an adjacent coding block for selecting adjacent pixel points of the target coding block) is 4x4, thus satisfying the minimum pixel point requirement necessary to determine the parameters in the target model.
[0082] Using the first equation above as an example, by differentiating it with respect to the parameter values α0, α1, α2, and α3 with respect to the Loss function, the following results are obtained.
number
[0083] Using the second example above, by differentiating with respect to the parameter values α0, α1, α2, and α3 with respect to the loss function, the following results are obtained.
number
number
number
number
[0084] In the above formulas 6.1 to 6.4 or formulas 7.1 to 7.4,
number
[0085] In one embodiment, considering that the function values of some activation function terms need to be jointly determined by multiple parameters, initial estimates of these multiple parameters can be determined first through relational equations that do not introduce activation function terms.
[0086] In practice, the values of each parameter in the third equation representing the first model (i.e., initial estimates) can be determined based on the true chromaticity and luminance downsampled reconstruction values of the target coded block. For example, a third mapping relationship can be determined between the error estimate corresponding to the third equation and the values of each parameter in the third equation, based on the true chromaticity and luminance downsampled reconstruction values of the target coded block (for example, by determining the third mapping relationship through a reference function), and then the values of each parameter in the third equation can be determined based on the third mapping relationship. As can be understood, the third equation can then be considered as the target model with the activation function related term removed, and the variable parameters of the third equation should include at least all parameters associated with that activation function related term.
[0087] After obtaining the values of each parameter in the third equation, the final values of each parameter in the target model can then be determined based on the true chromaticity and luminance downsampling reconstruction values of the target coded block, as well as the values of each parameter in the third equation. For example, the values of the target parameter items in the target model can first be determined through the initial estimates of each parameter, i.e., the values of the parameter items after modification by the activation function, and then the values of each parameter in the target model can be determined based on the true chromaticity and luminance downsampling reconstruction values of the target coded block, as well as the values of the target parameter items.
[0088] In one possible embodiment, a second mapping relationship can be determined between an error estimate corresponding to a target model and the values of each parameter in the target model, based on the true chromaticity and luminance downsampling reconstruction values of the target coding block, as well as the values of the target parameter items (for example, by substituting the values of the target parameter items into a reference function corresponding to the target model to determine the second mapping relationship), and further, the values of each parameter in the target model can be determined based on the second mapping relationship. Optionally, the error estimate corresponding to the target model may be a linear least mean squares error estimate, in which case the second mapping relationship can be determined based on LMMSE.
[0089] As an example, the determination of parameter values in the target model will be explained using the nonlinear mapping relationships 1 and 2 according to Embodiment 1 described above as an example.
[0090] For nonlinear mapping relationship 1, based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, a third mapping relationship can be determined by the fourth equation between the error estimate corresponding to the third equation associated with the nonlinear mapping relationship 1 and the values of each parameter in the third equation, where, The third equation associated with the nonlinear mapping relationship 1 is as follows:
number
number
number
number
number
number
[0091] For the nonlinear mapping relationship 2, based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, a third mapping relationship can be determined by the fifth equation between the error estimate corresponding to the third equation associated with the nonlinear mapping relationship 2 and the values of each parameter in the third equation, where, The third equation associated with the aforementioned nonlinear mapping relationship 2 is as follows:
number
number
number
number
number
number
[0092] Taking the fourth equation as an example, by differentiating it with respect to α0, α1, α2, and α3 based on the loss function, we obtain the following results.
number
number
number
[0093] In equations (8.1) to (8.4),
number
[0094] In this embodiment, the two parameter values α2 and α3 have the same coefficients in equations (8.1) to (8.4), and equations (8.3) and (8.4) are equivalent.
number
number
number
[0095] Using equations (9.1) to (9.3) above, we can obtain α0, α1, and α23. Furthermore, by setting α2 = (α23 >> 1) and α3 = α23 - α2, we can find α0, α1, α2, and α3 (i.e., the initial estimates). Next, we substitute α1 and α2 into equations (6.1) and (6.4) to approximate and find α0' and α3', and then find α2' using α2' = α23 + α3'.
[0096] Nonlinear function-related terms, for example, the ReLU function term.
number
number
number
[0097] Based on the determined α0', α1', α2', α3' and the constructed target model, that is, by substituting α0', α1', α2', α3' into the target model representing the nonlinear mapping relationship 1, chromaticity prediction of the target coded block in the video image can be achieved. This increases the applicability range of chromaticity prediction and the compression ratio of coded chromaticity blocks when a small amount of computational power is used.
[0098] In a second aspect, as shown in Figure 12, an embodiment of the present application provides an encoding method, which comprises at least: Step S201: A step of constructing a target model based on an activation function and a first model, wherein the first model is for representing a mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values. Step S202 determines the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, Step S203 determines the chromaticity prediction value of the target coding block based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block. Step S204, which determines the target difference value based on the chromaticity prediction value of the target coding block and the chromaticity true value of the target coding block, The process includes step S205, which generates coding information corresponding to the target coding block based on the index information of the target model, the values of each variable parameter, the luminance downsampling reconstruction value of the target coding block, and the target difference value.
[0099] In one possible embodiment, the first model is intended to represent a nonlinear mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values.
[0100] In one possible embodiment, the target model represents the following mapping relationship:
number
number
number
number
[0101] In one possible embodiment, determining the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block is: The process includes determining the values of each parameter in the target model using a first formula, based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block. The first equation above is as follows:
number
number
number
number
[0102] In one possible embodiment, the target model represents the following mapping relationship:
number
number
number
number
[0103] In one possible embodiment, determining the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block is: The process includes determining the values of each parameter in the target model using a second formula, based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block. The second equation above is as follows:
number
number
number
number
[0104] In one possible embodiment, the first model is represented by a third equation, and based on the activation function and the first model, the target model can be constructed as follows: This includes correcting at least one parameter term in the third equation using the activation function to obtain the target model.
[0105] In one possible embodiment, determining the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block is: Based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, the values of each parameter in the third equation are determined, This includes determining the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, and the values of each parameter in the third equation.
[0106] In one possible embodiment, the values of each parameter in the target model are determined based on the true chromaticity and luminance downsampling reconstruction values of the target coding block, and the values of each parameter in the third equation. The value of the target parameter item in the target model is determined by the value of each parameter in the third equation described above, wherein the target parameter item is a parameter item modified by an activation function, This includes determining the value of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, and the value of the target parameter item in the target model.
[0107] In one possible embodiment, the values of each parameter in the target model are determined based on the true chromaticity and luminance downsampling reconstruction values of the target coding block, and the values of the target parameter items in the target model. Based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, and the values of the target parameter items in the target model, a second mapping relationship is determined between the error estimate corresponding to the target model and the values of each parameter in the target model. This includes determining the values of each parameter in the target model based on the second mapping relationship described above.
[0108] In one possible embodiment, the error estimate corresponding to the target model is a linear least mean squares error estimate.
[0109] In one possible embodiment, determining the values of each parameter in the third equation based on the true chromaticity value and the luminance downsampling reconstruction value of the target coding block is: Based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, a third mapping relationship is determined between the error estimate corresponding to the third equation and the values of each parameter in the third equation. This includes determining the values of each parameter in the third equation based on the third mapping relationship.
[0110] In one possible embodiment, a third mapping relationship between the error estimate corresponding to the third equation and the values of each parameter in the third equation is determined based on the chromaticity true value and the luminance downsampling reconstruction value of the target coding block. The process includes determining a third mapping relationship between the error estimate corresponding to the third equation and the values of each parameter in the third equation, based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, using a fourth equation. Here, the third equation is as follows:
number
number
number
number
Number
[0111] As one possible embodiment, determining the third mapping relationship between the error estimation value corresponding to the third formula and the values of each parameter in the third formula based on the true chromaticity value and the luminance downsampling reconstruction value of the target encoding block includes determining the third mapping relationship between the error estimation value corresponding to the third formula and the values of each parameter in the third formula according to the fifth formula based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel points corresponding to the target encoding block, where the third formula is as follows
Number
Number
Number
Number
number
number
[0112] In one possible embodiment, the activation function is: ReLU function, Sigmoid function, Tanh function, Leaky ReLU function, ELU function, PReLU function, Softmax function, Swish function, Maxout function, It is one of the Softplus functions.
[0113] As can be seen from the steps above, by introducing an activation function into the construction of the target model, the target model can have high chromaticity prediction accuracy for both target coding blocks with good and bad linear correlations between luminance and chromaticity. This improves the applicability of the chromaticity prediction method and also improves the compression effect of video coding.
[0114] To ensure understanding, the encoding method and decoding method according to the embodiment of the present application may be implemented by an encoding device and a decoding device. In the embodiment of the present application, the encoding device and decoding device will be described as an example in which the encoding device and decoding device each execute the encoding method and decoding method.
[0115] In the third aspect, the embodiment of the present application provides a decoding device. As shown in FIG. 13, the decoding device 100 includes an information acquisition module 101 for acquiring encoded information corresponding to a target encoded block and information for obtaining a luminance downsampling reconstruction value of the target encoded block, where the encoded information includes at least index information of a target model, values of each parameter, the luminance downsampling reconstruction value of the target encoded block, and a target difference value; a model determination module 102 for determining a target model based on the index information of the target model, where the target model is constructed based on an activation parameter and a first model, and the first model is for representing a mapping relationship between a luminance downsampling reconstruction value and a chrominance prediction value; a chrominance prediction module 103 for determining a chrominance prediction value of the target encoded block based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target encoded block; and a chrominance reconstruction module 104 for determining a chrominance reconstruction value of the target encoded block based on the chrominance prediction value of the target encoded block and the target difference value.
[0116] Optionally, the first model is for representing a non-linear mapping relationship between a luminance downsampling reconstruction value and a chrominance prediction value.
[0117] Optionally, the target model is for representing the following mapping relationship,
Equation
Equation
number
number
[0118] Selectively, the value of each parameter in the target model is determined by a first equation based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, The first equation above is as follows:
number
number
number
number
[0119] Selectable, the target model represents the following mapping relationship:
number
number
number
number
[0120] Selectively, the values of each parameter in the target model are determined by a second equation based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block. The second equation above is as follows:
number
number
number
number
[0121] Selectively, the first model is represented by a third equation, and the target model is obtained by modifying at least one parameter term in the third equation with the activation function.
[0122] Selectively, the values of each parameter in the target model are determined based on the true chromaticity and luminance downsampling reconstruction values of the target coding block, and the values of each parameter in the third equation, the values of each parameter in the third equation are determined based on the true chromaticity and luminance downsampling reconstruction values of the target coding block.
[0123] Selectively, the values of each parameter in the target model are determined based on the true chromaticity and luminance downsampling reconstruction values of the target coding block, and the values of the target parameter items in the target model, the values of the target parameter items in the target model are determined by the values of each parameter in the third equation, and the target parameter items are parameter items modified by an activation function.
[0124] Selectively, the values of each parameter in the target model are determined based on a second mapping relationship, the second mapping relationship being a mapping relationship between the error estimate corresponding to the target model and the values of each parameter in the target model, determined based on the chromaticity true value and luminance downsampling reconstruction value of the target coding block, and the values of the target parameter items in the target model.
[0125] Selectively, the error estimate corresponding to the target model is a linear least mean squares error estimate.
[0126] Selectively, the values of each parameter in the third equation are determined based on a third mapping relationship, which is a mapping relationship between the error estimate corresponding to the third equation and the values of each parameter in the third equation, determined based on the chromaticity true value and luminance downsampling reconstruction value of the target coding block.
[0127] Selectively, the third mapping relationship is determined by a fourth equation based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, Here, the third equation is as follows:
number
number
number
number
number
number
[0128] Selectively, the third mapping relationship is determined by a fifth equation based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block. Here, the third equation is as follows:
number
number
number
number
number
number
[0129] Selectively, the activation function is ReLU function, Sigmoid function, Tanh function, Leaky ReLU function, ELU function, PReLU function, Softmax function, Swish function, Maxout function, It is one of the Softplus functions.
[0130] The decoding device according to the embodiment of the present application can implement various processes realized by the embodiment of the decoding method described in the first aspect and achieve the same technical effects, and to avoid duplication, it will not be described further here.
[0131] In a fourth aspect, an embodiment of the present application provides an encoding device, and as shown in Figure 14, the encoding device 200 is: A first construction module 201 for constructing a target model based on an activation function and a first model, wherein the first model represents a mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values, A first processing module 202 for determining the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, A second processing module 203 for determining the chromaticity prediction value of the target coding block based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block, A third processing module 204 for determining a target difference value based on the chromaticity prediction value of the target coding block and the chromaticity true value of the target coding block, The system includes a first generation module 205 for generating coding information corresponding to the target coding block based on the index information of the target model, the values of each of the variable parameters, the luminance downsampling reconstruction value of the target coding block, and the target difference value.
[0132] Selectively, the first model is intended to represent a nonlinear mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values.
[0133] Selectable, the target model represents the following mapping relationship:
number
number
number
number
[0134] Selectively, the first processing module 202 is: The system includes a first processor module for determining the values of each parameter in the target model by a first formula, based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, The first equation above is as follows:
number
number
number
number
[0135] Selectable, the target model represents the following mapping relationship:
number
number
number
number
[0136] Selectively, the first processing module 202 is: The system includes a second processing submodule for determining the values of each parameter in the target model by a second equation, based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, The second equation above is as follows:
number
number
number
number
[0137] Selectively, the first model is represented by the third equation, and the first construction module 201 is, The system includes a first constructor module that corrects at least one parameter term in the third equation using the activation function to obtain the target model.
[0138] Selectively, the first processing module 202 is: A third processing submodule for determining the values of each parameter in the third equation based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, The system includes a fourth processing submodule for determining the values of each parameter in the target model based on the true chromaticity and luminance downsampling reconstruction values of the target coding block, and the values of each parameter in the third equation.
[0139] Selectively, the fourth processing submodule is: A fifth processing submodule that determines the value of a target parameter item in the target model based on the values of each parameter in the third equation, wherein the target parameter item is a parameter item modified by an activation function, The system includes a sixth processing submodule for determining the values of each parameter in the target model based on the true chromaticity and luminance downsampling reconstruction values of the target coding block, and the values of the target parameter items in the target model.
[0140] Selectively, the sixth processing submodule is: A seventh processing submodule for determining a second mapping relationship between an error estimate corresponding to the target model and the values of each parameter in the target model, based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, and the values of the target parameter items in the target model, It includes an eighth processing submodule for determining the values of each parameter in the target model based on the second mapping relationship described above.
[0141] Selectively, the error estimate corresponding to the target model is a linear least mean squares error estimate.
[0142] Selectively, the third processing submodule is: A ninth processing submodule for determining a third mapping relationship between an error estimate corresponding to the third equation and the values of each parameter in the third equation, based on the true chromaticity value and the luminance downsampling reconstruction value of the target coding block, It includes a tenth processing submodule for determining the values of each parameter in the third equation based on the third mapping relationship.
[0143] Selectively, the ninth processing submodule is: The system includes an eleventh processing submodule for determining a third mapping relationship between the error estimate corresponding to the third equation and the values of each parameter in the third equation, based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, using a fourth equation. Here, the third equation is as follows:
number
number
number
number
number
number
[0144] Selectively, the ninth processing submodule is: A twelfth processing submodule is included for determining a third mapping relationship between the error estimate corresponding to the third equation and the values of each parameter in the third equation, based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, using a fifth equation. Here, the third equation is as follows:
number
number
number
number
number
number
[0145] Selectively, the activation function is ReLU function, Sigmoid function, Tanh function, Leaky ReLU function, ELU function, PReLU function, Softmax function, Swish function, Maxout function, It is one of the Softplus functions.
[0146] The encoding apparatus according to the embodiment of the present application can realize various processes realized by the encoding method embodiment described in the second embodiment and achieve the same technical effects, and to avoid duplication, it will not be described further here.
[0147] The apparatus described above is merely illustrative, and in it, the units described as separating members may or may not be physically separated, and the units shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. The objective of this embodiment can be realized by selecting some or all of its modules as needed. Those skilled in the art will understand and implement it without any creative work.
[0148] Each component of the present disclosure can be implemented in hardware, or in software modules running on one or more processors, or in combination thereof. Those skilled in the art will understand that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all components of the computing devices relating to embodiments of the present disclosure. The present disclosure can be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for performing some or all of the methods described herein. Such a program implementing the present disclosure may be stored on a computer-readable medium or may be in the form of one or more signals. Such signals may be downloaded from an internet site, provided on a carrier signal, or provided in any other form.
[0149] For example, Figure 15 shows a computing device capable of implementing the method of the present disclosure. This computing device conventionally includes a processor 1010 and a computer program product or computer-readable medium in the form of a memory 1020. The memory 1020 may be electronic memory such as flash memory, EEPROM (electrically erasable and programmable read-only memory), EPROM, hard disk, or ROM. The memory 1020 has a storage space 1030 for program code 1031 to perform any method step of the method. For example, the storage space 1030 for program code may include each program code 1031 to implement each step of the method. These program codes can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards, or floppy disks. Such computer program products are typically portable or fixed storage units, as shown in Figure 16. The storage unit may have storage segments, storage spaces, etc., with a configuration similar to that of memory 1020 in the computing device shown in Figure 15. The program code may be compressed in an appropriate format, for example. Typically, the storage unit contains computer-readable code 1031', i.e., code readable by a processor such as 1010, and when this code is executed by the computing device, it causes the computing device to perform each step in the manner described above.
[0150] Embodiments of the present application further provide a decoding device comprising a processor and memory, wherein the memory stores a program or instruction executable on the processor, and when the program or instruction is executed by the processor, each process of the embodiment of the decoding method described above is realized, achieving the same technical effect. To avoid duplication, no further explanation is provided here.
[0151] Embodiments of the present application further provide an encoding device comprising a processor and memory, wherein the memory stores a program or instruction executable on the processor, and when the program or instruction is executed by the processor, each process of the embodiment of the encoding method described above is realized, achieving the same technical effect. To avoid duplication, no further explanation is provided here.
[0152] Embodiments of the present application further provide a readable storage medium in which a program or instruction is stored, and when the program or instruction is executed by a processor, the processes of the above-described embodiment of the encoding method or the above-described embodiment of the decoding method can be realized, achieving the same technical effect. To avoid duplication, no further explanation is provided here.
[0153] Here, the processor is the processor in the terminal device described in the above embodiment. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disk, or optical disk.
[0154] Embodiments of the present application further provide a chip comprising a processor and a communication interface, wherein the communication interface and the processor are coupled, and the processor executes a program or instructions to realize each process of the embodiment of the encoding method or the embodiment of the decoding method, thereby achieving the same technical effect. To avoid duplication, no further explanation is provided here.
[0155] It should be understood that the chip according to the embodiment of this application may also be called a system-level chip, system chip, chip system, or on-chip system chip.
[0156] Embodiments of the present application further provide a computer program / program product which is stored in a storage medium and executed by at least one processor to implement each process of the above-described embodiment of the encoding method or the above-described embodiment of the decoding method, thereby achieving the same technical effect. To avoid duplication, no further explanation is provided here.
[0157] Embodiments of the present application further provide an encoding / decoding system comprising an encoding device and a decoding device, wherein the encoding device is for performing the steps of the encoding method described in the second embodiment, and the decoding device is for performing the steps of the decoding method described in the first embodiment.
[0158] In this specification, the terms “includes,” “incorporates,” or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus containing a set of elements includes not only those elements but also other elements not expressly listed, or elements specific to such process, method, article, or apparatus. Unless otherwise limited, an element defined by the phrase “includes one” does not preclude the presence of another identical element in a process, method, article, or apparatus containing that element. Furthermore, the scope of methods and apparatus in embodiments of this application is not limited to performing functions in the order illustrated or discussed, but may include performing functions essentially concurrently or in reverse order based on the relevant functions; for example, a described method may be performed in an order different from the described order, and various steps may be added, omitted, or combined. Furthermore, features described with reference to some examples may be combined in other examples.
[0159] From the above description of the embodiments, those skilled in the art will clearly understand that the methods of the above embodiments can be implemented by adding a necessary general-purpose hardware platform in addition to the software, and of course by hardware alone, although in many cases the former is a better embodiment. Based on this understanding, the essential or prior art contribution of the technical means of the present application can be embodied in the form of a computer software product, which is stored in a single storage medium (e.g., ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a single terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to perform the methods described in each embodiment of the present application.
[0160] Although embodiments of the present application have been described above with reference to the drawings, the present application is not limited to the embodiments described above. The embodiments described above are merely illustrative and not limiting. A person skilled in the art can take many forms that fall within the scope of protection of the present application without departing from the scope protected by the spirit and claims of the present application, under the disclosure of the present application.
[0161] In this specification, “one example,” “example,” or “one or more examples” means that a particular feature, structure, or property described with reference to an example is included in at least one example of this disclosure. However, the phrase “in one example” does not necessarily refer to the same example.
[0162] The instructions provided herein describe many specific details. However, it is understood that the embodiments of this disclosure can be implemented without these specific details. In some examples, known methods, structures, and techniques are not described in detail so as not to obscure the understanding of this specification.
[0163] In the claims, no reference numerals between parentheses limit the claims. The term “including” does not preclude the existence of any element or step not described in the claims. The term “one” or “one” preceding an element does not preclude the existence of multiple such elements. This disclosure may be implemented by hardware comprising several different elements and a appropriately programmed computer. In unit claims that enumerate several devices, some of these devices may be embodied by the same hardware item. The use of terms such as first, second, and third does not indicate any order. These words can be interpreted as names.
[0164] Finally, the above embodiments are for illustrative purposes only and not limiting to the technical solutions of the present disclosure. While the present disclosure has been described in detail with reference to the above embodiments, those skilled in the art should understand that they may modify the technical solutions described in each of the above embodiments or replace some of their technical features with equivalent ones. Such modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present disclosure. [Explanation of Symbols]
[0165] 100 Decoders 101 Information Acquisition Module 102 Model Decision Module 103 Chromaticity Prediction Module 104 Chromaticity Reconstruction Module 200 Encoding device 201 First Construction Module 202 First Processing Module 203 Second processing module 204 Third Processing Module 205 First generation module
Claims
1. The method involves obtaining coding information corresponding to a target coding block, wherein the coding information includes at least the index information of the target model, the values of each parameter, the luminance downsampling reconstruction value of the target coding block, and the target difference value. The method involves determining a target model based on the index information of the target model, wherein the target model is constructed based on activation parameters and a first model, and the first model is intended to represent the mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values. Based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block, the chromaticity prediction value of the target coding block is determined. A decoding method characterized by determining a chromaticity reconstruction value of the target coding block based on the chromaticity prediction value of the target coding block and the target difference value.
2. The method according to claim 1, characterized in that the first model is for representing a nonlinear mapping relationship between a luminance downsampling reconstruction value and a chromaticity prediction value.
3. The aforementioned target model is intended to represent the following mapping relationships: [Math 1] Here, [Math 2] is the activation function, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Math 3] This is the predicted chromaticity value of the pixel point at coordinate (i, j), [Math 4] The method according to claim 2, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, and >> is a bit shift operation.
4. The values of each parameter in the target model are determined by the first equation, based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block. The first equation is as follows: [Math 5] Here, Loss is the linear least mean squared error estimate. [Math 6] is the activation function, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 7] This is the true chromaticity value of the pixel point at coordinate (i, j), [Number 8] The method according to the previous version, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, the value range of coordinate (i, j) is the coordinate value range of a reference pixel point corresponding to the target coding block, and >> is a bit shift operation.
5. The aforementioned target model is intended to represent the following mapping relationships: [Number 9] Here, [Number 10] is the activation function, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Math 11] This is the predicted chromaticity value of the pixel point at coordinate (i, j), [Math 12] The method according to claim 2, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, and >> is a bit shift operation.
6. The values of each parameter in the target model are determined by a second equation based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block. The second equation above is as follows: [Number 13] Here, Loss is the linear least mean squared error estimate. [Number 14] is the activation function, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 15] This is the true chromaticity value of the pixel point at coordinate (i, j), [Number 16] The method according to 5, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, the value range of coordinate (i, j) is the coordinate value range of a reference pixel point corresponding to the target coding block, and >> is a bit shift operation.
7. The method according to claim 1, characterized in that the first model is represented by a third equation, and the target model is obtained by modifying at least one parameter in the third equation with the activation function.
8. The method according to 7, characterized in that the values of each parameter in the target model are determined based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, and the values of each parameter in the third equation, and the values of each parameter in the third equation are determined based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block.
9. The method according to 8, characterized in that the value of each parameter in the target model is determined based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, and the value of the target parameter item in the target model, the value of the target parameter item in the target model is determined by the value of each parameter in the third equation, and the target parameter item is a parameter item modified by an activation function.
10. The method according to 9, characterized in that the values of each parameter in the target model are determined based on a second mapping relationship, the second mapping relationship being a mapping relationship between an error estimate corresponding to the target model and the values of each parameter in the target model, determined based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, and the values of the target parameter items in the target model.
11. The method according to 10, characterized in that the error estimate corresponding to the target model is a linear least mean squares error estimate.
12. The method according to 8, characterized in that the values of each parameter in the third equation are determined based on a third mapping relationship, the third mapping relationship being a mapping relationship between an error estimate corresponding to the third equation, which is determined based on the true chromaticity value and the luminance downsampling reconstruction value of the target coding block, and the values of each parameter in the third equation.
13. The third mapping relationship is determined by the fourth equation based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block. Here, the third equation is as follows: [Number 17] Here, [Number 18] is the predicted chromaticity value of the pixel point at coordinate (i, j), bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 19] is the luminance downsampling reconstruction value of the pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, and >> is a bit shift operation. The fourth equation is as follows: [Number 20] Here, Loss is the linear least mean squared error estimate, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Math 21] This is the true chromaticity value of the pixel point at coordinate (i, j), [Number 22] The method according to 12, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, the value range of coordinate (i, j) is the coordinate value range of a reference pixel point corresponding to the target coding block, and >> is a bit shift operation.
14. The third mapping relationship is determined by the fifth equation based on the true chromaticity value and luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block. Here, the third equation is as follows: [Number 23] Here, [Number 24] is the predicted chromaticity value of the pixel point at coordinate (i, j), bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 25] is the luminance downsampling reconstruction value of the pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, and >> is a bit shift operation. The fifth equation is as follows: [Number 26] Here, Loss is the linear least mean squared error estimate, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 27] This is the true chromaticity value of the pixel point at coordinate (i, j), [Number 28] The method according to 12, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, the value range of coordinate (i, j) is the coordinate value range of a reference pixel point corresponding to the target coding block, and >> is a bit shift operation.
15. The aforementioned activation function is, ReLU function, Sigmod function, Tanh function, Leaky ReLU function, ELU function, PReLU function, Softmax function, Swish function, Maxout function, The method according to any one of claims 1 to 14, characterized in that it is one of the Softplus functions.
16. The method involves constructing a target model based on an activation function and a first model, wherein the first model represents a mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values. Based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, the values of each parameter in the target model are determined, Based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block, the chromaticity prediction value of the target coding block is determined. Based on the predicted chromaticity value of the target coded block and the true chromaticity value of the target coded block, the target difference value is determined. An encoding method characterized by generating encoding information corresponding to the target encoding block based on the index information of the target model, the values of each of the variable parameters, the luminance downsampling reconstruction value of the target encoding block, and the target difference value.
17. The method according to 16, characterized in that the first model is for representing a nonlinear mapping relationship between a luminance downsampling reconstruction value and a chromaticity prediction value.
18. The aforementioned target model is intended to represent the following mapping relationships: [Number 29] Here, [Number 30] is the activation function, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 31] This is the predicted chromaticity value of the pixel point at coordinate (i, j), [Number 32] The method according to 17, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, and >> is a bit shift operation.
19. Determining the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block is: The process includes determining the values of each parameter in the target model using a first formula, based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block. The first equation is as follows: [Number 33] Here, Loss is the linear least mean squared error estimate. [Number 34] is the activation function, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 35] This is the true chromaticity value of the pixel point at coordinate (i, j), [Number 36] The method according to 18, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, the value range of coordinate (i, j) is the coordinate value range of a reference pixel point corresponding to the target coding block, and >> is a bit shift operation.
20. The aforementioned target model is intended to represent the following mapping relationships: [Number 37] Here, [Number 38] is the activation function, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 39] This is the predicted chromaticity value of the pixel point at coordinate (i, j), [Number 40] The method according to 17, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, and >> is a bit shift operation.
21. Determining the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block is: The process includes determining the values of each parameter in the target model using a second equation, based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, The second equation above is as follows: [Number 41] Here, Loss is the linear least mean squared error estimate. [Number 42] is the activation function, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 43] This is the true chromaticity value of the pixel point at coordinate (i, j), [Number 44] The method according to 20, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, the value range of coordinate (i, j) is the coordinate value range of a reference pixel point corresponding to the target coding block, and >> is a bit shift operation.
22. The first model is represented by the third equation, and based on the activation function and the first model, the target model can be constructed as follows: The method according to 16, characterized in that it includes correcting at least one parameter item in the third equation with the activation function to obtain the target model.
23. Determining the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block is: Based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, the values of each parameter in the third equation are determined, The method according to 22, characterized in that it includes determining the value of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, and the value of each parameter in the third formula.
24. Determining the values of each parameter in the target model based on the true chromaticity and luminance downsampling reconstruction values of the target coding block, and the values of each parameter in the third equation, The value of the target parameter item in the target model is determined by the value of each parameter in the third equation described above, wherein the target parameter item is a parameter item modified by an activation function. The method according to 23, characterized in that it includes determining the value of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, and the value of the target parameter item in the target model.
25. Determining the values of each parameter in the target model based on the true chromaticity and luminance downsampling reconstruction values of the target coding block, and the values of the target parameter items in the target model, Based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, and the values of the target parameter items in the target model, a second mapping relationship is determined between the error estimate corresponding to the target model and the values of each parameter in the target model. The method according to 24, characterized in that it includes determining the values of each parameter in the target model based on the second mapping relationship.
26. The method according to 25, characterized in that the error estimate corresponding to the target model is a linear least mean squares error estimate.
27. Determining the values of each parameter in the third equation based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block is: Based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, a third mapping relationship is determined between the error estimate corresponding to the third equation and the values of each parameter in the third equation. The method according to 23, characterized in that it includes determining the values of each parameter in the third equation based on the third mapping relationship.
28. Determining a third mapping relationship between the error estimate corresponding to the third equation and the values of each parameter in the third equation, based on the true chromaticity value and the luminance downsampling reconstruction value of the target coding block, is: The process includes determining a third mapping relationship between the error estimate corresponding to the third equation and the values of each parameter in the third equation, based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, using a fourth equation. Here, the third equation is as follows: [Number 45] Here, [Number 46] is the predicted chromaticity value of the pixel point at coordinate (i, j), bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 47] is the luminance downsampling reconstruction value of the pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, and >> is a bit shift operation. The fourth equation is as follows: [Number 48] Here, Loss is the linear least mean squared error estimate, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 49] This is the true chromaticity value of the pixel point at coordinate (i, j), [Number 50] The method according to 27, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, the value range of coordinate (i, j) is the coordinate value range of a reference pixel point corresponding to the target coding block, and >> is a bit shift operation.
29. Determining a third mapping relationship between the error estimate corresponding to the third equation and the values of each parameter in the third equation, based on the true chromaticity value and the luminance downsampling reconstruction value of the target coding block, is: The process includes determining a third mapping relationship between the error estimate corresponding to the third equation and the values of each parameter in the third equation, based on the true chromaticity value and the luminance downsampling reconstruction value of the reference pixel point corresponding to the target coding block, using a fifth equation. Here, the third equation is as follows: [Number 51] Here, [Number 52] is the predicted chromaticity value of the pixel point at coordinate (i, j), bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 53] is the luminance downsampling reconstruction value of the pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, and >> is a bit shift operation. The fifth equation is as follows: [Number 54] Here, Loss is the linear least mean squared error estimate, bitDepth is the bit depth, and midValue is the midpoint of the bit depth. [Number 55] This is the true chromaticity value of the pixel point at coordinate (i, j), [Number 56] The method according to 27, characterized in that is the luminance downsampling reconstruction value of a pixel point at coordinate (i, j), α0, α1, α2, α3 are parameter values, the value range of coordinate (i, j) is the coordinate value range of a reference pixel point corresponding to the target coding block, and >> is a bit shift operation.
30. The aforementioned activation function is, ReLU function, Sigmod function, Tanh function, Leaky ReLU function, ELU function, PReLU function, Softmax function, Swish function, Maxout function, The method according to any one of claims 16 to 29, characterized in that it is one of the Softplus functions.
31. An information acquisition module for obtaining encoding information corresponding to a target encoding block and a luminance downsampling reconstruction value of the target encoding block, wherein the encoding information includes at least the index information of the target model, the values of each parameter, the luminance downsampling reconstruction value of the target encoding block, and the target difference value. A model determination module for determining a target model based on the index information of the target model, wherein the target model is constructed based on activation parameters and a first model, and the first model represents a mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values. A chromaticity prediction module for determining the chromaticity prediction value of the target coding block based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block, A decoding device comprising a chromaticity reconstruction module for determining the chromaticity reconstruction value of the target coding block based on the chromaticity prediction value of the target coding block and the target difference value.
32. A first construction module for constructing a target model based on an activation function and a first model, wherein the first model represents a mapping relationship between luminance downsampling reconstruction values and chromaticity prediction values. A first processing module for determining the values of each parameter in the target model based on the true chromaticity value and luminance downsampling reconstruction value of the target coding block, A second processing module for determining the chromaticity prediction value of the target coding block based on the target model, the values of each parameter, and the luminance downsampling reconstruction value of the target coding block, A third processing module for determining a target difference value based on the chromaticity prediction value of the target coding block and the chromaticity true value of the target coding block, An encoding device comprising a first generation module for generating encoding information corresponding to the target encoding block based on the index information of the target model, the values of each of the variable parameters, the luminance downsampling reconstruction value of the target encoding block, and the target difference value.
33. A decoding device comprising a processor and memory, wherein the memory stores a program or instruction executable on the processor, and when the program or instruction is executed by the processor, the steps of the decoding method described in any one of claims 1 to 15 are realized.
34. An encoding device comprising a processor and memory, wherein the memory stores a program or instruction executable on the processor, and when the program or instruction is executed by the processor, it realizes a step of the encoding method according to any one of claims 16 to 30.
35. Memory in which computer-readable code is stored, Includes one or more processors, A computing device characterized in that, when the computer-readable code is executed by one or more processors, it executes the decoding method according to any one of claims 1 to 15 or the encoding method according to any one of claims 16 to 30.
36. A computer program including a computer-readable code, characterized in that when the computer-readable code is executed on a computing device, the computing device causes the computing device to execute the decoding method described in any one of claims 1 to 15 or the encoding method described in any one of claims 16 to 30.
37. A computer-readable medium characterized by storing the computer program described in claim 36.