Coding method, coding device and computer readable storage medium

By learning the dependence between the current template and the dependency template using pre-trained target neural network, the problem of prediction inaccurate in the prior art is solved, and more efficient video encoding accuracy and compression efficiency are achieved.

CN115988200BActive Publication Date: 2025-08-29ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211471302.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-08-29
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

In the existing video encoding technology, linear prediction methods lead to inaccurate predictions, making it difficult to effectively improve the accuracy and compression efficiency of the encoding process.

Method used

Using a pre-trained target neural network, a complex linear or nonlinear relationship is established through the dependency learning between the current template and the dependency template, and the target prediction value of the pixel to be encoded is determined.

Benefits of technology

It improves the accuracy and encoding efficiency of the prediction results, and improves the compression effect of video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115988200B_ABST
    Figure CN115988200B_ABST
Patent Text Reader

Abstract

The present application discloses an encoding method, encoding device, and computer-readable storage medium. The encoding method includes: obtaining a dependent block, a current template, and a dependent template of a current block, wherein the current template includes multiple reconstructed pixel points around the current block, and the dependent template includes multiple reconstructed pixel points around the dependent block; inputting the dependent block, the current template, and the dependent template into a pre-trained target neural network to obtain a target prediction value of the pixel point to be encoded in the current block; wherein the target neural network determines the dependency relationship between the dependent template and the current template based on the reconstructed pixel values ​​of the pixel points in the current template and the dependent template, and determines the target prediction value of the pixel point to be encoded based on the reconstructed pixel values ​​of the pixel points in the dependent block and the dependency relationship. The encoding method provided by the present application can improve the accuracy of prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of video coding technology, and in particular relates to a coding method, a coding device, and a computer-readable storage medium. Background Art

[0002] Video image data is relatively large and usually needs to be encoded and compressed before transmission or storage. The encoded data is called a video stream, which is transmitted to the user end via a wired or wireless network and then decoded and viewed by the user end.

[0003] Currently, when encoding video image data, simple linear prediction methods are usually used directly, such as cross-component linear prediction technology and local illumination compensation technology, which can easily lead to inaccurate predictions. Therefore, the current prediction process needs to be further improved. Summary of the Invention

[0004] The present application provides a coding method, a coding device and a computer-readable storage medium, which can improve the accuracy of prediction.

[0005] A first aspect of an embodiment of the present application provides an encoding method, which includes: obtaining a dependent block, a current template, and a dependent template of a current block, wherein the current template includes multiple reconstructed pixel points around the current block, and the dependent template includes multiple reconstructed pixel points around the dependent block; inputting the dependent block, the current template, and the dependent template into a pre-trained target neural network to obtain a target prediction value of the pixel point to be encoded in the current block; wherein the target neural network determines the dependency relationship between the dependent template and the current template based on the current template and the reconstructed pixel values ​​of the pixel points in the dependent template, and determines the target prediction value of the pixel point to be encoded based on the reconstructed pixel values ​​of the pixel points in the dependent block and the dependency relationship.

[0006] A second aspect of an embodiment of the present application provides a decoding method, which includes: receiving encoded data sent by an encoder; decoding the encoded data to obtain a predicted value of a current pixel in a current decoding block; wherein the predicted value of the current pixel in the current decoding block is obtained by processing using any of the above-mentioned encoding methods.

[0007] A third aspect of an embodiment of the present application provides an encoding device, which includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit, respectively. Program data is stored in the memory. The processor implements the steps of any of the above methods by executing the program data in the memory.

[0008] A fourth aspect of an embodiment of the present application provides a decoding device, which includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit, respectively. Program data is stored in the memory. The processor implements the steps of any of the above methods by executing the program data in the memory.

[0009] A fifth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program can be executed by a processor to implement the steps in the above method.

[0010] Beneficial effects: This application uses a pre-trained target neural network to learn the dependency relationship between the current template and the dependent template, which can establish a more complex linear or nonlinear relationship, so that the established dependency relationship is more in line with the actual situation of the image, thereby making the prediction results accurate, improving the accuracy of the prediction results, and improving compression efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which:

[0012] Figure 1 This is a flow chart of an implementation method of the present application;

[0013] Figure 2 It is a schematic diagram of the current block and dependent blocks in an application scenario;

[0014] Figure 3 It is a schematic diagram of the current block and dependent blocks in another application scenario;

[0015] Figure 4 This is a schematic diagram of the current block and dependent blocks in another application scenario;

[0016] Figure 5 It is a relative diagram of the current block and the current template;

[0017] Figure 6 This is a schematic diagram of the structure of an embodiment of the target neural network of the present application;

[0018] Figure 7 yes Figure 6 A schematic diagram of the specific structure of the target neural network in an example;

[0019] Figure 8 yes Figure 7Schematic diagram of the structure of the residual block in the enhancement unit;

[0020] Figure 9 This is a partial flow chart of another embodiment of the encoding method of the present application;

[0021] Figure 10 This is a partial flow chart of another embodiment of the encoding method of the present application;

[0022] Figure 11 This is a flowchart of an implementation method of the present invention;

[0023] Figure 12 This is a schematic structural diagram of an embodiment of an encoder of the present application;

[0024] Figure 13 It is a structural diagram of another embodiment of the encoder of the present application;

[0025] Figure 14 This is a schematic structural diagram of an embodiment of a decoder of the present application;

[0026] Figure 15 It is a structural diagram of another embodiment of the decoder of the present application;

[0027] Figure 16 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0028] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0029] It should be noted that the terms "first" and "second" in this application are only used for descriptive purposes and should not be understood as indicating or suggesting relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.

[0030] See Figure 1 , Figure 1 : is a flow chart of an embodiment of the encoding method of the present application, the encoding method comprising:

[0031] S110: Obtain the dependent blocks, current template, and dependent templates of the current block.

[0032] Specifically, the current block refers to the coding block currently to be encoded, which can also be called the current coding block. The current block can be a luminance block or a chrominance block, wherein the video frame in which the current block is located is defined as the current frame. The current template includes multiple reconstructed pixels around the current block.

[0033] The dependent block includes multiple reconstructed pixel points, and the dependent block and the current block can be in the same video frame or in different video frames. Specifically, when the current block is predicted using inter-frame prediction, the dependent block is in the reference frame; when the current block is predicted using intra-frame prediction, the dependent block is in the current frame.

[0034] The dependent template includes a plurality of reconstructed pixels surrounding the dependent block, and the distribution of the reconstructed pixels included in the current template relative to the current block is the same as the distribution of the reconstructed pixels included in the dependent template relative to the dependent block. For example, when encoding the coding blocks in the image frame sequentially from left to right and from top to bottom, if the reconstructed pixels included in the current template are distributed to the left and above the current block, then the reconstructed pixels included in the dependent template are also distributed to the left and above the dependent block.

[0035] In this embodiment, the dependent block is a cross-component reconstructed block corresponding to the current block, a reconstructed block adjacent to the current block in the current frame, a target reference block corresponding to the current block in the reference frame, or a cross-component reconstructed block corresponding to the target reference block.

[0036] Specifically, the cross-component reconstructed block corresponding to the current block corresponds to the corresponding area in the image of the current block, but the color components corresponding to the cross-component reconstructed block and the current block are different. Figure 2 When the dependent block is a cross-component reconstructed block corresponding to the current block, if the current block is a chroma block, the dependent block may be a luminance block corresponding to the chroma block or another chroma block corresponding to the chroma block; or if the current block is a luminance block, the dependent block may be a chroma block corresponding to the luminance block, wherein the chroma block may be a Cb chroma block or a Cr chroma block. It will be understood that when the dependent block is a cross-component reconstructed block corresponding to the current block, the dependent block is also in the current frame.

[0037] When the dependent block is a reconstructed block related to the current block in the current frame, combined with Figure 3If the coding blocks in the image frame are encoded in the order from left to right and from top to bottom, the dependent block can be Figure 3 However, in other implementations, when the dependent block is a reconstructed block related to the current block in the current frame, the dependent block may also be another reconstructed block in the current frame that has spatial correlation with the current block, and this is not limited here.

[0038] Combine Figure 4 The target reference block corresponding to the current block in the reference frame is a reconstructed block in the reference frame that has temporal correlation with the current block. The process of determining the target reference block includes determining the reconstructed block closest to the current coding block in the reference frame through methods such as motion search. This reconstructed block is the target reference block. The color components of the target reference block are the same as those of the current block. That is, if the current block is a luminance block, the target reference block is also a luminance block. If the current block is a Cb chrominance block, the target reference block is also a Cb chrominance block.

[0039] The cross-component reconstructed block corresponding to the target reference block is identical to the cross-component reconstructed block corresponding to the current block, and the cross-component reconstructed block corresponding to the target reference block corresponds to a corresponding region in the image, except that the cross-component reconstructed block and the target reference block correspond to different color components. For example, when the target reference block is a luma block, the cross-component reconstructed block is the chroma block corresponding to the luma block; when the target reference block is a Cb chroma block, the cross-component reconstructed block can be either the corresponding luma block or the corresponding Cr chroma block.

[0040] Among them, the dependent block obtained in step S110 can be a cross-component reconstructed block corresponding to the current block, a reconstructed block related to the current block in the current frame, a target reference block corresponding to the current block in the reference frame, or a cross-component reconstructed block corresponding to the target reference block. The specific setting can be made according to needs and is not limited here.

[0041] It can be understood that when it is necessary to combine the inter-frame prediction technology to predict the current block, the dependent block and the current frame are in different video frames respectively. When it is necessary to combine the intra-frame prediction technology to predict the current block, the dependent block and the current block are in the same video frame.

[0042] In this embodiment, step S110 includes:

[0043] S111: Determine the dependent blocks of the current block.

[0044] S112: Determine a current template and a dependent template based on reconstructed pixel points in the same target area outside the current block and outside the dependent block respectively.

[0045] Among them, the target area outside the current block includes at least one of the first sub-area, the second sub-area, the third sub-area, the fourth sub-area and the fifth sub-area outside the current block, the first sub-area is located on the first side of the current block and its two ends are flush with the two ends of the current block, the second sub-area is located on the second side of the current block and its two ends are flush with the two ends of the current block, the third sub-area connects the first sub-area and the second sub-area, the fourth sub-area is located on the side of the first sub-area away from the third sub-area, and the fifth sub-area is located on the side of the second sub-area away from the third sub-area.

[0046] Specifically, the area where the current template is distributed relative to the current block is the same as the area where the dependent template is distributed relative to the dependent block, so only the current template is introduced below:

[0047] Combine Figure 5 The current template includes reconstructed pixels of at least one of the first, second, third, fourth, and fifth subregions outside the current block. When encoding the coding blocks in the image frame sequentially from left to right and from top to bottom, the first and fourth subregions are located above the current block, and the second and fifth subregions are located to the left of the current block.

[0048] The width and height of the first sub-region and the fourth sub-region may be equal to each other, and the width and height of the second sub-region and the fifth sub-region may be equal to each other, which is not limited here.

[0049] At the same time, this application does not limit the height of the first sub-region and the fourth sub-region, which may include multiple rows of reconstructed pixel points, and does not limit the width of the second sub-region and the fifth sub-region, which may include multiple columns of reconstructed pixel points.

[0050] In the prior art, the current template usually only includes a row of pixels and / or a column of pixels adjacent to the current block, and cannot utilize more spatial information, resulting in inaccurate determined dependencies. However, the present application sets the target area in the above manner, which can make full use of the spatial information and ensure the accuracy of the final prediction result.

[0051] S120: Input the dependent block, the current template, and the dependent template into a pre-trained target neural network to obtain a target prediction value of the pixel to be encoded in the current block.

[0052] Among them, the target neural network determines the dependency relationship between the dependent template and the current template based on the reconstructed pixel values ​​of the pixels in the current template and the dependent template, and determines the target prediction value of the pixel to be encoded based on the reconstructed pixel values ​​of the pixels in the dependent block and the dependency relationship.

[0053] Specifically, the dependency relationship between the dependent template and the current template is similar to the dependency relationship between the dependent block and the current block. Therefore, after receiving the dependent block, the current template and the dependent template, the target neural network learns the dependency relationship between the dependent template and the current template, and then uses the dependency relationship as the dependency relationship between the dependent block and the current block. Finally, based on the reconstructed pixel value of the pixel point in the dependent block and the dependency relationship, the target prediction value of the pixel point to be encoded can be determined.

[0054] In this embodiment, a pre-trained target neural network is used to learn the dependency relationship between the current template and the dependent template, so that a more complex linear or nonlinear relationship can be established, so that the established dependency relationship is more consistent with the actual situation of the image, thereby making the prediction result accurate and improving compression efficiency.

[0055] It should be noted that if the size of the current block is different from that of the dependent block, the dependent block needs to be preprocessed before being input into the target neural network, or the result needs to be post-processed after the target neural network outputs the result. For example, if the dependent block is a luminance block, the current block is a chrominance block, and the size of the luminance block is larger than the size of the chrominance block, then after the target neural network outputs the result, the output result needs to be downsampled to finally obtain the target prediction value of each pixel to be encoded in the current block. The downsampling processing method includes but is not limited to bilinear interpolation filters, bilinear cubic interpolation filters, neural network-based downsampling filters, etc., which are not limited here.

[0056] To better understand the above solution, let's illustrate it with an example:

[0057] In the first instance, the current block is a luminance block, the dependent block is the left dependent luminance block of the current block, and the current template includes reconstructed pixel points of the second sub-area, the third sub-area and the fifth sub-area outside the current block. After the dependent block, the current template and the dependent template are input into the target neural network, the target neural network outputs the target prediction value of each pixel to be encoded in the current block.

[0058] In the second example, the current block is a chroma block, the dependent block is a luminance block corresponding to the cross-component of the chroma block, and the current template includes reconstructed pixels in the first sub-region, the second sub-region, the third sub-region, the fourth sub-region, and the fifth sub-region outside the current block, and the image format is YUV420. After the dependent block, the current template, and the dependent template are input into the target neural network, after the target neural network outputs the result, the result output by the target neural network is also downsampled to finally obtain the target prediction value of each pixel in the chroma block. The downsampling method can be a bilinear cubic interpolation filter.

[0059] In the third example, the current block is a luminance block, the dependent block is a target reference block (also a luminance block) corresponding to the current block in the reference frame, and the current template includes reconstructed pixels in the first sub-region, the second sub-region, the third sub-region, the fourth sub-region, and the fifth sub-region outside the current block. After the dependent block, the current template, and the dependent template are input into the target neural network, the target neural network outputs the target prediction value of each pixel to be encoded in the current block.

[0060] In this embodiment, step S120 specifically includes:

[0061] S121: Input the dependent block, current template, dependent template, and target side information into the target neural network to obtain a target prediction value for the pixel to be encoded. The target side information includes at least one of the following: a quantization parameter of the current block, a quantization parameter of the dependent block, and a correspondence between the current block and the dependent block. The target neural network determines the dependency relationship between the dependent template and the current template based on the current template, the reconstructed pixel values ​​of the pixels in the dependent template, and the target side information.

[0062] Specifically, the quantization parameter of the current block includes at least one of the sequence-level quantization parameter, slice-level quantization parameter, and CU-level quantization parameter of the current block, and the quantization parameter of the dependent block includes at least one of the sequence-level quantization parameter, slice-level quantization parameter, and CU-level quantization parameter of the dependent block.

[0063] The corresponding relationship between the current block and the dependent block refers to the dependency type of the dependent block and the current block. For example, if the current block is a luminance block, the dependent block is Figure 3 If the current block is the upper left dependent block of the current block, the correspondence between the current block and the dependent block is: the current luminance block depends on the upper left luminance block; if the current block is a chrominance block, the dependent block is the cross-component luminance block corresponding to the chrominance block, then the correspondence between the current block and the dependent block is: the current chrominance block depends on the corresponding luminance block; if the current block is a luminance block, the dependent block is the target reference block corresponding to the current luminance block in the reference frame, then the correspondence between the current block and the dependent block is: the current luminance block depends on the reference luminance block, and so on, which will not be introduced one by one here.

[0064] The target side information includes one or more of the following: a quantization parameter of the current block, a quantization parameter of the dependent block, and a correspondence between the current block and the dependent block, which is not limited here.

[0065] Specifically, at this time, the target neural network further combines the target side information to determine the dependency relationship between the dependent template and the current template, which can further improve the accuracy of determining the dependency relationship and ultimately improve the prediction results.

[0066] It should be noted that in other implementations, the target side information may not be input into the target neural network. In this case, the target neural network can determine the dependency relationship between the dependent template and the current template only based on the reconstructed pixel values ​​of the pixels in the current template and the dependent template.

[0067] In this embodiment, when the target side information is set to include the correspondence between the current block and the dependent block, the following beneficial effects are achieved:

[0068] If the target neural network input does not include target side information, or the target side information does not include the correspondence between the current block and the dependent block, then during prediction, the target neural network cannot know the correspondence between the dependent block and the current block. Therefore, in order to ensure that the trained target neural network can predict the current block, when training the target neural network, the dependency between the second sample block and the first sample block in the sample group used must be consistent with the dependency between the current block and the dependent block during the prediction process. The sample group used for training includes a first sample block, a second sample block, a first sample template, and a second sample template. The first sample template includes multiple reconstructed pixel points around the first sample block, and the second sample template includes multiple reconstructed pixel points around the second sample block. When using the sample group to train the target neural network, the first sample block, the first sample template, and the second sample template are used as input, and the second sample block is used as a label to train the target neural network.

[0069] When the target side information is set to include the correspondence between the current block and the dependent block, since the target neural network can know the correspondence between the current block and the dependent block during the prediction process, the sample groups used when training the target neural network may have the following situations: the dependency relationship between the second sample block and the first sample block in some sample groups is consistent with the dependency relationship between the current block and the dependent block in the prediction process, while the dependency relationship between the second sample block and the first sample block in other sample groups is inconsistent with the dependency relationship between the current block and the dependent block in the prediction process.

[0070] Simply put, during the training process, the correspondence between the second sample block and the first sample block in the sample group can be diverse, so that after the training is completed, the target neural network can predict multiple current blocks, and the correspondence between the multiple current blocks and their respective corresponding dependent blocks can be different, thereby improving the applicability of the target neural network.

[0071] See Figure 6 , Figure 6 It is a structural diagram of an embodiment of the target neural network of the present application. The target neural network 100 includes a relationship fitting module 110, a feature extraction module 120, a prediction module 130, a channel conversion module 140 and a residual connection line 150.

[0072] The relationship fitting module 110 is used to determine the dependency relationship according to the reconstructed pixel values ​​of the pixels in the current template and the dependent template.

[0073] Specifically, the relationship fitting module 110 can fit the dependency relationship between the current template and the dependent template based on the reconstructed pixel values ​​of the pixels in the current template and the dependent template through convolution, full connection and other processing methods. In order to improve the accuracy of fitting, the relationship fitting module 110 can further combine the target side information for fitting during fitting. That is to say, at this time, the input of the relationship fitting module 110 includes not only the current template and the dependent template, but also the target side information. In particular, the relationship fitting module 110 can be a convolutional neural sub-network, a fully connected neural sub-network or a hybrid neural sub-network of convolution and full connection. The present application does not limit the specific structure of the relationship fitting module 110.

[0074] The feature extraction module 120 is used to extract features from the dependency blocks to obtain dependency features.

[0075] Specifically, the feature extraction module 120 can map the dependency block from the pixel domain to the feature domain through convolution, full connection, etc. The feature extraction module 120 can specifically be a convolutional neural sub-network, a fully connected neural sub-network, or a hybrid neural sub-network of convolution and full connection. This application does not limit the specific structure of the feature extraction module 120.

[0076] The prediction module 130 is connected to the relationship fitting module 110 and the feature extraction module 120 at the same time, and is used to fuse the dependency relationship with the dependency feature to obtain a fused feature.

[0077] Specifically, in this embodiment, the prediction module 130 includes a prediction unit 131 and an enhancement unit 132. The prediction unit 131 is connected to the relationship fitting module 110 and the feature extraction module 120 at the same time, and is used to fuse the dependency relationship with the dependency feature to obtain a fused feature. The enhancement unit 132 is connected to the prediction unit 131 and is used to perform quality enhancement processing on the fused feature to improve the accuracy of subsequent predictions. Among them, the enhancement unit 132 can perform quality enhancement processing on the fused feature by convolution, full connection, etc. The enhancement unit 132 can be a convolutional neural sub-network, a fully connected neural sub-network, or a hybrid neural sub-network of convolution and full connection. This application does not impose any specific restrictions on the enhancement unit 132.

[0078] It should be noted that, in other implementations, the prediction module 130 may only include the prediction unit 131 but not the enhancement unit 132 .

[0079] The channel conversion module 140 is connected to the prediction module 130 and is used to perform channel conversion processing on the fused features to obtain predicted features, wherein the dimension of the predicted features is the same as the dimension of the dependent block.

[0080] Specifically, the dimension of the predicted feature is the same as the dimension of the dependent block. For example, assuming that the dimension of the dependent block is W×H, the dimension of the predicted feature is also W×H, that is, it includes the feature values ​​of W×H feature points.

[0081] Among them, the channel conversion module 140 can perform channel conversion processing on the fusion features through operations such as convolution, and the channel conversion module 140 can be a convolutional neural sub-network, a fully connected neural sub-network, or a mixed neural sub-network of convolution and fully connected. This application does not impose specific restrictions on the channel conversion module 140.

[0082] The residual connection line 150 connects the input of the feature extraction module 120 and the output of the channel conversion module 140 to add the predicted features and the reconstructed pixel values ​​of the pixels in the dependent block as the output of the target neural network.

[0083] Specifically, through the residual connection line 150, the predicted features and the reconstructed pixel values ​​of the pixels in the dependent block are added together to form the output of the target neural network, and the output is the target prediction value of the current block.

[0084] Among them, the setting of the residual connection line 150 can make each module in the target neural network 100 only need to learn the difference between the dependent block and the prediction feature, thereby reducing the training difficulty of the target neural network 100 and improving the training effect.

[0085] It should be noted that in other implementations, the residual connection line 150 may not be provided.

[0086] In this embodiment, the dependency relationship includes multiple weights, and the dependent features include feature values ​​of multiple feature points. The multiple feature points are distributed on multiple channels, and the feature points distributed on different channels correspond one to one. The prediction template 130 is specifically used to: fuse the feature value of each feature point with the weight corresponding to the feature point to obtain a fused feature. The multiple weights correspond one to one with the multiple feature points, or the feature points corresponding to different channels correspond to the same weight, or the feature points distributed on the same channel correspond to the same weight.

[0087] Specifically, if the dimension of the dependent feature is W×H×C, it means that the dependent feature includes the eigenvalues ​​of W×H×C feature points, and these eigenvalues ​​are distributed on C channels, each channel is distributed with W rows and H columns of pixels, and the pixels at the same position on different channels correspond one to one. For example, the pixels in the first row and the first column on different channels correspond one to one, the pixels in the second row and the third column on different channels correspond one to one, and so on.

[0088] The prediction module 130 may select the following fusion methods when performing fusion processing:

[0089] The first is that each feature point corresponds to a weight, and the number of weights is the same as the number of feature points. If the dimension of the dependent feature is W×H×C, the number of weights is W×H×C. Then, during the fusion process, for each feature point, its corresponding eigenvalue and the corresponding weight are fused. Specifically, the eigenvalue of the feature point can be multiplied by the corresponding weight.

[0090] The second is that the weights corresponding to the feature points on different channels are the same, that is, the weights corresponding to the feature points at the same position on different channels are the same. For example, the weights corresponding to the pixels at position (1, 2) on each channel are the same.

[0091] That is, if the dimension of the dependent feature is W×H×C, the number of weights is W×H.

[0092] The third type is that one channel corresponds to one weight, and the weights of multiple feature points on the same channel are the same. For any feature point, its corresponding weight is the weight corresponding to the channel it is in.

[0093] That is, if the dimension corresponding to the dependent feature is W×H×C, the number of weights is C.

[0094] Among them, when performing fusion processing, the prediction module 130 can choose any one of the above three methods for fusion processing, or it can simultaneously adopt two or three of the above fusion processing methods. The simultaneous adoption of two fusion processing methods is used as an example for introduction: first, the two fusion processing methods are respectively adopted for fusion to obtain respective fusion features, and then the obtained fusion features are further fused to obtain the final fusion features.

[0095] For example, the prediction module 13 adopts the first fusion process and the second fusion process respectively to obtain two fusion features, and then further fuses the two fusion features to obtain the final fusion feature.

[0096] In order to better understand the above target neural network 100, the following is an illustration with reference to an example:

[0097] See Figure 7 In this example, in addition to the current template and the dependent template, the input of the target neural network 100 also includes the quantization parameter of the current block (specifically, the slice-level quantization parameter of the current block) and the quantization parameter of the dependent block (specifically, the slice-level quantization parameter of the dependent block). In this example, the current block is a chroma block, and the dependent block is the cross-color luminance block corresponding to the chroma block.

[0098] In this example, after receiving the current template, the dependent template, the quantization parameter of the current block, and the quantization parameter of the dependent block, the relationship fitting module 110 performs fitting processing to obtain the dependency relationship between the current template and the dependent template. In this example, the relationship fitting module 110 includes 3 cascaded fully connected layers.

[0099] Meanwhile, in this example, the feature extraction module 120 includes four [first convolutional layers, first activation layers] and one second convolutional layer cascaded in sequence.

[0100] At the same time, the prediction unit 131 fuses the dependency relationship output by the relationship fitting module 110 with the dependency feature output by the feature extraction module 120 to obtain a fused feature.

[0101] The fused features are then fed into the enhancement unit 132, which includes four cascaded residual blocks and a local residual link A. Figure 8 In this example, each residual block includes a [fourth convolution layer, second activation layer] and a fifth convolution layer, both of which are cascaded, and also includes a local residual connection line B.

[0102] At the same time, the channel conversion module 140 includes a third convolutional layer.

[0103] It should be noted that the present application does not impose any limitation on the specific structure of the target neural network 100. For example, in other instances, the relationship fitting module 110 may include 1 fully connected layer, 2 fully connected layers, 4 fully connected layers, or even more fully connected layers, or the feature extraction module 120 may include 2 [first convolution layer, first activation layer] and a second convolution layer cascaded in sequence, or 1 [first convolution layer, first activation layer] and a second convolution layer cascaded, or the enhancement unit 132 includes 1 residual block, 2 residual blocks, 3 residual blocks, 5 residual blocks, or even more residual blocks, or the channel conversion module 140 may include more than one third convolution layer, and may include 2 third convolution layers, 3 third convolution layers, or even more third convolution layers.

[0104] In an application scenario of this embodiment, when training a target neural network, multiple sample groups are used to train the target neural network respectively; wherein the quantization parameters for encoding the multiple sample groups are not completely the same.

[0105] Specifically, each sample group includes a first sample block, a second sample block, a first sample template and a second sample template. The first sample template includes multiple reconstructed pixel points around the first sample block, and the second sample template includes multiple reconstructed pixel points around the second sample block. When using the sample group to train the target neural network, the first sample block, the first sample template and the second sample template are used as input, and the second sample block is used as a label to train the target neural network.

[0106] In the process of training the target neural network, the quantization parameters of encoding multiple sample groups are not exactly the same, which can improve the generalization of the target neural network. In subsequent use, the target neural network can be used to predict the predicted value of the current block of each quantization parameter.

[0107] For ease of understanding, the quantization parameter is explained here as a slice-level quantization parameter: during training, the target neural network is trained with multiple sample groups with slice-level quantization parameters equal to 22, 27, 32, 37, etc., and subsequently during prediction, the target neural network can be used for training regardless of whether the slice-level quantization parameter of the current block is equal to 22, 27, 32, 37 or other values.

[0108] In another application scenario, see Figure 9 , before step S120, further comprising:

[0109] S130: Determine the difference between the first quantization parameters corresponding to the multiple first neural networks and the second quantization parameter corresponding to the current block.

[0110] S140: Determine the first neural network corresponding to the smallest difference as the target neural network.

[0111] When training the first neural network, the corresponding sample group is used to train the first neural network, wherein the quantization parameter of the encoded sample group is the first quantization parameter corresponding to the first neural network.

[0112] Specifically, in this application scenario, multiple first neural networks are pre-trained, wherein the first quantization parameters corresponding to the multiple first neural networks are different. For any first neural network, a corresponding sample group is used to train it during the training process. The sample group includes a first sample block, a second sample block, a first sample template, and a second sample template. The first sample template includes multiple reconstructed pixel points around the first sample block, and the second sample template includes multiple reconstructed pixel points around the second sample block. When using the sample group to train the target neural network, the first sample block, the first sample template, and the second sample template are used as input, and the second sample block is used as a label to train the target neural network. The quantization parameter of the sample group corresponding to the encoded first neural network is the first quantization parameter corresponding to the first neural network.

[0113] At the same time, the quantization parameter of the current block is defined as the second quantization parameter. When the current block needs to be predicted, among multiple first neural networks, the first neural network whose corresponding first quantization parameter is closest to the second quantization parameter is found, and then the first neural network is used as the target neural network to predict the current block.

[0114] For ease of understanding, the quantization parameters are described here as slice-level quantization parameters:

[0115] Four first neural networks are pre-trained: first neural network 1 is trained with multiple first sample groups, wherein the quantization parameter for encoding the multiple first sample groups is 22; first neural network 2 is trained with multiple second sample groups, wherein the quantization parameter for encoding the multiple second sample groups is 27; first neural network 3 is trained with multiple third sample groups, wherein the quantization parameter for encoding the multiple third sample groups is 32; first neural network 4 is trained with multiple fourth sample groups, wherein the quantization parameter for encoding the multiple fourth sample groups is 37.

[0116] During prediction, for the current block with a slice-level quantization parameter of 35, since 37 is closest to 35, the first neural network 4 is determined as the target neural network, and the target neural network is used to predict the current block.

[0117] In another application scenario, see Figure 10 , the method of the present application further includes:

[0118] S150: After the target neural network is sequentially determined as a plurality of second neural networks, a target prediction value of each pixel to be encoded is obtained.

[0119] S160: Determine the cost value corresponding to each second neural network according to the target prediction value of all the pixels to be encoded under each second neural network.

[0120] S170: Determine the second neural network with the smallest cost as the final neural network.

[0121] S180: Determine the target prediction value of each pixel to be encoded under the final neural network as the final prediction value of each pixel to be encoded.

[0122] The quantization parameters corresponding to the multiple second neural networks are different, and when training the second neural network, the second neural network is trained using the corresponding sample group, and the quantization parameter of the encoded sample group is the quantization parameter corresponding to the second neural network; or the target relative positions corresponding to the multiple second neural networks are different, and when training the second neural network, the second neural network is trained using the corresponding sample group, the sample group includes a first sample block and a first sample template, the first sample template includes multiple reconstructed pixel points around the first sample block, and the relative position of the first sample template and the first sample block is the target relative position corresponding to the second neural network.

[0123] Specifically, in this application scenario, multiple second neural networks are used as target neural networks in sequence, and step S120 is performed respectively to obtain the target prediction value of each pixel to be encoded in the current block under each second neural network.

[0124] Therefore, for each second neural network, the corresponding cost value of the second neural network can be obtained based on the target prediction values ​​of all pixels to be encoded under the second neural network. The cost value is specifically a rate-distortion cost value. The process of determining the cost value is related to existing technology and will not be detailed here.

[0125] After obtaining the cost value corresponding to each second neural network, the second neural network with the smallest cost value can be found. Among them, the cost value corresponding to the second neural network is the smallest, which means that the second neural network is the most accurate for predicting the current block. Therefore, the target prediction value of the pixel to be encoded under the second neural network is finally used as the final prediction value of the pixel.

[0126] That is, in this embodiment, multiple second neural networks are used to compete, and the best second neural network is finally selected.

[0127] To facilitate understanding, let's explain this with examples:

[0128] Four second neural networks are pre-trained: second neural network A, second neural network B, second neural network C, and second neural network D.

[0129] Then, during prediction, the current template, dependent template and current block are all input into the second neural network A, the second neural network B, the second neural network C and the second neural network D to obtain the predicted value of each pixel to be encoded under the second neural network A, the second neural network B, the second neural network C and the second neural network D.

[0130] Then, for the second neural network A, the cost value corresponding to the second neural network A is determined according to the target prediction values ​​of all the pixels to be encoded in the current block under the second neural network A; for the second neural network B, the cost value corresponding to the second neural network B is determined according to the target prediction values ​​of all the pixels to be encoded in the current block under the second neural network B; for the second neural network C, the cost value corresponding to the second neural network C is determined according to the target prediction values ​​of all the pixels to be encoded in the current block under the second neural network C; for the second neural network D, the cost value corresponding to the second neural network D is determined according to the target prediction values ​​of all the pixels to be encoded in the current block under the second neural network D.

[0131] If the cost values ​​corresponding to the second neural network A, the second neural network B, the second neural network C, and the second neural network D are 1000, 750, 900, and 1200 respectively, the target prediction value of the pixel to be encoded under the second neural network B will be determined as the final prediction value of the pixel.

[0132] In this application scenario 1 example, the same as the multiple first neural networks mentioned above, the multiple second neural networks correspond to different quantization parameters. For ease of understanding, this is explained with an example:

[0133] Four second neural networks are pre-trained: second neural network A is trained with multiple second sample groups, wherein the quantization parameters for encoding the multiple second sample groups are all 22; second neural network B is trained with multiple second sample groups, wherein the quantization parameters for encoding the multiple second sample groups are all 27; second neural network C is trained with multiple third sample groups, wherein the quantization parameters for encoding the multiple third sample groups are all 32; second neural network D is trained with multiple fourth sample groups, wherein the quantization parameters for encoding the multiple fourth sample groups are all 37.

[0134] In another example of this application scenario, the target relative positions corresponding to the multiple second neural networks are different. The target relative position corresponding to the second neural network is the position of the first sample template relative to the first sample block in the corresponding sample group, and the position of the second sample template relative to the second sample block. When the second neural network is trained using the sample group, the first sample block, the first sample template, and the second sample template are used as input, and the second sample block is used as a label to train the target neural network.

[0135] For example, the number of the multiple second neural networks is set to three. When training the second neural network A, the reconstructed pixel points included in the first sample template are distributed in the upper area of ​​the first sample block. When training the second neural network B, the reconstructed pixel points included in the first sample template are distributed in the left area of ​​the first sample block. When training the second neural network C, the reconstructed pixel points included in the first sample template are partially distributed in the upper area of ​​the first sample block and partially distributed in the left area of ​​the first sample block.

[0136] In the above application scenario, in order to indicate the final neural network to the decoding end, corresponding syntactic elements need to be generated.

[0137] Specifically, a first syntax element is first generated, and the first syntax element is used to indicate whether to use the target neural network to predict the current block, that is, whether to execute step S120.

[0138] If the first syntax element indicates that the target neural network is not used to predict the current block, step S120 is not performed, and multiple second neural networks are not used for competition, so there is no need to generate other syntax elements.

[0139] However, if the first syntax element indicates that the target neural network is used to predict the current block, that is, step S120 is executed, then in order to indicate the final neural network to the decoding end, a second syntax element is generated. In this case, multiple second neural networks can be represented in advance with different labels, and then the second syntax element is set to be equal to the label corresponding to the final target neural network to indicate the final target neural network to the decoding end. For ease of understanding, this is explained here with reference to an example:

[0140] For example, in the intra-frame prediction technology, the first syntax element CCLM_NN is set to indicate whether the target neural network is enabled to predict the current block. When CCLM_NN=0, it means that the target neural network is not enabled to predict the current block, and when CCLM_NN=1, it means that the target neural network is enabled to predict the current block, and when CCLM_NN=1, the second syntax element CCLM_NN_IDX is set to indicate the final neural network to the decoding end. For example, assuming that the four second neural networks are represented by 0, 1, 2, and 3 respectively, when the target prediction value of the pixel to be encoded under the second neural network labeled 0 is determined as the final prediction value, CCLM_NN_IDX=0 is set.

[0141] Similarly, in the inter-frame prediction technology, the first syntax element LIC_NN is set to indicate whether the target neural network is enabled to predict the current block. When LIC_NN=0, it means that the target neural network is not enabled to predict the current block, and when LIC_NN=1, it means that the target neural network is enabled to predict the current block, and when LIC_NN=1, the second syntax element LIC_NN_IDX is set to indicate the final neural network to the decoding end. For example, assuming that the four second neural networks are represented by 0, 1, 2, and 3 respectively, when the target prediction value of the pixel to be encoded under the second neural network numbered 0 is determined as the final prediction value, LIC_NN_IDX=0 is set.

[0142] When the dependent block is in the current frame, the method of this embodiment further includes:

[0143] S210: Predict the current block using a traditional intra-frame prediction technology to obtain a first prediction value of a pixel to be encoded in the current block.

[0144] S220: Determine a first generation value according to the first prediction values ​​of all pixels to be encoded.

[0145] S230: Determine the second generation value based on the target prediction values ​​of all pixels to be encoded.

[0146] S240: In response to the first generation value being less than the second generation value, the first prediction value of the pixel to be encoded is determined as the final prediction value of the pixel to be encoded; otherwise, the target prediction value of the pixel to be encoded is determined as the final prediction value of the pixel to be encoded.

[0147] Specifically, when intra-frame prediction technology is needed to predict the current block, that is, when the dependent block is in the current frame, the traditional intra-frame prediction technology competes with the technology of using the target neural network for prediction in this application:

[0148] The intra-frame prediction technology in the existing technology is used to predict the current block to obtain the first prediction value of each pixel to be encoded, and the first generation value is determined based on the first prediction value of each pixel to be encoded. At the same time, the second generation value is determined based on the target prediction value of each pixel to be encoded.

[0149] If the first-generation value is less than the second-generation value, it means that the accuracy of prediction using the intra-frame prediction technology in the existing technology is higher than the accuracy of prediction of the current block using the target neural network, and the result of prediction using the intra-frame prediction technology in the existing technology is used as the final result. Otherwise, the result of prediction of the current block using the target neural network is used as the final result.

[0150] In the above scheme, the scheme using the target neural network for prediction competes with the scheme using the intra-frame prediction technology for prediction in the prior art; however, in other schemes, the scheme using the target neural network for prediction can also directly replace the intra-frame prediction scheme in the prior art, that is, the target prediction value of the pixel point to be encoded is directly used as the final prediction value.

[0151] When the dependent block is a cross-component reconstructed block of the current block, the above-mentioned traditional intra prediction technology may specifically be a cross-component linear prediction technology, namely, CCLM technology.

[0152] In another embodiment, when the dependent block is in a reference frame, the method of the present application further includes:

[0153] S310: Predict the current block using a traditional inter-frame prediction technology to obtain a first prediction value of a pixel to be encoded in the current block.

[0154] S320: Determine a first generation value according to the first prediction values ​​of all pixels to be encoded.

[0155] S330: Determine the second generation value based on the target prediction values ​​of all pixels to be encoded.

[0156] S340: In response to the first generation value being less than the second generation value, the first prediction value of the pixel to be encoded is determined as the final prediction value of the pixel to be encoded; otherwise, the target prediction value of the pixel to be encoded is determined as the final prediction value of the pixel to be encoded.

[0157] Specifically, when inter-frame prediction technology is needed to predict the current block, that is, when the dependent block is in the reference frame, the traditional inter-frame prediction technology competes with the technology of this application using the target neural network for prediction:

[0158] The current block is predicted using the inter-frame prediction technology in the existing technology to obtain the first prediction value of each pixel to be encoded, and the first generation value is determined based on the first prediction value of each pixel to be encoded. At the same time, the second generation value is determined based on the target prediction value of each pixel to be encoded.

[0159] If the first-generation value is less than the second-generation value, it means that the accuracy of prediction using the inter-frame prediction technology in the existing technology is higher than the accuracy of prediction of the current block using the target neural network, and the result of prediction using the inter-frame prediction technology in the existing technology is used as the final result. Otherwise, the result of prediction of the current block using the target neural network is used as the final result.

[0160] In the above scheme, the scheme using the target neural network for prediction competes with the scheme using the inter-frame prediction technology for prediction in the prior art; however, in other schemes, the scheme using the target neural network for prediction can also directly replace the inter-frame prediction scheme in the prior art, that is, the target prediction value of the pixel point to be encoded is directly used as the final prediction value.

[0161] When the dependent block is a target reference block corresponding to the current block in a reference frame, the above-mentioned traditional inter-frame prediction technology may be a local illumination compensation technology, namely, a LIC technology.

[0162] See Figure 11 , Figure 11 This is a flowchart of an embodiment of a decoding method of the present application, which includes:

[0163] S410: Receive encoded data sent by the encoder.

[0164] S420: Obtain a predicted value of a current pixel in a current decoding block by decoding the encoded data.

[0165] Among them, the final prediction value of the current pixel point in the current decoding block is obtained by using the encoding method in any of the above implementation methods. The detailed steps can be found in the relevant content and will not be repeated here.

[0166] See Figure 12 , Figure 12 3 is a schematic diagram of the structure of an embodiment of an encoder of the present application. The encoder 300 includes a processor 310, a memory 320, and a communication circuit 330. The processor 310 is coupled to the memory 320 and the communication circuit 330, respectively. The memory 320 stores program data. The processor 310 executes the program data in the memory 320 to implement the steps of the encoding method in any of the above-mentioned embodiments. The detailed steps can be found in the above-mentioned embodiments and will not be repeated here.

[0167] The encoder 300 can be any device with algorithm processing capabilities, such as a computer or a mobile phone, and is not limited here.

[0168] See Figure 13 , Figure 13 4 is a schematic diagram of another embodiment of the encoder of the present application. The encoder 400 includes an acquisition module 410 and a prediction module 420 connected to each other.

[0169] The acquisition module 410 is used to acquire a dependent block of the current block, a current template, and a dependent template. The current template includes multiple reconstructed pixel points around the current block, and the dependent template includes multiple reconstructed pixel points around the dependent block.

[0170] The prediction module 420 is used to input the dependent block, the current template and the dependent template into a pre-trained target neural network to obtain a target prediction value of the pixel to be encoded in the current block.

[0171] Among them, the target neural network determines the dependency relationship between the dependent template and the current template based on the reconstructed pixel values ​​of the pixels in the current template and the dependent template, and determines the target prediction value of the pixel to be encoded based on the reconstructed pixel values ​​of the pixels in the dependent block and the dependency relationship.

[0172] Among them, the encoder 400 executes the steps of the encoding method in any of the above-mentioned implementation methods when working. The detailed steps can be found in the above-mentioned related content and will not be repeated here.

[0173] The encoder 400 can be any device with algorithm processing capabilities, such as a computer or a mobile phone, and is not limited here.

[0174] See Figure 14 , Figure 14 1 is a schematic diagram of the structure of one embodiment of a decoder of the present application. The decoder 500 includes a processor 510, a memory 520, and a communication circuit 530. The processor 510 is coupled to the memory 520 and the communication circuit 530, respectively. The memory 520 stores program data. The processor 510 executes the program data in the memory 520 to implement the steps of the decoding method in any of the above-mentioned embodiments. The detailed steps can be found in the above-mentioned embodiments and will not be repeated here.

[0175] The decoder 500 can be any device with algorithm processing capabilities, such as a computer or a mobile phone, and is not limited here.

[0176] See Figure 15 , Figure 15 FIG. 6 is a schematic diagram of another embodiment of a decoder of the present invention. The decoder 600 includes an acquisition module 610 and a decoding module 620 .

[0177] The acquisition module 610 is used to receive the encoded data sent by the encoder.

[0178] The decoding module 620 is connected to the acquisition module 610 and is configured to obtain a predicted value of a current pixel in a current decoding block by decoding the encoded data.

[0179] Among them, the predicted value of the current pixel point in the current decoding block is obtained by using the encoding method in any of the above implementation methods. The specific process can be found in the above content and will not be repeated here.

[0180] Among them, the decoder 600 executes the steps of the decoding method in any of the above-mentioned implementation modes when working. The detailed steps can be found in the above-mentioned related content and will not be repeated here.

[0181] The decoder 600 can be any device with algorithm processing capabilities, such as a computer or a mobile phone, and is not limited here.

[0182] See Figure 16 , Figure 16 The computer-readable storage medium 700 stores a computer program 710, which can be executed by a processor to implement the steps of any of the above methods.

[0183] Among them, the computer-readable storage medium 700 can specifically be a device that can store the computer program 710, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or it can also be a server that stores the computer program 710. The server can send the stored computer program 710 to other devices for execution, or it can also run the stored computer program 710 itself.

[0184] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A coding method, characterized in that: The method comprises: Obtaining a dependent block of a current block, a current template, and a dependent template, wherein the current template includes a plurality of reconstructed pixel points around the current block, and the dependent template includes a plurality of reconstructed pixel points around the dependent block; Inputting the dependent block, the current template, and the dependent template into a pre-trained target neural network to obtain a target prediction value of a pixel to be encoded in the current block; The target neural network determines a dependency relationship between the dependent template and the current template based on the reconstructed pixel values ​​of the pixels in the current template and the dependent template, and determines the target prediction value of the pixel to be encoded based on the reconstructed pixel values ​​of the pixels in the dependent block and the dependency relationship; The method further comprises: respectively determining differences between first quantization parameters corresponding to the plurality of first neural networks and a second quantization parameter corresponding to the current block; Determine the first neural network corresponding to the smallest difference as the target neural network; When training the first neural network, the corresponding sample group is used to train the first neural network, wherein the quantization parameter encoding the sample group is the first quantization parameter corresponding to the first neural network.

2. The method according to claim 1, characterized in that The dependent block is a cross-component reconstructed block corresponding to the current block, a reconstructed block related to the current block in the current frame, a target reference block corresponding to the current block in the reference frame, or a cross-component reconstructed block corresponding to the target reference block.

3. The method according to claim 1, characterized in that The step of obtaining the dependent blocks, current template, and dependent templates of the current block includes: Determining the dependent block of the current block; Determine the current template and the dependent template according to reconstructed pixel points in the same target area outside the current block and outside the dependent block respectively; Among them, the target area outside the current block includes at least one of the first sub-area, the second sub-area, the third sub-area, the fourth sub-area and the fifth sub-area outside the current block, the first sub-area is located on the first side of the current block and its two ends are flush with the two ends of the current block, the second sub-area is located on the second side of the current block and its two ends are flush with the two ends of the current block, the third sub-area connects the first sub-area and the second sub-area, the fourth sub-area is located on the side of the first sub-area away from the third sub-area, and the fifth sub-area is located on the side of the second sub-area away from the third sub-area.

4. The method according to claim 1, wherein The step of inputting the dependent block, the current template, and the dependent template into a pre-trained target neural network to obtain a target predicted value of a pixel to be encoded in the current block includes: Inputting the dependent block, the current template, the dependent template, and the target side information into the target neural network to obtain the target prediction value of the pixel to be encoded; The target side information includes at least one of the quantization parameters of the current block, the quantization parameters of the dependent block, and the corresponding relationship between the current block and the dependent block; the target neural network determines the dependency relationship between the dependent template and the current template based on the current template, the reconstructed pixel values ​​of the pixels in the dependent template, and the target side information.

5. The method according to claim 1, wherein The target neural network includes: a relationship fitting module, configured to determine the dependency relationship based on the reconstructed pixel values ​​of the pixels in the current template and the dependent template; A feature extraction module, configured to extract features from the dependency block to obtain dependency features; A prediction module, connected to the relationship fitting module and the feature extraction module, for fusing the dependency relationship with the dependency feature to obtain a fused feature; a channel conversion module, connected to the prediction module, configured to perform channel conversion processing on the fused features to obtain prediction features, wherein the dimension of the prediction features is the same as the dimension of the dependent block; A residual connection line connects the input of the feature extraction module and the output of the channel conversion module to add the predicted features and the reconstructed pixel values ​​of the pixels in the dependent block as the output of the target neural network.

6. The method according to claim 5, characterized in that The dependency relationship includes a plurality of weights, the dependency feature includes feature values ​​of a plurality of feature points, the plurality of feature points are distributed on a plurality of channels, and the feature points distributed on different channels correspond to each other one to one; The prediction module is specifically configured to: fuse the feature value of each feature point with the weight corresponding to the feature point to obtain the fused feature; The multiple weights correspond to the multiple feature points one by one, or the feature points corresponding to different channels correspond to the same weight, or the feature points distributed in the same channel correspond to the same weight.

7. The method according to claim 1, characterized in that The method further comprises: Using multiple sample groups respectively to train the target neural network; The quantization parameters for encoding the plurality of sample groups are not completely the same.

8. The method according to claim 1, characterized in that The method further comprises: After sequentially determining the target neural network as a plurality of second neural networks, obtaining the target prediction value of each pixel to be encoded; Determining a cost value corresponding to each second neural network according to the target prediction value of all the pixels to be encoded under each second neural network; Determine the second neural network with the smallest cost value as the final neural network; Determine the target prediction value of each pixel to be encoded under the final neural network as the final prediction value of each pixel to be encoded; wherein the plurality of second neural networks correspond to different quantization parameters, wherein, when training the second neural network, the second neural network is trained using the corresponding sample group, wherein the quantization parameter for encoding the sample group is the quantization parameter corresponding to the second neural network; Alternatively, the target relative positions corresponding to multiple second neural networks are different, wherein, when training the second neural network, the second neural network is trained using a corresponding sample group, the sample group includes a first sample block and a first sample template, the first sample template includes multiple reconstructed pixel points around the first sample block, and the relative position of the first sample template and the first sample block is the target relative position corresponding to the second neural network.

9. The method according to claim 8, characterized in that After determining the target prediction value of each pixel to be encoded under the final neural network as the final prediction value of each pixel to be encoded, the method further includes: Generate the first syntactic element; In response to the first syntax element indicating the step of inputting the dependent block, the current template and the dependent template into a pre-trained target neural network to obtain a target prediction value of the pixel to be encoded in the current block, a second syntax element is generated, and the second syntax element indicates the final neural network.

10. The method according to claim 1, characterized in that The dependent block is in the current frame, and the method further includes: Predicting the current block using a traditional intra-frame prediction technique to obtain a first predicted value of the pixel to be encoded in the current block; Determining a first generation value based on the first predicted values ​​of all the pixels to be encoded; Determining a second generation value according to the target prediction values ​​of all the pixels to be encoded; In response to the first generation value being less than the second generation value, the first prediction value of the pixel to be encoded is determined as the final prediction value of the pixel to be encoded; otherwise, the target prediction value of the pixel to be encoded is determined as the final prediction value of the pixel to be encoded.

11. The method according to claim 1, wherein The dependent block is in a reference frame, and the method further includes: Predicting the current block using a traditional inter-frame prediction technique to obtain a first predicted value of the pixel to be encoded in the current block; Determining a first generation value based on the first predicted values ​​of all the pixels to be encoded; Determining a second generation value according to the target prediction values ​​of all the pixels to be encoded; In response to the first generation value being less than the second generation value, the first prediction value of the pixel to be encoded is determined as the final prediction value of the pixel to be encoded; otherwise, the target prediction value of the pixel to be encoded is determined as the final prediction value of the pixel to be encoded.

12. A decoding method, characterized in that: The method comprises: Receive the encoded data sent by the encoder; Decoding the encoded data to obtain a predicted value of a current pixel in a current decoding block; The predicted value of the current pixel in the current decoding block is obtained by processing using the encoding method according to any one of claims 1 to 11.

13. An encoder, characterized in that The encoder includes a processor, a memory and a communication circuit, the processor is coupled to the memory and the communication circuit respectively, the memory stores program data, and the processor implements the steps in the method according to any one of claims 1 to 11 by executing the program data in the memory.

14. A decoder, characterized in that: The decoder includes a processor, a memory and a communication circuit. The processor is coupled to the memory and the communication circuit respectively. The memory stores program data. The processor implements the steps in the method according to claim 12 by executing the program data in the memory.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program can be executed by a processor to implement the steps in the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Cross-component chrominance prediction method and device based on neural network

    CN115190312A