Image decoding and encoding method, apparatus, device, and storage medium
By performing feature enhancement and synthesis transformation on the feature reconstruction values after decoding the image bitstream, the problem of poor image reconstruction quality in deep learning image coding technology is solved, and higher quality image reconstruction is achieved.
Patent Information
- Application Number
- CN202310055970.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-13
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2043-01-13
AI Technical Summary
Deep learning-based image coding techniques often result in poor image quality after compression and reconstruction, making it difficult to improve.
By decoding the image bitstream, the feature reconstruction values are determined, and feature enhancement and synthesis transformation are performed on them, including obtaining the feature standard deviation and standard deviation representation value, determining the feature mask according to the preset threshold, performing feature enhancement on the feature reconstruction values, obtaining the enhanced feature values, and finally obtaining the reconstructed image patch.
It reduces the distortion of image features during the quantization process and improves the quality of the reconstructed image.
Smart Images

Figure CN118368434B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to an image decoding and encoding method, device, equipment and storage medium. BACKGROUND
[0002] Nowadays, deep learning and neural networks are constantly making breakthroughs in the field of video image compression. The image encoding technology based on deep learning has greatly led the traditional encoding standard in encoding performance, and how to improve the image quality of the image reconstructed after compression by the image encoding technology based on deep learning is a big problem.
[0003] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY
[0004] The main purpose of the present application is to provide an image decoding and encoding method, device, equipment and storage medium, which aims to improve the image quality of the image reconstructed after compression by the image encoding technology based on deep learning.
[0005] To achieve the above purpose, the present application provides an image decoding method, which comprises the following steps:
[0006] Decoding the image code stream and determining the feature reconstruction value corresponding to the current image feature obtained by decoding;
[0007] Feature enhancement is performed on the feature reconstruction value to obtain an enhanced feature value;
[0008] Synthetic transformation is performed on the enhanced feature value to obtain a reconstructed image block.
[0009] In one possible embodiment of the present application, the current image feature is a three-dimensional feature matrix;
[0010] The feature enhancement is performed on the feature reconstruction value to obtain an enhanced feature value, comprising:
[0011] Obtaining the feature standard deviation and the standard deviation representation value corresponding to each matrix element in the current image feature;
[0012] According to the feature standard deviation, the standard deviation representation value and the preset defined threshold, the feature mask corresponding to each matrix element in the current image feature is determined;
[0013] Based on the feature mask, the feature enhancement is performed on the feature reconstruction value to obtain an enhanced feature value.
[0014] In a possible implementation of the present application, the determining of the feature mask corresponding to each matrix element in the current image feature according to the feature standard deviation, the standard deviation representation value and the preset threshold value comprises:
[0015] If the feature standard deviation and the standard deviation representation value corresponding to the matrix element satisfy a preset enhancement condition, the feature mask corresponding to the matrix element is set to a first type value.
[0016] If the feature standard deviation and the standard deviation representation value corresponding to the matrix element do not satisfy the preset enhancement condition, the feature mask corresponding to the matrix element is set to a second type value.
[0017] In a possible implementation of the present application, before the feature mask corresponding to the matrix element is set to the first type value if the feature standard deviation and the standard deviation representation value corresponding to the matrix element satisfy the preset enhancement condition, the method further comprises:
[0018] If the feature standard deviation corresponding to the matrix element is greater than the preset threshold value and the standard deviation representation value is a first type representation value, it is determined that the feature standard deviation and the standard deviation representation value corresponding to the matrix element satisfy the preset enhancement condition.
[0019] Or,
[0020] If the feature standard deviation corresponding to the matrix element is less than the preset threshold value and the standard deviation representation value is a second type representation value, it is determined that the feature standard deviation and the standard deviation representation value corresponding to the matrix element satisfy the preset enhancement condition.
[0021] In a possible implementation of the present application, the feature enhancement is performed on the feature reconstruction value based on the feature mask to obtain an enhanced feature value, which comprises:
[0022] The matrix element with the first type value of the corresponding feature mask is taken as a target matrix element.
[0023] The feature reconstruction value corresponding to the target matrix element is enhanced to obtain an enhanced feature value.
[0024] In a possible implementation of the present application, the feature enhancement is performed on the feature reconstruction value corresponding to the target matrix element to obtain an enhanced feature value, which comprises:
[0025] The feature reconstruction value, the residual reconstruction value and the predicted feature value corresponding to the target matrix element are obtained.
[0026] A first enhancement value is determined according to a first scaling factor and the residual reconstruction value, and a second enhancement value is determined according to a second scaling factor and the predicted feature value.
[0027] determine an enhanced feature value according to the first enhanced value and the second enhanced value.
[0028] In a possible implementation of the present application, before the first enhanced value is determined according to the first scaling factor and the predicted feature value, and the second enhanced value is determined according to the second scaling factor and the residual reconstruction value, the method further comprises:
[0029] obtaining a feature channel corresponding to the target matrix element;
[0030] determining a first scaling factor and a second scaling factor according to the feature channel, wherein different feature channels correspond to different first scaling factors and second scaling factors.
[0031] In a possible implementation of the present application, the enhancing the feature reconstruction value corresponding to the target matrix element to obtain an enhanced feature value comprises:
[0032] obtaining a feature reconstruction value and a predicted feature value corresponding to the target matrix element;
[0033] determining a first enhanced value according to a first scaling factor and the feature reconstruction value, and determining a second enhanced value according to a second scaling factor and the predicted feature value;
[0034] determining an enhanced feature value according to the first enhanced value and the second enhanced value.
[0035] In a possible implementation of the present application, the enhancing the feature reconstruction value corresponding to the target matrix element to obtain an enhanced feature value comprises:
[0036] obtaining a feature reconstruction value and a residual reconstruction value corresponding to the target matrix element;
[0037] determining a first enhanced value according to a first scaling factor and the feature reconstruction value, and determining a second enhanced value according to a second scaling factor and the residual reconstruction value;
[0038] determining an enhanced feature value according to the first enhanced value and the second enhanced value.
[0039] In a possible implementation of the present application, the enhancing the feature reconstruction value corresponding to the target matrix element to obtain an enhanced feature value comprises:
[0040] obtaining a feature reconstruction value and a feature standard deviation corresponding to the target matrix element;
[0041] determining a first enhanced value according to a first scaling factor and the feature reconstruction value, and determining a second enhanced value according to a second scaling factor and the feature standard deviation;
[0042] determine an enhanced feature value according to the first enhanced value and the second enhanced value.
[0043] In a possible implementation of the present application, the feature reconstruction value includes a luma reconstruction value and a chroma reconstruction value, and the enhanced feature value includes a luma enhanced feature value and a chroma enhanced feature value.
[0044] The enhancing of the feature reconstruction value corresponding to the target matrix element to obtain an enhanced feature value includes:
[0045] enhancing the chroma reconstruction value corresponding to the target matrix element to obtain a chroma enhanced feature value;
[0046] enhancing the luma reconstruction value corresponding to the target matrix element to obtain a luma enhanced feature value.
[0047] In a possible implementation of the present application, after the enhancing of the chroma reconstruction value corresponding to the target matrix element to obtain a chroma enhanced feature value, the method further includes:
[0048] determining a first enhanced value according to the luma reconstruction value and a first scaling factor;
[0049] secondarily enhancing the chroma enhanced feature value according to the first enhanced value.
[0050] In a possible implementation of the present application, the determining of the first enhanced value according to the luma reconstruction value and a first scaling factor further includes:
[0051] extracting a component indication parameter from the image code stream;
[0052] if the component indication parameter is a chroma enhancement type parameter, determining a first enhanced value according to the luma reconstruction value and a first scaling factor.
[0053] In a possible implementation of the present application, the decoding of the image code stream and the determining of a feature reconstruction value corresponding to a current image feature obtained by decoding include:
[0054] decoding the image code stream and determining a residual reconstruction value corresponding to a current image feature obtained by decoding;
[0055] predicting according to the feature reconstruction value of the already reconstructed feature to obtain a predicted feature value;
[0056] determining the feature reconstruction value corresponding to the current image feature according to the residual reconstruction value and the predicted feature value.
[0057] In a possible implementation of the present application, after the decoding of the image code stream and the determining of a residual reconstruction value corresponding to a current image feature obtained by decoding, the method further includes:
[0058] According to the enhanced feature value of the reconstructed feature, a predicted feature value is obtained by prediction;
[0059] According to the residual reconstruction value and the predicted feature value, a feature reconstruction value corresponding to the current image feature is determined.
[0060] In a possible implementation of the present application, the feature enhancement mode of the enhanced feature value of the reconstructed feature and the enhanced feature value corresponding to the current image feature is the same or different.
[0061] In a possible implementation of the present application, whether the feature enhancement mode of the enhanced feature value of the reconstructed feature and the enhanced feature value corresponding to the current image feature is the same or different is determined by a syntax flag, and the syntax flag is read from an image code stream.
[0062] In a possible implementation of the present application, the feature enhancement of the feature reconstruction value to obtain an enhanced feature value comprises:
[0063] A syntax application interval parameter is extracted from an image code stream, and feature position information corresponding to the current image feature is obtained;
[0064] According to the feature position information and the syntax application interval parameter, an enhanced syntax parameter is determined;
[0065] According to the enhanced syntax parameter, the feature reconstruction value is subjected to feature enhancement to obtain an enhanced feature value.
[0066] In addition, to achieve the above object, the present application further provides an image decoding device, which comprises the following modules:
[0067] A code stream decoding module is configured to decode an image code stream, and determine a feature reconstruction value corresponding to a current image feature obtained by decoding;
[0068] A feature enhancement module is configured to perform feature enhancement on the feature reconstruction value to obtain an enhanced feature value;
[0069] An image reconstruction module is configured to perform synthesis transformation on the enhanced feature value to obtain a reconstructed image block.
[0070] In addition, to achieve the above object, the present application further provides an image encoding method, which comprises:
[0071] Features of a to-be-encoded image block are extracted, and the extracted features are taken as current image features;
[0072] According to a feature reconstruction value corresponding to a reconstructed feature, a predicted feature value is obtained by prediction;
[0073] determining an encoding residual coefficient corresponding to the current image feature according to the predicted feature value;
[0074] writing the encoding residual coefficient into an image bitstream corresponding to the image block to be encoded.
[0075] In a possible implementation of the present application, after the encoding residual coefficient is written into the image bitstream corresponding to the image block to be encoded, the method further comprises:
[0076] performing image decoding on the image bitstream corresponding to the image block to be encoded to obtain a reconstructed image block;
[0077] determining an image encoding efficiency according to the reconstructed image block and the image block to be encoded.
[0078] In addition, to achieve the above object, the present application further provides an image encoding device, which comprises:
[0079] a feature extraction module, configured to perform feature extraction on an image block to be encoded, and take the extracted features as current image features;
[0080] a feature prediction module, configured to perform prediction according to a feature reconstruction value corresponding to a reconstructed feature to obtain a predicted feature value;
[0081] a residual calculation module, configured to determine an encoding residual coefficient corresponding to the current image feature according to the predicted feature value;
[0082] a parameter writing module, configured to write the encoding residual coefficient into an image bitstream corresponding to the image block to be encoded.
[0083] In addition, to achieve the above object, the present application further provides a decoding device, which comprises a processor, a memory, and a decoding program stored in the memory and executable on the processor, and the decoding program, when executed by the processor, implements the image decoding method as described above.
[0084] In addition, to achieve the above object, the present application further provides an encoding device, which comprises a processor, a memory, and a decoding program and / or an encoding program stored in the memory and executable on the processor, and the decoding program, when executed by the processor, implements the image decoding method as described above, and the encoding program, when executed by the processor, implements the image encoding method as described above.
[0085] Furthermore, to achieve the above objectives, the present invention also proposes a computer-readable storage medium storing an image decoding program and / or an image encoding program, wherein the image decoding program, when executed, implements the image decoding method as described above, and the image encoding program, when executed, implements the image encoding method as described above.
[0086] This invention decodes the image bitstream and determines the feature reconstruction value corresponding to the current image features obtained from the decoding; it then enhances the feature reconstruction value to obtain enhanced feature values; finally, it performs a synthetic transformation on the enhanced feature values to obtain reconstructed image blocks. Because feature enhancement is performed on the feature reconstruction value before image reconstruction, and then image reconstruction is performed based on the enhanced feature values, distortion of image features during quantization and other processes is reduced, thereby improving the image quality of the reconstructed image. Attached Figure Description
[0087] Figure 1 This is a schematic diagram of the structure of an electronic device in the hardware operating environment involved in the embodiments of the present invention;
[0088] Figure 2 This is a flowchart illustrating the first embodiment of the image decoding method of the present invention;
[0089] Figure 3 This is a schematic diagram of a matrix structure according to an embodiment of the present invention;
[0090] Figure 4 This is a schematic diagram of an image encoding and decoding process according to an embodiment of the present invention;
[0091] Figure 5 This is a schematic diagram of an image encoding and decoding process according to an embodiment of the present invention;
[0092] Figure 6 This is a flowchart illustrating the second embodiment of the image decoding method of the present invention;
[0093] Figure 7 This is a flowchart illustrating the third embodiment of the image decoding method of the present invention;
[0094] Figure 8 This is a flowchart illustrating the fourth embodiment of the image decoding method of the present invention.
[0095] Figure 9 This is a schematic diagram of the feature reconstruction sequence according to an embodiment of the present invention;
[0096] Figure 10 This is a flowchart illustrating the first embodiment of the image encoding method of the present invention;
[0097] Figure 11 This is a flowchart illustrating the second embodiment of the image encoding method of the present invention.
[0098] Figure 12 Structure block diagram of the first embodiment of the image decoding apparatus of the present application.
[0099] Figure 13 Structure block diagram of the first embodiment of the image encoding apparatus of the present application.
[0100] The implementation, functional features and advantages of the present application will be further explained with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0101] It should be understood that the specific embodiments described herein merely exemplify the application and do not limit the application.
[0102] Reference Figure 1 , Figure 1 The decoding device or the structure schematic diagram of the decoding device related to the hardware running environment of the embodiment of the present application.
[0103] As shown in Figure 1 , the electronic device can include a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to realize the connection and communication between the components. The user interface 1003 can include a display, an input unit such as a keyboard, and can also include a standard wired interface, a wireless interface. The network interface 1004 can optionally include a standard wired interface, a wireless interface (such as a wireless fidelity (WIreless-FIdelity, WI-FI) interface). The memory 1005 can be a high-speed random access memory (RAM), and can also be a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. The memory 1005 can also be a storage device independent of the aforementioned processor 1001.
[0104] Those skilled in the art can understand that Figure 1 the structure shown in the figure does not constitute a limitation on the electronic device, and can include more or fewer components than the figure, or combine certain components, or different component arrangements.
[0105] As shown in Figure 1 , the memory 1005 as a storage medium can include an operating system, a network communication module, a user interface module, and a decoding and / or encoding program.
[0106] In Figure 1The electronic device shown in the embodiment of the present application, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the electronic device of the present application can be arranged in the decoding device or the decoding device, and the decoding program and / or the decoding program stored in the memory 1005 are called by the processor 1001, and the image decoding method or the image encoding method provided by the embodiment of the present application is executed.
[0107] The embodiment of the present application provides an image decoding method, which refers to Figure 2 , Figure 2 The flowchart of the first embodiment of the image decoding method of the present application is shown.
[0108] In the embodiment, the image decoding method comprises the following steps:
[0109] Step S10: decoding the image code stream, and determining the feature reconstruction value corresponding to the current image feature obtained by decoding.
[0110] It should be noted that the execution subject of the embodiment can be a decoding device for encoding processing of image data, and the decoding device can be a personal computer, a server or other electronic device, of course, it can also be other devices which can realize the same or similar functions, and the embodiment does not limit the decoding device, and the image decoding method of the present application is described by taking the decoding device as an example in the embodiment and the following embodiments.
[0111] In the image encoding process, the encoding device generally decodes the encoded image code stream after encoding, and determines whether the parameters used in the encoding need to be adjusted according to the image quality of the decoded image, so the execution subject of the embodiment can also be the encoding device.
[0112] It should be noted that the image code stream can be a code stream generated by the encoding device after encoding the image data to be compressed and encoded. The decoding device decodes the image code stream, extracts the image features of the encoded image data from the image code stream, and the current image features obtained in the decoding process are the current image features. The feature reconstruction value can be the image feature obtained by feature recovery of the current image feature in the decoding process.
[0113] In the specific processing process, the encoding device can divide the image data into an image block for processing when encoding the image data, of course, the image data can also be divided into multiple image blocks for processing, and the embodiment does not limit this.
[0114] The technical terms involved in encoding or decoding images include: JPEG (Joint Photographic Experts Group), JPEG-AI (Joint Photographic Experts Group Artificial Intelligence), entropy encoding, neural network (NN), convolutional neural network (CNN), feature, rate-distortion principle, etc., which are described herein.
[0115] Among them, JPEG (Joint Photographic Experts Group) is a standard for continuous tone still image compression, and the file suffix is.jpg or.jpeg. It is the most commonly used image file format. It mainly uses a combination of predictive coding (DPCM), discrete cosine transform (DCT) and entropy coding to remove redundant image and color data. It belongs to lossy compression format, which can compress images into very small storage space, and to a certain extent, it will cause damage to image data. Especially using too high compression ratio, will reduce the quality of the image recovered after final decompression, if you pursue high quality image, it is not appropriate to use too high compression ratio.
[0116] The scope of JPEGAI is to create a learning-based image coding standard that provides a single-stream, compact, compression-domain representation that is both human visualizable and significantly improves compression efficiency over commonly used image coding standards for both image processing and computer vision tasks. JPEGAI targets a wide range of applications, such as cloud storage, visual surveillance, autonomous cars and devices, image acquisition, storage and management, real-time monitoring of visual data, and media distribution. The goal is to design an encoding solution that significantly improves the compression efficiency of commonly used coding standards with the same subjective quality and provides efficient compression-domain processing for machine learning-based image processing and computer vision tasks. Other key requirements include hardware / software implementation-friendly encoding and decoding, support for 8-bit and 10-bit depth, efficient encoding of images using text and graphics, and progressive decoding.
[0117] Entropy encoding is an encoding process that does not lose any information according to the entropy principle. Information entropy is the average amount of information of the source (a measure of uncertainty). Common entropy encodings include: Shannon coding, Huffman coding and arithmetic coding.
[0118] The neural network in the present application refers to an artificial neural network, rather than a biological neural network. The neural network is an operation model composed of a large number of nodes (or neurons) connected with each other. In the artificial neural network, the neuron processing unit can represent different objects, such as features, letters, concepts, or some meaningful abstract patterns. The types of processing units in the network are divided into three categories: input units, output units and hidden units. The input units accept signals and data from the external world; the output units output the processing results of the system; the hidden units are the units between the input and output units, which cannot be observed from the outside of the system. The connection weights between neurons reflect the connection strength between units, and the information representation and processing are embodied in the connection relationship of the network processing units. The artificial neural network is a non-programmed, brain-like information processing method, which essentially obtains a parallel distributed information processing function through the transformation and dynamics of the network, and simulates the information processing function of the human brain neural system at different levels and levels. At present, in the field of video processing, commonly used neural networks include convolutional neural network (CNN), recurrent neural network (RNN), fully connected network, etc.
[0119] The convolutional neural network is a kind of feedforward neural network, which is one of the most representative network structures in deep learning technology. The artificial neurons of the convolutional neural network can respond to a part of the surrounding units in the coverage range, and have excellent performance for large image processing. Generally, the basic structure of CNN includes two layers, one of which is the feature extraction layer (also called the convolution layer), and the input of each neuron is connected with the local receptive field of the previous layer, and the local features are extracted. Once the local features are extracted, the positional relationship between them and other features is also determined; the second is the feature mapping layer (also called the activation layer), each calculation layer of the network is composed of multiple feature mappings, and each feature mapping is a plane, and all the neurons on the plane have equal weights. The feature mapping structure can use Sigmoid function, ReLU function, Leaky-ReLU function, PReLU function, GDN function, etc. as the activation function of the convolutional network. In addition, since the neurons on a mapping plane share weights, the number of free parameters of the network is reduced. One of the advantages of CNN compared with traditional image processing algorithms is that it avoids the complex pre-processing process of the image (extracting artificial features, etc.), and can directly input the original image for end-to-end learning. One of the advantages of CNN compared with traditional neural networks is that traditional neural networks use full connection, that is, all the neurons from the input layer to the hidden layer are connected, which will result in a large number of parameters, making the network training time-consuming and even difficult to train, while CNN avoids this difficulty through local connection and weight sharing.
[0120] The feature involved in the present application is a three-dimensional feature matrix of CxWxH (such asFigure 3 As shown, Figure 3 C represents the number of channels, H represents the feature height, and W represents the feature width. The feature matrix can be an input of a neural network or an output of a neural network.
[0121] There are two indicators for evaluating the coding efficiency: the code rate and the PSNR. The smaller the bit stream is, the greater the compression rate is; the greater the PSNR is, the better the image coding efficiency is. In the mode selection, the decision formula is essentially a comprehensive evaluation of the two. The cost corresponding to the mode is: J(mode) = D + λ * R. Wherein, D represents Distortion, and the SSE index is usually used to measure the Distortion, and the SSE is the sum of squares of the difference between the reconstructed block and the source image; λ is the Lagrange multiplier; R is the actual number of bits required for image block coding under the mode, including the total number of bits required for coding mode information, motion information, residual error, etc. In the mode selection, if the RDO principle is used to compare and decide the coding mode, the best coding performance can be usually guaranteed.
[0122] In a possible implementation of the present application, the value of the image feature can be large or even complex. In order to improve the coding efficiency, the encoding device can perform residual calculation on the image feature and then encode the obtained residual data when encoding the image feature. When decoding the image data, a corresponding recovery process needs to be performed to ensure that the feature reconstruction value close to the image feature before encoding can be obtained. Therefore, the step S10 in the present embodiment can include:
[0123] decoding the image code stream and determining the residual reconstruction value corresponding to the current image feature obtained by decoding;
[0124] predicting the feature reconstruction value of the already reconstructed feature to obtain a predicted feature value;
[0125] determining the feature reconstruction value corresponding to the current image feature according to the residual reconstruction value and the predicted feature value.
[0126] It should be noted that if the encoding device calculates the residual data and then encodes, the residual reconstruction value corresponding to the current image feature can be obtained only when the image code stream is decoded. The already reconstructed feature can be part of the image feature that has completed feature recovery. The predicted feature value can be obtained by predicting the feature reconstruction value of the already reconstructed feature using a mean prediction network.
[0127] For ease of understanding, the present application will be described below in combination with Figure 4 . Figure 4The figure is a schematic diagram of the image coding process of the embodiment. In the figure, Bitstream #1 is the auxiliary code stream, and Bitstream #2 is the image code stream. As shown in Figure 4 The original image block (i.e. the image data to be encoded by the encoding device) is subjected to feature extraction by the analysis transformation network, and the image feature y is obtained. At the same time, the auxiliary information z_hat is calculated by the hyperparameter encoding network, and the predicted feature value mu is calculated by the mean prediction network. Then, the encoding device performs residual processing to obtain the residual original value r of the current feature. After residual processing and quantization (Q&AE), the obtained encoding residual coefficient r_coef is written into the image code stream (Bitstream #2).
[0128] Then, when processing the image code stream (Bitstream #2), the decoding device extracts the encoding residual coefficient r_coef from the image code stream, and then performs inverse quantization and residual recovery (AD&IQ) on the encoding residual coefficient. Then, the residual reconstruction value r_hat corresponding to the current image feature is obtained. Then, the mean prediction network performs prediction based on the feature reconstruction value y_hat of the reconstructed feature, and obtains the predicted feature value mu. Based on the predicted feature value and the residual reconstruction value, the feature reconstruction value y_hat of the current image feature is determined. Then, the feature reconstruction value is subjected to feature enhancement to obtain the enhanced feature value y_hat_en. Finally, the enhanced feature value is subjected to synthesis coding by the synthesis transformation network to obtain the reconstructed image block x_hat. The parameters used in the processing of Q&AE and AD&IQ are obtained by processing the auxiliary information z_hat by the probability hyperparameter decoding network.
[0129] The analysis transformation network, the hyperparameter encoding network, the probability hyperparameter decoding network, and the synthesis transformation network can all be neural networks constructed based on deep learning.
[0130] In actual use, the determination of the feature reconstruction value corresponding to the current image feature based on the residual reconstruction value and the predicted feature value can be adding the residual reconstruction value and the predicted feature value, and taking the sum value obtained after the addition as the feature reconstruction value corresponding to the current image feature. In the prediction, the feature reconstruction value of the reconstructed feature near the position of the current image feature can be used for prediction.
[0131] In a possible implementation of the present application, in order to further improve the coding efficiency, in the process of encoding the image, after the residual data is calculated, the encoding device can further perform residual processing or quantization processing on the residual data, and then write the obtained encoding residual coefficients into the image code stream. At this time, the residual reconstruction value corresponding to the current image feature determined by decoding the obtained current image feature can be obtained by decoding the image code stream, extracting the encoding residual coefficients corresponding to the current image feature, performing inverse quantization and residual recovery on the encoding residual coefficients, and obtaining the residual reconstruction value corresponding to the current image feature.
[0132] In a possible implementation of the present application, when performing prediction, the enhanced feature value of the reconstructed feature can also be used for prediction. At this time, the step S10 of the present embodiment can include:
[0133] decoding the image code stream and determining the residual reconstruction value corresponding to the current image feature obtained by decoding;
[0134] performing prediction according to the enhanced feature value of the reconstructed feature to obtain a predicted feature value;
[0135] determining the feature reconstruction value corresponding to the current image feature according to the residual reconstruction value and the predicted feature value.
[0136] It should be noted that in some cases, the encoding end can use the enhanced feature value of the reconstructed feature to perform prediction when calculating the residual value in the process of encoding the image data, determine the predicted feature value corresponding to the current encoded image feature, and calculate the encoding residual coefficient of the current encoded image feature based on the predicted feature value. Therefore, in the decoding process, the same method is also needed, that is, the mean value prediction network is used to perform prediction according to the enhanced feature value of the reconstructed feature to obtain a predicted feature value, and then the residual reconstruction value and the predicted feature value are added to obtain the feature reconstruction value corresponding to the current image feature.
[0137] For ease of understanding, the present application will be described below in combination with Figure 5 . Figure 5 The image encoding and decoding process diagram of the present embodiment is shown in Figure 5 . In the process of encoding and decoding the image data, the processing flow is the same as the above Figure 4The basic similarity is that the first feature enhancement manner is used to perform feature enhancement on the feature reconstruction value after the feature reconstruction value of the current image feature is obtained in the decoding process (i.e., feature enhancement 1 in the figure), the obtained first enhanced feature value y_hat_en1 is input into the synthesis transformation network for synthesis transformation processing, and the reconstructed image block x_hat is obtained. Meanwhile, the second feature enhancement manner is also used to perform feature enhancement on the feature reconstruction value (i.e., feature enhancement 2 in the figure), and the obtained second enhanced feature value y_hat_en2 is input into the mean value prediction network. In the processing of other image features, the current image feature is used as the reconstructed feature, the second enhanced feature value is used for prediction, and the predicted feature value is calculated.
[0138] It should be noted that the purpose of the feature enhancement 2 is to generate the subsequent feature prediction value as shown in the figure, so it is often performed column by column (the column can be diagonal, not necessarily vertical), and the features of the entire image cannot be completed at the same time. When the feature enhancement 1 is performed, the features of the entire image have been reconstructed, and thus the features of the entire image can be executed in parallel. Figure 5
[0139] In actual use, the feature enhancement manner of the enhanced feature value of the reconstructed feature and the enhanced feature value corresponding to the current image feature can be the same or different. That is, the feature enhancement process of the first feature enhancement manner and the second feature enhancement manner can be the same process, and in this case, y_hat_en1 is equal to y_hat_en2, and the decoding device only needs to decode one set of related syntax parameters from the image code stream. Of course, according to actual needs, the first feature enhancement manner and the second feature enhancement manner can be set as different processes, and in this case, the decoding device needs to decode two sets of syntax parameters from the image code stream.
[0140] In specific applications, the management personnel of the decoding device or the encoding device can also set a configuration parameter to allow skipping the first feature enhancement manner and / or the second feature enhancement manner, i.e., directly setting y_hat_en1 and / or y_hat_en2 to be equal to y_hat. The configuration parameter can be set in the image code stream, and of course, other ways of setting the configuration parameter can also be tried, which are not limited in the embodiment.
[0141] In a possible implementation of the present application, the encoding device can write a syntax flag in the image code stream in advance, which is used to indicate whether the first feature enhancement mode and the second feature enhancement mode use the same parameters. In this case, the decoding device can determine whether the first feature enhancement mode and the second feature enhancement mode use the same parameters according to the syntax flag read from the image code stream. For example, the encoding device can use a 1-bit syntax flag useSameParaFlag to represent whether the first feature enhancement mode and the second feature enhancement mode use the same syntax parameters. If the decoding device reads useSameParaFlag = 1, it indicates that the first feature enhancement mode and the second feature enhancement mode use the same syntax parameters. If the decoding device reads useSameParaFlag = 0, it indicates that the first feature enhancement mode and the second feature enhancement mode use different syntax parameters.
[0142] In a possible implementation of the present application, the syntax parameters involved in the feature enhancement mode can be as shown in the following table.
[0143] Table 1 Syntax parameter semantic table
[0144] Name Coding length Semantics numFilters 8 bit Indicates how many sets of syntax parameters, if the value is N, then the range of idx below is 0~N-1 applicationList[idx] 2 bit Indicates which component the idx set of parameters is used for, 0 means used for luma and chroma components, 1 means used only for luma component, 2 means used only for chroma component, 3 means not used. preciseList[idx][0] 1 bit 0 means the precision is 1 / 100, 1 means the precision is 1 / 10000. preciseList[idx][0] is the precision of thrList. If the precision is 1 / 100, then the actual value of thrList[idx] is the decoded value of thrList[idx] divided by 100. preciseList[idx][1] 1 bit 0 means the precision is 1 / 100, 1 means the precision is 1 / 10000. preciseList[idx][1] is the precision of scaleList greaterList[idx] 1 bit 1 means greater than a certain threshold, 0 means less than a certain threshold scaleList[idx][0] 8 / 16 bit The above scaling factor Scale1[idx], the coding length depends on preciseList[idx][0]. If preciseList[idx][0] is 0, then the precision is 1 / 100, then 8 bits are used for coding; If preciseList[idx][0] is 1, then the precision is 1 / 10000, then 16 bits are used for coding scaleList[idx][1] 8 / 16 bit The above scaling factor Scale2[idx] (if any), the coding length depends on the precision indicated by preciseList[idx][1], as above thrList[idx] 8 / 16 bit Threshold parameter, the coding length depends on the precision indicated by preciseList[idx][0], as above
[0145] In a possible implementation of the present application, on the basis of the syntax parameters in Table 1, other parameters such as blockSizeList[idx] and modeList[idx] can be additionally added. The coding length of blockSizeList[idx] can be 8 bits, and the semantics are as follows: for the block size NxN, if the value of N is 1, it means that each pixel is based on. Specifically, if N is greater than 1, a representative value of the NxN block is obtained by taking the minimum value, the maximum value, the average value, and the like. The representative value is enhanced based on the feature enhancement method of the present application to obtain a new representative value. Then, the enhanced value of the NxN block is set to the new representative value. The coding length of modeList[idx] can be 3 bits, and the semantics are as follows: the values of 1-4 respectively represent min, avg, max, and max pool for up-sampling (for block size not equal to 1x1). The value of 5 indicates that a group of filters has two scales. In this case, the parameters are as shown in Table 2.
[0146] Table 2 Syntax parameter semantic table
[0147] Name Coding length Semantics numFilters 8 bit Indicates how many sets of adjustment parameters, if the value is N, then the range of idx below is 0 to N-1 applicationList[idx] 2 bit Indicates which component the idx set of parameters is used for, 0 means used for luma and chroma components, 1 means used only for luma component, 2 means used only for chroma component, 3 means not used. preciseList[idx][0] 1 bit 0 means precision is 1 / 100, 1 means precision is 1 / 10000, preciseList[idx][0] is the precision of thrList, preciseList[idx][1] is the precision of scaleList preciseList[idx][1] 1 bit same as table 1 greaterList[idx] 1 bit 1 means greater than a threshold, 0 means less than a threshold blockSizeList[idx] 8 bit block size NxN, if N is 1, it means based on each pixel. modeList[idx] 3 bit 1~4 means min, avg, max, max pool for up and down sampling (for block size not 1x1). 5 means a set of filter has 2 scales. scaleList[idx][0] 8 / 16 bit scaling factor, coding length depends on preciseList[idx][0]. If preciseList[idx][0] is 0, it means precision is 1 / 100, then coded with 8 bits; if preciseList[idx][0] is 1, it means precision is 1 / 10000, then coded with 16 bits scaleList[idx][1] 8 / 16 bit scaling factor coded when modeList[idx] is 5, coding length depends on precision indicated by preciseList[idx][1], same as above thrList[idx] 8 / 16 bit threshold parameter, coding length depends on precision indicated by preciseList[idx][0], same as above
[0148] Step S20: performing feature enhancement on the feature reconstruction value to obtain an enhanced feature value.
[0149] It should be noted that, since the image features are subjected to quantization and other processing when encoding the image data, the features may be partially distorted, and the feature reconstruction values can be enhanced first, and then the enhanced feature values are used for image reconstruction, so as to improve the image quality of the reconstructed image.
[0150] Step S30: performing a synthesis transformation on the enhanced feature values to obtain a reconstructed image block.
[0151] It should be noted that, the synthesis encoding of the enhanced feature values to obtain the reconstructed image block can be performed by a synthesis transformation network constructed in advance to perform synthesis transformation on the enhanced feature values, so as to obtain the reconstructed image block. The synthesis transformation network can be a network constructed based on deep learning or neural network.
[0152] It can be understood that, if the image data is only divided into one image block for processing by the encoding device during the encoding process, the reconstructed image block obtained at this time is the completed reconstructed image data corresponding to the image data. If the image data is divided into multiple image blocks for processing by the encoding device during the encoding process, the reconstructed image block obtained at this time is only the reconstructed image data corresponding to one image block in the image data.
[0153] The embodiment decodes the image code stream, determines the feature reconstruction values corresponding to the current image features obtained by decoding, enhances the feature reconstruction values to obtain enhanced feature values, and performs a synthesis transformation on the enhanced feature values to obtain a reconstructed image block. Since the feature reconstruction values are enhanced before image reconstruction, and then the enhanced feature values are used for image reconstruction, the distortion of the image features in the quantization process is reduced, and thus the image quality of the reconstructed image is improved.
[0154] Reference Figure 6 , Figure 6 FIG. 2 is a flowchart of a second embodiment of an image decoding method according to the present application.
[0155] Based on the above first embodiment, the step S20 of the image decoding method of the present embodiment comprises:
[0156] Step S201: obtaining the feature standard deviation and the standard deviation representation value corresponding to each matrix element in the current image features.
[0157] It should be noted that, the current image features can be a three-dimensional feature matrix, and the three-dimensional dimensions thereof can be represented by c, i and j, wherein c is a channel identifier, and i and j represent the feature height and the feature width, respectively. The feature standard deviation corresponding to the matrix element can be the standard deviation of the feature corresponding to the matrix element and the feature mean value, and the standard deviation representation value is a representation value for representing whether the standard deviation corresponding to the matrix element is a large standard deviation.
[0158] Wherein, the standard deviation representation value can be divided into a first type representation value and a second type representation value (for the sake of concise representation, true and false can be used to represent, true is the first type representation value, and false is the second type representation value), if the standard deviation representation value corresponding to the matrix element is the first type representation value, it indicates that the standard deviation corresponding to the matrix element is a large standard deviation; if the standard deviation representation value corresponding to the matrix element is the second type representation value, it indicates that the standard deviation corresponding to the matrix element is not a large standard deviation.
[0159] Step S202: determining the feature mask corresponding to each matrix element in the current image feature according to the feature standard deviation, the standard deviation representation value, and a preset threshold.
[0160] It should be noted that the feature mask can be a representation value for identifying whether to perform feature enhancement on the feature reconstruction value corresponding to the matrix element, wherein the feature mask can be divided into a first type value and a second type value (for the sake of concise representation, true and false can be used to represent, true is the first type value, and false is the second type value), if the feature mask corresponding to the matrix element is the first type value, it indicates that the feature enhancement needs to be performed on the feature reconstruction value corresponding to the matrix element; if the feature mask corresponding to the matrix element is the second type value, it indicates that the feature enhancement does not need to be performed on the feature reconstruction value corresponding to the matrix element.
[0161] In actual use, a preset enhancement condition can be set to determine whether the feature standard deviation and the standard deviation representation value corresponding to the matrix element satisfy the preset enhancement condition to set the feature mask of the matrix element, and at this time, the step S202 of the embodiment can include:
[0162] If the feature standard deviation and the standard deviation representation value corresponding to the matrix element satisfy the preset enhancement condition, the feature mask corresponding to the matrix element is set to the first type value.
[0163] If the feature standard deviation and the standard deviation representation value corresponding to the matrix element do not satisfy the preset enhancement condition, the feature mask corresponding to the matrix element is set to the second type value.
[0164] It can be understood that if the feature standard deviation and the standard deviation representation value corresponding to the matrix element satisfy the preset enhancement condition, it indicates that the feature reconstruction value corresponding to the matrix element needs to be enhanced, and therefore, the feature mask corresponding to the matrix element can be set to the first type value.
[0165] And if the feature standard deviation and the standard deviation representation value corresponding to the matrix element do not satisfy the preset enhancement condition, it indicates that the feature reconstruction value corresponding to the matrix element does not need to be enhanced, and therefore, the feature mask corresponding to the matrix element can be set to the second type value.
[0166] In a specific implementation, the matrix elements with large standard deviation and corresponding feature standard deviation greater than a certain threshold value can be enhanced, and the matrix elements with non-large standard deviation and corresponding feature standard deviation less than a certain threshold value can be enhanced. The preset threshold value can be set by the administrator of the encoding device or the decoding device. Then, before the step of setting the feature mask corresponding to the matrix element to the first type value, if the feature standard deviation and the standard deviation representation value of the matrix element satisfy the preset enhancement condition, the step can further include:
[0167] If the feature standard deviation of the matrix element is greater than the preset threshold value, and the standard deviation representation value is the first type representation value, it is determined that the feature standard deviation and the standard deviation representation value of the matrix element satisfy the preset enhancement condition.
[0168] Or,
[0169] If the feature standard deviation of the matrix element is less than the preset threshold value, and the standard deviation representation value is the second type representation value, it is determined that the feature standard deviation and the standard deviation representation value of the matrix element satisfy the preset enhancement condition.
[0170] In actual use, when the feature reconstruction value is enhanced, at least one filter can be used. The enhancement of each filter is sequentially performed according to the order of the filter (the execution order of the filter can be set by the administrator of the encoding device or the decoding device in advance).
[0171] Different filters can be distinguished by setting different indexes (idx). When determining the feature mask, different filters can set different feature masks for the same matrix element. For example, for the filter with index idx, the corresponding feature mask can be:
[0172]
[0173] Wherein, mask[idx,c,i,j] is the feature mask set by the filter with index idx for the matrix element with matrix coordinates (c, i, j), Threshold[idx] is the preset threshold value corresponding to the filter with index idx, GreaterFlag[idx] is the standard deviation representation value set by the filter with index idx for the matrix element with matrix coordinates (c, i, j), and σ[c,i,j] is the feature standard deviation corresponding to the matrix element with matrix coordinates (c, i, j).
[0174] In a specific implementation, if mask[idx, c, i, j] is true (true and false can also be replaced by 1 and 0), that is, mask[idx, c, i, j] = 1, it indicates that the filter with the index idx performs feature enhancement on the matrix element with the matrix coordinates (c, i, j). If the filter is multiple sets, the matrix element with the same matrix coordinates can be subjected to multiple feature enhancements.
[0175] Step S203: performing feature enhancement on the feature reconstruction value based on the feature mask to obtain an enhanced feature value.
[0176] It should be noted that performing feature enhancement on the feature reconstruction value based on the feature mask to obtain an enhanced feature value can be performing feature enhancement on the feature reconstruction value corresponding to the matrix element in the current image feature that needs to be enhanced according to the feature mask, thereby obtaining the enhanced feature value.
[0177] In the embodiment, the feature standard deviation and the standard deviation representation value corresponding to each matrix element in the current image feature are obtained, the feature mask corresponding to each matrix element in the current image feature is determined according to the feature standard deviation, the standard deviation representation value and the preset threshold, and the feature reconstruction value is enhanced based on the feature mask to obtain an enhanced feature value. Since the feature mask corresponding to each matrix element is set in advance according to the feature standard deviation and the standard deviation representation value corresponding to each matrix element and the preset threshold, the matrix element in the feature matrix corresponding to the current image feature that needs to be enhanced is marked by setting the feature mask, so that the matrix element that needs to be enhanced can be quickly determined when the feature enhancement is performed, and the processing efficiency is improved.
[0178] Reference Figure 7 , Figure 7 The flowchart of the third embodiment of the image decoding method of the application is shown.
[0179] Based on the second embodiment, the step S203 of the image decoding method of the embodiment includes:
[0180] Step S2031: taking the matrix element with the corresponding feature mask being the first type value as a target matrix element.
[0181] It should be noted that if the feature mask corresponding to the matrix element is the first type value, it indicates that the feature reconstruction value corresponding to the matrix element needs to be enhanced, and therefore, the matrix elements in the three-dimensional matrix corresponding to the current image feature can be screened according to the feature mask, and the matrix element with the corresponding feature mask being the first type value is taken as a target matrix element.
[0182] Step S2032: enhancing the feature reconstruction value corresponding to the target matrix element to obtain an enhanced feature value.
[0183] It should be noted that when the feature reconstruction value corresponding to the target matrix element is enhanced, different ways can be used for enhancement to obtain an enhanced feature value.
[0184] In a possible implementation of the present application, when the feature reconstruction value corresponding to the target matrix element is enhanced, the preset scaling factor and the feature reconstruction value, the residual reconstruction value and the predicted feature value corresponding to the target matrix element can be combined for enhancement. At this time, the step S2032 in the present embodiment can include:
[0185] obtaining the feature reconstruction value, the residual reconstruction value and the predicted feature value corresponding to the target matrix element;
[0186] determining a first enhancement value according to the first scaling factor and the residual reconstruction value, and determining a second enhancement value according to the second scaling factor and the predicted feature value;
[0187] determining an enhanced feature value according to the feature reconstruction value, the first enhancement value and the second enhancement value.
[0188] It should be noted that the first scaling factor and the second scaling factor can be preset scaling factors, wherein different filters can correspond to different first scaling factors and second scaling factors.
[0189] In actual use, the enhanced feature value can be determined according to the feature reconstruction value, the first enhancement value and the second enhancement value based on the first feature enhancement formula.
[0190] The first feature enhancement formula is:
[0191] y_hat_en[c,i,j]= y_hat[c,i,j]+ mean_hat [c,i,j] * Scale2[idx]+residual_hat [c,i,j]*Scale1[idx]
[0192] In the formula, (c, i, j) is the matrix coordinate of the target matrix element, idx is the index of the filter, y_hat_en[c, i, j] is the enhanced feature value corresponding to the target matrix element, y_hat[c, i, j] is the feature reconstruction value corresponding to the target matrix element, residual_hat[c, i, j] * Scale1[idx] is the first enhanced value, mean_hat[c, i, j] * Scale2[idx] is the second enhanced value, mean_hat[c, i, j] is the predicted feature value corresponding to the target matrix element, residual_hat[c, i, j] is the residual reconstruction value corresponding to the target matrix element, Scale1[idx] is the first scaling factor, and Scale2[idx] is the second scaling factor.
[0193] In a possible implementation of the present application, different scaling factors can also be set for the same filter when the channels corresponding to the target matrix elements are different, and before the step of determining the first enhanced value according to the first scaling factor and the predicted feature value and determining the second enhanced value according to the second scaling factor and the residual reconstruction value, the embodiment can further include:
[0194] obtaining the feature channel corresponding to the target matrix element;
[0195] determining the first scaling factor and the second scaling factor according to the feature channel.
[0196] It should be noted that the feature channel corresponding to the target matrix element can be the matrix coordinate (c, i, j) of the target matrix element, and c is extracted, and the feature channel corresponding to the target matrix element is determined according to the value of c.
[0197] In actual use, determining the first scaling factor and the second scaling factor according to the feature channel can be obtaining the factor channel mapping table corresponding to the filter currently used, and the first scaling factor and the second scaling factor corresponding to the feature channel are found in the factor channel mapping table. The factor channel mapping table contains the mapping relationship between the feature channel and the scaling factor, and different feature channels in the factor channel mapping table can correspond to different first scaling factors and second scaling factors. The factor channel mapping table can be set in advance by the manager of the encoding device or the decoding device.
[0198] At this time, the enhanced feature value can be determined according to the feature reconstruction value, the first enhanced value and the second enhanced value according to the second feature enhancement formula, and the second feature enhancement formula can be:
[0199] y_hat_en[c, i, j] = y_hat[c, i, j] + mean_hat[c, i, j] * Scale2[c, idx] + residual_hat[c, i, j] * Scale1[c, idx]
[0200] In the formula, (c, i, j) is the matrix coordinates of the target matrix element, idx is the index of the filter, y_hat_en[c, i, j] is the enhanced feature value corresponding to the target matrix element, y_hat[c, i, j] is the feature reconstruction value corresponding to the target matrix element, residual_hat[c, i, j] * Scale1[c, idx] is the first enhanced value, mean_hat[c, i, j] * Scale2[c, idx] is the second enhanced value, mean_hat[c, i, j] is the predicted feature value corresponding to the target matrix element, residual_hat[c, i, j] is the residual reconstruction value corresponding to the target matrix element, Scale1[c, idx] is the first scaling factor corresponding to the feature channel of the target matrix element, and Scale2[c, idx] is the second scaling factor corresponding to the feature channel of the target matrix element.
[0201] The management personnel of the encoding device or the decoding device can also set the corresponding control switch for each channel in advance, and turn off the feature enhancement of a certain channel by turning off the corresponding control switch. Of course, the feature enhancement of a certain channel can also be turned off by modifying the scaling factor of the factor channel mapping table corresponding to the channel to 0.
[0202] In a possible implementation of the present application, when the feature reconstruction value corresponding to the target matrix element is enhanced, the scaling factor, the feature reconstruction value corresponding to the target matrix element, and the predicted feature value can be combined for enhancement. At this time, the step S2032 in the present embodiment can include:
[0203] obtaining the feature reconstruction value and the predicted feature value corresponding to the target matrix element;
[0204] determining a first enhanced value according to the first scaling factor and the feature reconstruction value, and determining a second enhanced value according to the second scaling factor and the predicted feature value;
[0205] determining an enhanced feature value according to the first enhanced value and the second enhanced value.
[0206] It should be noted that the first scaling factor and the second scaling factor can be preset scaling factors, wherein different filters can correspond to different first scaling factors and second scaling factors. Similarly, for the same filter, different first scaling factors and second scaling factors can also be set according to different characteristic channels of the target matrix elements.
[0207] In actual use, the enhanced characteristic value can be determined according to the first enhancement value and the second enhancement value based on a third characteristic enhancement formula, and the third characteristic enhancement formula is:
[0208] y_hat_en[c,i,j]= y_hat[c,i,j]* Scale1[idx]+ mean_hat [c,i,j]* Scale2[idx]
[0209] In the formula, (c, i, j) is the matrix coordinates of the target matrix element, idx is the index of the filter, y_hat_en[c, i, j] is the enhanced characteristic value, y_hat[c, i, j]*Scale1[idx] is the first enhancement value, mean_hat[c, i, j]*Scale2[idx] is the second enhancement value, y_hat[c, i, j] is the characteristic reconstruction value corresponding to the target matrix element, mean_hat[c, i, j] is the predicted characteristic value corresponding to the target matrix element, Scale1[idx] is the first scaling factor, and Scale2[idx] is the second scaling factor.
[0210] In a possible implementation of the present application, when the characteristic reconstruction value corresponding to the target matrix element is enhanced, the preset scaling factor and the characteristic reconstruction value and the residual reconstruction value corresponding to the target matrix element can be combined for enhancement. At this time, the step S2032 in the present embodiment can include:
[0211] Obtaining the characteristic reconstruction value and the residual reconstruction value corresponding to the target matrix element;
[0212] Determining a first enhancement value according to the first scaling factor and the characteristic reconstruction value, and determining a second enhancement value according to the second scaling factor and the residual reconstruction value;
[0213] Determining an enhanced characteristic value according to the first enhancement value and the second enhancement value.
[0214] It should be noted that the first scaling factor and the second scaling factor can be preset scaling factors, wherein different filters can correspond to different first scaling factors and second scaling factors. Similarly, for the same filter, different first scaling factors and second scaling factors can also be set according to different characteristic channels of the target matrix elements.
[0215] In actual use, the enhanced feature value can be determined according to the first enhancement value and the second enhancement value based on a fourth feature enhancement formula, the fourth feature enhancement formula being:
[0216] y_hat_en[c,i,j]= y_hat[c,i,j]* Scale1[idx]+ residual_hat [c,i,j]*Scale2[idx]
[0217] In the formula, (c, i, j) is the matrix coordinate of the target matrix element, idx is the index of the filter, y_hat_en[c, i, j] is the enhanced feature value, y_hat[c, i, j]*Scale1[idx] is the first enhancement value, residual_hat[c, i, j]*Scale2[idx] is the second enhancement value, y_hat[c, i, j] is the feature reconstruction value corresponding to the target matrix element, residual_hat[c, i, j] is the residual reconstruction value corresponding to the target matrix element, Scale1[idx] is the first scaling factor, and Scale2[idx] is the second scaling factor.
[0218] In a possible implementation of the present application, when the feature reconstruction value corresponding to the target matrix element is enhanced, the feature reconstruction value corresponding to the target matrix element and the feature standard deviation can be enhanced in combination with the preset scaling factor. In this case, the step S2032 in the present embodiment can include:
[0219] obtaining the feature reconstruction value corresponding to the target matrix element and the feature standard deviation;
[0220] determining the first enhancement value according to the first scaling factor and the feature reconstruction value, and determining the second enhancement value according to the second scaling factor and the feature standard deviation;
[0221] determining the enhanced feature value according to the first enhancement value and the second enhancement value.
[0222] It should be noted that the first scaling factor and the second scaling factor can be preset scaling factors, wherein different filters can correspond to different first scaling factors and second scaling factors. Similarly, for the same filter, different first scaling factors and second scaling factors can also be set according to different feature channels of the target matrix element.
[0223] In actual use, the enhanced feature value can be determined according to the first enhancement value and the second enhancement value based on a fifth feature enhancement formula, the fifth feature enhancement formula being:
[0224] y_hat_en[c, i, j] = y_hat[c, i, j] * Scale1[idx] + σ[c, i, j] * Scale2[idx]
[0225] wherein (c, i, j) is the matrix coordinate of the target matrix element, idx is the index of the filter, y_hat_en[c, i, j] is the enhanced feature value, y_hat[c, i, j] * Scale1[idx] is the first enhanced value, σ[c, i, j] * Scale2[idx] is the second enhanced value, y_hat[c, i, j] is the feature reconstruction value corresponding to the target matrix element, σ[c, i, j] is the feature standard deviation corresponding to the target matrix element, Scale1[idx] is the first scaling factor, and Scale2[idx] is the second scaling factor.
[0226] In a possible implementation of the present application, the feature reconstruction value includes reconstruction values corresponding to different components, such as luminance reconstruction values and chrominance reconstruction values, and correspondingly, the enhanced feature value can also include enhanced values of different components, such as luminance enhanced feature values and chrominance enhanced feature values. When the feature reconstruction value corresponding to the target matrix element is enhanced, the feature enhancement process of different components can be relatively independent and not affect each other. In this case, the step S2032 of the present embodiment can include:
[0227] enhancing the chrominance reconstruction value corresponding to the target matrix element to obtain a chrominance enhanced feature value;
[0228] enhancing the luminance reconstruction value corresponding to the target matrix element to obtain a luminance enhanced feature value.
[0229] It should be noted that when the feature is enhanced, different parameters, such as different scaling factors, can be set for the enhancement of different components in the same filter.
[0230] In a possible implementation of the present application, the enhanced feature value of a certain component can also be enhanced again by using the feature reconstruction value of another component, for example, the chrominance enhanced feature value is enhanced again by using the luminance reconstruction value. In this case, the step of enhancing the chrominance reconstruction value corresponding to the target matrix element to obtain a chrominance enhanced feature value in the present embodiment can be followed by:
[0231] determining a first enhanced value according to the luminance reconstruction value and a first scaling factor;
[0232] enhancing the chrominance enhanced feature value again according to the first enhanced value.
[0233] It should be noted that the first scaling factor can be a preset scaling factor, and different filters can correspond to different first scaling factors.
[0234] In actual use, the chroma enhancement feature value can be secondarily enhanced according to the first enhancement value based on a sixth feature enhancement formula, which can be:
[0235] y_hat_chroma_en2[c,i,j]=y_hat_chroma_en[c,i,j]+Scale1[idx]*y_hat_luma[c,i,j]
[0236] In the formula, y_hat_chroma_en2[c,i,j] is the secondarily enhanced chroma enhancement feature value, y_hat_chroma_en[c,i,j] is the chroma enhancement feature value, Scale1[idx]*y_hat_luma[c,i,j] is the first enhancement value, Scale1[idx] is the first scaling factor, and y_hat_luma[c,i,j] is the luminance reconstruction value corresponding to the target matrix element.
[0237] Of course, in a specific implementation, the luminance enhancement feature value can also be secondarily enhanced by using the chroma reconstruction value corresponding to the target matrix element.
[0238] In a possible implementation of the present application, in order to enable the decoding device to determine whether secondary enhancement is needed, the step of determining the first enhancement value according to the luminance reconstruction value and the first scaling factor can include:
[0239] extracting a component indication parameter from the image code stream;
[0240] If the component indication parameter is a chroma enhancement type parameter, the first enhancement value is determined according to the luminance reconstruction value and the first scaling factor.
[0241] It should be noted that the component indication parameter can be an indication parameter for indicating whether secondary enhancement is needed and the component enhancement feature value that needs to be secondarily enhanced. For example, the component indication parameter can have a value of 0-3. If the component indication parameter is 0, it means that secondary enhancement is not needed. If the component indication parameter is 1, it means that the chroma enhancement feature value needs to be secondarily enhanced. If the component indication parameter is 2, it means that the luminance enhancement feature value needs to be secondarily enhanced. If the component indication parameter is 3, it means that both the luminance enhancement feature value and the chroma enhancement feature value need to be secondarily enhanced.
[0242] It can be understood that if the component indication parameter is a chroma enhancement type parameter, it indicates that secondary enhancement is needed, and the component to be enhanced is a chroma enhancement feature value, at this time, a first enhancement value can be determined according to the luminance reconstruction value and the first scaling factor, and then the chroma enhancement feature value is enhanced according to the first enhancement value.
[0243] In a possible implementation of the present application, the filter can also use the same scaling factor to enhance the feature of all matrix elements in the current image, at this time, the feature reconstruction value corresponding to the target matrix element can be enhanced according to the seventh feature enhancement formula, the seventh feature enhancement formula is:
[0244] y_hat_en[c,i,j]= y_hat[c,i,j]* Scale1[idx]
[0245] In the formula, (c, i, j) is the matrix coordinates of the target matrix element, idx is the index of the filter, y_hat[c, i, j] is the feature reconstruction value corresponding to the target matrix element, y_hat_en[c, i, j] is the enhanced feature value, and Scale1[idx] is the scaling factor corresponding to the filter with index idx.
[0246] In a possible implementation of the present application, different scaling factors can also be set to enhance the features of the matrix elements in different channels of the current image (of course, the scaling factors of some same channels can also be the same), at this time, the feature reconstruction value corresponding to the target matrix element can be enhanced according to the eighth feature enhancement formula, the eighth feature enhancement formula is:
[0247] y_hat_en[c,i,j]= y_hat[c,i,j]* Scale1[c]
[0248] In the formula, (c, i, j) is the matrix coordinates of the target matrix element, idx is the index of the filter, y_hat[c, i, j] is the feature reconstruction value corresponding to the target matrix element, y_hat_en[c, i, j] is the enhanced feature value, and Scale1[c] is the scaling factor corresponding to the c channel.
[0249] The embodiment takes the matrix element corresponding to the first type value of the feature mask as the target matrix element, enhances the feature reconstruction value corresponding to the target matrix element, and obtains the enhanced feature value. Since part of the matrix elements that need to be enhanced are marked as target matrix elements according to the feature mask, and then the feature reconstruction value of the target matrix element is enhanced, the number of matrix elements that need to be processed in the processing process is reduced, and the execution efficiency of the image decoding method is improved.
[0250] ReferenceFigure 8 , Figure 8 Figure 2 is a flowchart illustrating a method for image decoding according to a fourth embodiment of the present application.
[0251] Based on the first embodiment, the step S20 of the image decoding method of the present embodiment comprises:
[0252] Step S201': extracting syntax application interval parameters from the image code stream and obtaining feature position information corresponding to the current image feature.
[0253] It should be noted that if the same enhancement method is used for feature enhancement of the image features of the entire image block (i.e., the syntax parameters used in the enhancement process are completely the same), the parameters (such as the first scaling factor and the second scaling factor) used in the enhancement process may only be optimal for part of the image features in the image block, but not optimal for other subsequent image features, and may even have a negative effect. In order to avoid this phenomenon, the encoding device can set multiple sets of syntax parameters when encoding, and use different syntax parameters to enhance the image features at different positions in the image block.
[0254] In actual use, the syntax application interval parameters can be parameters for indicating the application range of each set of syntax parameters, and the feature position information corresponding to the current image feature can include the position of the current image feature in the image block, such as the row number, the column number, and the like.
[0255] Step S202': determining the enhancement syntax parameters according to the feature position information and the syntax application interval parameters.
[0256] It should be noted that determining the enhancement syntax parameters according to the feature position information and the syntax application interval parameters can be determining the application range of each syntax parameter according to the syntax application interval parameters, comparing the feature position information with the application range of each syntax parameter, determining the application range in which the feature position information is located, and taking the syntax parameter corresponding to the application range in which the feature position information is located as the enhancement syntax parameter.
[0257] In actual use, a first syntax parameter and a second syntax parameter can be set, i.e., two sets of syntax parameters, and then a syntax application interval parameter is set to distinguish which part of the image features the first syntax parameter and the second syntax parameter are used for feature enhancement. The syntax application interval parameter limitation can be to limit the application range of the syntax parameters in the dimensions of channels, rows, columns, diagonal columns, and the like.
[0258] For example, assuming that there are two sets of syntax parameters, a first syntax parameter and a second syntax parameter, and the syntax application interval parameter extracted from the bitstream is applyLineNum=K, then for image features in columns 1 to K (or rows or diagonal columns), the first syntax parameter is used for feature enhancement, and for image features in columns K to the last column, the second syntax parameter is used for feature enhancement.
[0259] Of course, in actual use, more than two sets of syntax parameters can also be set, and in this case, a syntax application interval parameter can be set for each set of syntax parameters, and then the application range of each set of syntax parameters is determined according to the syntax application interval parameter.
[0260] For example, assuming that there are N sets of syntax parameters in total, there are N syntax application interval parameters, which can be represented as applyLineNum(i) (i=1~N), applyLineNum(i) is the syntax application interval parameter corresponding to the ith set of syntax parameters, and in this case, the application range of the first set of syntax parameters can be determined as image features in columns 1 to applyLineNum(1) (or rows or diagonal columns), the application range of the second set of syntax parameters can be determined as image features in columns applyLineNum(1)+1 to applyLineNum(1)+applyLineNum(2) (or rows or diagonal columns), the application range of the third set of syntax parameters can be determined as image features in columns applyLineNum(2)+1 to applyLineNum(2)+applyLineNum(3) (or rows or diagonal columns), and so on.
[0261] The image features in the row or column are relatively easy to understand, but the diagonal column is relatively complex. In order to facilitate understanding, the Figure 9 is described below, but the present solution is not limited thereto, Figure 9 is a feature reconstruction order diagram of the present application. As shown in Figure 9 , the reconstruction order used when performing feature reconstruction is from the top left to the bottom right, and the diagonal column is used for reconstruction, Figure 9 where Row represents the row in which the image feature is located, column represents the column in which the image feature is located, the virtual box circle represents a feature that has already been reconstructed (Samplesthat are already processed), the black circle represents an image feature that is currently being reconstructed (Current sample), and T represents the feature offset (Wave) of each reconstruction, Figure 9 indicates that the image feature in the 8th diagonal column is being reconstructed, and then the diagonal column in which the current image feature is located can be determined according to the feature position information (i.e., the row and the column) of the current image feature, and then the set of syntax parameters used for feature enhancement can be determined according to the diagonal column.
[0262] Step S203': performing feature enhancement on the feature reconstruction value according to the enhanced syntax parameter to obtain an enhanced feature value.
[0263] It can be understood that the feature enhancement on the feature reconstruction value according to the enhanced syntax parameter to obtain an enhanced feature value can be performed in the manner of feature enhancement provided by any of the above embodiments of the image coding method using the enhanced syntax parameter, which will not be described herein again.
[0264] The present embodiment extracts a syntax application interval parameter from an image code stream and obtains feature position information corresponding to the current image feature; determines an enhanced syntax parameter according to the feature position information and the syntax application interval parameter; and performs feature enhancement on the feature reconstruction value according to the enhanced syntax parameter to obtain an enhanced feature value. Since the enhanced syntax parameter used when performing feature enhancement is determined according to the feature position information of the current image feature and the syntax application interval parameter extracted from the image code stream, different syntax parameters can be applied when performing feature enhancement on image features at different positions, so that the effect of feature enhancement can be ensured as much as possible.
[0265] Reference Figure 10 , Figure 10 FIG. 1 is a flowchart of an image coding method according to a first embodiment of the present application.
[0266] In the present embodiment, the image coding method comprises the following steps:
[0267] Step S910: performing feature extraction on a to-be-coded image block and taking the extracted feature as a current image feature.
[0268] It should be noted that the execution subject of the present embodiment can be the coding device, which can be a personal computer, a server, or other electronic devices that can realize the same or similar functions, and the present embodiment does not limit the same. In the present embodiment and the following embodiments, the coding device is taken as an example to describe the image coding method of the present application.
[0269] It should be noted that the to-be-coded image block can be an image block obtained by dividing image data to be coded, wherein the image data can be divided into only one image block, or the image data can be divided into multiple image blocks.
[0270] In actual use, the feature extraction on the to-be-coded image block and the taking of the extracted feature as a current image feature can be performed by using an analysis transformation network to extract features from the to-be-coded image block, and then taking the current extracted image feature as a current image feature.
[0271] Step S920: predicting the feature reconstruction value corresponding to the reconstructed feature to obtain a predicted feature value.
[0272] In actual use, the prediction of the feature reconstruction value corresponding to the reconstructed feature to obtain a predicted feature value can be a feature prediction of the feature reconstruction value corresponding to the reconstructed feature by using a mean prediction network, so as to obtain a predicted feature value. When predicting by using the mean prediction network, the auxiliary information calculated by the hyperparameter coding network can also be input into the mean prediction network, so that the mean prediction network predicts in combination with the feature reconstruction value corresponding to the reconstructed feature and the auxiliary information.
[0273] Step S930: determining the coding residual coefficient corresponding to the current image feature according to the predicted feature value.
[0274] In actual use, the residual original value can be obtained by subtracting the predicted feature value from the feature value of the current image feature, and then the residual original value is subjected to residual processing and quantization processing, so as to obtain the coding residual coefficient.
[0275] Step S940: writing the coding residual coefficient into the image code stream corresponding to the to-be-encoded image block.
[0276] It should be noted that the coding residual coefficient is written into the image code stream corresponding to the to-be-encoded image block, so that when image decoding is needed, the coding residual coefficient can be directly read from the image code stream, and then the coding residual coefficient is subjected to inverse quantization processing and residual recovery processing, so as to obtain the residual reconstruction value, and finally the feature reconstruction value corresponding to the to-be-encoded image block can be determined in combination with the predicted feature value of the mean prediction network and the residual reconstruction value, so as to facilitate image reconstruction.
[0277] In actual use, the encoding device can also perform operations such as parameter calculation, parameter setting, syntax setting, flag setting, etc. The specific implementation can be derived by referring to the contents of any of the above image encoding methods.
[0278] The present embodiment extracts features from the to-be-encoded image block and takes the extracted features as the current image features; predicts the feature reconstruction value corresponding to the reconstructed feature to obtain a predicted feature value; determines the coding residual coefficient corresponding to the current image feature according to the predicted feature value; and writes the coding residual coefficient into the image code stream corresponding to the to-be-encoded image block. Since the coding residual coefficient of the current image feature is calculated and written into the image code stream during encoding, the coding efficiency of the image data is reduced.
[0279] Reference Figure 11 , Figure 11 is a flowchart of a second embodiment of an image encoding method of the present application.
[0280] Based on the first embodiment of the image encoding method described above, the present embodiment can further include, after the step S940:
[0281] Step S950: image decoding the image bitstream corresponding to the to-be-encoded image block to obtain a reconstructed image block.
[0282] It should be noted that after the encoding device completes the encoding of the image data, it also needs to verify the encoding efficiency to ensure that the encoding efficiency is high. At this time, the encoding device can image decode the image bitstream corresponding to the to-be-encoded image block to obtain a reconstructed image block.
[0283] Step S960: determining the image encoding efficiency according to the reconstructed image block and the to-be-encoded image block.
[0284] In actual use, determining the image encoding efficiency according to the reconstructed image block and the to-be-encoded image block can be comparing the reconstructed image block and the to-be-encoded image block, calculating the corresponding code rate and PSNR according to the comparison result, and thus determining the image encoding efficiency.
[0285] In a possible implementation of the present application, a preset efficiency threshold can also be set in advance. After the image encoding efficiency is obtained, the image encoding efficiency is compared with the preset efficiency threshold. If the image encoding efficiency is less than the preset efficiency threshold, it indicates that the image encoding efficiency at this time is low. At this time, the parameters in the network or model used in the encoding process can be adjusted to try to improve the image encoding efficiency.
[0286] In actual use, when image decoding the image bitstream corresponding to the to-be-encoded image block, the image decoding method provided by any embodiment of the image decoding method described above can be used, and the present embodiment does not limit this. The encoding device can also perform operations such as parameter calculation, parameter setting, syntax setting, flag setting, etc. The specific implementation can be derived by referring to the content of any embodiment of the image encoding method described above.
[0287] The present embodiment obtains a reconstructed image block by image decoding the image bitstream corresponding to the to-be-encoded image block; and determines the image encoding efficiency according to the reconstructed image block and the to-be-encoded image block. Since the image bitstream is image decoded after the encoding is completed, the reconstructed image block obtained by image decoding is compared with the to-be-encoded image block to determine the image encoding efficiency. This makes it possible to adjust the parameters in various networks used in the encoding process according to the image encoding efficiency, thereby improving the image encoding efficiency.
[0288] Furthermore, the embodiment of the present application further provides a storage medium, wherein the storage medium stores an image decoding program and / or an image encoding program, the image decoding program is used to implement the image decoding method, and the image encoding program is used to implement the image encoding method.
[0289] With reference to Figure 12 , Figure 12 FIG. 1 is a structural block diagram of an image decoding device according to a first embodiment of the present application.
[0290] As shown in Figure 12 , the image decoding device according to the embodiment of the present application comprises:
[0291] a code stream decoding module 10, configured to decode an image code stream and determine a feature reconstruction value corresponding to a feature of a current image obtained by decoding;
[0292] a feature enhancement module 20, configured to perform feature enhancement on the feature reconstruction value to obtain an enhanced feature value;
[0293] an image reconstruction module 30, configured to perform synthetic transformation on the enhanced feature value to obtain a reconstructed image block.
[0294] The embodiment decodes the image code stream, determines a feature reconstruction value corresponding to a feature of a current image obtained by decoding, performs feature enhancement on the feature reconstruction value to obtain an enhanced feature value, and performs synthetic transformation on the enhanced feature value to obtain a reconstructed image block. Since the feature reconstruction value is enhanced before image reconstruction, and the image is reconstructed according to the enhanced feature value, the distortion of the image feature in the quantization process is reduced, and thus the image quality of the reconstructed image is improved.
[0295] In a possible implementation of the present application, the feature of the current image is a three-dimensional feature matrix.
[0296] The feature enhancement module 20 is further configured to obtain a feature standard deviation and a standard deviation representation value corresponding to each matrix element in the feature of the current image, determine a feature mask corresponding to each matrix element in the feature of the current image according to the feature standard deviation, the standard deviation representation value and a preset threshold, and perform feature enhancement on the feature reconstruction value based on the feature mask to obtain an enhanced feature value.
[0297] In a possible implementation of the present application, the feature enhancement module 20 is further configured to set the feature mask corresponding to a matrix element as a first type of value if the feature standard deviation and the standard deviation representation value corresponding to the matrix element satisfy a preset enhancement condition, and set the feature mask corresponding to the matrix element as a second type of value if the feature standard deviation and the standard deviation representation value corresponding to the matrix element do not satisfy the preset enhancement condition.
[0298] In a possible implementation of the present application, the feature enhancement module 20 is further configured to: if the feature standard deviation corresponding to the matrix element is greater than a preset threshold, and the standard deviation representation value is a first type representation value, determine that the feature standard deviation corresponding to the matrix element and the standard deviation representation value satisfy a preset enhancement condition; or, if the feature standard deviation corresponding to the matrix element is less than the preset threshold, and the standard deviation representation value is a second type representation value, determine that the feature standard deviation corresponding to the matrix element and the standard deviation representation value satisfy the preset enhancement condition.
[0299] In a possible implementation of the present application, the feature enhancement module 20 is further configured to: take, as a target matrix element, a matrix element corresponding to a feature mask of a first type value; and enhance a feature reconstruction value corresponding to the target matrix element to obtain an enhanced feature value.
[0300] In a possible implementation of the present application, the feature enhancement module 20 is further configured to: obtain a feature reconstruction value, a residual reconstruction value, and a predicted feature value corresponding to the target matrix element; determine a first enhancement value according to a first scaling factor and the residual reconstruction value, and determine a second enhancement value according to a second scaling factor and the predicted feature value; and determine an enhanced feature value according to the feature reconstruction value, the first enhancement value, and the second enhancement value.
[0301] In a possible implementation of the present application, the feature enhancement module 20 is further configured to: obtain a feature channel corresponding to the target matrix element; and determine the first scaling factor and the second scaling factor according to the feature channel, wherein different feature channels correspond to different first scaling factors and second scaling factors.
[0302] In a possible implementation of the present application, the feature enhancement module 20 is further configured to: obtain a feature reconstruction value and a predicted feature value corresponding to the target matrix element; determine a first enhancement value according to a first scaling factor and the feature reconstruction value, and determine a second enhancement value according to a second scaling factor and the predicted feature value; and determine an enhanced feature value according to the first enhancement value and the second enhancement value.
[0303] In a possible implementation of the present application, the feature enhancement module 20 is further configured to: obtain a feature reconstruction value and a residual reconstruction value corresponding to the target matrix element; determine a first enhancement value according to a first scaling factor and the feature reconstruction value, and determine a second enhancement value according to a second scaling factor and the residual reconstruction value; and determine an enhanced feature value according to the first enhancement value and the second enhancement value.
[0304] In a possible implementation of the present application, the feature enhancement module 20 is further configured to obtain a feature reconstruction value corresponding to the target matrix element and a feature standard deviation; determine a first enhancement value according to a first scaling factor and the feature reconstruction value, and determine a second enhancement value according to a second scaling factor and the feature standard deviation; and determine an enhanced feature value according to the first enhancement value and the second enhancement value.
[0305] In a possible implementation of the present application, the feature reconstruction value includes a luminance reconstruction value and a chroma reconstruction value, and the enhanced feature value includes a luminance enhanced feature value and a chroma enhanced feature value.
[0306] The feature enhancement module 20 is further configured to enhance the chroma reconstruction value corresponding to the target matrix element to obtain a chroma enhanced feature value, and enhance the luminance reconstruction value corresponding to the target matrix element to obtain a luminance enhanced feature value.
[0307] In a possible implementation of the present application, the feature enhancement module 20 is further configured to determine a first enhancement value according to the luminance reconstruction value and a first scaling factor, and perform secondary enhancement on the chroma enhanced feature value according to the first enhancement value.
[0308] In a possible implementation of the present application, the feature enhancement module 20 is further configured to extract a component indication parameter from an image code stream, and if the component indication parameter is a chroma enhancement type parameter, determine a first enhancement value according to the luminance reconstruction value and a first scaling factor.
[0309] In a possible implementation of the present application, the code stream decoding module 10 is further configured to decode an image code stream, and determine a residual reconstruction value corresponding to a current image feature obtained by decoding; predict an enhanced feature value of a reconstructed feature to obtain a predicted feature value; and determine a feature reconstruction value corresponding to the current image feature according to the residual reconstruction value and the predicted feature value.
[0310] In a possible implementation of the present application, the code stream decoding module 10 is further configured to predict an enhanced feature value of a reconstructed feature to obtain a predicted feature value; and determine a feature reconstruction value corresponding to the current image feature according to the residual reconstruction value and the predicted feature value.
[0311] In a possible implementation of the present application, the feature enhancement manner of the enhanced feature value of the reconstructed feature and the enhanced feature value corresponding to the current image feature is the same or different.
[0312] In a possible implementation of the present application, the feature enhancement manner of the enhanced feature value of the reconstructed feature and the enhanced feature value corresponding to the current image feature is determined by a syntax flag, and the syntax flag is read from an image code stream.
[0313] In one possible implementation of this application, the feature enhancement module 20 is further configured to extract syntax application interval parameters from the image bitstream and obtain feature location information corresponding to the current image feature; determine enhanced syntax parameters based on the feature location information and the syntax application interval parameters; and perform feature enhancement on the feature reconstruction value based on the enhanced syntax parameters to obtain enhanced feature values.
[0314] Reference Figure 13 , Figure 13 This is a structural block diagram of the first embodiment of the image encoding device of the present invention.
[0315] like Figure 13 As shown, the image encoding device proposed in this embodiment of the invention includes:
[0316] Feature extraction module 110 is used to extract features from the image block to be encoded and to use the extracted features as the current image features;
[0317] The feature prediction module 120 is used to predict feature values based on the feature reconstruction values corresponding to the reconstructed features;
[0318] The residual calculation module 130 is used to determine the coding residual coefficients corresponding to the current image features based on the predicted feature values;
[0319] The parameter writing module 140 is used to write the encoding residual coefficients into the image bitstream corresponding to the image block to be encoded.
[0320] This embodiment extracts features from the image block to be encoded and uses the extracted features as the current image features; it then predicts the predicted feature values based on the reconstructed feature values; it determines the coding residual coefficients corresponding to the current image features based on the predicted feature values; and it writes the coding residual coefficients into the image bitstream corresponding to the image block to be encoded. Because the coding residual coefficients of the current image features are calculated and written into the image bitstream during encoding, the encoding efficiency of the image data is reduced.
[0321] In one possible implementation of this application, the parameter writing module 140 is further configured to perform image decoding on the image bitstream corresponding to the image block to be encoded to obtain a reconstructed image block; and determine the image encoding efficiency based on the reconstructed image block and the image block to be encoded.
[0322] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.
[0323] It should be noted that the above-described workflow is merely illustrative and does not limit the scope of protection of the present application. In actual applications, a person skilled in the art can select part or all of the above-described workflow to achieve the purpose of the embodiment according to actual needs, which is not limited herein.
[0324] In addition, technical details not described in detail in the present embodiment can be found in the image decoding or image decoding method provided by any embodiment of the present application, which will not be described here.
[0325] In addition, it should be noted that in this document, the terms "comprise", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or system. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of another identical element in the process, method, article or system that includes the element.
[0326] The above-mentioned embodiment numbers of the present application are only for description, not representing the advantages and disadvantages of the embodiments.
[0327] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and necessary general hardware platforms, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory (ROM) / RAM, a magnetic disk, an optical disk) and includes a number of instructions to make a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) execute the methods described in various embodiments of the present application.
[0328] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. An image decoding method, characterized in that, The image decoding method includes the following steps: Decode the image bitstream and determine the feature reconstruction value corresponding to the current image features obtained from the decoding; The reconstructed feature values are enhanced to obtain enhanced feature values; The enhanced feature values are synthesized and transformed to obtain reconstructed image patches; The step of decoding the image bitstream and determining the feature reconstruction value corresponding to the current image features obtained through decoding includes: The image bitstream is decoded, and the residual reconstruction value corresponding to the current image features obtained from the decoding is determined. Based on the reconstructed feature values, prediction is performed to obtain the predicted feature values; The feature reconstruction value corresponding to the current image feature is determined based on the residual reconstruction value and the predicted feature value.
2. The image decoding method as described in claim 1, characterized in that, The current image features are a three-dimensional feature matrix; The step of enhancing the reconstructed feature values to obtain enhanced feature values includes: Obtain the feature standard deviation and standard deviation representation value corresponding to each matrix element in the current image features; The feature mask corresponding to each matrix element in the current image feature is determined based on the feature standard deviation, the standard deviation representation value, and the preset threshold. Based on the feature mask, feature enhancement is performed on the reconstructed feature values to obtain enhanced feature values.
3. The image decoding method as described in claim 2, characterized in that, The step of determining the feature mask corresponding to each matrix element in the current image feature based on the feature standard deviation, the standard deviation representation value, and a preset threshold includes: If the feature standard deviation and standard deviation representation value corresponding to the matrix element meet the preset enhancement conditions, then the feature mask corresponding to the matrix element is set to the first type value. If the feature standard deviation and standard deviation representation value corresponding to the matrix element do not meet the preset enhancement conditions, then the feature mask corresponding to the matrix element is set to the second type value.
4. The image decoding method as described in claim 3, characterized in that, Before setting the feature mask corresponding to the matrix element to the first type value if the feature standard deviation and standard deviation representation value corresponding to the matrix element satisfy the preset enhancement condition, the method further includes: If the feature standard deviation corresponding to a matrix element is greater than a preset threshold, and the standard deviation representation value is a first type representation value, then it is determined that the feature standard deviation and the standard deviation representation value corresponding to the matrix element satisfy a preset enhancement condition. or, If the feature standard deviation corresponding to a matrix element is less than a preset threshold, and the standard deviation representation value is a second type representation value, then it is determined that the feature standard deviation and standard deviation representation value corresponding to the matrix element satisfy the preset enhancement condition.
5. The image decoding method as described in claim 2, characterized in that, The step of enhancing the reconstructed feature values based on the feature mask to obtain enhanced feature values includes: Use the matrix elements whose corresponding feature masks are of the first type as the target matrix elements; The feature reconstruction values corresponding to the elements of the target matrix are enhanced to obtain enhanced feature values.
6. The image decoding method as described in claim 5, characterized in that, The step of enhancing the feature reconstruction values corresponding to the elements of the target matrix to obtain enhanced feature values includes: Obtain the feature reconstruction value, residual reconstruction value, and predicted feature value corresponding to the elements of the target matrix; A first enhancement value is determined based on a first scaling factor and the residual reconstruction value, and a second enhancement value is determined based on a second scaling factor and the predicted feature value. The enhanced feature value is determined based on the reconstructed feature value, the first enhanced value, and the second enhanced value.
7. The image decoding method as described in claim 6, characterized in that, Before determining the first enhancement value based on the first scaling factor and the predicted feature value, and determining the second enhancement value based on the second scaling factor and the residual reconstruction value, the method further includes: Obtain the feature channels corresponding to the elements of the target matrix; The first scaling factor and the second scaling factor are determined based on the feature channels, and different feature channels correspond to different first scaling factors and second scaling factors.
8. The image decoding method as described in claim 5, characterized in that, The step of enhancing the feature reconstruction values corresponding to the elements of the target matrix to obtain enhanced feature values includes: Obtain the feature reconstruction values and predicted feature values corresponding to the elements of the target matrix; A first enhancement value is determined based on a first scaling factor and the reconstructed feature value, and a second enhancement value is determined based on a second scaling factor and the predicted feature value. The enhancement feature value is determined based on the first enhancement value and the second enhancement value.
9. The image decoding method as described in claim 5, characterized in that, The step of enhancing the feature reconstruction values corresponding to the elements of the target matrix to obtain enhanced feature values includes: Obtain the feature reconstruction values and residual reconstruction values corresponding to the elements of the target matrix; A first enhancement value is determined based on a first scaling factor and the reconstructed feature value, and a second enhancement value is determined based on a second scaling factor and the reconstructed residual value. The enhancement feature value is determined based on the first enhancement value and the second enhancement value.
10. The image decoding method as described in claim 5, characterized in that, The step of enhancing the feature reconstruction values corresponding to the elements of the target matrix to obtain enhanced feature values includes: Obtain the feature reconstruction values and feature standard deviations corresponding to the elements of the target matrix; A first enhancement value is determined based on a first scaling factor and the reconstructed feature value, and a second enhancement value is determined based on a second scaling factor and the feature standard deviation. The enhancement feature value is determined based on the first enhancement value and the second enhancement value.
11. The image decoding method as described in claim 5, characterized in that, The reconstructed feature values include luminance reconstructed values and chrominance reconstructed values, and the enhanced feature values include luminance enhancement feature values and chrominance enhancement feature values; The step of enhancing the feature reconstruction values corresponding to the elements of the target matrix to obtain enhanced feature values includes: The chromaticity reconstruction values corresponding to the elements of the target matrix are enhanced to obtain chromaticity enhancement feature values; The brightness reconstruction values corresponding to the target matrix elements are enhanced to obtain brightness enhancement feature values.
12. The image decoding method as described in claim 11, characterized in that, After enhancing the chroma reconstruction values corresponding to the elements of the target matrix to obtain chroma enhancement feature values, the method further includes: The first enhancement value is determined based on the reconstructed brightness value and the first scaling factor; The chromaticity enhancement feature value is further enhanced based on the first enhancement value.
13. The image decoding method as described in claim 12, characterized in that, The step of determining the first enhancement value based on the reconstructed brightness value and the first scaling factor further includes: Extract component indicator parameters from the image bitstream; If the component indication parameter is a chroma enhancement type parameter, then the first enhancement value is determined based on the luminance reconstruction value and the first scaling factor.
14. The image decoding method as described in claim 1, characterized in that, After decoding the image bitstream and determining the residual reconstruction value corresponding to the current image features obtained from the decoding, the method further includes: Predictive feature values are obtained by making predictions based on the enhanced feature values of the reconstructed features; The feature reconstruction value corresponding to the current image feature is determined based on the residual reconstruction value and the predicted feature value.
15. The image decoding method as described in claim 1, characterized in that, The enhanced feature value of the reconstructed feature may be enhanced in the same way or in a different way than the enhanced feature value corresponding to the current image feature.
16. The image decoding method as described in claim 15, characterized in that, The syntax flags determine whether the enhanced feature value of the reconstructed feature is the same as or different from the enhanced feature value corresponding to the current image feature, and the syntax flags are read from the image bitstream.
17. The image decoding method as described in claim 1, characterized in that, The step of enhancing the reconstructed feature values to obtain enhanced feature values includes: Extract the syntax application interval parameters from the image bitstream and obtain the feature location information corresponding to the current image feature; The enhanced syntax parameters are determined based on the feature location information and the syntax application interval parameters. The feature reconstruction value is enhanced according to the enhanced syntax parameters to obtain enhanced feature values.
18. An image decoding apparatus, characterized in that, The image decoding device includes the following modules: The bitstream decoding module is used to decode the image bitstream and determine the feature reconstruction value corresponding to the current image features obtained by decoding; The feature enhancement module is used to enhance the feature reconstruction values to obtain enhanced feature values; The image reconstruction module is used to perform synthetic transformation on the enhanced feature values to obtain reconstructed image blocks; The bitstream decoding module is further configured to decode the image bitstream and determine the residual reconstruction value corresponding to the current image feature obtained by decoding; make predictions based on the feature reconstruction values of the reconstructed features to obtain predicted feature values; and determine the feature reconstruction value corresponding to the current image feature based on the residual reconstruction value and the predicted feature value.
19. An image encoding method, characterized in that, The image encoding method includes: Feature extraction is performed on the image block to be encoded, and the extracted features are used as the current image features; Predict the predicted feature value based on the feature reconstruction value corresponding to the reconstructed feature; The coding residual coefficients corresponding to the current image features are determined based on the predicted feature values; Write the coding residual coefficients into the image bitstream corresponding to the image block to be encoded; The image bitstream is decoded, and the residual reconstruction value corresponding to the current image features obtained from the decoding is determined. Based on the reconstructed feature values, prediction is performed to obtain the predicted feature values; The feature reconstruction value corresponding to the current image feature is determined based on the residual reconstruction value and the predicted feature value; The reconstructed feature values are enhanced to obtain enhanced feature values; The enhanced feature values are synthesized and transformed to obtain reconstructed image patches; The image coding efficiency is determined based on the reconstructed image block and the image block to be encoded.
20. An image encoding device, characterized in that, The image encoding device includes: The feature extraction module is used to extract features from the image blocks to be encoded and to use the extracted features as the current image features. The feature prediction module is used to predict feature values based on the feature reconstruction values corresponding to the reconstructed features. The residual calculation module is used to determine the coding residual coefficients corresponding to the current image features based on the predicted feature values; The parameter writing module is used to write the encoding residual coefficients into the image bitstream corresponding to the image block to be encoded. The parameter writing module is further configured to decode the image bitstream and determine the residual reconstruction value corresponding to the current image feature obtained by decoding; predict the predicted feature value based on the feature reconstruction value of the reconstructed feature; determine the feature reconstruction value corresponding to the current image feature based on the residual reconstruction value and the predicted feature value; perform feature enhancement on the feature reconstruction value to obtain enhanced feature value; perform synthetic transformation on the enhanced feature value to obtain reconstructed image block; and determine the image coding efficiency based on the reconstructed image block and the image block to be encoded.
21. A decoding device, characterized in that, The decoding device includes: a processor, a memory, and a decoding program stored in the memory and executable on the processor. When the decoding program is executed by the processor, it implements the image decoding method as described in any one of claims 1-17.
22. An encoding device, characterized in that, The encoding device includes: a processor, a memory, and a decoding program and / or an encoding program stored in the memory and executable on the processor. When the decoding program is executed by the processor, it implements the image decoding method as described in any one of claims 1-17, and when the encoding program is executed by the processor, it implements the image encoding method as described in claim 19.
23. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an image decoding program and / or an image encoding program, wherein when the image decoding program is executed by a processor, it implements the image decoding method as described in any one of claims 1-17, and when the image encoding program is executed by a processor, it implements the image encoding method as described in claim 19.
Citation Information
Patent Citations
Depth map end-to-end intelligent compression coding method and device
CN115278246A
A method and an apparatus for encoding / decoding images and videos using artificial neural network based tools
WO2022221374A1