Coded video image compression distortion restoration method and device, equipment and storage medium

By decoding the code stream of the video image encoded code stream and parsing the code stream parameters, structured data is obtained as prior information, and after fusing it with the decoded image, it is used for distortion repair, which solves the problem of encoding distortion and improves the repair effect.

CN120017865AActive Publication Date: 2025-05-16HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510464093.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-16
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

Existing video image encoders have encoding distortion during the encoding process, resulting in problems such as artifact noise, tailing, loss of details, color distortion, mosaic, color level and breathing effects.

Method used

By obtaining the encoded code stream of the current frame video image, decoding is obtained, and the code stream parameter set is parsed to obtain structured data as code stream prior information, perform data preprocessing and fusing with the decoded image, input to the pre-trained neural network for distortion repair.

Benefits of technology

By introducing prior information data, the neural network's ability to repair and generalize encoding distortions is improved, and the distortion repair effect is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017865A_ABST
    Figure CN120017865A_ABST
Patent Text Reader

Abstract

The invention provides a coding video image compression distortion restoration method and device, equipment and a storage medium. The method comprises the following steps: acquiring a coding code stream of a current frame video image; decoding the coding code stream of the current frame video image to obtain a decoded image; analyzing a code stream parameter set in the coding code stream to obtain parameter type structured data, taking the parameter type structured data as code stream prior information, and performing data preprocessing on the code stream prior information to obtain prior information data; and fusing the prior information data and the decoded image, inputting fusion information into a pre-trained neural network, and performing distortion repair on the decoded image by using the pre-trained neural network. Through the scheme of the invention, the compression distortion restoration performance of the coded video image can be optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of video image post-processing, and in particular to a method, device, equipment and storage medium for repairing compression distortion of encoded video images. Background Art

[0002] In order to save space, video images are encoded before transmission. Complete video encoding may include prediction, transformation, quantization, entropy coding, filtering and other processes.

[0003] Existing mainstream video image encoders all operate based on blocks (such as coding units, prediction units, transform units, filtering units, etc.), and coding distortion may occur in the above process.

[0004] For example, in the intra-frame prediction stage, there will be inaccurate predictions and artifact noise. In the inter-frame prediction stage, there will be inaccurate motion estimation and tailing. In the transform quantization stage, there will be residual coefficient quantization loss, resulting in detail loss (blurring), color distortion, and checkerboard. In the filtering stage, the similarity of the original pixels will be destroyed, resulting in mosaics, color levels, etc. The coding quality between different frames varies greatly, resulting in a breathing effect. Summary of the invention

[0005] In view of this, the present application provides a method, apparatus, device and storage medium for repairing compression distortion of encoded video images.

[0006] Specifically, the present application is implemented through the following technical solutions: According to a first aspect of an embodiment of the present application, a method for repairing compression distortion of a coded video image is provided, comprising: Get the encoded code stream of the current frame video image; Decoding the coded bitstream of the current frame video image to obtain a decoded image; and, Parsing a code stream parameter set in the coded code stream to obtain parameter-type structured data, using the parameter-type structured data as code stream prior information, and performing data preprocessing on the code stream prior information to obtain prior information data; The prior information data is fused with the decoded image, and the fused information is input into a pre-trained neural network, and the pre-trained neural network is used to perform distortion repair on the decoded image.

[0007] According to a second aspect of an embodiment of the present application, a device for repairing compression distortion of a coded video image is provided, comprising: An acquisition unit, used to acquire a coded bit stream of a current frame of video image; A decoding unit, configured to decode the coded bit stream of the current frame video image to obtain a decoded image; and A preprocessing unit, configured to parse a code stream parameter set in the coded code stream to obtain parameter-type structured data, use the parameter-type structured data as code stream prior information, and perform data preprocessing on the code stream prior information to obtain prior information data; A restoration unit is used to fuse the prior information data with the decoded image, input the fused information into a pre-trained neural network, and use the pre-trained neural network to perform distortion restoration on the decoded image.

[0008] According to a third aspect of an embodiment of the present application, there is provided an electronic device, the electronic device comprising: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; The processor is used to execute machine executable instructions to implement the method provided by the first aspect.

[0009] According to a fourth aspect of an embodiment of the present application, a machine-readable storage medium is provided, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the method provided in the first aspect is implemented.

[0010] It can be seen from the above technical scheme that in the embodiment of the present application, by obtaining the encoded code stream of the current frame video image, on the one hand, the encoded code stream of the current frame video image is decoded to obtain a decoded image; on the other hand, the code stream parameter set in the encoded code stream is parsed to obtain parameter-type structured data, and the parameter-type structured data is used as code stream prior information, and the code stream prior information is preprocessed to obtain prior information data, and then, the prior information data can be fused with the decoded image, and the fused information can be input into a pre-trained neural network, and the pre-trained neural network is used to perform distortion repair on the decoded image. By introducing the prior information data in the process of using the neural network to repair the distortion, and using the prior information data to assist the neural network in repairing the distortion, the neural network's ability to repair coding distortion and the generalization ability of the neural network are improved, and the distortion repair effect is optimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a flowchart of a method for repairing compression distortion of a coded video image shown in an exemplary embodiment of the present application; Figure 2 It is a schematic diagram of an implementation flow of a neural network-based coded video image compression distortion repair solution shown in an exemplary embodiment of the present application; Figure 3 is a schematic diagram of a priori information data shown in an exemplary embodiment of the present application; Figure 4AIt is a schematic diagram of a direct fusion of prior information data and LQ shown in an exemplary embodiment of the present application; Figure 4B It is a schematic diagram of a direct fusion of prior information data and LQ under a U-Net network shown in an exemplary embodiment of the present application; Figure 5 It is a schematic diagram of feature fusion of a priori information data and LQ shown in an exemplary embodiment of the present application; Figure 6 It is a schematic diagram of a prior information data and LQ multi-scale fusion shown in an exemplary embodiment of the present application; Figure 7 It is a schematic diagram of a fusion of prior information data and LQ spatial domain attention shown in an exemplary embodiment of the present application; Figure 8 It is a schematic diagram of a prior information data and LQ hierarchical fusion shown in an exemplary embodiment of the present application; Fig. 9 It is a structural schematic diagram of a device for repairing compression distortion of a coded video image shown in an exemplary embodiment of the present application; Fig.10 It is a schematic diagram of the hardware structure of an electronic device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0012] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, some technical terms involved in the embodiments of the present application are explained below.

[0013] Post-processing: Video image post-processing refers to the further processing of video image data after encoding, transmission and decoding, including de-coding distortion, denoising, super-resolution, etc. Post-processing can be implemented through software or hardware devices, and is usually used to improve subjective quality.

[0014] Prior information: information that can be obtained after the encoded video image is decoded by the decoder, including the prediction value in the prediction stage, the division of transform units (TU), the division of coding units (CU), the QP value of each CU unit, the bit value (BIT) required for encoding each CU unit, the value (SKIP) for skipping during inter-frame prediction of each CU unit, the reconstructed value before filtering (the prediction value plus the residual coefficient after inverse transformation), the boundary strength value (BS) of deblocking filtering, etc.

[0015] In order to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application are further described in detail below with reference to the accompanying drawings.

[0016] See also Figure 1, is a flow chart of a method for repairing compression distortion of a coded video image provided in an embodiment of the present application, wherein the method for repairing compression distortion of a coded video image can be applied to a decoding end device, such as Figure 1 As shown, the method for repairing compression distortion of coded video images may include the following steps: Step S100: Obtain the encoded bit stream of the current frame of video image.

[0017] In the embodiment of the present application, the current frame video image refers to the video frame image currently being decoded by the decoding end device.

[0018] Step S110: Decode the coded bitstream of the current frame video image to obtain a decoded image.

[0019] Step S120: parsing the bitstream parameter set in the coded bitstream to obtain parameter-type structured data, using the parameter-type structured data as bitstream prior information, and performing data preprocessing on the bitstream prior information to obtain prior information data.

[0020] In an embodiment of the present application, for the current frame video image, when the decoding end device obtains the frame video image, on the one hand, it can decode the encoded code stream of the current frame video image to obtain a decoded image; wherein, the decoded image is a low-quality image (Low-Quality Image, abbreviated as LQ), and the image quality needs to be improved through distortion repair.

[0021] On the other hand, in order to optimize the distortion repair effect and improve the distortion repair performance, the decoding end device can also obtain the code stream prior information based on the encoded code stream of the current frame video image, and perform data preprocessing on the code stream prior information to obtain the prior information data.

[0022] Exemplarily, the prior information data can be used to assist the neural network in performing distortion repair on the decoded image so that the neural network can better perform distortion repair.

[0023] Exemplarily, the parameter-type structured data may be obtained by parsing the bitstream parameter set in the coded bitstream, and the parameter-type structured data may be used as the bitstream prior information.

[0024] In one example, parameter-type structured data may include some or all of the following parameters: Structured data of video image level parameters, coding unit parameters, transform unit parameters, prediction mode parameters, quantization parameters, number of bits, and filtering parameters.

[0025] Step S130: fuse the prior information data with the decoded image, input the fused information into a pre-trained neural network, and use the pre-trained neural network to perform distortion repair on the decoded image.

[0026] In an embodiment of the present application, when a decoded image and prior information data can be obtained, the input features of a pre-trained neural network (a neural network used for repairing compression distortion of encoded video images) can be determined based on the prior information data and the decoded image, and based on the determined input features, the pre-trained neural network can be used to repair the distortion of the decoded image.

[0027] Exemplarily, the prior information data can be integrated into the input features of the neural network at different locations of the neural network, at different scales, and / or using different fusion methods to enhance the neural network's ability to repair coding distortion.

[0028] Exemplarily, the prior information data obtained in step S120 may be fused with the decoded image to obtain fused information.

[0029] Exemplarily, the above-mentioned fusion may include splicing in the channel dimension, and the splicing in the channel dimension may include splicing the prior information data and the decoded image in the channel dimension; and / or, extracting features from the prior information data, and splicing the obtained prior information features with the decoded image in the channel dimension.

[0030] Exemplarily, when the fusion information is obtained, the fusion information can be input into a pre-trained neural network, and the pre-trained neural network is used to perform distortion repair on the decoded image.

[0031] Exemplarily, the prior information may include mask-type prior information and / or non-mask-type prior information.

[0032] Among them, mask-type prior information can guide the network to focus on the distorted areas in the decoded image; non-mask-type prior information can provide additional detail information, such as unfiltered prior information can provide more original detail information, thereby optimizing the distortion repair effect of the neural network.

[0033] Exemplarily, different types of prior information data may be fused with the decoded image in different fusion modes.

[0034] Exemplarily, the width and height of the prior information data may be respectively consistent with the width and height of the decoded image.

[0035] Since the prior information needs to be fed into the network to be learned together with the decoded image, it is usually necessary to ensure that the width / height of the prior information is consistent with the width / height of the decoded image.

[0036] In an example, the number of channels of the prior information data may also be consistent with the number of channels of the decoded image.

[0037] It can be seen that in Figure 1In the method flow shown, by obtaining the encoded code stream of the current frame video image, on the one hand, the encoded code stream of the current frame video image is decoded to obtain a decoded image; on the other hand, the code stream parameter set in the encoded code stream is parsed to obtain parameter-type structured data, the parameter-type structured data is used as code stream prior information, and the code stream prior information is preprocessed to obtain prior information data, and then, the prior information data can be fused with the decoded image, and the fused information can be input into a pre-trained neural network, and the pre-trained neural network is used to perform distortion repair on the decoded image. By introducing the prior information data in the process of using the neural network to repair the distortion, and using the prior information data to assist the neural network in performing distortion repair, the neural network's ability to repair coding distortion and the generalization ability of the neural network are improved, and the distortion repair effect is optimized.

[0038] In some embodiments, the above-mentioned fusion of the prior information data and the decoded image may include: The prior information data and the decoded image are concatenated in the channel dimension.

[0039] Exemplarily, when the prior information data is obtained in the above manner, the prior information data and the decoded image can be spliced ​​in the channel dimension to obtain splicing features, and the splicing features are input into a pre-trained neural network as fusion information, and the pre-trained neural network is used to repair the distortion of the decoded image.

[0040] In another example, the above-mentioned fusing of the prior information data with the decoded image may include: Feature extraction is performed on the prior information data to obtain prior information features, the prior information features and the decoded image are spliced ​​in the channel dimension, and channel dimensionality reduction is performed on the spliced ​​features.

[0041] Exemplarily, when the prior information data is obtained in the above manner, feature extraction can be performed on the obtained prior information data. For example, feature extraction can be performed on the prior information data using a convolution method to obtain prior information features, and the prior information features and the decoded image can be spliced ​​in the channel dimension to obtain spliced ​​features.

[0042] Exemplarily, in order to reduce the processing complexity of the neural network, the above-mentioned splicing features can also be aligned to perform channel dimensionality reduction processing, and the splicing features after dimensionality reduction processing are input into a pre-trained neural network, and the pre-trained neural network is used to repair the distortion of the decoded image.

[0043] For example, assuming that the type of prior information data is k (k ≥ 1), and the number of channels of the decoded image is C, when the spliced ​​feature is obtained by the above-mentioned feature splicing method, the channel dimension of the spliced ​​feature can be reduced to (k+1)*C channels.

[0044] It should be noted that in the embodiments of the present application, in the process of fusing the prior information data with the decoded image, feature extraction can also be performed on the decoded image. For example, feature extraction is performed on the decoded image using a convolution method, and feature splicing is performed on the decoded image features and the prior information features to obtain spliced ​​features. The specific implementation will not be described here.

[0045] In addition, in the embodiment of the present application, when the splicing features are obtained, it is not necessary to reduce the channel dimension of the splicing features, but the splicing features are used as input features of the pre-trained neural network.

[0046] As an example, the above-mentioned inputting the fusion information into the pre-trained neural network and using the pre-trained neural network to perform distortion repair on the decoded image may include: Downsampling the fused information at at least one scale to obtain fused information at at least one downsampling scale; The fused information and the fused information of at least one down-sampling scale are respectively used as input features of corresponding scales of a pre-trained neural network, and the pre-trained neural network is used to perform distortion repair on the decoded image.

[0047] Exemplarily, in order to enhance feature diversity, improve the receptive field of the network, enhance the network's sensitivity to details, and improve the robustness of the network, the fused information obtained by fusing the prior information data with the decoded image can be downsampled at at least one scale, for example, the fused information is downsampled by x2 or x4 to obtain fused information of at least one downsampled scale (in the case of including fused information of the original scale, there are at least two fused information of different scales).

[0048] The fusion information of the original scale and the at least one down-sampling scale can be respectively input into different scales of a pre-trained neural network as input features of the corresponding scale of the pre-trained neural network, and the pre-trained neural network is used to perform distortion repair on the decoded image.

[0049] As an example, the prior information data may include at least one type of target prior information data, and the target prior information data is used to indicate an area that needs to be focused on during the distortion restoration process.

[0050] The above-mentioned inputting the fusion information into the pre-trained neural network and using the pre-trained neural network to perform distortion repair on the decoded image may include: The fused information is input into a pre-trained neural network, and, at the decoding end of the pre-trained neural network, the target prior information data is integrated into the network in a spatial domain attention manner, and the pre-trained neural network is used to repair the distortion of the decoded image.

[0051] Exemplarily, the prior information data includes at least one type of target prior information data, and the target prior information data is used to indicate an area that needs to be focused on during the distortion restoration process.

[0052] For example, the target prior information data may include deblocking filter boundary strength prior information (BS for short), where the BS includes a boundary strength value of each deblocking filter unit saved when decoding a video image.

[0053] Based on the boundary strength values ​​of each deblocking filter unit in the decoded image, for a region with a larger boundary strength value, it indicates that the region in the decoded image has been strongly filtered and needs to be paid special attention to during the distortion repair process.

[0054] For another example, the target prior information data may include quantization parameter prior information (QP for short) and bit prior information (BIT) required for the coding unit, where QP includes the mapped quantization parameter of each coding unit saved when decoding the video image, and BIT includes the bits required to encode each pixel under each coding unit saved when decoding the video image.

[0055] For any coding unit in the decoded image, if the quantization parameter of the coding unit is large, it indicates that the distortion degree of the area corresponding to the coding unit is large. At the same time, if the number of bits required for encoding the coding unit is also large, it indicates that the area needs special attention.

[0056] Accordingly, in the process of using a pre-trained neural network to repair the distortion of the decoded image, on the one hand, the prior information data and the decoded image can be fused in the above manner to obtain fused information, and the fused information can be input into the pre-trained neural network; on the other hand, the target prior information data can be integrated into the network in a spatial attention manner, and the pre-trained neural network can be used to repair the distortion of the decoded image.

[0057] In some embodiments, the prior information data may include mask-type prior information data and non-mask-type prior information data; the width, height and number of channels of the prior information data are respectively consistent with the width, height and number of channels of the decoded image.

[0058] The above-mentioned fusion of the prior information data and the decoded image may include: The non-masked prior information data and the decoded image are spliced ​​in the channel dimension; or, feature extraction is performed on the non-masked prior information data to obtain prior information features, and the prior information features and the decoded image are spliced ​​in the channel dimension; The splicing features and the mask-like prior information data are fused by using pixel-by-pixel multiplication and / or pixel-by-pixel addition.

[0059] Exemplarily, the prior information data may include mask-type prior information data and non-mask-type prior information data.

[0060] Exemplarily, the visualization effect of the non-mask prior information data is close to the visualization effect of the decoded image, that is, the visual effect of the non-mask prior information data is closer to the image.

[0061] For example, the non-mask prior information may include part or all of the prediction prior information (PRED for short), the reconstruction prior information before filtering (REC for short), and the transform unit division prior information (TU for short).

[0062] The visualization effect of mask-like prior information data is similar to the mask of the decoded image.

[0063] For example, mask-type prior information may include part or all of the prior information such as deblocking filter boundary strength prior information (BS for short), quantization parameter prior information (QP for short), coding unit required bit prior information (BIT for short), and whether the coding unit is skipped prior information (SKIP for short).

[0064] Exemplarily, in the process of fusing the prior information data with the decoded image, for non-mask prior information data, it can be fused with the decoded image by splicing in the channel dimension; or, feature extraction can be performed first, and then the extracted features and the decoded image are spliced ​​in the channel dimension to obtain spliced ​​features.

[0065] For mask-like prior information data, the splicing features and the mask-like prior information data may be fused using pixel-by-pixel multiplication and / or pixel-by-pixel addition.

[0066] In some embodiments, the prior information data may include part or all of the following prior information: Quantization parameter prior information and bit prior information at the image level, image block level or pixel level; Image block-level prediction prior information, deblocking filter boundary strength prior information, transform unit division prior information, and whether to skip inter-frame prediction prior information; Reconstructed pixel prior information before pixel-level filtering.

[0067] In some embodiments, the prior information data includes part or all of the following prior information: Prediction prior information (PRED for short); wherein the prediction prior information includes the prediction value under each transform unit saved when decoding the video image.

[0068] Reconstruction prior information before filtering (REC for short); wherein the reconstruction prior information before filtering includes the reconstruction value before filtering under each coding unit saved when decoding the video image.

[0069] Exemplarily, the reconstructed value before filtering is the residual coefficient obtained by adding the predicted pixel to the inverse quantization and inverse transformation.

[0070] Deblocking filter boundary strength prior information (BS for short); wherein the deblocking filter boundary strength prior information includes a boundary strength value of each deblocking filter unit saved when decoding a video image.

[0071] Exemplarily, one filter unit stores one boundary strength value.

[0072] For example, assuming that the boundary strength values ​​include 0, 1, and 2, corresponding to no filtering, weak filtering, and strong filtering, respectively, then in the process of decoding the video image, the saved boundary strength value can be mapped to 0~255, and the corresponding mapping values ​​(that is, the saved boundary strength values) can be 0, 127, and 255, respectively, representing no filtering, weak filtering, and strong filtering, respectively.

[0073] Quantization parameter prior information (QP for short); wherein the quantization parameter prior information includes the mapped quantization parameter of each coding unit saved when decoding the video image.

[0074] Exemplarily, one filtering unit stores one quantization parameter.

[0075] Exemplarily, the quantization parameter range of the encoder may be uniformly mapped to 0-255 to obtain the mapped quantization parameter.

[0076] Transformation unit partition prior information (TU for short); wherein the transformation unit partition prior information includes an average value of predicted pixel values ​​under each transformation unit saved when decoding a video image, or the size of each transformation unit.

[0077] Exemplarily, the average value of the predicted pixel values ​​under the transformation unit may be the sum of the predicted pixel values ​​of each pixel in the transformation unit divided by the area of ​​the transformation unit (that is, the number of pixels in the transformation unit).

[0078] Exemplarily, the size of the transformation unit corresponding to the transformation unit division prior information may be the width / height of the transformation unit.

[0079] For example, for an 8*8 transform unit, the value corresponding to the transform unit in the transform unit division prior information is 8.

[0080] A priori information of bits required for coding units (BIT for short); wherein the a priori information of bits required for coding units includes the bits required for coding each pixel under each coding unit saved when decoding a video image.

[0081] Exemplarily, the bits required to encode each pixel in each coding unit may be a value of bits used to encode the coding unit divided by the area of ​​the coding unit, and then the value is mapped to 0-255.

[0082] Whether the coding unit skips the prior information (SKIP for short); wherein, whether the coding unit skips the prior information includes saving whether each coding unit is skipped when decoding the video image.

[0083] Exemplarily, the skip mode only exists in inter-frame prediction. During the decoding of video images, each coding unit has only one situation, skipped or not skipped, which is mapped to 0 or 255 to save whether each coding unit is skipped.

[0084] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, the technical solutions provided by the embodiments of the present application are described below with reference to specific examples.

[0085] The embodiment of the present application provides a scheme for repairing compression distortion of coded video images based on a neural network, such as Figure 2 As shown, the implementation process of the neural network-based coded video image compression distortion repair solution can be as follows: S1. By modifying the decoder in the decoding device, data preprocessing is performed on the code stream prior information to obtain a variety of prior information data that are consistent with the width, height, and number of channels of the encoded video image.

[0086] Exemplarily, the input of data preprocessing may include structured data of video image level parameters, coding unit parameters, transform unit parameters, prediction mode parameters, quantization parameters, number of bits, and filtering parameters decoded by the decoder, and the output data is prior information data with a size consistent with the width, height, and number of channels of the encoded video image.

[0087] Exemplarily, the coverage of the prior information data may include: Quantization parameter prior information and bit prior information at the image level, image block level or pixel level; Image block-level prediction prior information, deblocking filter boundary strength prior information, transform unit division prior information, and whether to skip inter-frame prediction prior information; Reconstructed pixel prior information before pixel-level filtering.

[0088] S2. Integrate the prior information data at different positions and scales of the neural network by addition, multiplication, connection or cross product to enhance the neural network's ability to repair coding distortion.

[0089] The prior information data and the integration of the prior information data are explained below respectively.

[0090] 1. Prior information data.

[0091] like Figure 3 As shown, the prior information data includes the following prior information: Prediction prior information (PRED for short); wherein the prediction prior information includes the prediction value under each transform unit saved when decoding the video image.

[0092] Reconstruction prior information before filtering (REC for short); wherein the reconstruction prior information before filtering includes the reconstruction value before filtering under each coding unit saved when decoding the video image.

[0093] Exemplarily, the reconstructed value before filtering is the residual coefficient obtained by adding the predicted pixel to the inverse quantization and inverse transformation.

[0094] Deblocking filter boundary strength prior information (BS for short); wherein the deblocking filter boundary strength prior information includes a boundary strength value of each deblocking filter unit saved when decoding a video image.

[0095] Exemplarily, one filter unit stores one boundary strength value.

[0096] For example, assuming that the boundary strength values ​​include 0, 1, and 2, corresponding to no filtering, weak filtering, and strong filtering, respectively, then in the process of decoding the video image, the saved boundary strength value can be mapped to 0~255, and the corresponding mapping values ​​(that is, the saved boundary strength values) can be 0, 127, and 255, respectively, representing no filtering, weak filtering, and strong filtering, respectively.

[0097] Quantization parameter prior information (QP for short); wherein the quantization parameter prior information includes the mapped quantization parameter of each coding unit saved when decoding the video image.

[0098] Exemplarily, one filtering unit stores one quantization parameter.

[0099] Exemplarily, the quantization parameter range of the encoder may be uniformly mapped to 0-255 to obtain the mapped quantization parameter.

[0100] Transformation unit partition prior information (TU for short); wherein the transformation unit partition prior information includes an average value of predicted pixel values ​​under each transformation unit saved when decoding a video image, or the size of each transformation unit.

[0101] Exemplarily, the average value of the predicted pixel values ​​under the transformation unit may be the sum of the predicted pixel values ​​of each pixel in the transformation unit divided by the area of ​​the transformation unit (that is, the number of pixels in the transformation unit).

[0102] Exemplarily, the size of the transformation unit corresponding to the transformation unit division prior information may be the width / height of the transformation unit.

[0103] For example, for an 8*8 transform unit, the value corresponding to the transform unit in the transform unit division prior information is 8.

[0104] A priori information of bits required for coding units (BIT for short); wherein the a priori information of bits required for coding units includes the bits required for coding each pixel under each coding unit saved when decoding a video image.

[0105] Exemplarily, the bits required to encode each pixel in each coding unit may be a value of bits used to encode the coding unit divided by the area of ​​the coding unit, and then the value is mapped to 0-255.

[0106] Whether the coding unit skips the prior information (SKIP for short); wherein, whether the coding unit skips the prior information includes saving whether each coding unit is skipped when decoding the video image.

[0107] Exemplarily, the skip mode only exists in inter-frame prediction. During the decoding of video images, each coding unit has only one situation, skipped or not skipped, which is mapped to 0 or 255 to save whether each coding unit is skipped.

[0108] 2. Integration of prior information data.

[0109] Exemplarily, the prior information data can be input into the neural network using different strategies to enhance the neural network's ability to repair coding distortion.

[0110] For example, taking the U-Net network as an example, at the loss function level, the method of integrating prior information data can be determined in combination with the main sources of coding distortion.

[0111] 2.1. Direct fusion.

[0112] For example, Figure 4A As shown in FIG. 1 , all prior information data are concatenated (CONCAT) with LQ (i.e., decoded image) in the channel dimension to become 8C channels, and the concatenated features are input into the neural network; where C represents the number of channels of LQ.

[0113] For example, the schematic diagram of the case where the neural network is a U-Net network can be as follows: Figure 4B shown.

[0114] like Figure 4B As shown, H and W are the width and height of the decoded image, and F is the number of channels of the spliced ​​features.

[0115] 2.2. Feature fusion.

[0116] like Figure 5 As shown, convolution is used to extract features from the prior information data, and the extracted prior information features are concatenated with LQ in the channel dimension. Convolution is then used to perform channel dimension change processing, processed into 8C channels, and input into the neural network.

[0117] 2.3. Multi-scale fusion.

[0118] Exemplarily, all prior information data are first subjected to feature extraction using convolution and then concatenated with LQ in the channel dimension; or, all prior information data are concatenated with LQ in the channel dimension.

[0119] The concatenated features are downsampled at multiple scales and used as inputs of different scales of the neural network.

[0120] like Figure 6 As shown, taking the U-Net network as an example, the concatenated features can be downsampled by x2 and x4 and used as inputs of the corresponding scales of the neural network respectively.

[0121] 2.4. Spatial Attention Fusion.

[0122] Exemplarily, all prior information data are first subjected to feature extraction using convolution and then concatenated with LQ in the channel dimension; or, all prior information data are concatenated with LQ in the channel dimension, and the concatenated features are input into the neural network.

[0123] At the neural network decoding end, QP, BIT and BS (i.e., the target prior information data mentioned above) are integrated into the neural network in the form of spatial attention, guiding the neural network to pay more attention to the repair of the distorted areas of the video image.

[0124] For example, Figure 7 As shown, taking the U-Net network as an example, the schematic diagram of integrating QP, BIT and BS into the neural network in a spatial domain attention manner can be shown as follows Figure 7 shown.

[0125] 2.5. Hierarchical integration.

[0126] Exemplarily, non-mask prior information data (such as PRED, REC and TU) are subjected to convolution for feature extraction and are concatenated with LQ in the channel dimension. They are then fused with mask prior information data (such as BS, SKIP, QP and BIT) by pixel-by-pixel multiplication and / or pixel-by-pixel addition, and then input into the neural network.

[0127] like Figure 8 As shown, PRED, REC and TU are convolved for feature extraction and concatenated with LQ in the channel dimension; then pixel-by-pixel product is performed with BS and SKIP; then pixel-by-pixel addition is performed with QP and BIT; and then input into the neural network.

[0128] The method provided by the present application is described above. The device provided by the present application is described below: like Fig. 9 FIG. 1 is a flow chart of a device for repairing compression distortion of a coded video image provided by an embodiment of the present application, wherein the device for repairing compression distortion of a coded video image can be deployed in a decoding end device, such as Fig. 9 As shown, the coded video image compression distortion repair device may include: The acquisition unit 910 is used to acquire the encoded bit stream of the current frame video image; The decoding unit 920 is used to decode the coded bit stream of the current frame video image to obtain a decoded image; and The preprocessing unit 930 is used to parse the code stream parameter set in the coded code stream to obtain parameter-type structured data, use the parameter-type structured data as code stream prior information, and perform data preprocessing on the code stream prior information to obtain prior information data; The restoration unit 940 is used to fuse the prior information data with the decoded image, input the fused information into a pre-trained neural network, and use the pre-trained neural network to perform distortion restoration on the decoded image.

[0129] Exemplarily, the specific process of the acquisition unit 910, the decoding unit 920, the preprocessing unit 930 and the repair unit 940 to implement the repair of the compression distortion of the encoded video image can be referred to the relevant description in the above embodiment, and the embodiments of the present application are not repeated here.

[0130] See also Fig.10, is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device may include a processor 1001 and a machine-readable storage medium 1002 storing machine-executable instructions. The processor 1001 and the machine-readable storage medium 1002 may communicate via a system bus 1003. In addition, by reading and executing the machine-executable instructions corresponding to the encoded video image compression distortion repair logic in the machine-readable storage medium 1002, the processor 1001 may execute the encoded video image compression distortion repair method described above.

[0131] The machine-readable storage medium 1002 mentioned herein may be any electronic, magnetic, optical or other physical storage device that may contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium may be: RAM (Radom Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disk, DVD, etc.), or similar storage medium, or a combination thereof.

[0132] In some embodiments, a machine-readable storage medium is also provided, wherein the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the above-described method for repairing compression distortion of coded video images is implemented. For example, the machine-readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0133] An embodiment of the present application also provides a computer application, which, when executed by a processor, can implement the method for repairing compression distortion of encoded video images disclosed in the above example of the present application.

Claims

1. A method for repairing compression distortion of coded video images, characterized in that: include: Get the encoded code stream of the current frame video image; Decoding the encoded code stream of the current frame video image to obtain a decoded image; as well as, Parsing a code stream parameter set in the coded code stream to obtain parameter-type structured data, using the parameter-type structured data as code stream prior information, and performing data preprocessing on the code stream prior information to obtain prior information data; The prior information data is fused with the decoded image, and the fused information is input into a pre-trained neural network, and the pre-trained neural network is used to perform distortion repair on the decoded image.

2. The method according to claim 1, characterized in that The width, height and number of channels of the prior information data are respectively consistent with the width and height of the decoded image; The fusing the prior information data with the decoded image comprises: Splicing the prior information data and the decoded image in a channel dimension; or, Feature extraction is performed on the prior information data to obtain prior information features, the prior information features and the decoded image are spliced ​​in the channel dimension, and channel dimension reduction processing is performed on the spliced ​​features.

3. The method according to claim 2, characterized in that The step of inputting the fusion information into a pre-trained neural network and using the pre-trained neural network to perform distortion repair on the decoded image comprises: Downsampling the fused information at at least one scale to obtain fused information at at least one downsampling scale; The fusion information and the fusion information of the at least one down-sampling scale are respectively used as input features of the corresponding scale of the pre-trained neural network, and the pre-trained neural network is used to perform distortion repair on the decoded image.

4. The method according to claim 2, characterized in that: The prior information data includes at least one type of target prior information data, and the target prior information data is used to indicate an area that needs to be focused on during the distortion restoration process; The step of inputting the fusion information into a pre-trained neural network and using the pre-trained neural network to perform distortion repair on the decoded image comprises: The fusion information is input into a pre-trained neural network, and at the decoding end of the pre-trained neural network, the target prior information data is integrated into the network in a spatial domain attention manner, and the pre-trained neural network is used to repair the distortion of the decoded image.

5. The method according to claim 1, characterized in that The prior information data includes mask-type prior information data and non-mask-type prior information data; the width and height of the prior information data are respectively consistent with the width and height of the decoded image; The fusing the prior information data with the decoded image comprises: The non-mask prior information data and the decoded image are spliced ​​in the channel dimension; or, feature extraction is performed on the non-mask prior information data to obtain prior information features, and the prior information features and the decoded image are spliced ​​in the channel dimension; The splicing features are fused with the mask-like prior information data by using pixel-by-pixel multiplication and / or pixel-by-pixel addition.

6. The method according to any one of claims 1 to 5, characterized in that: The parameter-type structured data includes some or all of the following parameters: Structured data of video image level parameters, coding unit parameters, transform unit parameters, prediction mode parameters, quantization parameters, number of bits, and filtering parameters.

7. The method according to any one of claims 1 to 5, characterized in that: The prior information data includes part or all of the following prior information: Quantization parameter prior information and bit prior information at the image level, image block level or pixel level; Image block-level prediction prior information, deblocking filter boundary strength prior information, transform unit division prior information, and whether to skip inter-frame prediction prior information; Reconstructed pixel prior information before pixel-level filtering.

8. The method according to any one of claims 1 to 5, characterized in that: The prior information data includes part or all of the following prior information: Prediction prior information; wherein the prediction prior information includes a prediction value under each transform unit saved when decoding the video image; Reconstruction prior information before filtering; wherein the reconstruction prior information before filtering includes the reconstruction value before filtering under each coding unit saved when decoding the video image; Deblocking filter boundary strength prior information; wherein the deblocking filter boundary strength prior information includes a boundary strength value of each deblocking filter unit saved when decoding the video image; Quantization parameter prior information; wherein the quantization parameter prior information includes the mapped quantization parameter of each coding unit saved when decoding the video image; Transformation unit division prior information; wherein the transformation unit division prior information includes an average value of predicted pixel values ​​under each transformation unit saved when decoding the video image, or the size of each transformation unit; A priori information on bits required for the coding unit; wherein the a priori information on bits required for the coding unit includes the bits required for encoding each pixel under each coding unit saved when decoding the video image; Whether the coding unit skips the priori information; wherein, whether the coding unit skips the priori information includes saving whether each coding unit is skipped when decoding the video image.

9. A device for repairing compression distortion of coded video images, characterized in that: include: An acquisition unit, used to acquire a coded bit stream of a current frame of video image; A decoding unit, used for decoding the coded bit stream of the current frame video image to obtain a decoded image; as well as, A preprocessing unit, configured to parse a code stream parameter set in the coded code stream to obtain parameter-type structured data, use the parameter-type structured data as code stream prior information, and perform data preprocessing on the code stream prior information to obtain prior information data; A restoration unit is used to fuse the prior information data with the decoded image, input the fused information into a pre-trained neural network, and use the pre-trained neural network to perform distortion restoration on the decoded image.

10. An electronic device, characterized in that: The electronic device comprises: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; The processor is used to execute machine executable instructions to implement the method described in any one of claims 1 to 8.

11. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores a plurality of computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Single-image super-resolution analysis method matched with natural degradation conditions

    CN112163998A

  • Image processing method and device

    CN115375909A

  • Image processing method and related device

    CN115409697A

  • Variable code rate image compression method, system and device, terminal and storage medium

    CN115988215A

  • Image super-division processing method and device, computer equipment and medium

    CN117689539A

Cited By

  • Image quality enhancement method and device, model training method and device, storage medium and program product

    CN121660908A