Video compression method, video decompression method and related devices thereof

By enhancing motion estimation and error information based on the current frame and reference frame during video encoding, the problem of inaccurate motion compensation information is solved, and the prediction accuracy and effect of video compression are improved.

CN120676157APending Publication Date: 2025-09-19ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410317760.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing video coding technologies have difficulty ensuring the accuracy of motion compensation information during the motion information compression process of motion estimation, resulting in poor prediction accuracy and video compression effects.

Method used

By performing motion estimation based on the current frame and the reference frame, first motion information is obtained, and the first motion information is enhanced using error information to obtain second motion information. Finally, the current frame is predicted and compressed based on the enhanced motion information.

Benefits of technology

The accuracy of motion information is improved, thereby improving the prediction accuracy of the current frame and the video compression effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676157A_ABST
    Figure CN120676157A_ABST
Patent Text Reader

Abstract

The invention provides a video compression method, a video decompression method and related devices. The video compression method comprises the following steps: performing motion estimation based on a current frame and a reference frame to obtain first motion information; performing motion compensation based on the first motion information and the reference frame to obtain a preliminary prediction frame of the current frame; calculating first error information of the preliminary prediction frame and the current frame; enhancing the first motion information based on the first error information to obtain second motion information; and performing prediction and compression of the current frame based on the second motion information. The prediction accuracy and the video compression effect can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of coding technology, and in particular to a video compression method, a video decompression method and related devices. Background Art

[0002] Video encoding and decoding systems primarily consist of three main components: encoding, transmission, and decoding. Since video images are relatively large, the primary function of video encoding is to compress the video pixel data (RGB, YUV, etc.) into a video stream, thereby reducing the video data volume and ultimately reducing network bandwidth and storage space during transmission.

[0003] Traditional video coding systems primarily involve video acquisition, prediction (including intra-frame prediction and inter-frame prediction, which remove spatial and temporal redundancy in video images, respectively), transform quantization, and entropy coding. These processes compress and encode video images to produce compressed data, which the decoder then reconstructs based on. With the development of traditional video coding standards, computational complexity has increased. Furthermore, traditional video coding technologies typically optimize for objective metrics like PSNR, making them less adaptable to subjective quality requirements.

[0004] In recent years, deep neural networks have made significant progress in video processing applications such as video detection, video super-resolution, video denoising, and video enhancement. Due to their powerful nonlinear representation capabilities and the advantages of joint training, deep neural networks have shown great potential in the image and video fields. However, related neural network video compression schemes directly compress motion information from motion estimation, making it difficult to ensure the accuracy of the compensation information derived from this motion information. As a result, these methods may suffer from poor prediction accuracy and video compression performance. Summary of the Invention

[0005] The present application provides a video compression method, a video decompression method and related devices, which can improve the prediction accuracy and video compression effect.

[0006] To achieve the above objectives, the present application provides a video compression method, which includes:

[0007] Performing motion estimation based on the current frame and the reference frame to obtain first motion information;

[0008] Performing motion compensation based on the first motion information and the reference frame to obtain a preliminary prediction frame of the current frame;

[0009] Calculating first error information between the preliminary predicted frame and the current frame;

[0010] enhancing the first motion information based on the first error information to obtain second motion information;

[0011] The current frame is predicted and compressed based on the second motion information.

[0012] To achieve the above objectives, the present application provides a video decompression method, the method comprising:

[0013] decoding a motion information code stream of a current frame to obtain third motion information, where the motion information code stream is obtained by compressing second motion information, where the second motion information is obtained by enhancing the first motion information based on first error information, where the first motion information is obtained by motion estimation based on the current frame and a reference frame, where the first error information is error information between the preliminary predicted frame and the current frame, where the preliminary predicted frame is obtained by motion compensation based on the first motion information and the reference frame;

[0014] The context code stream of the current frame is decoded using the third motion information to obtain a reconstructed image of the current frame.

[0015] To achieve the above-mentioned purpose, the present application provides an electronic device, which includes a memory and a processor; the memory stores a computer program, and the processor is used to execute the computer program to implement the steps of the above-mentioned method.

[0016] To achieve the above objectives, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above method when the computer program is executed by a processor.

[0017] The method of the present application is: performing motion estimation based on the current frame and the reference frame to obtain first motion information, using the first error information between the current frame and the preliminary predicted frame of the current frame determined by the first motion information to enhance the first motion information to obtain second motion information, and then predicting and compressing the current frame based on the second motion information. In this way, using the first error information between the current frame and the preliminary predicted frame of the current frame determined by the first motion information to enhance and optimize the first motion information can improve the accuracy of the motion information, thereby improving the accuracy of the current frame prediction, and thus improving the video compression effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 This is a flow chart of the first embodiment of the video compression method of the present application;

[0020] Figure 2 This is a flow chart of an embodiment of the video compression method of the present application;

[0021] Figure 3 This is a flow chart of an embodiment of motion information enhancement in the video compression method of the present application;

[0022] Figure 4 This is a flow chart of another embodiment of the motion information enhancement method in the video compression method of the present application;

[0023] Figure 5 This is a flow chart of an example of motion information enhancement in the video compression method of the present application;

[0024] Figure 6 This is a flow chart of another embodiment of the motion information enhancement method in the video compression method of the present application;

[0025] Figure 7 This is a flowchart of another example of motion information enhancement in the video compression method of the present application;

[0026] Figure 8 This is a flow chart of an embodiment of the predictive frame enhancement method in the video compression method of the present application;

[0027] Figure 9 This is a processing diagram of an embodiment of a context coding network in the video compression method of the present application;

[0028] Figure 10 This is a processing diagram of an embodiment of a context decoding network in the video compression method of the present application;

[0029] Figure 11 This is a flow chart of an implementation method of a video decompression method of the present application;

[0030] Figure 12 This is a schematic diagram of the structure of an implementation method of an end-to-end video compression network of the present application;

[0031] Figure 13 It is a schematic diagram of the structure of the electronic device of the present application;

[0032] Figure 14 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0033] In order to enable those skilled in the art to better understand the technical solution of the present application, the video compression method, video decompression method and related devices provided by the present application are further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0034] In related technologies, the motion information obtained by motion estimation is directly used to predict, compress, or reconstruct the current frame; or the motion information is enhanced by combining the reference frame, the current frame, and the context information of the reference frame, and then the enhanced motion information is used to predict, compress, or reconstruct the current frame. The corresponding motion information will be used at the decoding end to perform a warp operation (also called compensation or alignment) on the reference frame to obtain the current predicted frame. Because this motion information is not compensated before encoding in related technologies, the accuracy of the predicted frame obtained after compensation at the decoding end cannot be guaranteed, that is, the accuracy of the predicted frame cannot be perceived.

[0035] See also Figure 1 and Figure 2 , Figure 1 This is a flow chart of the first embodiment of the video compression method of the present application. Figure 2 This is a flow chart of an embodiment of the video compression method of the present application. Optionally, the video compression method of this embodiment can be applied to an electronic device. The electronic device can be an encoding device, which is not particularly limited here. The video compression method of this embodiment can include the following steps.

[0036] S101: Perform motion estimation based on a current frame and a reference frame to obtain first motion information.

[0037] Motion estimation can be performed based on the current frame and the reference frame to obtain first motion information, so that the first error information between the current frame and the preliminary predicted frame of the current frame determined by the first motion information can be subsequently used to enhance the first motion information to obtain second motion information, and then the current frame can be predicted and compressed based on the second motion information. In this way, the first error information between the current frame and the preliminary predicted frame of the current frame determined by the first motion information is used to enhance and optimize the first motion information, which can improve the accuracy of the motion information, and then improve the accuracy of the current frame prediction, thereby improving the video compression effect.

[0038] The current frame and the reference frame may be processed by a motion estimation module to obtain inter-frame motion information, that is, first motion information.

[0039] Optionally, the motion estimation module may be an optical flow estimation module, and the optical flow estimation module may perform motion estimation processing on the current frame and the reference frame to obtain optical flow information between frames, that is, obtain the first motion information.

[0040] S102: Perform motion compensation based on the first motion information and the reference frame to obtain a preliminary prediction frame of the current frame.

[0041] After performing motion estimation based on the current frame and the reference frame to obtain first motion information, motion compensation can be performed based on the first motion information and the reference frame to obtain a preliminary prediction frame of the current frame.

[0042] Optionally, a motion compensation network may be used to process the reference frame based on the first motion information to obtain a preliminary prediction frame for the current frame. The motion compensation network may be used to perform a compensation operation such as warping or deformable convolution on the reference frame based on the first motion information to extract values ​​from the reference frame to obtain the preliminary prediction frame for the current frame. Furthermore, the compensation operation may be a warping operation based on bilinear interpolation.

[0043] S103: Calculate first error information between the preliminary prediction frame and the current frame.

[0044] After the preliminary prediction frame of the current frame is obtained through motion compensation, first error information between the preliminary prediction frame and the current frame may be calculated so as to subsequently enhance the first motion information using the first error information.

[0045] In one embodiment, the preliminary prediction frame and the current frame may be directly subtracted to obtain the first error information.

[0046] In another embodiment, the preliminary prediction frame and the current frame are processed by an error calculation network to obtain first error information.

[0047] Among them, the preliminary prediction frame and the current frame can be spliced ​​first to obtain splicing features; then the splicing features can be input into the error calculation network to use the error calculation network to process the preliminary prediction frame and the current frame in the splicing features to obtain first error information.

[0048] Alternatively, the first motion information, the preliminary prediction frame, and the current frame may be processed by an error calculation network to obtain first error information. Thus, the first motion information, the preliminary prediction frame, and the current frame may be concatenated to obtain a concatenated feature. The error calculation network may then be used to process the first motion information, the preliminary prediction frame, and the current frame in the concatenated feature to obtain first error information.

[0049] Furthermore, in order to improve the processing effect of the error calculation network, the variables generated by the reference frame decoding process (i.e., the first cache information) can also be used as auxiliary information for the error calculation network calculation processing, that is, the first cache information, the first motion information, the preliminary prediction frame and the current frame can be used as joint inputs and input into the error calculation network to allow the error calculation network to learn the error information of the preliminary prediction frame and the current frame, thereby obtaining the first error information.

[0050] The first cache information may include at least one of a motion context (motion information of a reference frame), a reference frame reconstruction feature, a reference frame reconstructed image, and the like.

[0051] Optionally, the first error information output by the error calculation network may be information of one scale or information of multiple scales, which is not limited here.

[0052] The structure of the error calculation network is not limited as long as it has the function of determining the error information between the preliminary prediction frame and the current frame. For example, it may include but is not limited to a Unet network, a convolutional network with downsampling function, and a commonly used convolutional network.

[0053] S104: Enhance the first operation information based on the first error information to obtain second motion information.

[0054] After calculating the first error information between the preliminary prediction frame and the current frame, the first error information can be used to enhance the first motion information to obtain second motion information, which can then be used to perform operations such as prediction and compression on the current frame. This allows the first error information to be used as feedback and optimize the first motion information, making the motion information of the current frame more accurate. This also allows for the accuracy of the prediction result obtained by motion compensation at the decoder using this motion information.

[0055] Optionally, an addition method, a cascade splicing method, or a discriminative addition method may be used to enhance the first motion information based on the first error information to obtain the second motion information.

[0056] For the addition method, in one embodiment, as Figure 3 As shown, the first error information and the first motion information can be directly added to obtain the second motion information.

[0057] For the addition method, in another embodiment, as Figure 4 As shown, the first error information can be range-adjusted to obtain the second error information; then the second error information and the first motion information are added to obtain the second motion information, so as to limit the value range of the second error information through range adjustment, thereby preventing the error information from destroying the main information in the motion information.

[0058] Optionally, assuming that the range adjustment is an F function, the second error information x1=F(x), where x is the first error information.

[0059] The aforementioned F function includes but is not limited to tanh, sigmoid activation function, a combination of activation function and range factor (customized maximum range value), etc.

[0060] For example, Figure 5 As shown, the range adjustment function F is the Tanh process multiplied by the range factor. For example, when the range factor is set to 10, Figure 5The range adjustment function shown limits the second error information value to the range of [-10, 10].

[0061] For the discriminative additive method, such as Figure 6 As shown, the first error information x can be divided to obtain a scaling factor x2 and residual information x1. The scaling product is added to the residual information x1 to obtain the second motion information m1. The scaling product is the product of the scaling factor x2 and the first motion information m. Thus, the second motion information m1 = x2*m+x1. The scaling factor x2 is used to control the information content of the first motion information m. In other words, the multiplication operation discriminately selects important information in the first motion information for enhancement.

[0062] like Figure 6 As shown, the first error information can be divided into channels to obtain first sub-information x12 and second sub-information x11; optionally, the first sub-information x12 is the scaling factor x2, or the first sub-information x12 is range-adjusted to obtain the scaling factor x2; optionally, the second sub-information x11 is the residual information x1, or the second sub-information x11 is range-adjusted to obtain the residual information x1; that is, for the discriminative addition method, the range adjustment of the scaling factor branch and the residual information branch are both optional, that is, the range functions F11 and F12 can be used simultaneously, or not used at the same time, or only one of them can be used.

[0063] As described above, the range adjustment function for the first sub-information x12 is F12, and the range adjustment function for the second sub-information x11 is F11, so the residual information x1=F11(x11), and the scaling factor x2=F12(x12).

[0064] The aforementioned F11 function and F12 function include but are not limited to tanh, sigmoid activation functions, and combinations of activation functions and range factors (customized maximum range values).

[0065] When using the F11 function and the F12 function at the same time, the F11 function and the F12 function can be the same or different. Figure 7 As shown, the range adjustment function F11 is tanh processing multiplied by the range factor. For example, when the range factor is set to 10, the numerical range of the residual information x1 is [-10, 10]; F12 selects the sigmoid function.

[0066] For the cascade splicing method, the first error information and the first motion information may be directly spliced ​​to obtain the second motion information.

[0067] Optionally, in some scenarios, the first motion information may be subjected to information adjustment and / or dimension adjustment to obtain intermediate motion information. This intermediate motion information is then enhanced based on the first error information using an addition method, a cascaded concatenation method, or a discriminative addition method. The specific steps of the addition method, the cascaded concatenation method, and the discriminative addition method can be found in the above description.

[0068] Exemplarily, for the addition method, the intermediate motion information and the first error information may be added to obtain the second motion information; or the intermediate motion information and the second error information may be added to obtain the second motion information.

[0069] Illustratively, for the cascade splicing method, the intermediate motion information and the first error information may be spliced ​​together to obtain the second motion information.

[0070] For example, for the discriminative addition method, the intermediate motion information and the scaling factor may be multiplied and then the residual information may be added to obtain the second motion information.

[0071] It is understandable that the dimension of the intermediate motion information is consistent with the dimension of the first error information. When the adjustment network is used to adjust the first motion information to obtain the intermediate motion information, the sampling multiple of the adjustment network is the same as the sampling multiple of the error calculation network, that is, the adjustment network and the error calculation network are both dimensionality reduction, dimensionality increase, or dimensionality maintenance. Corresponding to the error calculation network, the intermediate motion information output by the adjustment network can be information of one scale or information of multiple scales. In addition, the intermediate motion information output by the adjustment network and the first error information output by the error calculation network are of the same scale, that is, the first error information and the intermediate motion information of the same scale are used to enhance the motion information.

[0072] There is no restriction on adjusting the structure of the network, which can include but not limited to Unet networks, convolutional networks with downsampling functions, common convolutional networks and interpolation, etc.

[0073] The structure of the adjustment network can be the same as or different from that of the error calculation network, and there is no limitation here. For example, the error calculation network can have a strided convolution (downsampling by a factor of 2) + 1 residual block structure, and the adjustment network can have a bilinear downsampling by a factor of 2 (the factor remains the same as that of the error calculation network) and then divided by 2 (division by 2 means reducing the numerical range).

[0074] S105: Predict and compress the current frame based on the second motion information.

[0075] After the motion information of the current frame is enhanced using the above steps, the enhanced second motion information can be used to predict and compress the current frame.

[0076] Optionally, motion compensation can be performed based on the second motion information and in combination with the reference frame to obtain an enhanced prediction frame of the current frame; and the enhanced prediction frame of the current frame can be used to perform operations such as compression of the current frame. It can be understood that the use of the enhanced prediction frame of the current frame to compress the current frame here can refer to: using the current frame and the enhanced prediction frame to calculate context information, and then quantizing and arithmetic coding the context information to obtain the context code stream of the current frame. Among them, the use of the current frame and the enhanced prediction frame to calculate context information can refer to: using the current frame and the enhanced prediction frame to calculate residual information. Of course, other calculation methods can also be used.

[0077] In another implementation scenario, the second motion information can be input into a motion vector encoding network, which performs dimensionality reduction processing on the second motion information to obtain a first motion feature. A motion vector entropy model performs quantization and arithmetic coding on the first motion feature to obtain a second motion feature. A motion vector decoding network performs dimensionality increase processing on the second motion feature to obtain third motion information. The current frame is predicted and compressed based on the third motion information. That is, motion compensation is performed based on the third motion information and in combination with a reference frame to obtain an enhanced prediction frame of the current frame. The enhanced prediction frame of the current frame is used to perform operations such as compression of the current frame.

[0078] Among them, the motion vector encoding network can reduce the dimension of the second motion information and extract a compact feature representation to obtain the first motion feature; the motion vector entropy model can quantize and perform arithmetic coding on the first motion feature to obtain the probability of each character appearing in the quantized first motion feature, and perform arithmetic coding to output the second motion feature; the motion vector decoding network can increase the dimension of the second motion feature and reconstruct it to obtain the third motion information.

[0079] The second motion information may be subjected to dimensionality reduction processing based on the variables generated in the reference frame decoding process (ie, the second cache information) to obtain the first motion feature.

[0080] The second cache information may include at least one of a motion context, a reference frame reconstruction feature, and a reference frame reconstructed image.

[0081] The second cache information may be the same as or different from the first cache information.

[0082] like Figure 8 As shown, motion compensation is performed in conjunction with the reference frame (i.e. Figure 8After obtaining the enhanced prediction frame of the current frame by performing the alignment in the image (alignment in the image), the enhanced prediction frame can be further enhanced to obtain a fused prediction frame; and then the fused prediction frame can be used to perform operations such as compression of the current frame. Here, using the fused prediction frame of the current frame to compress the current frame can refer to: using the current frame and the fused prediction frame to calculate the context information, and then quantizing and arithmetic coding the context information to obtain the context code stream of the current frame. Among them, using the current frame and the fused prediction frame to calculate the context information can refer to: using the current frame and the fused prediction frame to calculate the residual information. Of course, other calculation methods can also be used.

[0083] Optionally, the step of performing enhancement processing on the enhanced prediction frame may be: fusing the reference frame and the enhanced prediction frame to fully consider neighborhood information in the reference frame, improve prediction accuracy, and obtain a fused prediction frame.

[0084] Among them, the fusion method of the reference frame and the enhanced prediction frame may include but is not limited to using at least one network of a residual block network based on 2D convolution, an attention network, and a convolution network based on 3D convolution to fuse the reference frame and the enhanced prediction frame.

[0085] It is more preferred to use a 3D convolutional network to perform time domain fusion of the reference frame and the enhanced prediction frame, so as to make full use of the reference frame information by using the 3D convolutional network that can process time domain information to improve the accuracy of the motion compensated prediction frame, thereby obtaining a more accurate prediction frame.

[0086] The motion compensation based on the second motion information or the third motion information in combination with the reference frame may be: performing a warp or deformable convolution operation on the reference frame using the second motion information or the third motion information to obtain an enhanced prediction frame of the current frame.

[0087] Optionally, operations such as compressing the current frame using the enhanced prediction frame or the fused prediction frame of the current frame may include: processing the enhanced prediction frame or the fused prediction frame of the current frame using the context network to obtain preliminary reconstruction information of the current frame; and processing the preliminary reconstruction information using the frame reconstruction network to obtain a reconstructed image and reconstruction features of the current frame.

[0088] The context network may include a context coding network, a context entropy model, and a context decoding network. The context coding network is used to reduce the dimension and compress the context information (which may be referred to as the first context information) between the current frame and the predicted frame (i.e., the enhanced predicted frame or the fused predicted frame) to obtain the first context feature. The context information may include, but is not limited to, the difference information or cascade information between the current frame and the predicted frame. The purpose of the context coding network is to remove the temporal domain correlation between frames and the spatial domain correlation within frames.

[0089] The context entropy model is used to obtain the probability of occurrence of each character in the quantized first context feature, perform arithmetic coding, and output a context information code stream; the context information code stream is then decoded to obtain the second context feature.

[0090] The context decoding network is used to upgrade the low-dimensional context features (i.e., the second context features) obtained by decoding the above-mentioned code stream and reconstruct the second context information, and combine the second context information with the predicted frame (i.e., the enhanced predicted frame or the fused predicted frame) to obtain the preliminary reconstruction information of the current frame. For example, if the first context information of the context coding network is the difference between the current frame and the predicted frame, the context decoding network will add the predicted frame and the second context information to obtain the preliminary reconstruction information of the current frame. For another example, if the first context information of the context coding network is the cascade information (i.e., splicing information) between the current frame and the predicted frame, the context decoding network will fuse the predicted frame and the second context information to obtain the preliminary reconstruction information of the current frame.

[0091] In related technologies, context coding networks and context decoding networks typically perform dimensionality reduction and dimensionality increase on single-scale context information to obtain corresponding features. This application introduces a multi-scale structure to fully remove temporal and spatial domain correlations, thereby further context-compressing the bitstream. Optionally, both the multi-scale context coding network and the context decoding network can include parallel main and auxiliary branches, each containing information at multiple scales. Information at the same scale in the two branches is fused to enhance the features of the main branch.

[0092] like Figure 9As shown, the multi-scale context coding network can include a main branch and an auxiliary branch. The main branch includes a first context module and at least one first downsampling module connected to the first context module. The first context module is used to calculate context information for the current frame and the predicted frame, and the at least one first downsampling module is used to downsample the context information output by the first context module. The auxiliary branch includes at least one second downsampling module, which is used to downsample the current frame and the predicted frame to obtain larger-scale current frame and predicted frame respectively. The auxiliary branch may also include a second context module connected to each second downsampling module. The first downsampling module corresponding to the second context module is connected to a fusion layer, and the second context module is connected to the fusion layer provided after the corresponding first downsampling module to fuse the output results of the corresponding second context module and the first downsampling module through the fusion layer. Optionally, the output results of the corresponding second context module and the first downsampling module can be fused by splicing, splicing plus convolution network, addition, or weighted averaging. It is understood that the first downsampling module corresponding to the second context module is a first downsampling module whose output scale is the same as the output scale of the second context module.

[0093] This multi-scale context encoding network contains context information at multiple scales (i, where i = 2, 3, ... N). Each i-scale is sequentially downsampled to the i+1 scale, which has a smaller spatial dimension. The structure of each scale is similar. We use scales i to i+1 as an example (other scales are similar). Figure 9 The dotted box represents the main branch of the context module, and the remaining N scales (from top to bottom on the left) are multi-scale branches (i.e., auxiliary branches) that compensate for the main branch. The multi-scale branches compensate for the same-scale context of the main branch through fusion.

[0094] like Figure 9 As shown, the specific processing flow of the context coding network is as follows:

[0095] 1. The first context module obtains level i context information of the level i current frame and the level i predicted frame. Methods for obtaining this context information include, but are not limited to: the difference between the current frame and the predicted frame, the concatenation of the current frame and the predicted frame, or the concatenation of the difference and the concatenation result.

[0096] 2. First downsampling module: i-level context information is input into the first downsampling module to obtain low-scale main branch context. The first downsampling module includes but is not limited to strided convolution downsampling, bicubic interpolation, and convolutional networks including downsampling operations, and its downsampling multiple can be a multiple of 2;

[0097] 3. The second downsampling module downsamples the current frame at level i and the predicted frame at level i (the sampling multiple is the same as in step 2) to obtain the current frame at level i+1 and the predicted frame at level i+1;

[0098] 4. Obtain i+1 level context: Input the i+1 level current frame and the i+1 level predicted frame from step 3 into the second context module similar to step 1 to obtain the i+1 level scale context.

[0099] 5. Fusion: The low-dimensional main branch context from step 2 and the i+1 scale context from step 4 are fused to obtain the final context at level i+1. Fusion methods include but are not limited to concatenation, concatenation plus convolutional networks, etc.

[0100] As shown above, both the first context module and the second context module are referred to as context modules. The context modules are used to obtain context information for the current frame and the predicted frame at their corresponding scales. The context information can be obtained by, but is not limited to, calculating the difference between the current frame and the predicted frame at the corresponding scale; or by concatenating the current frame and the predicted frame at the corresponding scale; or by concatenating the difference between the current frame and the predicted frame at the corresponding scale and the concatenated structure of the current frame and the predicted frame at the corresponding scale.

[0101] The processing methods of the first context module and the second context module can be the same or different, and this is not limited here. For example, the processing method of the first context module is difference calculation, and the processing method of the second context module is splicing. For another example, the processing methods of the first context module and the second context module are both difference calculation. In addition, if the context coding network includes multiple second context modules, the processing methods of two second context modules can be the same or different, and this is not limited here. For example, the processing method of one second context module is difference calculation, and the processing method of another second context module is splicing.

[0102] The first downsampling module and the second downsampling module may include, but are not limited to, strided convolution downsampling, bicubic interpolation, and a convolutional network including downsampling operations. Optionally, the downsampling multiples of the first downsampling module and the second downsampling module may be multiples of 2.

[0103] Corresponding to the context encoding network, the context decoding network can also adopt a multi-scale network structure. The context decoding network mainly consists of multiple upsampling modules. Each upsampling will obtain features with larger spatial dimensions.

[0104] like Figure 10As shown, the context decoding network with a multi-scale structure can include a main branch and an auxiliary branch. The main branch in the context decoding network includes a first fusion layer and at least one first upsampling module. The first fusion layer is used to fuse the input context features of the context decoding network and the predicted frames of the same scale to obtain initial fused features. Optionally, the input context features and the predicted frames of the same scale can be fused by splicing, splicing plus convolutional network, addition, etc. The at least one first upsampling module is used to upsample the initial fused features. The auxiliary branch includes at least one second upsampling module. The second upsampling module is used to upsample the predicted frames of the corresponding scale to obtain a predicted frame of a smaller scale. The second upsampling module is connected to the first upsampling module corresponding to the second upsampling module, and the second upsampling module is connected to the second fusion layer set after the corresponding first upsampling module to fuse the output results of the corresponding second upsampling module and the first upsampling module through the second fusion layer. Optionally, the output results of the corresponding second upsampling module and the first upsampling module can be fused by splicing, splicing plus convolutional network, addition, etc. Both the first fusion layer and the second fusion layer can be referred to as fusion layers. The fusion method of the fusion layer in the context decoding network can be determined based on the processing method of the context modules of the corresponding scale in the context encoding network. For example, the processing method of the first context module is difference calculation, and the fusion method of the last second fusion layer is addition. It is understood that the output results of the second upsampling module and its corresponding first upsampling module have the same scale.

[0105] like Figure 10 As shown in Figure 1, the context decoding network consists of a main branch and a multi-scale branch, which contains N similar multi-scale feature processing structures. Taking the i+2 level scale to the i+1 level scale as an example (i=0,1,2…N), the specific process is introduced:

[0106] Main branch (context branch): The i+2 level decoded features are passed through the first upsampling module to obtain the i+1 level context. Upsampling methods include but are not limited to deconvolution, interpolation, and a combination of deconvolution / interpolation + convolution networks.

[0107] Auxiliary branch: The i+2 level prediction frame is passed through the second upsampling module to obtain the i+1 level prediction frame. Upsampling methods include but are not limited to deconvolution, interpolation, and a combination of deconvolution / interpolation + convolution networks.

[0108] Fusion: The i+1 scale context and the i+1 scale prediction frame are fused. Fusion methods include but are not limited to addition (when the main branch context on the encoder side is difference (subtraction), the decoder needs to add it back), splicing, splicing plus convolutional network, etc., and output the i+1 scale fusion result, that is, the i+1 level decoding feature;

[0109] Repeat the above steps until a preliminary reconstructed frame of the desired dimension is obtained.

[0110] The first upsampling module and the second upsampling module may include but are not limited to deconvolution, interpolation, a combination of deconvolution / interpolation+convolution networks, etc. Optionally, the upsampling multiples of the first upsampling module and the second upsampling module may be multiples of 2.

[0111] The scales of the context coding network and the context decoding network may be the same, for example, both the context coding network and the context decoding network may adopt a three-scale structure. In other embodiments, the scales of the context coding network and the context decoding network may be different.

[0112] For encoding devices, decoding network solutions of different scales can be selected based on the device's computing power. Furthermore, a syntactic element P (P = 0, 1, ... M) can be stored in the bitstream to control the selection of decoding networks of different scales, where M represents M multi-scale decoding networks. That is, the encoding device evaluates the device's computing power and transmits a P value, which can be multiple, representing various computing powers. Accordingly, the decoding device can select the decoding network solution corresponding to the transmitted P value. If P has multiple values, the device can select an appropriate multi-scale decoding network solution based on the computing power of the device. If the computing power is low, a single small-scale structure can be selected.

[0113] In a specific example, the context network includes a multi-scale context encoding network and a multi-scale context decoding network. The multi-scale context encoding network has the following configuration: a four-scale structure, downsampling by 2 times each time, i.e., the downsampling multiple of each first downsampling module and each second downsampling module is 2 times, and each first downsampling module and each second downsampling module uses a combination of strided convolution and convolutional network for downsampling; the first context module and the second context module both use a difference calculation method, i.e., output the difference between the current frame and the predicted frame of the corresponding scale; the fusion layer fuses the output results of the second context module and the first downsampling module of the corresponding scale through a fusion method combining splicing and convolutional network, i.e., the output results of the second context module and the first downsampling module of the corresponding scale are spliced ​​and then input into the convolutional network to obtain a fusion result. The multi-scale context decoding network matches the context encoding network. The structure of the multi-scale context decoding network is configured as follows: a four-scale structure, upsampling by 2 times each time, that is, the upsampling multiple of each first upsampling module and each second upsampling module is 2 times, and each first upsampling module and each second upsampling module uses a combination of deconvolution + convolution network with upsampling function for upsampling; the first fusion layer and the second fusion layer all fuse the corresponding prediction frames and context information by adding, that is, adding the corresponding prediction frames and context information to obtain the fusion result. The setting scheme P of the syntactic elements of the context decoding network can have M = 4 modes: P = 0, indicating scheme 1 of the decoding network, the context decoding network only contains one scale structure, for example, only contains Figure 10 N-level scale; P = 1, indicating the decoding network scheme 2: that is, the context decoding network contains 2 scale structures, such as Figure 10 N levels and i+2 scales; P = 2, indicating the decoding network scheme 3, that is, the context decoding network contains 3 scale structures, such as Figure 10 N levels, i+2 and i+1 scales; P = 3, indicating the scheme 4 of the decoding network, that is, the context decoding network contains all scale structures, that is, including Figure 10 The encoder evaluates the computing power of the device and inserts the appropriate syntactic element P value into the bitstream. For example, if the computing power is low, P = 0. The decoder selects the context decoding network solution 1 with P = 0 based on the decoded P value.

[0114] In another specific example, the context network includes a multi-scale context encoding network and a multi-scale context decoding network. The multi-scale context encoding network has the following configuration: a four-scale structure, downsampling by 2 times each time, i.e., the downsampling multiple of each first downsampling module and each second downsampling module is 2 times, and each first downsampling module and each second downsampling module uses a combination of strided convolution and convolutional network for downsampling; the first context module extracts context information by splicing, and all second context modules extract context information of corresponding scales by difference calculation; the fusion layer fuses the output results of the second context module and the first downsampling module of corresponding scales by splicing, i.e., the output results of the second context module and the first downsampling module of corresponding scales are spliced ​​to obtain a fusion result. The multi-scale context decoding network matches the context encoding network. The structure of the multi-scale context decoding network is configured as follows: a four-scale structure, with upsampling of 2 times each time, that is, the upsampling multiple of each first upsampling module and each second upsampling module is 2 times, and each first upsampling module and each second upsampling module uses a combination of deconvolution + convolution networks with upsampling functions for upsampling; the last second fusion layer fuses the corresponding predicted frame and context information by splicing, and the first fusion layer and the remaining second fusion layers fuse the corresponding predicted frames and context information by adding.

[0115] The multi-scale scheme in the above-mentioned context encoding network and decoding network can also be extended to the motion vector encoding network and motion vector decoding network.

[0116] This application also provides a video decompression method, such as Figure 11 As shown, the video decompression method includes the following steps.

[0117] S201: Decode the motion information code stream of the current frame to obtain third motion information.

[0118] The motion information code stream is obtained by compressing the second motion information, the second motion information is obtained by enhancing the first motion information based on the first error information, the first motion information is obtained by motion estimation based on the current frame and the reference frame, the first error information is the error information between the preliminary prediction frame and the current frame, and the preliminary prediction frame is obtained by motion compensation based on the first motion information and the reference frame.

[0119] The compression processing of the second motion information may include: performing dimensionality reduction processing on the second motion information to obtain the first motion feature; and performing quantization and arithmetic coding on the first motion feature to obtain a motion information code stream.

[0120] S202: Decode the context code stream of the current frame using the third motion information to obtain a reconstructed image of the current frame.

[0121] Decoding the context code stream of the current frame using the third motion information may include: decoding the context code stream of the current frame, performing dimensionality upscaling processing, and fusing the context code stream with the third motion information to obtain a reconstructed image of the current frame.

[0122] Optionally, an end-to-end video compression network may be used to compress and decompress the video.

[0123] like Figure 12 As shown, the end-to-end video compression network includes a motion estimation module, a motion vector encoding network, a motion vector entropy model, a motion vector decoding network, a motion compensation and time domain prediction network, a context network and / or a frame reconstruction module.

[0124] The video compression process using an end-to-end video compression network may include the following parts:

[0125] A. Motion information compression and reconstruction part:

[0126] (1) The motion estimation module obtains inter-frame motion information, namely, first motion information, based on the current frame and the reference frame.

[0127] (2) The motion vector coding network reduces the dimension of the first motion information and extracts compact information to obtain the motion information to be encoded (i.e., the first motion feature). At the front end, middle module, or tail end of the motion vector coding network, the first motion information can be enhanced and then input into the remaining coding dimension reduction networks. The above steps S102, S103, and S104 can be used to enhance the first motion information to obtain the second motion information, and then the second motion information is input into the remaining structures in the motion vector coding network.

[0128] (3) The motion vector entropy model obtains the probability of each character appearing in the quantized motion information to be encoded (i.e., the first motion feature), performs arithmetic encoding, and outputs a motion information code stream; the motion vector entropy model decodes the motion information code stream to obtain the second motion feature.

[0129] (4) The motion vector decoding network upgrades the decoded low-dimensional motion features (i.e., the second motion features) and reconstructs the third motion information.

[0130] B. Motion compensation and time domain prediction part:

[0131] The motion compensation and time domain prediction network performs motion compensation operations such as warping on the reference frame features through the third motion information to extract relevant information from the reference frame to obtain an enhanced prediction frame or a fused prediction frame of the current frame.

[0132] C. Context compression and reconstruction part:

[0133] The context coding network reduces and compresses the first context information between the current frame and the predicted frame (i.e., the enhanced predicted frame or fused predicted frame) to obtain first context features. This context information includes, but is not limited to, the difference or concatenation of the two frames. The goal is to use this first context information to remove temporal and spatial correlations between frames.

[0134] The context entropy model obtains the probability of each character appearing in the quantized context information to be encoded (i.e., the first context feature), performs arithmetic encoding, and outputs a context code stream; the context code stream is decoded to obtain the second context feature.

[0135] The context decoding network upscales the low-dimensional context features (i.e., second context features) obtained by decoding the context bitstream and reconstructs the second context information. During this process, the second context information is combined with the predicted frame (i.e., the enhanced predicted frame or fused predicted frame described above) to obtain preliminary reconstruction information for the current frame. For example, if the context of the context encoding network is a difference value, the context decoding network will add the predicted frame information to obtain preliminary reconstruction information for the current frame.

[0136] The frame reconstruction module further processes the preliminary reconstruction information to obtain the reconstructed image and reconstruction features of the current frame, both of which will be saved in the cache information as a reference frame for the next frame.

[0137] Among them, using the end-to-end video compression network to perform video decompression processing may include the following parts:

[0138] A. Motion information code stream reconstruction part:

[0139] (1) The motion vector entropy model decodes the motion information code stream to obtain the second motion feature.

[0140] (2) The motion vector decoding network upgrades the decoded low-dimensional motion features (i.e., the second motion features) and reconstructs the third motion information.

[0141] B. Motion compensation and time domain prediction part:

[0142] The motion compensation and time domain prediction network performs motion compensation operations such as warping on the reference frame features through the third motion information to extract relevant information from the reference frame to obtain an enhanced prediction frame or a fused prediction frame of the current frame.

[0143] C. Context information decompression and image reconstruction part:

[0144] The context entropy model decodes the context code stream to obtain a second context feature.

[0145] The context decoding network upscales the low-dimensional context features (i.e., second context features) obtained by decoding the context bitstream and reconstructs the second context information. During this process, the second context information is combined with the predicted frame (i.e., the enhanced predicted frame or fused predicted frame described above) to obtain preliminary reconstruction information for the current frame. For example, if the context of the context encoding network is a difference value, the context decoding network will add the predicted frame information to obtain preliminary reconstruction information for the current frame.

[0146] The frame reconstruction module further processes the preliminary reconstruction information to obtain the reconstructed image and reconstruction features of the current frame, both of which will be saved in the cache information as a reference frame for the next frame.

[0147] See also Figure 13 , Figure 13 The electronic device 10 includes a memory 11 and a processor 12 coupled to each other, wherein the memory 11 is used to store program instructions, and the processor 12 is used to execute the program instructions to implement the method of any of the above-mentioned embodiments.

[0148] The logical process of the above method is presented in a computer program. In terms of the computer program, if it is sold or used as an independent software product, it can be stored in a computer storage medium. Therefore, this application proposes a computer-readable storage medium. Figure 14 , Figure 14 This is a structural diagram of an embodiment of a computer-readable storage medium of the present application. In this embodiment, a computer program 21 is stored in the computer-readable storage medium 20. When the computer program 21 is executed by the processor, the steps in the above-mentioned video encoding method are implemented.

[0149] The computer-readable storage medium 20 may specifically be a medium capable of storing a computer program, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Alternatively, the computer-readable storage medium 20 may be a server storing the computer program, which may send the stored computer program to other devices for execution, or may execute the stored computer program itself. Physically, the computer-readable storage medium 20 may be a combination of multiple entities, such as multiple servers, a server plus memory, or a memory plus a mobile hard disk, among other combinations.

[0150] The above is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A video compression method, characterized in that: The method comprises: Performing motion estimation based on the current frame and the reference frame to obtain first motion information; Performing motion compensation based on the first motion information and the reference frame to obtain a preliminary prediction frame of the current frame; Calculating first error information between the preliminary predicted frame and the current frame; enhancing the first motion information based on the first error information to obtain second motion information; The current frame is predicted and compressed based on the second motion information.

2. The video compression method according to claim 1, wherein: The calculating first error information between the preliminary predicted frame and the current frame includes: Calculating the first error information using at least one of the first motion information and the first buffer information, the preliminary prediction frame, and the current frame; The first cache information includes at least one of a motion context, a reference frame reconstruction feature, and a reference frame reconstructed image.

3. The video compression method according to claim 1, wherein: The enhancing the first motion information based on the first error information to obtain second motion information includes: The first error information and the first motion information are processed by an addition method, a cascade splicing method, or a discriminative addition method to obtain the second motion information.

4. The video compression method according to claim 3, wherein: The adopting an addition method, a cascade splicing method, or a discriminative addition method to process the first error information and the first motion information to obtain the second motion information includes: Adding the first error information to the first motion information, or adding the second error information to the first motion information to obtain the second motion information, where the second error information is obtained by adjusting the range of the first error information; or The first error information is divided to obtain a scaling factor and residual information; a scaling product is added to the residual information to obtain the second motion information, where the scaling product is the product of the scaling factor and the first motion information.

5. The video compression method according to claim 4, wherein: The dividing the first error information to obtain a scaling factor and residual information includes: performing channel segmentation on the first error information to obtain first sub-information and second sub-information; The first sub-information is the scaling factor, or the scaling factor is obtained by performing range adjustment on the first sub-information; The second sub-information is the residual information, or the residual information is obtained by performing range adjustment on the second sub-information.

6. The video compression method according to claim 1, wherein: The enhancing the first motion information based on the first error information to obtain second motion information includes: performing information adjustment and / or dimension adjustment on the first motion information to obtain intermediate motion information; enhancing the intermediate motion information based on the first error information to obtain second motion information; The intermediate motion information and the first error information have the same dimension and scale.

7. The video compression method according to claim 1, wherein: The predicting and compressing the current frame based on the second motion information includes: Performing motion compensation based on the second motion information and in combination with the reference frame to obtain an enhanced prediction frame of the current frame; Fusing the reference frame and the enhanced prediction frame to obtain a fused prediction frame; compressing the current frame using the fused prediction frame; Optionally, the fusing of the reference frame and the enhanced prediction frame includes: fusing the reference frame and the enhanced prediction frame using at least one of a residual block network based on 2D convolution, an attention network, and a convolutional network based on 3D convolution.

8. The video compression method according to claim 7, wherein: The performing motion compensation based on the second motion information and in combination with the reference frame to obtain an enhanced prediction frame of the current frame includes: performing dimensionality reduction processing on the second motion information based on the second cache information to obtain a first motion feature; quantizing, arithmetic encoding, and decoding the first motion feature to obtain a second motion feature; performing dimensionality upgrading processing on the second motion feature to obtain third motion information; Performing motion compensation based on the third motion information and in combination with the reference frame to obtain an enhanced prediction frame of the current frame; The second cache information includes at least one of a motion context, a reference frame reconstruction feature, and a reference frame reconstructed image.

9. The video compression method according to claim 1, wherein: The predicting and compressing the current frame based on the second motion information includes: Based on a multi-scale context coding network, dimensionality reduction compression is performed on at least one of the enhanced prediction frame and the fused prediction frame of the current frame and context information of the current frame to obtain a first context feature; quantizing and arithmetically encoding the first context features based on a context entropy model to obtain a context code stream; The method further includes: decoding the context code stream based on a context entropy model to obtain a second context feature; performing dimension upscaling and reconstructing the second context feature based on a multi-scale context decoding network to obtain second context information; and combining at least one of the enhanced prediction frame and the fused prediction frame of the current frame with the second context information to obtain a reconstructed image of the current frame; The enhanced prediction frame is obtained by performing motion compensation based on the second motion information and in combination with the reference frame, and the fused prediction frame is obtained by fusing the reference frame and the enhanced prediction frame.

10. The video compression method according to claim 9, wherein: The main branch of the context coding network includes a first context module, at least one first downsampling module connected to the first context module, and a fusion layer connected to the first downsampling module; the auxiliary branch of the context coding network includes at least one second downsampling module and a second context module connected to each second downsampling module, the second context module is connected to the fusion layer provided after its corresponding first downsampling module, the scale of the output result of the second context module and its corresponding first downsampling module is the same, and the inputs of the first context module and the at least one second downsampling module are at least one of the enhanced prediction block and the fused prediction block of the current frame and the current frame; and / or, The main branch of the context decoding network includes a first fusion layer, at least one first upsampling module connected after the first fusion layer, and a second fusion layer connected after the first upsampling module. The auxiliary branch of the context decoding network includes at least one second upsampling module, the second upsampling module is connected to the second fusion layer provided after its corresponding first upsampling module, the second upsampling module has the same scale as the output result of its corresponding first upsampling module, the input of the first fusion layer is the second context feature, and at least one of the enhanced prediction frame and the fused prediction frame having the same scale as the second context feature, and the input of the at least one second upsampling module is at least one of the enhanced prediction frame and the fused prediction frame having the same scale as the second context feature; Optionally, the first context module and / or the second context module is processed by calculating differences; the first fusion layer and / or the second fusion layer is processed by addition.

11. The video compression method according to claim 9, characterized in that: The multi-scale structure-based context decoding network performs dimension upgrading on the second context feature and reconstructs the second context information, which includes: Determining, based on device computing power, a context decoding network solution for the current frame from context decoding network solutions of multiple scales; The multi-scale structure-based context decoding network performs dimension upgrading and reconstruction on the second context feature to obtain the second context information, including: utilizing the context decoding network solution of the current frame, performing dimension upgrading and reconstruction on the second context feature to obtain the second context information; The method further includes: setting a decoding network scheme syntax value in the code stream of the current frame to a value corresponding to the context decoding network scheme of the current frame.

12. The video compression method according to claim 1, wherein: The video compression method is a method for compressing video using an end-to-end video compression network, wherein the end-to-end video compression network includes a motion estimation module, a pre-compensation module, an error calculation network, a feedback module, a motion vector encoding network, a motion vector entropy model, a motion vector decoding network, a motion compensation and time domain prediction network, and a context network; The performing motion estimation based on the current frame and the reference frame to obtain the first motion information includes: using the motion estimation module and performing motion estimation based on the current frame and the reference frame to obtain the first motion information; The performing motion compensation based on the first motion information and the reference frame to obtain a preliminary predicted frame of the current frame includes: using the pre-compensation module and performing motion compensation based on the first motion information and the reference frame to obtain a preliminary predicted frame of the current frame; The calculating the first error information between the preliminary predicted frame and the current frame includes: calculating the first error information between the preliminary predicted frame and the current frame using the error calculation network; The enhancing the first motion information based on the first error information to obtain second motion information includes: utilizing the feedback module and enhancing the first motion information based on the first error information to obtain second motion information; The predicting and compressing the current frame based on the second motion information includes: utilizing the motion vector encoding network, the motion vector entropy model, the motion vector decoding network, the motion compensation and time domain prediction network, and the context network, and predicting and compressing the current frame based on the second motion information.

13. A video decompression method, characterized in that: The method comprises: decoding a motion information code stream of a current frame to obtain third motion information, where the motion information code stream is obtained by compressing second motion information, where the second motion information is obtained by enhancing first motion information based on first error information, where the first motion information is obtained by motion estimation based on the current frame and a reference frame, where the first error information is error information between a preliminary prediction frame and the current frame, where the preliminary prediction frame is obtained by motion compensation based on the first motion information and the reference frame; The context code stream of the current frame is decoded using the third motion information to obtain a reconstructed image of the current frame.

14. An electronic device, characterized in that: The electronic device includes a memory and a processor; the memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of the method according to any one of claims 1 to 13.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.