A motion compensated video compression method and system based on neighboring feature guidance
Patent Information
- Application Number
- CN202510841020.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-06-23
AI Technical Summary
然而,当前主流研究多局限于基于光流估计的单帧参考特征运动补偿机制,在面对非刚性形变、遮挡等非稳态运动模式时,易引发残差域信息熵激增,导致码流冗余度上升与重建帧出现运动伪影等
[0027]本发明提出的一种基于相邻特征引导的运动补偿视频压缩方法,一方面,引入了当前帧相邻的多个帧的特征值进行特征生成,不仅增加了信息来源的多样性,还能够更准确地捕捉到运动细节,从而减少了因根据单一参考帧进行特征生成引起的参考信息不足和误差累积的问题,改善复杂运动场景下的重建质量;另一方面,通过协同滤波和帧重建挖掘时空域的非局部相关性得到重建帧与当前帧之间的误差,进一步优化压缩效率与提高重建帧的视觉保真度。
Smart Images

Figure CN120547349B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of video encoding and decoding technology, and specifically relates to a motion-compensated video compression method and system based on adjacent feature guidance. Background Technology
[0002] In recent years, the increasing demand for high-resolution video has posed significant challenges to data storage. Video compression methods are considered crucial for balancing high-quality video with limited storage costs. Currently, academia and engineering have established recognized video compression standards such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC. However, these methods follow traditional hybrid video coding frameworks, relying on manually designed motion compensation and transform quantization modules. This presents challenges in achieving end-to-end optimization for large-scale video data, limiting further improvements in algorithm performance.
[0003] Video frames are typically divided into I-frames and P-frames. An I-frame (keyframe) is usually the first frame in a video sequence; it is self-contained, meaning the information within an I-frame does not depend on data from other frames and can therefore be decoded independently. A P-frame (predicted frame) is obtained by encoding the previous frame (reference frame) of the current frame. A P-frame only stores the differences compared to the reference frame (i.e., motion compensation) to achieve efficient compression. In this process, the current frame refers to the frame being processed; the reconstructed frame is the fully compressed frame regenerated by combining information from the reference frame and the encoded information.
[0004] In recent years, deep learning-based video compression methods have gained widespread attention, exploring a new direction for processing large-scale video sets. However, current mainstream research is mostly limited to single-frame reference feature motion compensation mechanisms based on optical flow estimation. When faced with non-rigid deformation, occlusion, and other non-steady motion modes, these mechanisms are prone to causing a surge in residual domain information entropy, leading to increased bitstream redundancy and motion artifacts in the reconstructed frames. Summary of the Invention
[0005] The purpose of this invention is to provide a motion-compensated video compression method and system based on adjacent feature guidance, which solves the problems of cumulative error and bitstream redundancy caused by single reference features and improves the reconstruction quality in complex motion scenes.
[0006] To achieve the above objectives, the present invention provides a motion-compensated video compression method based on adjacent feature guidance, comprising:
[0007] Step 1: Preprocess the original video sequence, analyze and extract the key frame, reference frame and current frame of the current sequence, and extract the initial optical flow motion vector from the current frame and reference frame;
[0008] Step 2: Input the initial optical flow motion vector obtained in Step 1, along with the reconstructed features of the previous three frames of the current frame, into the motion compensation network of the motion compensation module to generate the features of the current frame.
[0009] Step 3: Generate context information using the features of the current frame, concatenate the context information with the current frame, compress the concatenated data, and obtain the implicitly represented video parameters.
[0010] Step 4: Perform collaborative filtering on the video parameters and reconstruct the current frame to obtain the reconstructed frame;
[0011] Step 5: Calculate the error between the reconstructed frame and the current frame, evaluate the bit overhead of the reconstructed frame, construct a joint optimization loss function, and dynamically adjust the video compression network parameters based on the loss function. Train the video compression model under PSNR and MS-SSIM constraints to obtain an optimized video compression model.
[0012] Step 6: Perform independent data compression on the key frames in the video sequence to be compressed, and input the remaining frames (excluding the key frames) into the optimized video compression model to complete the compression of the entire video sequence.
[0013] In addition, step 1 extracts the initial optical flow motion vector, including: modeling the motion between the current frame and the reference frame through optical flow estimation, calculating the motion field between frames, and extracting the initial optical flow motion vector; using the motion vector codec of the video preprocessing module to compress and reconstruct the initial optical flow motion vector to obtain the reconstructed initial optical flow motion vector.
[0014] In addition, step 2 specifically includes: performing multi-scale convolution extraction on the reconstructed features of the first frame before the current frame to obtain local detail features at different scales; performing spatial joint optimization on the reconstructed features of the second and third frames before the current frame to extract global spatial features; using a hybrid feature fusion mechanism based on iterative optimization to fuse local detail features with global spatial features to obtain fused intermediate features; and performing optical flow warping operation on the initial optical flow motion vector and intermediate features to obtain the features of the current frame.
[0015] In addition, step 3 specifically includes: processing the features of the current frame obtained in step 2 to construct context information; performing temporal feature extraction on the context information through the temporal prior encoder of the feature context compression module to obtain temporal prior information; using a context coding architecture to perform cross-frame feature interaction between the features of the current frame and the current context information to obtain hierarchical prior information; and fusing the temporal prior information and the hierarchical prior information to obtain the implicit representation of the video parameters of the current frame.
[0016] In addition, step 4 specifically includes: performing several combination operations to generate channel features based on the implicitly represented video parameters, including convolution and residual operations; performing pooling and convolution processing on the channel features, and segmenting the channel features after convolution and pooling processing to obtain segmentation features, and obtaining encoding information based on the segmentation features; using two U-Net networks to reconstruct the current frame based on the encoding information to obtain the reconstructed frame of the current frame.
[0017] Additionally, in step 5, the loss function L all The formula is as follows:
[0018]
[0019] Where T is the frame number in the video segment, t represents the time index of the frame in the current video segment, and R t λ represents the bit overhead of encoding the entire frame, λ is the Lagrange multiplier used to control the tradeoff between rate R and distortion D, and d(·) represents the calculation of the current frame x. t and the reconstructed frame of the current frame The distortion function between the two frames, where L represents the loss per frame.
[0020] The present invention also provides a motion-compensated video compression system based on adjacent feature guidance, which performs the method described above, including: a video preprocessing module, an adjacent feature-guided motion compensation module, a feature context compression module, a collaborative filtering module, a frame reconstruction module, a model training module, and a compression module;
[0021] The video preprocessing module is used to preprocess the original video sequence, analyze and extract the keyframes, reference frames and current frame of the current sequence, and extract the initial optical flow motion vector from the current frame and the reference frame;
[0022] The adjacent feature-guided motion compensation module is used to input the initial optical flow motion vector and the reconstructed features of the previous three frames of the current frame into the motion compensation network to generate the features of the current frame;
[0023] The feature context compression module is used to generate context information using the features of the current frame, concatenate the context information with the current frame, compress the concatenated data, and obtain implicitly represented video parameters.
[0024] The collaborative filtering module and the frame reconstruction module are used to perform collaborative filtering on the video parameters and reconstruct the current frame to obtain the reconstructed frame.
[0025] The model training module is used to calculate the error between the reconstructed frame and the current frame, evaluate the bit overhead of the reconstructed frame, construct a joint optimization loss function, and dynamically adjust the video compression network parameters based on the loss function. The video compression model is trained under PSNR and MS-SSIM metric constraints to obtain an optimized video compression model.
[0026] The compression module is used to independently compress key frames in the video sequence to be compressed, and input each remaining frame into the optimized video compression model to complete the compression of the entire video sequence.
[0027] This invention proposes a motion-compensated video compression method based on adjacent feature guidance. On the one hand, it introduces feature values from multiple adjacent frames of the current frame for feature generation, which not only increases the diversity of information sources but also captures motion details more accurately. This reduces the problems of insufficient reference information and error accumulation caused by feature generation based on a single reference frame, thus improving the reconstruction quality in complex motion scenes. On the other hand, it obtains the error between the reconstructed frame and the current frame by mining nonlocal correlations in the spatiotemporal domain through collaborative filtering and frame reconstruction, further optimizing compression efficiency and improving the visual fidelity of the reconstructed frame.
[0028] This invention proposes a motion-compensated video compression system based on adjacent features, which successfully solves the problem of insufficient reference information often encountered in video compression when processing fast-moving objects. By introducing feature information from adjacent frames, it not only effectively reduces distortion and artifacts caused by dynamic changes, but also improves temporal consistency and spatial fidelity during video compression, thereby effectively improving the visual quality and compression efficiency of video compression. Attached Figure Description
[0029] Figure 1 This is a flowchart of the motion compensation video compression method based on adjacent feature guidance of the present invention;
[0030] Figure 2 A flowchart of the motion compensation network guided by adjacent features according to the present invention;
[0031] Figure 3 This is a structural diagram of the global-local hybrid enhancement module of the present invention;
[0032] Figure 4 This is a schematic diagram of the motion compensation video compression system based on adjacent feature guidance of the present invention. Detailed Implementation
[0033] The technical solution and beneficial effects of the present invention will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to explain the present invention and should not be construed as limiting the present invention.
[0034] The motion-compensated video compression method based on adjacent feature guidance proposed in this invention employs a core architecture that combines deep learning intra-frame (I-frame) coding and inter-frame (P-frame) coding. In terms of coding hierarchy design, the system uses an alternating IP frame sequence structure (IPPP periodic cyclic mode), where the keyframe spacing parameter (GOP, Group of Pictures) is defined as the frame number interval between consecutive keyframes. This parameter configuration supports adaptive adjustment based on the spatiotemporal characteristics of the video stream, such as motion complexity and texture features, thereby achieving dynamic optimization of coding efficiency.
[0035] In one embodiment of the present invention, the flow of the motion-compensated video compression method guided by adjacent features is as follows: Figure 1 As shown. In step 1, the original video sequence is preprocessed to analyze and extract the keyframes, reference frames, and current frame of the current sequence, and the initial optical flow motion vector is extracted from the current frame and the reference frame. The reference frame is the frame preceding the current frame. If the current frame is the first frame of the original video sequence, then its reference frame is the key reconstructed frame obtained by compressing and reconstructing the keyframes.
[0036] Specifically, step 1 includes the following: First, construct a sequence from all frames in the original video, denoted as X = {x1, x2, ..., x...} t-1 ,x t}, where each frame x t This represents the image data of frame t in the original video. The keyframes (I-frames) and reference frames x in the current sequence are analyzed and extracted. t-1 and current frame x t Next, the motion between the current frame and the reference frame is modeled using optical flow estimation, the inter-frame motion field is calculated, and the preliminary optical flow motion vector v is extracted. t In this invention, the compression method for I-frames includes, but is not limited to, traditional image compression methods and deep learning-based image compression methods. Optical flow estimation methods also include, but are not limited to, SpyNet optical flow network estimation. After obtaining the initial optical flow motion vector v... t Then, the motion vector codec of the video preprocessing module is used to process the initial optical flow motion vector v. t Compression reconstruction was performed to obtain the initial optical flow motion vectors at three different scales after reconstruction. (Full resolution scale) (half-resolution scale), (Quarter-scale resolution). This motion vector codec includes the following functions: First, it encodes the padded motion vectors through a hyperprior autoencoder network to obtain hyperprior motion information mv. tThen, by quantizing and decoding the motion vector encoded information and reference motion vector using the prior parameter decoder in the video preprocessing module, a prior fusion network is constructed to fuse temporal prior information and hierarchical prior information to estimate the mean μ of the hidden coding space. t and scale σ t Finally, the motion vectors are encoded into a bitstream using quadtree entropy coding, which involves multiple spatial prior adapters and spatial prior parameters. Finally, a motion vector codec is used to decode the initial motion vector predictions, outputting the final reconstructed initial optical flow motion vectors at three different scales.
[0037] In step 2, the initial optical flow motion vector obtained in step 1, along with the reconstructed features of the previous three frames of the current frame, are input into the motion compensation network to generate the features of the current frame.
[0038] Specifically, in step 2, the motion compensation network structure is as follows: Figure 2 As shown. In the local domain, the reconstructed features F from the previous frame t-1 Multi-scale hierarchical features are extracted and spatially transformed using a network composed of convolutions and residual blocks. This process generates local detail features at three different scales: That is, local detail features at full resolution scale, local detail features at half resolution scale, and local detail features at quarter resolution scale. In the global domain, the reconstructed features F of the third frame immediately preceding the nearest neighbor. t-3 Applying the same operation, the optimized long-term spatial features are obtained as follows: The initial global spatial features are given by the following formula:
[0039]
[0040] Among them, operators `conv` represents element-wise multiplication, `resBlock` represents convolution, and `resBlock` represents residual operation. Represents global spatial features, Representing long-term global spatial features, Reconstructing features F from the second-to-last frame using ResBlock operation t-2 Extracted from [source]. Through joint modeling. and By fusing the two approaches, supervising and independent optimization are applied at each layer, and feature fusion and updating are performed through convolution and residual connections to ultimately obtain global spatial features.
[0041] Furthermore, when the current frame lacks the reconstructed features of the previous three frames, a method of copying the nearest neighbor features is employed. That is, if the reconstructed features of the previous three frames cannot be obtained at the beginning of the video sequence and input into the motion compensation network, the missing feature parameters are supplemented by copying the most recently available feature parameters. For example, if the first frame is being processed, data from the previous three frames is missing. In this case, the first frame will have its features extracted through a convolutional layer and will be copied three times to construct a virtual three-frame feature set. If the second frame is being processed and the reconstructed features of the first frame are the most recently available reconstructed features, then the feature parameters of the first frame will be copied twice and combined with itself to form a three-frame feature set. When processing the third frame, the most recently available reconstructed features of the second frame will be copied once, forming a three-frame feature set together with the reconstructed features of the first and second frames.
[0042] Specifically, a global-local hybrid enhancement module is used to achieve the collaborative interaction and efficient fusion of global spatial context and local detail features to obtain intermediate features. Specifically, the global-local hybrid enhancement module is further optimized in each iteration based on the results of the previous iteration, as detailed in the following process: Figure 3 As shown, the formula is as follows:
[0043]
[0044] Among them, operators This represents addition, where the time step exponent n represents the number of recursions, f represents the recursive function, and ReLU represents the activation function. Represents global spatial features, Representing local detail information, i = 1, 2, 3. Setting n = 3 to recursively perform 3 iterations corresponds to the three reconstructed features used in the motion compensation network framework. In the first recursive time step (n = 1), only the features at the current time step are processed. Perform convolution operations.
[0045] Then, an optical flow warping operation is performed on the initial optical flow motion vector and intermediate features. The optical flow warping operation is a transformation of spatial alignment and remapping to obtain the features of the current frame. In this embodiment, It is by It is obtained by gradually using bilinear interpolation downsampling, which can maintain fine spatial accuracy. This coarse-to-fine alignment strategy can effectively make up for the lack of global information.
[0046] In step 3, context information is generated using the features of the current frame, and the context information is concatenated with the current frame. The concatenated data is then compressed to obtain implicitly represented video parameters.
[0047] Specifically, step 3 includes, from input frame xt and the features of the current frame generated in step 2 Extracting multi-level contextual features. This process first involves the highest-level global features. Upsampling is performed using a convolutional block, two residual blocks, and one convolutional block. Then, the processed data is... and The data is concatenated and further processed using two convolutional blocks and two residual blocks. Finally, the processed data is... and The concatenation is repeated and a convolution operation is performed to obtain the final context information.
[0048] During the encoding and decoding process, the context information is converted into temporal prior information by the temporal prior encoder in the feature context compression module. Then, quantization and arithmetic coding are performed to determine the super-prior probability distribution information. The temporal prior encoder is composed of one convolutional block, one GDN analysis block, one convolutional block, one GDN analysis block, one convolutional block, one GDN analysis block, one convolutional block, and one convolutional block connected in series.
[0049] Then, the temporal prior information and the super-prior probability distribution information are processed by the context encoder of the feature context compression module to generate hierarchical prior information. The context encoder is composed of one convolutional block, one GDN analysis block, one residual block, one convolutional block, one GDN analysis block, one residual block, one convolutional block, one GDN analysis block, and one convolutional block connected in series.
[0050] This method utilizes a prior fusion network to combine temporal prior information with hierarchical prior feature information to estimate video parameters of implicit representations in the hidden coding space.
[0051] In step 4, the video parameters are subjected to collaborative filtering, and the reconstruction of the current frame is completed to obtain the reconstructed frame;
[0052] Specifically, in step 4, the implicitly represented video parameters obtained in step 3 are first processed. Perform several combination operations to generate channel features Combination operations include convolution operations and residual operations, as shown in the following formulas:
[0053]
[0054] in, The subscript ×z in the text indicates the implicit representation of the video parameters. The process of "performing convolution operation followed by residual operation" is performed z times. In this embodiment, z can be 2, meaning that the combination operation is performed twice.
[0055] Furthermore, the dependencies between channels are explicitly constructed to enhance the model's sensitivity to information channels. In particular, the use of global average pooling helps the model capture global information, which is lacking in convolutional layers. Therefore, as shown in the following formula, pooling is used to segment the convolutionally enhanced features.
[0056]
[0057] Where MaxPool(·) and AvgPool(·) represent the max pooling layer and average pooling layer with a stride of 2, respectively, and the parameter σ is obtained through a simple convolutional layer. It is obtained through the ReLU activation function, where * indicates channel concatenation and Split indicates feature segmentation operation. It is a segmentation feature.
[0058] Subsequently, segmentation features Encoded information is generated through linear layers and convolutional layers. The reconstructed frames are then generated using a frame reconstruction module consisting of two U-Nets, resulting in pixel-domain reconstructed frames. Additionally, the reconstructed features from the last convolutional layer This feature is used as a reference for compressing the next frame.
[0059] In step 5, the error between the reconstructed frame and the current frame is calculated, and the bit overhead of the reconstructed frame is evaluated. A joint optimization loss function is constructed, and the video compression network parameters are dynamically adjusted based on the loss function. The video compression model is trained under the constraints of PSNR (Peak Signal-to-Noise Ratio) and MS-SSIM (Multi-Scale Structural Similarity) indicators to obtain an optimized video compression model.
[0060] Specifically, in step 5, the loss function formula is as follows:
[0061]
[0062] Here, T represents the number of frames in the video segment, which is usually set to 1. However, to address error propagation and accumulation, a cascaded loss strategy is employed in the final training phase, setting T=7 to accumulate the loss of each frame within the video segment. Furthermore, t represents the t-th frame in the video, and Rt... t λ represents the bit overhead of encoding the entire frame, λ is the Lagrange multiplier used to control the tradeoff between rate R and distortion D, and d(·) represents the calculation of the current frame x. t and reconstructed frames The distortion function between, where L represents the loss per frame, L allThe formula for calculating the entire loss function.
[0063] In step 6, key frames in the video sequence to be compressed are processed independently, and each remaining frame is input into the optimized video compression model to complete the compression of the entire video sequence.
[0064] Specifically, keyframes undergo independent data compression processing, including: using existing image coding methods, such as a hybrid coding architecture combining discrete cosine transform and entropy coding, to compress keyframes; or using a deep neural network-driven intelligent compression algorithm to compress keyframes, wherein the deep neural network is trained through an end-to-end framework constructed by a differentiable quantization module and a contextual entropy model.
[0065] Another embodiment of the present invention includes a motion-compensated video compression system guided by adjacent features, performing the method as described above, such as... Figure 4 As shown, it includes: a video preprocessing module, a motion compensation module guided by adjacent features, a feature context compression module, a collaborative filtering module, a frame reconstruction module, a model training module, and a compression module;
[0066] The video preprocessing module is used to preprocess the original video sequence, analyze and extract the keyframes, reference frames and current frame of the current sequence, and extract the initial optical flow motion vector from the current frame and the reference frame;
[0067] The adjacent feature-guided motion compensation module is used to input the initial optical flow motion vector and the reconstructed features of the previous three frames of the current frame into the motion compensation network to generate the features of the current frame;
[0068] The feature context compression module is used to generate context information using the features of the current frame, concatenate the context information with the current frame, compress the concatenated data, and obtain implicitly represented video parameters.
[0069] The collaborative filtering module and the frame reconstruction module are used to perform collaborative filtering on the video parameters and reconstruct the current frame to obtain the reconstructed frame.
[0070] The model training module is used to calculate the error between the reconstructed frame of the current frame and the current frame, evaluate the bit overhead of the reconstructed frame, construct a jointly optimized loss function, and dynamically adjust the video compression network parameters based on the loss function. The video compression model is trained under the constraints of the PSNR peak signal-to-noise ratio standard and the MS-SSIM multi-scale structural similarity index to obtain an optimized video compression model.
[0071] The compression module is used to independently compress key frames in the video sequence to be compressed, and input each remaining frame into the optimized video compression model to complete the compression of the entire video sequence.
[0072] This invention uses the Vimeo-90k dataset for training, employs gradient descent to adjust hyperparameters, and uses an optimizer for parameter updates. All video sequences are randomly cropped into 256×256 segments. In this embodiment, the batch size is set to 4. The model is first trained for 31 epochs, and then fine-tuned for an additional 3 epochs using a cascading method.
[0073] In summary, this invention proposes a motion-compensated video compression system based on adjacent feature guidance, which solves the problems of accumulated error and bitstream redundancy caused by single reference features and improves the reconstruction quality in complex motion scenes. The method includes the following steps: preprocessing the original video sequence to extract the I-frame, reference frame, and current frame of the current sequence; performing optical flow estimation on the current frame and reference frame to calculate the inter-frame motion field and extract the initial optical flow motion vector; inputting the motion vector obtained in step 1 and the reconstructed features of the previous three frames into the motion compensation network in the adjacent feature-guided motion compensation module, using a feature fusion mechanism to generate the reconstructed features of the current frame; if the cardinality of the current feature set does not meet the requirements of three-frame collaborative processing, the missing feature parameters are obtained by copying the most recently available feature parameters; performing context compression on the reconstructed features from step 2 to obtain implicitly represented video parameters; performing collaborative filtering on the video parameters from step 3 and completing the frame reconstruction; calculating the error between the reconstructed frame and the current frame, simultaneously evaluating the bit overhead of the reconstructed frame, constructing a jointly optimized loss function, and dynamically adjusting the video compression network parameters based on this loss function. The model is trained under the constraints of PSNR and MS-SSIM metrics to obtain an optimized video compression model. Each frame in the video sequence to be compressed is input into the video compression model trained and optimized in step 5 to compress the P frames. This process is iterated to complete the compression of the entire video sequence.
[0074] The above embodiments are merely illustrative of the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of this invention.
Claims
1. A motion-compensated video compression method based on adjacent feature guidance, characterized in that, Includes the following steps: Step 1: Preprocess the original video sequence, analyze and extract the key frame, reference frame and current frame of the current sequence, and extract the initial optical flow motion vector from the current frame and reference frame; Step 2: The initial optical flow motion vector obtained in Step 1, along with the reconstructed features of the previous three frames, are input into the motion compensation network of the motion compensation module to generate features for the current frame; specifically including: Multi-scale convolution is performed on the reconstructed features of the first frame before the current frame to extract local detail features at different scales; Spatial joint optimization is performed on the reconstructed features of the first second and first third frames to extract global spatial features; A hybrid feature fusion mechanism based on iterative optimization is adopted to fuse local detail features with global spatial features to obtain fused intermediate features; The initial optical flow motion vector and intermediate features are subjected to optical flow warping operation to obtain the features of the current frame; Step 3: Generate context information using the features of the current frame, concatenate the context information with the current frame, compress the concatenated data, and obtain the implicitly represented video parameters. Step 4: Perform collaborative filtering on the video parameters and reconstruct the current frame to obtain the reconstructed frame; Step 5: Calculate the error between the reconstructed frame and the current frame, evaluate the bit overhead of the reconstructed frame, construct a joint optimization loss function, and dynamically adjust the video compression network parameters based on the loss function. Train the video compression model under PSNR and MS-SSIM constraints to obtain an optimized video compression model. Step 6: Perform independent data compression on the key frames in the video sequence to be compressed, and input the remaining frames (excluding the key frames) into the optimized video compression model to complete the compression of the entire video sequence.
2. The motion-compensated video compression method based on adjacent feature guidance as described in claim 1, characterized in that, Step 1 extracts the initial optical flow motion vector, including: The motion between the current frame and the reference frame is modeled by optical flow estimation, the motion field between frames is calculated, and the preliminary optical flow motion vector is extracted. The initial optical flow motion vector is compressed and reconstructed using the motion vector codec of the video preprocessing module to obtain the reconstructed initial optical flow motion vector.
3. The motion-compensated video compression method based on adjacent feature guidance as described in claim 1, characterized in that, Step 3 specifically includes: The features of the current frame obtained in step 2 are processed to construct context information; Temporal feature extraction is performed on the context information by the temporal prior encoder of the feature context compression module to obtain temporal prior information; A context coding architecture is used to perform cross-frame feature interaction between the features of the current frame and the current context information to obtain hierarchical prior information. By fusing temporal prior information with hierarchical prior information, the implicit representation of video parameters for the current frame is obtained.
4. The motion-compensated video compression method based on adjacent feature guidance as described in claim 1, characterized in that, Step 4 specifically includes: Based on the implicitly represented video parameters, several combination operations are performed to generate channel features. The combination operations include convolution operations and residual operations. The channel features are subjected to pooling and convolution processing, and the channel features after convolution and pooling processing are segmented to obtain segmentation features. Encoding information is obtained based on the segmentation features. Two U-Net networks are used to reconstruct the current frame based on the encoded information, resulting in a reconstructed frame.
5. The motion-compensated video compression method based on adjacent feature guidance as described in claim 1, characterized in that, In step 5, the loss function L all The formula is as follows: , Where T is the number of frames in the video segment. This represents the time index of a frame in the current video segment. This represents the bit overhead of encoding the entire frame. It is the Lagrange multiplier used to control the tradeoff between rate R and distortion D. Indicates the calculation of the current frame and the reconstructed frame of the current frame Distortion function between This represents the loss per frame.
6. A motion-compensated video compression system guided by adjacent features, comprising performing the method as described in any one of claims 1-5, including: The system includes a video preprocessing module, a motion compensation module guided by adjacent features, a feature context compression module, a collaborative filtering module, a frame reconstruction module, a model training module, and a compression module. The video preprocessing module is used to preprocess the original video sequence, analyze and extract the key frame, reference frame and current frame of the current sequence, and extract the initial optical flow motion vector from the current frame and the reference frame. The adjacent feature-guided motion compensation module is used to input the initial optical flow motion vector and the reconstructed features of the previous three frames of the current frame into the motion compensation network to generate the features of the current frame. The feature context compression module is used to generate context information using the features of the current frame, concatenate the context information with the current frame, compress the concatenated data, and obtain implicitly represented video parameters. The collaborative filtering module and the frame reconstruction module are used to perform collaborative filtering on the video parameters and complete the reconstruction of the current frame to obtain the reconstructed frame. The model training module is used to calculate the error between the reconstructed frame and the current frame, evaluate the bit overhead of the reconstructed frame, construct a jointly optimized loss function, and dynamically adjust the video compression network parameters based on the loss function. The video compression model is trained under PSNR and MS-SSIM metric constraints to obtain an optimized video compression model. The compression module is used to independently compress key frames in the video sequence to be compressed, and input each remaining frame into the optimized video compression model to complete the compression of the entire video sequence.
Citation Information
Patent Citations
Video compression method based on deep learning feature space
CN113298894A
Feature space context video compression method and system based on optical flow guidance
CN116684622A