Multi-granularity generative video compression method

By decomposing the video signal into potential subspaces of low-frequency space, space-time motion and high-frequency details and encoding them separately, the problem of static background repeated encoding and insufficient fidelity of motion details is solved, and the efficient video compression effect is achieved, and the compression performance of complex motion scenes is improved.

CN120343255BActive Publication Date: 2025-08-22ZHONGKE FANGCUN ZHIWEI (NANJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510797399.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-08-22
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

When existing video compression technology deals with complex motion scenarios, it is difficult to explicitly separate static backgrounds from dynamic motion, resulting in imbalance in bit rate allocation, limited multi-scale feature extraction capabilities, and traditional methods have shortcomings in motion detail fidelity and compression efficiency.

Method used

The multi-grained generated video compression method is used to decompose and map the video signal to three potential subspaces, namely low-frequency space information, spatiotemporal motion characteristics and high-frequency detail information, and each subspace is targeted to encode and generate a compressed code stream.

Benefits of technology

It realizes efficient content adaptability, improves compression performance, reduces bit rate, improves visual quality and compression efficiency, especially in complex motion scenarios, which significantly improves peak signal-to-noise ratio and reduces storage usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343255B_ABST
    Figure CN120343255B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-granularity generative video compression method, comprising: obtaining a video frame sequence, performing spatiotemporal feature decomposition, decomposing and mapping the video signal into three latent subspaces, obtaining first, second, and third latent subspace representations, wherein the first latent subspace represents the low-frequency spatial information and temporal fading components in the video, the second latent subspace represents the spatiotemporal motion features in the video, and the third latent subspace represents the high-frequency detail information in the video; encoding each of these latent subspaces to generate corresponding first, second, and third coded data, which are combined to generate a compressed bitstream. The present invention effectively solves the problem of unbalanced bitrate distribution and achieves content-adaptive and efficient compression. It also enhances multi-scale feature extraction capabilities, avoids subband matching mismatch issues, and improves compression performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of video compression, in particular to a multi-granularity generation type video compression method. Background Art

[0002] Video compression technology is crucial in today's information age. Its primary purpose is to reduce video data volume, thereby lowering storage costs and transmission bandwidth requirements. Effective video compression technology improves video storage efficiency and accelerates video transmission speeds, and is widely used in a wide range of fields, including video streaming services, video surveillance, and high-definition television. Technological innovations that achieve more efficient video compression will drive further development in these areas and enhance user experience.

[0003] Currently, traditional methods face an inherent conflict between redundancy elimination and motion preservation due to their spatiotemporal coupled representation. Mainstream video compression standards such as HEVC / AV1 utilize block-based motion compensation and transform coding, which can reduce spatial redundancy to a certain extent. However, repeated encoding of static backgrounds wastes bits, while rigid block partitioning in fast-moving areas leads to blocking artifacts. The spatiotemporal coupled transform domain representation disrupts the continuity of motion trajectories, resulting in rate-distortion optimization margins lower than theoretical values. Furthermore, existing frameworks fail to orthogonally decompose video components, causing dynamic motion and static background to interfere with each other during quantization, resulting in loss of high-frequency detail. The multi-scale feature extraction capabilities of existing coding systems are limited by the time-frequency analysis limitations of transform methods. While the traditional DCT transform is computationally efficient, its global basis functions produce Gibbs oscillations in regions with abrupt motion changes. The fixed decomposition structure of the 3D wavelet transform leads to mismatched subband allocation. The mismatch between the parameters of the spatial directional filter and the temporal wavelet distorts the motion phase, resulting in a significant decrease in PSNR. Entropy coding efficiency is limited by the inadequate modeling of spatiotemporal nonstationarity in traditional models. CABAC's local correlation model, based on the Markov assumption, struggles to capture long-term motion dependencies, and the motion vector entropy coding in HEVC suffers from high redundancy. While deep learning models improve spatial correlation modeling, their autoregressive structures have limited temporal receptive fields, and their computational complexity increases cubically with resolution. While the Transformer model can model long-term dependencies, direct application to entropy coding can lead to video decoding delays. Reconstruction enhancement techniques based on generative adversarial networks face the dual challenges of motion consistency and training stability. Existing approaches decouple optical flow estimation from the compression process, resulting in a high incidence of motion artifacts. Furthermore, mode collapse in adversarial training reduces the SSIM in flat areas, while overfitting leads to the generation of spurious edges. The root cause lies in the lack of physical constraints on compressed domain features in GANs. For example, a mismatch between the quantization step size and the generator's receptive field can amplify frequency domain distortion.

[0004] Existing technologies have difficulty in handling complex motion scenes in terms of multi-component decoupling, and do not explicitly separate static background from dynamic motion, resulting in an imbalance in bit rate distribution. The multi-scale feature extraction capability is limited by the time-frequency analysis defects of the transformation method, such as the Gibbs oscillation generated by the traditional DCT transform and the mismatch in sub-band allocation of the three-dimensional wavelet transform. Summary of the Invention

[0005] The purpose of the invention is to provide a multi-granularity generative video compression method, in order to solve at least one technical problem existing in the prior art.

[0006] The technical solution, a multi-granularity generative video compression method, includes:

[0007] Obtain a video frame sequence, perform spatiotemporal feature decomposition, and decompose the video signal into three latent subspaces to obtain the first, second, and third latent subspace representations. The first latent subspace represents the low-frequency spatial information and temporal fading components in the video, the second latent subspace represents the spatiotemporal motion features in the video, and the third latent subspace represents the high-frequency detail information in the video.

[0008] The first, second and third latent subspace representations are respectively encoded to generate corresponding first, second and third encoded data, which are combined to generate a compressed code stream.

[0009] Beneficial effects: The present invention explicitly separates static background and dynamic motion information, effectively solves the problem of unbalanced bit rate distribution, and realizes content-adaptive and efficient compression; enhances the multi-scale feature extraction capability, avoids the sub-band matching mismatch problem, and improves compression performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 A flowchart of the steps of a multi-granularity generative video compression method provided in an embodiment of the present application.

[0011] Figure 2 A flowchart of the steps for obtaining first, second and third latent subspace representations provided in an embodiment of the present application.

[0012] Figure 3 A flowchart of the steps for generating the first latent subspace representation provided in an embodiment of the present application.

[0013] Figure 4 A flowchart of the steps for generating the second latent subspace representation provided in an embodiment of the present application. DETAILED DESCRIPTION

[0014] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0016] Research has found that the efficiency of entropy coding is limited by traditional models' inadequate modeling of spatiotemporal nonstationarity. Furthermore, deep learning and Transformer models suffer from limited temporal receptive fields, high computational complexity, and decoding delays. In reconstruction enhancement techniques, approaches based on generative adversarial networks face the dual challenges of motion consistency and training stability, leading to motion artifacts, decreased SSIM in flat areas, and the generation of false edges. Fixed three-dimensional wavelet structures are unable to adapt to the multi-granularity characteristics of motion; traditional entropy models and generative reconstruction methods lack spatiotemporal synergy.

[0017] This paper aims to address the inherent contradiction between existing video compression technologies in efficiently removing spatiotemporal redundancy and accurately preserving motion details. It proposes a multi-granularity generative video compression method based on hierarchical latent spaces and spatiotemporal decoupling. While current mainstream video compression standards, such as HEVC and AV1, have achieved significant improvements in coding efficiency, their block-based hybrid coding frameworks often couple static background information with dynamic motion information. This results in redundant encoding of background information in videos containing large static areas (such as the sky or fixed background in surveillance scenes), reducing compression efficiency. Furthermore, for complex or nonlinear motion (such as rapid rotation, scale changes, and non-rigid deformations), the reliance on prediction residuals and simplified motion vector representations makes it difficult to accurately reconstruct details of moving objects, often resulting in motion blur or artifacts. Furthermore, existing entropy coding models, such as CABAC, while highly efficient, lack adaptability when handling the complex spatiotemporal non-stationary correlations in video signals. This is especially true for latent representations transformed by deep learning models, whose statistical properties can differ significantly from those of traditional prediction residuals, resulting in a failure to fully exploit compression potential. This is particularly true at low bitrates, where the compression efficiency of residual detail information is low. The present invention aims to improve compression performance, especially motion fidelity and visual quality at low bit rates, by introducing a spatiotemporal decoupled hierarchical latent representation, a hybrid entropy coding mechanism, and a generative reconstruction network.

[0018] like Figure 1 As shown, a multi-granularity generative video compression method is proposed, comprising the following steps:

[0019] Obtain a video frame sequence, perform spatiotemporal feature decomposition on the video frame sequence, and decompose the video signal into three latent subspaces to obtain the first, second, and third latent subspace representations. The first latent subspace represents the low-frequency spatial information and temporal fading components in the video, the second latent subspace represents the spatiotemporal motion features in the video, and the third latent subspace represents the high-frequency detail information in the video. Specifically, the static background flow, dynamic motion flow, and residual detail flow correspond to the first, second, and third latent subspaces.

[0020] The first, second and third latent subspace representations are respectively encoded to generate corresponding first, second and third encoded data; and the first, second and third encoded data are combined to generate a compressed code stream.

[0021] Specifically, an input video frame sequence is obtained, which can be video data of any resolution and frame rate. Spatiotemporal feature decomposition is performed on the video frame sequence, decomposing the video signal into three latent subspaces. Specifically, the first latent subspace represents low-frequency spatial information and temporally varying components (such as static background); the second latent subspace represents spatiotemporal motion features (such as object trajectories); and the third latent subspace represents high-frequency detail information (such as texture edges). Traditional video compression mixes all video components, resulting in repeated encoding of static background and loss of motion details. This embodiment, based on the physical properties of video content, decomposes it into three orthogonal components: the first latent subspace represents low-frequency static information (such as fixed background); the second latent subspace represents mid-frequency motion information (such as object trajectories); and the third latent subspace represents high-frequency detail information (such as texture edges). This decomposition allows for independent optimization of compression for each component. The three latent subspace representations are encoded separately, using different encoding strategies tailored to the characteristics of each subspace, to generate corresponding first, second, and third encoded data. These three encoded data are then combined to generate the final compressed bitstream.

[0022] Through the above steps, this embodiment achieves efficient and content-adaptive compression by decomposing the video into three semantically distinct latent subspaces and encoding them separately. This addresses the technical issues of repeated encoding of static backgrounds and insufficient fidelity of motion details in traditional methods, achieving efficient and content-adaptive compression. Experimental data shows that, while maintaining the same visual quality, the bit rate is approximately 45% lower than that of High-Efficiency Video Coding (HEVC), improving compression efficiency; the peak signal-to-noise ratio (PSNR) for complex motion scenes is improved by approximately 2.1dB; and storage usage is reduced.

[0023] like Figure 2 As shown, according to one aspect of the present application, obtaining the first, second and third latent subspace representations includes:

[0024] Perform motion estimation on the video frame sequence to obtain motion information;

[0025] Perform background modeling on the video frame sequence to obtain background information;

[0026] Extract spatiotemporal features from video frame sequences based on motion information and background information;

[0027] The spatiotemporal features are tensor decomposed and the tensor decomposition results are mapped into three approximately orthogonal latent subspaces, where the correlation between the latent subspaces is minimized by a preconfigured decomposition matrix to obtain the first, second and third latent subspace representations.

[0028] In one embodiment of the present application, a RAFT (Recurrent All-Pairs Field Transforms) optical flow network is used to perform motion estimation on an input video frame sequence, and a dense optical flow field between adjacent frames is calculated as motion information. The motion information (optical flow field) F m = {(u(x, y), v(x, y)) | all (x, y) ∈ Frame}, where u(x, y) is the horizontal velocity component of the pixel (x, y); v(x, y) is the vertical velocity component of the pixel (x, y), and Frame is a single frame image. Applied to a 1920×1080 resolution surveillance video, the RAFT network outputs an optical flow field F of the same size. m For example, for the pixel point (100, 200) in the static area, u(100, 200)≈0, v(100, 200)≈0; for the pixel point (500, 600) in the moving vehicle area, u(500, 600)=5.2 pixels / frame, v(500, 600)=-1.3 pixels / frame. Background modeling is performed on the video frame sequence in parallel, and stable background information F is extracted through temporal median filtering or Gaussian mixture model (GMM). b . The background modeling of the above 1920×1080 resolution surveillance video is performed using the time domain median filtering method. 32 consecutive frames are selected as the time window, and the median of each pixel position in the time dimension is calculated. For example, for the pixel point (100, 200) in the static background area, its pixel value sequence in 32 frames is [120, 118, 121, 119, 120, 122, 119, 121,...], and the stable background value F is obtained by median filtering. b (100, 200) = 120. For the pixel point (500, 600) in the road area where vehicles occasionally pass by, the pixel value sequence is [85, 87, 85, 180, 182, 178, 86, 85, ...], where 180-182 are the pixel values ​​when the vehicle passes by. After median filtering, the background value F is obtained. b (500, 600) = 86, effectively eliminating the influence of temporary occlusions. A Gaussian mixture model (GMM) was used in parallel to verify background modeling. A mixture model of three Gaussian components was established for each pixel position. Taking the pixel point (300, 400) as an example, the background component parameters obtained after training were: weight π1 = 0.85, mean μ1 = 95, variance σ1 2 =4 (main background); weight π2 = 0.12, mean μ2 = 180, variance σ2 2 =25 (shadow variation); weight π3=0.03, mean μ3=220, variance σ3 2= 100 (accidental foreground). Extract the component with the largest weight as background information and get F b (300, 400) = 95.

[0029] Based on motion information F m and background information F b , extracting spatiotemporal features from video frame sequences; specifically by fusing information such as the difference between the original frame and the background, and the marking of the motion area. Based on motion information F m and background information F b , extract spatiotemporal features from the video frame sequence. Calculate the difference map F between the original frame and the background diff = |F original - F b For the static area pixel (100, 200), the original frame value is 119, the background value is 120, and the difference value F diff (100, 200) = |119-120| = 1, indicating that the area is basically still. For the pixel point (500, 600) in the moving vehicle area, the original frame value is 182, the background value is 86, and the difference value F diff (500, 600) = |182-86| = 96, indicating significant motion. Combined with the motion amplitude ||F m (x, y)||2 calculates the motion area mark. For the pixel point (500, 600), the motion amplitude||F m (500, 600)||2= sqrt(5.2 2 + (-1.3) 2 ) = 5.36 pixels / frame, set the threshold τ = 1.0 pixels / frame, since 5.36>1.0, the point is marked as the motion area M mask (500, 600) = 1; while the motion amplitude of the pixel point (100, 200) in the static area is close to 0 and is marked as M mask (100, 200) = 0. Spatiotemporal feature matrix F st Constructed as four-channel features: F st = [F original , F b , F diff , M mask ], with a dimension of 1920×1080×4×T (T is the time length of 16 frames). The spatiotemporal features are decomposed into tensors, and the decomposition matrix is ​​designed to minimize the correlation between the three subspaces and achieve approximately orthogonal decomposition. stThe three-mode tensor decomposition (dimension 1920×1080×4×16) is performed, and the decomposition matrix is ​​designed to achieve approximate orthogonal decomposition. The feature tensor is reshaped into a two-dimensional matrix form for singular value decomposition (SVD) preprocessing, and the principal components are calculated along the spatial dimension (1920×1080 is reshaped into 2073600×1), the feature dimension (4×1), and the time dimension (16×1). The SVD decomposition of the 2073600×64 matrix (4×16=64) is performed, and the first 512 principal components are retained to obtain the spatial projection matrix U spatial (2073600×512). Feature-time decomposition: Perform SVD on the 64×2073600 matrix, retain the first 32 principal components, and obtain the feature-time projection matrix U feat_time (64×32). Orthogonalization is used to ensure that the correlation between the three subspaces is minimized. Calculate the correlation matrix: the correlation coefficient between the first subspace (static background) and the second subspace (dynamic motion) is ρ12 = 0.08; the correlation coefficient between the first subspace and the third subspace (residual details) is ρ13 = 0.12; the correlation coefficient between the second subspace and the third subspace is ρ23 = 0.15; apply the Gram-Schmidt orthogonalization process to reduce the correlation coefficient to below 0.05. The first latent subspace (static background flow): generated by spatial low-frequency filtering and temporal stability weighting. For the pixel point (100, 200), the temporal stability weight w(100, 200) = exp(-||F m (100, 200)||2 / σ) = exp(-0.1 / 2.0) = 0.95, and this point is mainly retained in the first subspace. The first subspace dimension is compressed to 240×135×8 (1 / 8 of the original). The second latent subspace (dynamic motion flow): generated based on the motion vector field decomposition. For the pixel point (500, 600) in the moving vehicle area, its motion component is mainly mapped to the second subspace. Five main motion modes are extracted through CP decomposition. The weight of the first mode is α1=0.6 (horizontal motion), the weight of the second mode is α2=0.3 (vertical motion), and the weights of the other modes are relatively small. The third latent subspace (residual detail flow): calculate the predicted frame F pred = F b +Motion Compensate (F m ), for pixel (500, 600), the predicted value is 86 + 5.2×cos(θ) + (-1.3)×sin(θ) = 91 (where θ is the motion angle). The residual between the actual value 182 and the predicted value 91 is 91. After 3D wavelet transform, high-frequency coefficients are mainly distributed in the HHH subband. The energy threshold is set to τ = 10, and coefficients with energy greater than 10 are retained.

[0030] like Figure 3As shown, according to one aspect of the present application, generating the first latent subspace representation includes:

[0031] Perform time domain filtering on background information to extract temporally stable regions;

[0032] Extracting spatial low-frequency components from a video frame sequence;

[0033] Calculate the motion amplitude of each pixel position based on the motion information, and mark the area where the motion amplitude is less than a preset threshold as a static area mask;

[0034] The temporally stable region, the spatial low-frequency component and the static region mask are weightedly fused to generate a first latent subspace representation, wherein the weight of the weighted fusion is adaptively determined according to the temporal stability of each region.

[0035] In one embodiment of the present application, the background information is filtered in the time domain to retain the temporally stable region: for each pixel position, the median of 16 consecutive frames is taken as the stable background; the spatial low-frequency component is extracted from the video frame sequence (e.g., by low-pass filtering); the motion amplitude of each pixel || F is calculated based on the motion information. m (x, y)||2= sqrt(u 2 (x, y) + v 2 (x, y)), the motion amplitude || F m The area where (x, y)||2<ε (ε=0.5 pixels / frame) is marked as a static area; the weight is adaptively determined according to the temporal stability, and the adaptive weight w(x, y) = exp(-||F m (x, y)||2 / σ), where σ is the scale parameter. The above information is weighted and fused to generate the first latent subspace representation (static background flow). Specifically, for the sky region, it remains almost unchanged for 16 consecutive frames, retaining a stable value after temporal filtering, with a motion amplitude close to 0 and a weight close to 1, completely preserved in the first subspace.

[0036] like Figure 4 As shown, according to one aspect of the present application, generating the second latent subspace representation includes:

[0037] Decompose the motion information into motion vector fields and extract the main motion components;

[0038] Perform temporal difference calculation on video frame sequences to identify motion areas;

[0039] Based on the main motion components and motion regions, mid-frequency spatiotemporal features are extracted, where motion patterns of different speeds and directions are captured through multi-scale motion analysis;

[0040] The mid-frequency spatiotemporal features are mapped to a second latent subspace to generate a second latent subspace representation, where the mapping preserves the temporal continuity of the motion trajectory.

[0041] In one embodiment of the present application, the optical flow field is decomposed into a motion vector field, and the main motion mode is extracted by principal component analysis (PCA): the complex motion is decomposed into K main modes: F m ≈ Σ k=1 K α k ·M k ; where K is the number of main motion components (typical value 5); α k is the weight coefficient of the kth motion mode; M k The kth main motion mode is represented by M1 and M2, respectively. In pedestrian scenes, these modes are primarily translation (M1) and sway (M2), with weights α1 = 0.8 and α2 = 0.2. Temporal differences between frames are calculated to identify regions of significant motion. Multi-scale analysis (e.g., pyramid structure) is used to capture motion at varying speeds. Temporal continuity of the motion trajectory is maintained to generate a second latent subspace representation (dynamic motion flow).

[0042] In one embodiment of the present application, a temporal difference operation is performed on a pedestrian walking video sequence with a resolution of 1920×1080. The pixel difference between adjacent frames is calculated: diff (x,y,t) = |F(x,y,t) - F(x,y,t-1)|. For the pixel point (800,400) in the pedestrian body area, the pixel value of the t-th frame is 145, and the pixel value of the t-1st frame is 118. The difference value T diff (800,400,t) = |145-118| = 27. Set the motion detection threshold T th = 15. Since 27>15, the pixel is marked as a motion region. For the background static region pixel (200,300), the pixel values ​​of consecutive frames are 92, 94, and 93 respectively, and the maximum difference value is 2, which is less than the threshold of 15, and is marked as a static region. Morphological operations are applied to remove noise: an opening operation of a 3×3 structural element is performed to remove isolated noise points, and a closing operation of a 5×5 structural element is performed to fill the internal holes of the motion region. The motion region mask R is obtained. mask , where the pedestrian contour area R mask = 1, background area R mask = 0. Further calculate the second-order derivative of the time difference to identify the motion boundary: T diff2 (x,y,t) = |T diff (x,y,t) - T diff(x,y,t-1)|. At the pedestrian edge pixel point (820,380), the first-order difference sequence is [5, 23, 18, 25, 8], and the second-order difference is 18, 15, 7, 17, indicating that there is a continuous motion boundary change in this area and it is marked as a key motion edge area. Based on the main motion component {M1, M2} and the motion area R mask , extracting mid-frequency spatiotemporal features. A multi-scale pyramid structure was constructed, sequentially downsampling the original 1920×1080 frames to 960×540 (first layer), 480×270 (second layer), and 240×135 (third layer). Medium-speed motion was analyzed at the first layer (960×540): for the pedestrian torso pixel (400,200) (corresponding to original coordinates 800,400), the main motion component M1 had a horizontal component u1 = 3.2 pixels / frame and a vertical component v1 = 0.8 pixels / frame, representing overall translational motion. The motion consistency index Consistency1 for this layer was calculated to be 0.85, indicating relatively stable motion. At the second layer (480×270), fast motion is analyzed: for the pixel point (200,110) in the pedestrian leg swing area, the swing amplitude of the main motion component M2 is Amplitude = 4.5 pixels / frame, and the swing frequency is f = 2.1Hz, which corresponds to the normal walking frequency. This layer captures the periodic swing pattern, and the motion pattern weight α2 in this area reaches 0.35. At the third layer (240×135), the overall motion trend is analyzed: the overall motion direction of the pedestrian is calculated as θ = arctan(v avg / u avg ) = arctan(0.9 / 3.8) = 13.3°, indicating that the pedestrian is moving to the upper right at a small angle. The velocity amplitude is ||V|| = sqrt(3.8 2 + 0.9 2 ) = 3.9 pixels / frame. Construct the intermediate frequency spatiotemporal feature tensor F mid The dimensions are 240×135×8×16 (space×feature channels×time). The eight feature channels include: channels 1-2: first-layer motion components (u1, v1); channels 3-4: second-layer motion components (u2, v2); channels 5-6: third-layer motion components (u3, v3); channel 7: motion amplitude ||V||; and channel 8: motion direction angle θ.

[0043] In one embodiment of the present application, the temporal continuity of the motion trajectory is maintained by a Kalman filter. The state vector State(t) = [x(t), y(t), vx(t), vy(t)] is established for the pedestrian center point trajectory. T, where (x, y) are the position coordinates and (vx, vy) are the velocity components. The state transition matrix A = [[1,0,Δt,0], [0,1,0,Δt], [0,0,1,0], [0,0,0,1]], where Δt = 1 / 30 second (30fps video). The observation matrix H = [[1,0,0,0], [0,1,0,0]]. In frame t, the pedestrian's center position is observed to be (805,405), the predicted position is (803,404), and the corrected position calculated using the Kalman gain K is (804,404.5). The velocity update is vx(t) = 3.5 pixels / frame and vy(t) = 1.0 pixels / frame, which is consistent with the main motion component M1 obtained by PCA decomposition. Third-order B-spline interpolation is applied to smooth the trajectory to eliminate trajectory jumps caused by occlusion or detection errors. During the 12th-14th frame, the detection position is offset due to partial occlusion. The original trajectory is [(798,402), (820,398), (810,406)], and after smoothing, it is [(798,402), (806,401), (810,405)], which maintains the continuity of the trajectory. mid Mapping to the second latent subspace. A nonlinear mapping function f(·) is used to compress the 240×135×8×16 feature tensor into a 64×64×5×16 latent representation. The design of the mapping weight matrix W follows the motion preservation principle: weight matrix W1 (240×64): spatial dimension compression, maintaining the spatial layout of the motion area; weight matrix W2 (135×64): spatial dimension compression, maintaining motion continuity; weight matrix W3 (8×5): feature dimension compression, compressing 8 multi-scale features into 5 main motion modes; in the latent space, the main motion information of the pedestrian is encoded as follows: 1st dimension: overall translation motion, coefficient range [-2.5, 2.5]; 2nd dimension: periodic swing, coefficient range [-1.8, 1.8]; 3rd dimension: motion acceleration change, coefficient range [-0.8, 0.8]; 4th dimension: direction change, coefficient range [-0.5, 0.5]; 5th dimension: motion uncertainty, coefficient range [-0.3, 0.3]; the representation of the pedestrian torso area in the latent space is: L2(32,32,1,t) = 2.1 (strong translational motion); the leg swing region is expressed as: L2(28,35,2,t) = 1.6×sin(2π×2.1×t / 30) (periodic swing). Add a temporal smoothing regularization term L in the latent space smooth = Σt ||L2(·,·,·,t) - L2(·,·,·,t-1)|| 2 , weight λ smooth= 0.1, ensuring smooth changes in the latent representation in the time dimension and avoiding unnatural jumps in the motion trajectory.

[0044] According to one aspect of the present application, generating the third latent subspace representation includes:

[0045] a predicted frame reconstructed based on the first and second latent subspace representations;

[0046] Calculate the residual information between the video frame sequence and the predicted frame;

[0047] Perform three-dimensional wavelet transform on the residual information, perform multi-scale decomposition in time and space dimensions, and extract high-frequency wavelet coefficients;

[0048] High-frequency wavelet coefficients are mapped to a third latent subspace to generate a third latent subspace representation, where the mapping preserves detail features through a sparsity constraint.

[0049] In one embodiment of the present application, the predicted frame F_pred = F is reconstructed based on the first two subspace representations. b +Motion_Compensate(F m ), where Motion_Compensate is the motion compensation function; calculate the residual F between the original frame and the predicted frame r = F_original - F_pred, where F_original is the original video frame. A three-dimensional discrete wavelet transform (3D-DWT) is performed on the residual using the Daubechies-4 wavelet basis, with a three-layer decomposition: Layer 1 generates eight subbands (LLL1, LLH1, ..., HHH1); Layer 2 further decomposes LLL1; Layer 3 further decomposes LLL2. High-frequency wavelet coefficients are extracted and mapped to a third latent subspace using a sparsity constraint (such as L1 regularization). Coefficients with energy greater than a threshold τ are retained to obtain the third latent subspace representation (the residual detail flow). High-frequency coefficients are primarily distributed around moving edges and textured regions. The sparsity constraint removes unimportant noise coefficients, preserving visually significant details.

[0050] In another embodiment of the present application, the predicted frame is reconstructed based on the first latent subspace (static background flow) and the second latent subspace (dynamic motion flow). The static background flow is reconstructed by Tucker decomposition: b_recon = G ×1 U1 ×2 U2 ×3 U3, get the background prediction. For the pixel (200,300), reconstruct the background value F b (200,300) = 92. Motion compensation function Motion Compensate(Fm) is implemented using bilinear interpolation: for the motion region pixel (800,400), based on the optical flow vector (u=3.5, v=1.0), interpolation sampling is performed from the previous frame position (796.5,399). The four neighboring pixel values ​​in the previous frame are: I(796,399)=118, I(797,399)=120, I(796,400)=115, I(797,400)=119. Bilinear interpolation calculation: Motion Compensate (800,400) = (1-0.5)×(1-0)×118 + 0.5×(1-0)×120 + (1-0.5)×0×115 + 0.5×0×119 = 119. Prediction frame synthesis F pred (800,400) = F b (800,400) +Motion Compensate (800,400) = 92 + 119 = 211. Calculate the residual between the original frame and the predicted frame. For the pedestrian body pixel (800,400), the original frame value F original (800,400) = 145, predicted frame value F pred (800,400) = 211, residual F r (800,400) = 145 - 211 = -66. For the background pixel (200,300), the original frame value is 94, the predicted frame value is 92, and the residual F r (200,300) = 94 - 92 = 2, indicating that the residual in the background area is small. For the pedestrian edge pixel (820,380), due to the complexity of the motion boundary, the original value is 162, the predicted value is 135, and the residual F r (820,380) =162 - 135 = 27, indicating that there is a large prediction error in the edge area, which needs to be compensated by residual information.

[0051] In one embodiment of the present application, for 16 frames of 240×135 resolution residual sequence F rPerform 3D-DWT decomposition using the Daubechies-4 wavelet basis. First-level decomposition: Apply the wavelet filter bank {h0, h1, g0, g1}, where h0 and h1 are the low-pass and high-pass decomposition filters, and g0 and g1 are the reconstruction filters. The Daubechies-4 filter coefficients are: h0 = [0.683, 1.183, 0.317, -0.183] (low-pass); h1 = [-0.183, -0.317, 1.183, -0.683] (high-pass). Decomposition along the spatial dimension (height): Convolve each row of pixels and downsample by a factor of 2 to obtain the low-frequency L and high-frequency H components, reducing the spatial dimensions from 240×135 to 120×135. Decomposition along the spatial dimension (width): Decomposition continues along the width dimension, obtaining four spatial subbands: LL, LH, HL, and HH, reducing the spatial dimensions to 120×68. Decomposition along the time dimension: Decompose the 16-frame time series to obtain 8 sub-bands: LLL1: low-frequency space + low-frequency time, dimension 120×68×8; LLH1: low-frequency space + high-frequency time, dimension 120×68×8; LHL1: low-high-frequency space + low-frequency time, dimension 120×68×8; LHH1: low-high-frequency space + high-frequency time, dimension 120×68×8; HLL1: high-low-frequency space + low-frequency time, dimension 120×68×8; HLH1: high-low-frequency space + high-frequency time, dimension 120×68×8; HHL1: high-frequency space + low-frequency time, dimension 120×68×8; HHH1: high-frequency space + high-frequency time, dimension 120×68×8; Second-level decomposition: Continue to decompose the LLL1 sub-band, the spatial dimension becomes 60×34×4, and produce 8 secondary sub-bands: LLL2, LLH2, LHL2, The third level decomposition is to continue decomposing the LLL2 subband, and the spatial dimension becomes 30×17×2, generating 8 third-level subbands: LLL3, LLH3, LHL3, LHH3, HLL3, HLH3, HHL3, HHH3.

[0052] In one embodiment of the present application, significant coefficients are extracted from 21 high-frequency subbands. Taking the HHH1 subband as an example, this subband captures high-frequency variations in space and time, primarily distributed in the moving edge region. For the pedestrian edge position corresponding to the coordinates (41, 19, 3) in the HHH1 subband, the wavelet coefficient value is 23.7. The coefficient energy is calculated as: Energy = |23.7| 2= 561.69. An energy threshold of τ = 25 was set. Since 561.69 > 25, this coefficient was retained. For the background region corresponding to coordinates (10, 15, 3) in the HHH1 subband, the wavelet coefficient value is 1.8, and the energy is 3.24 < 25, so it is thinned to 0. The HHH1 subband retains 12.3% of the coefficients, primarily at moving edges; the HHL1 subband retains 8.7% of the coefficients, primarily at horizontal edges; the HLH1 subband retains 6.2% of the coefficients, primarily at vertical edges; and the LHH1 subband retains 9.1% of the coefficients, primarily at temporally varying regions. L1 regularization was applied to the sparsity constraint: minimize ||W·C||1+ λ||C||1, where C is the wavelet coefficient vector, W is the mapping weight matrix, and λ = 0.01 is the sparsity regularization parameter. The coefficients in the HHH1 subband are sorted by energy, retaining the top 5% of the coefficients. Taking the position (41,19,3) as an example, the coefficient 23.7 is soft-thresholded: if |c| > λ, then c' = sign(c) × (|c| - λ) = sign(23.7) × (23.7 - 0.01) = 23.69; if |c| ≤ λ, then c' = 0. The sparsified high-frequency coefficients are reorganized into the third latent subspace representation. The original 21 high-frequency subbands total approximately 6 million coefficients, and after the sparsity constraint, approximately 600,000 non-zero coefficients are retained (sparseness rate 90%). Run-length encoding is used to record the positions of non-zero coefficients: position index: [245, 1203, 1847, 2956, ...]; coefficient value: [23.69, -15.2,8.7, 31.4, ...]; the compact representation of the third latent subspace includes: a sparse position vector: 600,000 dimensions, recording the positions of non-zero coefficients; a coefficient magnitude vector: 600,000 dimensions, recording the values ​​of non-zero coefficients; a subband identifier vector: 600,000 dimensions, identifying the subband to which the coefficient belongs; through this sparse representation, the third latent subspace effectively captures high-frequency details in the video, including motion edges, texture changes, and prediction errors.

[0053] According to one aspect of the present application, generating first coded data includes:

[0054] Constructing the first latent subspace representation as a three-dimensional tensor;

[0055] Perform Tucker decomposition on the three-dimensional tensor to generate a core tensor and factor matrices along the spatial and temporal dimensions, where the dimensions of the core tensor are set to a predetermined compression ratio of each dimension of the original tensor;

[0056] Combining the core tensor and the factor matrix to form the compressed first latent subspace representation;

[0057] The compressed first latent subspace representation is quantized and entropy encoded to generate first encoded data.

[0058] In one embodiment of the present application, the first latent subspace representation is constructed as a three-dimensional tensor of Height×Width×Time (dimensions H×W×T); Tucker decomposition is performed: F b ≈ G ×1U1×2U2×3U3; the dimension of the core tensor G is set to 1 / 8 of the original dimension, that is, (H / 8)×(W / 8)×(T / 8); U1 is the height direction factor matrix, dimension H×(H / 8); U2 is the width direction factor matrix, dimension W×(W / 8); U3 is the time direction factor matrix, dimension T×(T / 8); the factor matrices U1, U2, and U3 are solved by alternating least squares (ALS); × n This is a modulo-n product operation; the core tensor and factor matrices are quantized, with a typical quantization step size of 4 bits. The original data size is H × W × T, and the compressed data size is approximately 1 / 512 + 3 / 8, which is approximately 0.88%. For a 1920 × 1080 × 16 background tensor (approximately 132 MB), after Tucker decomposition, the core tensor G is 240 × 135 × 2 (approximately 0.26 MB); the three factor matrices total approximately 12.4 MB; the total compressed size is 12.66 MB, for a compression ratio of 90.4%.

[0059] According to one aspect of the present application, generating the second encoded data includes:

[0060] constructing the second latent subspace representation as a motion tensor;

[0061] Perform CP decomposition on the motion tensor and represent it as a weighted sum of a predetermined number of one-dimensional factor vectors, where each group of one-dimensional factor vectors captures one main motion mode;

[0062] Extract the spatial factor vector, temporal factor vector and corresponding decomposition weights of the CP decomposition to form a compressed second latent subspace representation;

[0063] Adaptive quantization and entropy coding are performed on the compressed second latent subspace representation to generate second coded data.

[0064] In one embodiment of the present application, the motion flow is constructed as a motion tensor; CP decomposition is performed to convert the motion tensor F m Expressed as: F m ≈ Σ r=1 R λ r ·a r Θb r Θc r ; Where R is the decomposition rank (set to 5); λ r is the weight of the rth component; ar is the spatial height factor vector (length H); b r is the spatial width factor vector (length W); c r is the time factor vector (length T); Θ is the vector outer product; each rank-1 component represents a major motion pattern: the 1st component: horizontal translational motion (λ1 is the largest); the 2nd component: vertical translational motion; the 3rd component: rotational motion; the 4th component: scaling motion; the 5th component: complex local deformation. The spatial factor vector a r 、b r and the time factor vector c r are optimized by gradient descent; the factor vectors are adaptively quantized, and more bits (6 - 8 bits) are allocated to regions with intense motion.

[0065] According to one aspect of the present application, the compressed first latent subspace representation is quantized and entropy-encoded to generate the first encoded data, including:

[0066] An autoregressive model constructed using a convolutional neural network processes the compressed first latent subspace representation element by element in a predetermined scanning order;

[0067] For the current element to be encoded, its encoded spatial neighborhood is extracted as the context;

[0068] The conditional probability distribution of the current element to be encoded is predicted based on the context through the autoregressive model;

[0069] The arithmetic encoder is used to encode the current element to be encoded guided by the conditional probability distribution to generate the first encoded data.

[0070] In an embodiment of the present application, a 5-layer convolutional neural network is constructed as the autoregressive model: the input layer is a 3×3×C context window (C is the number of channels); the first hidden layer is 64 3×3 convolutional kernels with ReLU activation; the second to fourth hidden layers are 128 3×3 convolutional kernels with residual connections; the output layer is 256 probability values (8-bit quantization). Using the raster scan order, the compressed background stream is processed element by element; for the element at position (i, j), a 3×3 neighborhood is extracted as the spatial context. Conditional probability modeling is performed: P(x ij |Context) = Π_{c∈C} P(x ij c |x_{<ij}); where x ij is the coefficient to be encoded at position (i, j); Context is the encoded neighborhood {x mn |m < i or (m = i and n < j)}; C is the set of channels; the CNN predicts the conditional probability distribution N(μ ij , σ ij 2). Where P() is the conditional probability function, Π is the multiplication symbol; μ ij is the mean of the Gaussian distribution predicted by the neural network; σ ij 2 is the variance of the predicted Gaussian distribution. The arithmetic encoder encodes based on the predicted probability distribution, achieving compression close to the entropy limit. Compared to the fixed probability model, adaptive prediction reduces entropy by approximately 1.2 bits per coefficient.

[0071] According to one aspect of the present application, adaptive quantization and entropy coding are performed on the compressed second latent subspace representation to generate second coded data, including:

[0072] constructing the compressed second latent subspace representation as a spatiotemporal sequence;

[0073] Extract the spatiotemporal context of the current position to be encoded from the spatiotemporal sequence, including historical information in the time dimension and neighborhood information in the spatial dimension;

[0074] The Transformer encoder is used to process spatiotemporal context and the self-attention mechanism is used to capture the long-range dependencies of motion patterns.

[0075] The conditional probability distribution of the current position to be encoded is predicted based on the long-range dependency, and the conditional probability distribution is used to guide the arithmetic encoder to generate second encoded data.

[0076] In one embodiment of the present application, the compressed motion stream is constructed as a spatiotemporal sequence with a typical sequence length of 64. A 6-layer Transformer encoder is used, with each layer containing 8 attention heads. A windowed attention mechanism is used with a window size of 8×8×4 (space×time). The self-attention mechanism captures long-range dependencies of motion, such as periodic motion patterns. The output layer predicts the parameters of the Gaussian mixture distribution to guide arithmetic coding. The attention calculation formula is Attention(Q, K, V) = softmax(QK T / sqrt(d)) · V; where Q is the query matrix (current position features); K is the key matrix (context features); V is the value matrix (context information); and d is the feature dimension (512). Periodic motion (such as fan rotation) has a strong correlation between the tth frame and the tTth frame. The Transformer directly establishes this long-range connection through self-attention, while traditional context-adaptive binary arithmetic coding (CABAC) can only utilize the local Markov assumption. Since the Transformer framework, windowed attention, spatiotemporal transformer, conditional probability modeling and other technologies are existing technologies, they will not be described in detail here.

[0077] According to one aspect of the present application, generating the third encoded data includes:

[0078] classifying the wavelet coefficients in the third latent subspace representation to identify zero coefficient positions and non-zero coefficient values;

[0079] Perform sparse position coding on the zero coefficient positions and record the spatial distribution of non-zero coefficients;

[0080] The Gaussian mixture model is used to model the non-zero coefficient values, in which different mixture components are fitted to the positive and negative coefficients respectively to obtain the probability distribution parameters;

[0081] Non-zero coefficients are arithmetically encoded based on probability distribution parameters and combined with sparse position coding to generate third encoded data.

[0082] In one embodiment of the present application, the distribution of wavelet coefficients is analyzed. Typically, more than 90% of the coefficients are zero. Run-length encoding is used for the position of zero coefficients. A three-component Gaussian mixture model is used for non-zero coefficients: positive coefficients: two Gaussian components (peak and tail); negative coefficients: one Gaussian component. The wavelet coefficient distribution is modeled as P(h) = Σ k=1 3 π k · N(h|μ k , σ k 2 ); zero-peak component: π1≈0.9, μ1=0, σ1=0.1; positive tail component: π2≈0.07, μ2=2.5, σ2=1.2; negative tail component: π3≈0.03, μ3=-2.0, σ3=0.8; where P(h) is the probability density function of the wavelet coefficient h, N(h|μ k , σ k 2 ) represents a one-dimensional Gaussian distribution (normal distribution), with μ k is the mean, σ k 2 is the variance, corresponding to the kth component. GMM parameters are estimated using the expectation maximization (EM) algorithm. Arithmetic coding is performed based on the GMM probabilities, and sparse position coding and coefficient coding are combined in the output. Sparse representation encodes 90% of the coefficients with 1 bit (indicating zero), while the remaining 10% of non-zero coefficients use an average of 4 bits, for an overall average of 1.3 bits per coefficient.

[0083] Optionally, the third coded data may be generated as follows:

[0084] Analyze the statistical characteristics of wavelet coefficients in the third latent subspace representation and identify their sparse distribution patterns;

[0085] According to the distribution characteristics of non-zero coefficients, a Gaussian mixture model is used for fitting, in which asymmetric mixture components are established for positive and negative coefficients respectively;

[0086] Predict the probability distribution parameters of each coefficient based on the Gaussian mixture model;

[0087] The probability distribution parameters are used to guide the arithmetic encoder, and the sparsity constraint is combined to encode the third latent subspace representation to generate third encoded data.

[0088] According to one aspect of the present application, the method further includes decoding and reconstructing the compressed code stream, including:

[0089] Decoding the compressed code stream to recover a first decoded representation, a second decoded representation, and a third decoded representation;

[0090] extracting motion information from the second decoded representation, the motion information representing a direction and magnitude of motion in the video;

[0091] Generate spatial attention weights based on motion information, which indicate the motion areas that need to be focused on during reconstruction;

[0092] A deformable convolutional reconstruction network is used to dynamically adjust the sampling position of the convolution kernel according to the attention weight, and the first, second and third decoding representations are fused to generate a reconstructed video frame.

[0093] In one embodiment of the present application, the compressed bitstream is entropy decoded to recover three quantized latent representations. Motion information in the form of optical flow is extracted from the second decoded representation. A spatial attention weight map is generated based on the motion information, with higher weights assigned to areas of intense motion. A deformable convolution reconstruction network is employed: the network comprises five deformable convolution layers and three upsampling layers. The convolution kernel dynamically adjusts the sampling position based on the optical flow vector: offset = Conv(motion_flow) deform_conv(x, offset); where motion_flow is the optical flow information, Conv() is a standard convolution operation, offset is a dynamic offset vector, deform_conv is a deformable convolution operation, and x is the feature map input to the deformable convolution layer. The three decoded representations are cascaded and fused: first reconstructing the background, then overlaying the motion, and finally adding details. Generator loss function: L_total = L_adv + 0.1×L_perceptual + 0.05×L_temporal, where L_adv is the adversarial loss using the PatchGAN discriminator; L_perceptual is the VGG-19 feature matching loss; and L_temporal is the temporal consistency loss. Experiments show that motion-aware reconstruction improves PSNR by 2.1dB in complex motion scenes compared to traditional methods.

[0094] In another embodiment, the mathematical principle of deformable convolution: standard convolution is y(p) = Σ nw(n)·x(p+n); deformable convolution is y(p) = Σ n w(n) x(p+n+Δn); where p is the output position; n is the convolution kernel sampling position; Δn is the offset determined by the optical flow; w(n) is the convolution weight; calculate the offset Δn = Conv_offset(F m ); Conv_offset is a specialized offset prediction network. For fast-moving soccer balls (20 pixels / frame), traditional convolutions produce motion blur. Deformable convolutions adjust sampling points based on the direction of motion, sampling along the trajectory to maintain a clear outline of the ball. Experiments show that the Structural Similarity Index (SSIM) of moving regions improves from 0.82 to 0.91. Cascade fusion process: Background reconstruction: G bg = UpSample(Decode(first decoded representation)); Motion overlay: G motion = G bg + Warp(Decode(Second Decoding Representation), F m ); Detail enhancement: G_final = G motion + IDWT(Decode(third decoded representation)).

[0095] According to one aspect of the present application, before encoding the three potential subspace representations respectively, the method further includes:

[0096] Analyze the content complexity represented by the first, second, and third latent subspaces to obtain the corresponding first, second, and third complexities;

[0097] A rate-distortion prediction model is constructed based on three complexities to predict the distortion and bit rate under different quantization parameter combinations;

[0098] The optimal quantization parameter set that minimizes the weighted sum of total distortion and total bit rate is solved by Lagrangian optimization;

[0099] The corresponding parameters in the optimal quantization parameter set are applied to the encoding process of the three latent subspace representations respectively to achieve adaptive bit allocation.

[0100] In one embodiment of the present application, the spatial frequency and temporal stability of the background flow are calculated to obtain the first complexity; the motion vector variance and motion edge density of the motion flow are analyzed to obtain the second complexity; and the non-zero coefficient ratio and energy distribution of the residual flow are statistically analyzed to obtain the third complexity. Construct a rate-distortion model: D(Q) = α1·Q1 β1 + α2·Q2 β2 + α3·Q3 β3Bitrate model: R(Q) = γ1·log(1+δ1 / Q1) + γ2·log(1+δ2 / Q2) + γ3·log(1+δ3 / Q3); Model parameters are obtained by fitting the training set. α1, α2, and α3 are the distortion weight coefficients of the background, motion, and residual substreams, respectively; Q1, Q2, and Q3 are the quantization steps of the background, motion, and residual streams, respectively; β1, β2, and β3 are the sensitivity indices of the distortion of the background, motion, and residual substreams to the quantization step size, respectively; γ1, γ2, and γ3 are the bitrate weight coefficients of the three substreams, respectively; and δ1, δ2, and δ3 control the decay rate of the bitrate with respect to the quantization step size. Lagrangian optimization solves the problem: the objective function is min J = D(Q) + λ·R(Q); λ ranges from 0.3 to 1.2, adjusted according to the target bitrate. An interior point method is used to solve for the optimal quantization parameter set Q* = {Q1*, Q2*, Q3*}. Adaptive quantization parameters are applied: static background: Q1* typically has a value of 4-8 (high compression); dynamic motion: Q2* typically has a value of 2-4 (fidelity priority); residual detail: Q3* typically has a value of 8-16 (very high compression). This optimization reduces the overall bitrate by 15%-20% while maintaining the same quality, while avoiding the local quality degradation caused by fixed allocation.

[0101] In another embodiment of the present application, the first complexity (background): C1 = α·Var_spatial(F b ) + β·Var_temporal(F b ); Var_spatial is the spatial variance, Var_temporal is the temporal variance, α=0.7, β=0.3. Second complexity (motion): C2= γ·E[||F m ||2] + δ·Entropy(F m ); where E[·] is the expected value, Entropy(·) is the motion vector entropy, γ=0.6, δ=0.4. The third complexity (residual): C3= ε·NNZ(F r ) / Total(F r ); where NNZ is the non-zero coefficient, Total is the total number of coefficients, and ε=1.0. Optimization solution process: Lagrangian function: L = D(Q1, Q2, Q3) + λ·R(Q1, Q2, Q3); gradient descent method is used to iteratively solve Q i (t+1) = Q i (t) - η·ΨL / ΨQ iWhere η = 0.01 is the learning rate, and iterations are performed until convergence; Ψ is the partial derivative. When the target bitrate is 1 Mbps and λ = 0.7, the calculations yield: Q1* = 6 (background), Q2* = 3 (motion), and Q3* = 12 (residual). The resulting actual bitrate is 0.98 Mbps and PSNR = 38.5 dB. Dynamic optimization improves PSNR by 1.8 dB compared to a fixed allocation (Q1 = Q2 = Q3 = 5) at the same bitrate.

[0102] According to another aspect of the present application, the core of the multi-granularity generative video compression method is to decompose the video signal into multiple orthogonal or nearly orthogonal sub-streams that are easier to independently model and compress, and design a customized encoding and reconstruction strategy for each sub-stream. Specifically, it includes: performing preprocessing on the input video frame sequence, including motion estimation (for example, using advanced optical flow estimation algorithms) and background modeling, preliminarily separating the dynamic foreground and static background areas, and providing prior information for subsequent decomposition. Using a three-dimensional wavelet transform (3D-DWT) module, using fast algorithms such as Mallat, the video data block (space-time cube) is subjected to multi-scale and multi-directional time-frequency decomposition to effectively extract features across time and space dimensions. Based on tensor decomposition theory, these multi-scale time-frequency features are mapped into three semantically distinct latent subspaces: the static background stream, which primarily contains low-frequency spatial information and temporally varying components in the video. This stream can be extracted by temporally averaging, clustering, or combining low-frequency wavelet coefficients with background modeling results. This stream varies slowly and is suitable for high-compression encoding. The dynamic motion stream captures mid-frequency spatiotemporal features in the video, primarily characterizing the trajectory, velocity, and deformation of objects. This stream can be generated or constrained by combining the results of an optical flow estimation network. This stream contains the primary motion information and requires precise encoding to ensure motion fidelity. The residual detail stream, primarily reconstructed from high-frequency wavelet coefficients, contains information such as moving edges, texture details, and random noise that the model fails to capture. This stream is typically low in energy but significantly impacts visual quality. To efficiently compress the three separated substreams, a hybrid autoregressive-Transformer entropy model is constructed to adapt to the statistical characteristics and spatiotemporal correlations of the different substreams. For static background flows that are highly temporally correlated but spatially stable, an autoregressive model based on a convolutional neural network (CNN) is used to predict the spatial context of the current latent representation, leveraging its powerful local feature extraction capabilities to eliminate spatial redundancy. For dynamic motion flows that contain complex spatiotemporal dependencies, a Transformer-based entropy model is introduced. Specifically, it utilizes a windowed or local attention mechanism to effectively capture the long-range spatiotemporal dependencies between latent representations while controlling computational complexity. This model can more accurately estimate the probability distribution of motion information. For residual detail flows, which typically exhibit sparse characteristics, a sparse coding strategy is adopted. For example, by adding an L1 regularization constraint to the loss function to encourage the sparsity of the latent representation, it is combined with efficient entropy coding methods (such as arithmetic coding) for compression.The probability distribution of the motion flow entropy model can be mathematically expressed as: P(m_t |m_(<t))=∏_(i,j)N(μ_ij (m_(<t)),σ_ij (m_(<t))), where m_t represents the latent representation of the dynamic motion flow at the current time t (e.g., a feature map or a tensor block), m_(<t) represents its temporal context information (i.e., all or part of the motion flow latent representations before time t), and i, j are indices in the latent representation space dimension. N represents the Gaussian (normal) distribution. μ_ij (m_(<t)) and σ_ij (m_(<t)) are the mean and standard deviation (or the square root of the variance) of this distribution at position (i, j), respectively. These two parameters are predicted by a Transformer-based network based on the context information m_(<t)) and are used to guide the arithmetic encoder to encode the quantized value of m_t. At the decoding end, a motion-aware generative adversarial network (GAN) is adopted for high-quality video reconstruction, with particular attention paid to the restoration of motion details. The compressed code streams of each sub-flow are restored to the quantized latent representations through the corresponding entropy decoder (arithmetic decoder). These restored latent representations are fed into a cascaded decoder architecture. In particular, to enhance the reconstruction quality of the motion region, a part (or the entire decoding network) in the decoder is designed as the generator of the GAN. The generator receives the decoded information from the static background flow, the dynamic motion flow, and the residual detail flow as input or conditions. The key is to use the decoded dynamic motion flow information (or the accurate optical flow information transmitted at the encoding end) to guide the attention mechanism of the generator. Specifically, an optical flow-guided deformable attention module can be adopted, enabling the sampling positions of the convolutional kernels to be adaptively adjusted according to the motion direction and amplitude indicated by the optical flow vectors, so as to focus more on the contours and internal details of moving objects and effectively combat the blurring and artifacts introduced by compression. The cascaded decoder structure gradually fuses the information from the three sub-flows, first reconstructing the basic static scene using the background flow, then superimposing the motion flow information to restore the dynamic content, incorporating the information of the residual detail flow for refinement, and finally synthesizing high-quality output video frames.

[0103] Compared to existing technologies, this embodiment achieves improved compression efficiency. Experimental data shows that while achieving comparable subjective or objective visual quality (e.g., SSIM values) to HEVC, this embodiment reduces bitrate. This is due to efficient compression of static backgrounds and accurate modeling of dynamic information. Regarding motion fidelity, by explicitly separating and encoding dynamic motion streams and reconstructing them using an optical flow-guided deformable attention GAN, details in fast and complex motion scenes can be more accurately restored, resulting in improved peak signal-to-noise ratio (PSNR) in these scenarios. While a deep learning model is employed, the overall encoding speed is improved compared to some complex traditional or early learning encoders through sub-stream decomposition and adaptive quantization, as well as GPU parallel processing optimization. In engineering applications, the dynamic bit allocation strategy significantly reduces storage usage and can flexibly adapt to different application scenarios, such as significantly compressing static backgrounds in surveillance video storage or prioritizing motion streams to ensure an interactive experience in low-latency scenarios such as cloud gaming.

[0104] This embodiment proposes to map the time-frequency features of the video to three approximately orthogonal latent subspaces of static background, dynamic motion, and residual details through tensor decomposition, thereby realizing content-adaptive hierarchical representation and laying the foundation for subsequent differentiated and efficient coding. Based on the statistical characteristics of different substreams, a hybrid entropy model is constructed, combining the spatial modeling capability of CNN and the ability of Transformer to capture long-range spatiotemporal dependencies, as well as sparse coding, to improve the compression efficiency of the latent representation. A motion-aware generative adversarial network is introduced at the decoding end, and optical flow information is used to guide the deformable attention mechanism, which effectively enhances the reconstruction fidelity of motion details, especially improving the visual quality in complex motion scenes.

[0105] In a specific embodiment, based on the hierarchical latent space and multi-granularity generative framework, efficient video compression is achieved through a sophisticated processing flow. Specifically, the input video sequence is preprocessed, including GOP segmentation and normalization, and an optical flow network (such as RAFT) is used to initially extract dense motion flow F m , combined with background modeling technology to separate the static background flow F b The calculated residual information F r Apply three-dimensional discrete wavelet transform (3D-DWT) to perform multi-scale time-frequency decomposition and extract residual detail flow; use tensor decomposition technology (Tucker decomposition processing F b , CP decomposition processing F m) projects the background and motion streams into a compact, low-rank latent space. A hybrid autoregressive-Transformer entropy model is constructed to specifically compress the latent representations of each substream: the CNN autoregressive model handles spatial correlations in the background stream, the spatiotemporal Transformer model captures long-range dependencies in the motion stream, and the GMM model encodes the residual coefficients. Bits are dynamically allocated through global rate-distortion optimization (Lagrangian minD + λR). On the decoding side, after entropy decoding, a motion-aware generative adversarial network (GAN) is used for high-quality reconstruction. Its generator utilizes a deformable attention mechanism guided by optical flow to focus on recovering motion details, and the final frame is output by cascading the three streams. The entire process aims to achieve a balance between efficient compression and high-fidelity motion recovery.

[0106] In order to make the purpose, technical solutions and advantages of the present invention more clear, the specific implementation steps of the multi-granularity generative video compression method are described in detail below:

[0107] Step 1: Video sequence input, preprocessing and preliminary spatiotemporal separation.

[0108] The input raw video data is normalized, and a motion estimation algorithm is used to initially separate significant dynamic motion from relatively static background information, laying the foundation for subsequent, more refined hierarchical latent space decomposition. The system receives the raw video sequence and, to facilitate parallel processing and exploit temporal correlation, segments it into consecutive groups of pictures (Groups of Pictures (GOPs)) of a preset duration (e.g., typically 16 frames). Each frame within each GOP undergoes pixel normalization, for example, by linearly mapping it to the interval [-1, 1]. This improves the stability and convergence speed of subsequent deep learning model training. To specifically process color and brightness information, frames are typically converted from RGB color space to YUV color space, extracting the luma (Y) and chroma (U, V) components separately. Subsequent processing can focus on the luma component or on all components. Entering the critical motion estimation stage (which can be considered part of or immediately following the preprocessing module), advanced dense optical flow estimation algorithms such as the RAFT (Recurrent All-PairsField Transforms) network are employed. The network can calculate the displacement vector of each pixel between the current frame and the reference frame (usually the previous frame), forming a dense two-dimensional optical flow field F m Here F m ={(u(x, y), v(x, y))|all (x, y)∈Frame} represents a vector field with the same size as the original frame, where u(x, y) and v(x, y) are the instantaneous velocity estimates of the pixel (x, y) in the horizontal and vertical directions respectively. This optical flow field F mThe main dynamic information in the video is preliminarily characterized. At the same time, or based on the optical flow results, the static background flow F can be estimated by simple time domain filtering (such as pixel-level time domain median filtering) or more complex background modeling methods (such as Gaussian mixture model GMM or deep learning-based methods). b A simplified criterion is that for a pixel position (x, y), if its optical flow amplitude in multiple consecutive frames is \|(u, v)\|2=sqrt(u 2 +v 2 ) is less than a preset small threshold ε (for example, ε = 0.5 pixels / frame), the pixel is considered to belong to the static background area, forming a preliminary static background image F b This step outputs the initial separation of background information F b and dense optical flow field F m , as well as the original (or preprocessed) video frame sequence for use by subsequent modules.

[0109] Step 2: Spatiotemporal multi-scale decomposition of residual information based on three-dimensional wavelet transform.

[0110] After the initial background and motion separation, the remaining video information, namely the residual frame, is processed and time-frequency analysis is performed using the three-dimensional discrete wavelet transform (3D-DWT) to extract multi-scale and multi-directional features across time and space dimensions. These features mainly include motion edges, texture details, and complex dynamic components that the model fails to capture. Calculate the residual frame sequence F r , which is defined as the original frame (after normalization and color space conversion, denoted as F_original minus the initial estimated static background flow F b and the image content corresponding to the dynamic motion flow (which can be obtained through the optical flow F m Warp the previous frame to get the predicted frame F m_pred , then F r =F_original-(F b +F m_pred ), or simplify the process, such as directly considering F r =F_original-F b ). This residual sequence F r Represents the change information in the video except for the stable background and main motion. For this three-dimensional data block (spatial dimension × time dimension) F rA three-dimensional discrete wavelet transform (3D-DWT) is applied. This transform recursively performs filtering and downsampling operations along three dimensions (height, width, and time) through a separable filter group. In specific implementation, the fast Mallat algorithm can be used. The selected wavelet basis function has an important impact on the performance. In this embodiment, the Daubechies 4 (db4) wavelet basis is selected because it has good tight support and certain regularity, which is conducive to the localization of features. The decomposition level (Scale Level) is set to 3, which means that the video data block will be decomposed into three representations of different scales. In each layer of decomposition, a low-frequency approximate subband (LLL) and seven high-frequency detail subbands (LLH, LHL, LHH, HLL, HLH, HHL, HHH) are generated, where L represents low-pass filtering and H represents high-pass filtering. The three letters correspond to the operations in the height, width, and time dimensions respectively. After three layers of decomposition, an approximate subband L with the lowest frequency is finally obtained. (3) , and a total of 3×7=21 high-frequency detail subband sets of different scales and directions {H k (s) |s∈{1,2,3},k∈{1..7}}. Among them, L (3) It captures the coarsest spatiotemporal structure in the residual information, while the high-frequency subband H k (s) It contains detailed information such as edges, textures, and noise of different frequencies and directions. These wavelet coefficients constitute the main content of the residual detail stream, and their sparsity makes it possible to perform efficient compression.

[0111] Step 3: Latent space projection and low-rank representation of background flow and motion flow.

[0112] The static background flow F preliminarily separated in step 1 b and dynamic motion flow F m Further compression and structured representation are performed, and tensor decomposition technology is used to project it into a more compact potential space to extract its core structural information for subsequent entropy coding. For static background flow F that is highly correlated in time but has a relatively stable spatial structure b (Usually a three-dimensional tensor with dimensions of Height×Width×Time, compressed using the Tucker decomposition model. Tucker decomposition represents a high-dimensional tensor as the product of a core tensor and a series of factor matrices along each dimension. Mathematically, F b ≈F b core ×_1 U_1 ×_2 U_2 ×_3 U_3. Here, F b coreIt is a size much smaller than the original F b The core tensor of F is set to 1 / 8 of the original video dimensions (height, width, and time) in this embodiment, that is, (H / 8)×(W / 8)×(T / 8). U_1 (dimension H×H / 8), U_2 (dimension W×W / 8), and U_3 (dimension T×T / 8) are factor matrices along the height, width, and time dimensions, respectively (usually required to be orthogonal). They can be regarded as basis vectors of the corresponding dimensions. b core It captures the interactive information under different basis vector combinations and represents the most essential potential structure of the background flow. This decomposition can effectively remove the spatial and temporal redundancy of the background. m (i.e., dense optical flow field, also a three-dimensional tensor H×W×2, or regarded as two H×W×T tensors), its structure may be more complex and may not be suitable for the low-rank assumption. However, for compression, CP decomposition (Canonical Polyadic Decomposition, also known as PARAFAC) is chosen in this embodiment. CP decomposition represents the tensor as the sum of the outer products of several (rank R) one-dimensional vectors. Mathematically, F m ≈∑ r=1 R λ r a r Θb r Θc r Here, R is the decomposition rank, which is set to 5 in this embodiment. r is the weight of the rth component (usually incorporated into the factor vector), a r (dimension H), b r (dimension W) and c r (Dimension T or 2T, depending on F m ) are one-dimensional factor vectors along the three dimensions respectively. Θ represents the vector outer product. CP decomposition attempts to find the most important R spatiotemporal motion patterns that constitute the original motion flow. In this way, F b and F m are converted into a latent representation with reduced parameters (F b core , U_1, U_2, U_3 and λ r , a r , b r , c r for r=1 to 5), these parameters will be used as input to the subsequent entropy coding module.

[0113] Step 4: Hybrid entropy coding of static background flow and dynamic motion flow.

[0114] The latent representations of the static background flow and dynamic motion flow obtained in step 3 are efficiently compressed and encoded using a hybrid entropy model. This model adopts different strategies for the statistical characteristics and spatiotemporal correlations of different flows. A customized probability model is designed for each flow to accurately estimate the probability distribution of each element in its latent representation. An efficient entropy encoder such as arithmetic coding is used to perform lossless or near-lossless compression based on these probabilities. For the compressed static background flow core tensor F b core (The factor matrix U_i usually also needs to be encoded, but the core tensor is the main information carrier.) Considering that the background usually has strong local spatial correlation, this embodiment adopts an autoregressive model based on a convolutional neural network (CNN), such as PixelCNN or its variants. This type of model predicts the probability distribution of the coefficients (or quantization indexes) x_(i, j, k) of the current position (i, j, k) one by one according to a predetermined scanning order (such as raster scanning). The prediction condition is the neighborhood (context) C_(i, j, k) of the position that has been encoded in the scanning order. Its conditional probability can be expressed as P(x_(i, j, k) |C_(i, j, k);θ_AR), where θ_AR is the parameter of the autoregressive model (learned by CNN). The model uses the powerful local feature extraction capability of CNN to capture spatial redundancy. For the potential representation of dynamic motion flow (such as the factor vector sequence λ obtained by CP decomposition), the model can be used to extract local features. r , a r , b r , c r forr=1toR, or directly the original optical flow sequence F m encoding), considering that motion information usually contains complex long-range spatiotemporal dependencies (such as the continuous motion trajectory of an object, periodic motion, etc.), this embodiment introduces an entropy model based on Transformer. The self-attention mechanism of Transformer can effectively capture the long-range dependencies between elements in the sequence. In order to adapt to video data, a spatiotemporal Transformer architecture can be used. To reduce computational complexity, windowed attention or local attention mechanisms can be used to limit attention calculations to local spatiotemporal neighborhoods. The model predicts the probability distribution of the potential representation of the motion flow (or its quantized index) y_t at the current time t, conditional on its temporal context information Context_t={y_1,...,y_(t-1)} (and possible spatial context). For example, the model can predict the parameters of a Gaussian distribution: P(y_t |Context_t;θ_Transformer)=N(y_t |μ_t,σ t 2 ), where the mean μ_t and variance σ t2 These precise probability models P(x_(i, j, k) |...) and P(y_t |...) are then fed into the arithmetic encoder, which instructs it to represent the quantized values ​​of these latent variables with a number of bits close to the information entropy.

[0115] Step 5: Entropy coding and global rate-distortion optimization for the residual detail stream.

[0116] The residual detail flow (ie, wavelet coefficient H) obtained by 3D-DWT in step 2 k (s) ) strategy for entropy coding, and dynamically adjusts the quantization parameters of all three streams through global rate-distortion optimization (RDO) to maximize reconstruction quality within a given bitrate budget. The residual detail stream is primarily composed of high-frequency wavelet coefficients, which typically exhibit sparse, non-Gaussian statistical characteristics. That is, most coefficients are close to zero, with only a few having significant amplitudes, representing image details such as edges and textures. To address this characteristic, this embodiment employs an entropy model based on the Gaussian Mixture Model (GMM). A GMM can fit complex probability distributions as a weighted sum of multiple Gaussian components, effectively capturing the peaked and heavy-tailed characteristics of the wavelet coefficient distribution. Furthermore, an asymmetric GMM can be employed, using different mixture models for positive and negative coefficients to more accurately model their distributions. The model P(h |θ_GMM) (where h represents a wavelet coefficient and θ_GMM is a parameter of the GMM model) is used to guide the quantization and arithmetic coding of the wavelet coefficients. In order to further enhance the compression effect, sparsity constraints can be introduced during model training or quantization, such as adding an L1 regularization term to the loss function to encourage the sparsity of the potential representation (wavelet coefficients). In the global rate-distortion optimization link, simply assigning fixed quantization parameters to each stream (such as 4 bits for background, 6 bits for motion, and 2 bits for residual mentioned in the draft) is usually not optimal. The quantization strength needs to be dynamically adjusted according to the overall bitrate target and the specific characteristics of the video content. This embodiment uses the Lagrange multiplier method for optimization. The goal is to minimize the weighted sum of the total distortion D and the total bitrate R: min Q L=D(Q)+λR(Q). Where Q={Q b , Q m , Q r} represents the set of quantization parameters (e.g., quantization step size) assigned to the static background stream, dynamic motion stream, and residual detail stream. D(Q) is the total distortion of the final reconstructed video relative to the original video when the quantization parameter is Q (which can be measured using MSE, MS-SSIM, or combined perceptual loss). R(Q) is the corresponding total bit rate, which is composed of the sum of the number of bits encoded by each of the three streams (R(Q)=R b (Q b )+R m (Q m )+R r (Q r λ is the Lagrangian multiplier, which controls the trade-off between bitrate and distortion: a larger λ value favors lower bitrate (allowing greater distortion); a smaller λ value favors lower distortion (allowing higher bitrate). In practice, an appropriate λ value must be selected for the target bitrate range (in this example, the experimental range is 0.3 to 1.2). An iterative search or model-based approach is used to find the optimal quantization parameter set Q* that minimizes the Lagrangian cost L. This optimization process ensures the most efficient bitrate distribution among the three streams, maximizing overall compression performance.

[0117] Step 6: Decoding and high-quality reconstruction based on motion-aware generative adversarial network.

[0118] At the decoding end, the decoded information of each sub-stream is used to reconstruct high-quality video through a specially designed motion-aware generative adversarial network (GAN), with a particular focus on recovering motion details that may have been lost during the compression process. Three sub-streams corresponding to the static background stream, the dynamic motion stream, and the residual detail stream are separated from the received compressed bitstream. Using the entropy decoder corresponding to the encoder (such as an arithmetic decoder) and based on the transmitted probability model information (or re-inferred probability model at the decoder end), each sub-stream is decoded back to the quantized latent representation: the core tensor F of the background stream is recovered. b core (i.e., factor matrix), potential representation of motion flow (such as CP factor or optical flow information F m ), and the quantized wavelet coefficients H of the residual detail stream k (s). Next, these recovered potential representations are fed into a cascaded decoder architecture for reconstruction. The core of this architecture is a generator G of a generative adversarial network (GAN). This generator G is designed to be motion-aware, that is, it pays special attention to using the decoded motion information to improve the reconstruction quality. Specifically, the generator G receives decoded information from three streams as input or conditions. The key is that it contains or utilizes a flow-guided attention mechanism module (Flow-Guided Attention Module) inside. This module receives the decoded dynamic motion flow information (usually the optical flow field F m ) as a guiding signal. For example, a flow-guided deformable convolution or an attention weighting mechanism can be used to enable the reconstruction network (such as the sampling position of the convolution kernel or the weight of the feature channel) to be adaptively adjusted according to the motion direction and amplitude indicated by the optical flow vector. The network can focus more computing resources and attention on areas such as the contours and internal textures of moving objects, more effectively combat the blur and artifacts introduced by compression, and restore sharper and more realistic motion details. The discriminator D is used to train the generator G. It receives the reconstructed frame output by the generator and compares it with the real original frame (training stage) to determine its authenticity. This embodiment uses a discriminator D with a multi-scale PatchGAN structure, which can evaluate the authenticity of image blocks at different scales, helping to improve the generation quality of details and textures. The loss function during training includes the adversarial loss L_adv (making the generated frames difficult to distinguish by the discriminator) and the perceptual loss L_perceptual (for example, the feature matching loss based on the pre-trained VGG-19 network is used to ensure that the generated frames are similar to the original frames in terms of deep features, with a weight of 0.1). The total loss is L_GAN=L_adv+0.1L_perceptual. The structure of the cascade decoding is reflected in the possibility of first using the decoded background flow information (such as through F b core And the factor matrix reconstruction F b ) to reconstruct the basic static scene, and then superimpose (or fuse) the motion flow information (such as using F m The decoded residual wavelet coefficients are then used to generate a detail enhancement map using an inverse DWT. This map is then incorporated into the reconstruction result for refinement, resulting in a high-quality output video frame F*. The entire process is optimized through end-to-end training. Testing on an HEVC Class B sequence on an NVIDIA V100 GPU revealed that, when λ = 0.7, this implementation achieved a 0.02 improvement in the SSIM index while reducing the bit rate by 43.6% compared to HEVC, demonstrating its effectiveness.

[0119] The present invention can revolutionize the current video streaming transmission and distribution model. Online video platforms face multiple challenges, including massive data storage, high bandwidth costs, and users' continued pursuit of high-quality, low-latency experiences. Although existing compression standards (such as H.264 / AVC, HEVC, and AV1) continue to improve, it is still difficult to achieve both extreme compression and detail preservation when processing videos containing a large amount of static background (such as talk shows and landscape documentaries) or mixed complex motion (such as sports events and action movies). The present invention decomposes the video into three sub-streams of static background, dynamic motion, and residual details through spatiotemporal decoupling, which can achieve unprecedented content-adaptive compression. For slowly changing or static background areas in the video (such as the sky, walls, and fixed scenery), the static background stream can be encoded with an extremely high compression ratio, reducing the overall bit rate, which corresponds to a significant reduction in CDN (content distribution network) costs. Dynamic motion streams concentrate resources on accurately encoding key information such as the object's motion trajectory and deformation. Combined with a motion-aware GAN reconstruction network, especially the deformable attention mechanism guided by optical flow, they can accurately restore the edges and texture details of fast-moving objects, effectively suppressing the motion blur and blocking effects common in traditional encoding at low bit rates. This provides a clearer and smoother visual experience, especially in high-speed motion scenes such as live sports or action movies. The hybrid entropy model optimizes the statistical characteristics of different streams to further tap into the compression potential. Therefore, by applying the present invention, video service providers can provide higher resolution (such as promoting the popularization of 4K / 8K) or a more stable playback experience (reducing buffering) under the same bandwidth, or reduce the bit rate by approximately 45% while maintaining the existing quality, thereby reducing operating costs and improving user satisfaction in different network environments. The advantages are particularly evident in bandwidth-constrained scenarios such as mobile terminals.

[0120] In the field of video surveillance, the explosive growth of data volumes is placing enormous pressure on storage systems and network transmission. At the same time, clearly recording critical motion details during unusual events (such as intrusions, accidents, and suspicious behavior) is a core requirement for security systems. Traditional surveillance video encoding often redundantly encodes scenes containing large static background areas (such as the field of view of a fixed camera), wasting significant storage space and bandwidth. The layered compression method proposed in this paper offers a natural advantage in this regard. Through background modeling in preprocessing and subsequent tensor decomposition, it efficiently separates fixed background elements (such as buildings, roads, and the sky) from the surveillance footage into a static background stream. This is then deeply compressed using a CNN-based autoregressive entropy model, reducing storage usage (based on experimental data and combined with dynamic bit allocation). Furthermore, critical dynamic information in the surveillance scene, such as the movement trajectories and speed changes of pedestrians and vehicles, and even areas requiring high-resolution detail such as faces and license plates, are effectively mapped into the dynamic motion stream. Using a Transformer-based entropy model for encoding, this method captures complex spatiotemporal dependencies and ensures the integrity of motion information. More importantly, the motion-aware GAN reconstruction network at the decoding end uses optical flow information to guide the attention mechanism, specifically enhancing the reconstruction quality of moving areas. Even after compression, it can clearly restore the detailed features of moving targets, effectively avoiding the loss or blurring of details caused by over-compression. This is crucial for post-event forensics and real-time analysis (such as AI behavior recognition and target tracking). For example, in low-light conditions at night or in inclement weather, this method can better preserve the outlines and details of moving objects. Therefore, this invention can help security systems achieve longer video storage cycles, lower transmission bandwidth requirements, and more reliable key evidence acquisition capabilities, thereby improving the overall efficiency and cost-effectiveness of intelligent security systems.

[0121] Immersive applications such as VR / AR and the metaverse place extremely high demands on the transmission and rendering of video / image data, typically requiring high resolution (4K / 8K or even higher), high frame rates (above 90Hz), wide fields of view (such as 360° panoramic video), and low-latency interaction. The massive data volume poses a significant challenge to network bandwidth and terminal processing capabilities. Any compression artifacts or delays can disrupt user immersion and even induce motion sickness. The layered generative compression method of the present invention provides an effective approach to addressing these issues. The spatiotemporal decoupling mechanism is particularly well-suited for VR / AR scenarios. In VR, much of the environment within the user's field of view may be static or slowly changing, and can be efficiently compressed within the static background stream. However, the user's interactive objects, avatars, or other dynamic elements belong to the dynamic motion stream, requiring precise encoding to ensure realistic and smooth interaction. In AR, real-world background information and overlaid virtual dynamic information can be similarly separated and differentially encoded. The present invention's focus on motion fidelity is crucial. The precise encoding of dynamic motion streams and the reconstruction capabilities of motion-aware GANs can ensure the clear and natural presentation of details such as object movement, user gesture tracking, and AR overlay animations in the virtual world, avoiding blur, smearing, or artifacts, which is crucial for maintaining immersion and interaction accuracy. The reduction in bit rate directly translates into lower transmission latency, which is a core advantage for VR / AR applications that require real-time rendering and interaction (such as cloud VR and remote collaboration), and can achieve a smoother and more immediate user experience. The combination of hybrid entropy coding and generative reconstruction can maintain acceptable visual quality at extremely low bit rates, expanding the application possibilities of VR / AR in mobile networks or wireless environments. In summary, the present invention can effectively reduce the data burden of VR / AR applications, improve motion expressiveness, reduce latency, and bring users a higher quality and more comfortable immersive interactive experience, effectively promoting the development of related industries.

[0122] The preferred embodiments of the present invention are described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the scope of protection of the present invention.

Claims

1. A multi-granularity generative video compression method, characterized in that: include: Obtain a video frame sequence, perform spatiotemporal feature decomposition, and decompose the video signal into three latent subspaces to obtain the first, second, and third latent subspace representations. The first latent subspace represents the low-frequency spatial information and temporal fading components in the video, the second latent subspace represents the spatiotemporal motion features in the video, and the third latent subspace represents the high-frequency detail information in the video. Encoding the first, second, and third latent subspace representations respectively to generate corresponding first, second, and third coded data, and combining them to generate a compressed code stream; The first, second and third latent subspace representations are obtained, including: Perform motion estimation on the video frame sequence to obtain motion information; Perform background modeling on the video frame sequence to obtain background information; Extract spatiotemporal features from video frame sequences based on motion information and background information; The spatiotemporal features are tensor decomposed and mapped into three approximately orthogonal latent subspaces, where the correlation between the latent subspaces is minimized by a preconfigured decomposition matrix to obtain the first, second and third latent subspace representations.

2. The method according to claim 1, characterized in that The generation of the first latent subspace representation includes: Perform time domain filtering on background information to extract temporally stable regions; Extracting spatial low-frequency components from a video frame sequence; Calculate the motion amplitude of each pixel position based on the motion information, and mark the area where the motion amplitude is less than a preset threshold as a static area mask; The temporally stable region, the spatial low-frequency component and the static region mask are weightedly fused to generate a first latent subspace representation, wherein the weight of the weighted fusion is adaptively determined according to the temporal stability of each region.

3. The method according to claim 1, characterized in that The generation of the second latent subspace representation includes: Decompose the motion information into motion vector fields and extract the main motion components; Perform temporal difference calculation on video frame sequences to identify motion areas; Based on the main motion components and motion regions, mid-frequency spatiotemporal features are extracted, where motion patterns of different speeds and directions are captured through multi-scale motion analysis; The mid-frequency spatiotemporal features are mapped to a second latent subspace to generate a second latent subspace representation, where the mapping preserves the temporal continuity of the motion trajectory.

4. The method according to claim 1, wherein The generation of the third latent subspace representation includes: a predicted frame reconstructed based on the first and second latent subspace representations; Calculate the residual information between the video frame sequence and the predicted frame; Perform three-dimensional wavelet transform on the residual information, perform multi-scale decomposition in time and space dimensions, and extract high-frequency wavelet coefficients; High-frequency wavelet coefficients are mapped to a third latent subspace to generate a third latent subspace representation, where the mapping preserves detail features through a sparsity constraint.

5. The method according to claim 1, wherein Generating the first encoded data includes: Constructing the first latent subspace representation as a three-dimensional tensor; Perform Tucker decomposition on the three-dimensional tensor to generate a core tensor and factor matrices along the spatial and temporal dimensions, where the dimensions of the core tensor are set to a predetermined compression ratio of each dimension of the original tensor; Combining the core tensor and the factor matrix to form the compressed first latent subspace representation; The compressed first latent subspace representation is quantized and entropy encoded to generate first encoded data.

6. The method according to claim 1, characterized in that Generating the second encoded data includes: constructing the second latent subspace representation as a motion tensor; Perform CP decomposition on the motion tensor and represent it as a weighted sum of a predetermined number of one-dimensional factor vectors, where each group of one-dimensional factor vectors captures one main motion mode; Extract the spatial factor vector, temporal factor vector and corresponding decomposition weights of the CP decomposition to form a compressed second latent subspace representation; Adaptive quantization and entropy coding are performed on the compressed second latent subspace representation to generate second coded data.

7. The method according to claim 5, characterized in that Generating the first encoded data includes: An autoregressive model constructed using a convolutional neural network processes the compressed first latent subspace representation element by element in a predetermined scanning order; For the current element to be encoded, extract its encoded spatial neighborhood as context; Predict the conditional probability distribution of the current element to be encoded based on the context through an autoregressive model; The conditional probability distribution is used to guide the arithmetic encoder to encode the current element to be encoded to generate first encoded data.

8. The method according to claim 6, characterized in that Generating the second encoded data includes: constructing the compressed second latent subspace representation as a spatiotemporal sequence; Extract the spatiotemporal context of the current position to be encoded from the spatiotemporal sequence, including historical information in the time dimension and neighborhood information in the spatial dimension; The Transformer encoder is used to process spatiotemporal context and the self-attention mechanism is used to capture the long-range dependencies of motion patterns. The conditional probability distribution of the current position to be encoded is predicted based on the long-range dependency, and the arithmetic encoder is guided to generate the second encoded data accordingly.

9. The method according to claim 1, characterized in that Generating the third encoded data includes: classifying the wavelet coefficients in the third latent subspace representation to identify zero coefficient positions and non-zero coefficient values; Perform sparse position coding on the zero coefficient positions and record the spatial distribution of non-zero coefficients; The Gaussian mixture model is used to model the non-zero coefficient values, in which different mixture components are fitted to the positive and negative coefficients respectively to obtain the probability distribution parameters; Non-zero coefficients are arithmetically encoded based on probability distribution parameters and combined with sparse position coding to generate third encoded data.

Citation Information

Patent Citations

  • Deep learning-based compression method using frequency decomposition

    CN119072926A