Multi-granularity generation type video compression method
By decomposing the video signal into three potential subspaces for encoding, low-frequency space, space-time motion and high-frequency details, the problems of code rate allocation imbalance and subband matching mismatch in the prior art are solved, and efficient video compression and high-quality reconstruction in complex motion scenarios are achieved.
Patent Information
- Application Number
- CN202510797399.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-16
AI Technical Summary
When existing video compression technology deals with complex motion scenarios, it is difficult to explicitly separate static backgrounds from dynamic motion, resulting in imbalance in bit rate allocation, limited multi-scale feature extraction capabilities, and problems of subband matching mismatch and motion artifacts.
The video signal is decomposed and mapped to three potential subspaces: low-frequency spatial information, spatiotemporal motion characteristics and high-frequency detail information, and is respectively encoded to generate compressed code streams. The hybrid entropy encoding mechanism and generative reconstruction network are used to improve compression performance using tensor decomposition and deep learning models.
It realizes efficient compression of content adaptability, reduces bit rate, improves compression efficiency and motion fidelity, and improves visual quality especially in complex motion scenarios.
Smart Images

Figure CN120343255A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video compression, and in particular, to a multi-granularity generative video compression method. Background Art
[0002] Video compression technology is of crucial significance in today's information age. Its main purpose is to reduce the amount of video data in order to lower storage costs and transmission bandwidth requirements. Effective video compression technology can improve the storage efficiency of videos, accelerate the transmission speed of videos, and is widely applied in many fields such as video streaming services, video surveillance, high-definition televisions, etc. Achieving more efficient video compression through technological innovation can promote the further development of these fields and enhance the user experience.
[0003] Currently, traditional methods have an inherent contradiction between redundancy elimination and motion preservation due to spatio-temporal coupling representation. Mainstream video compression standards such as HEVC / AV1 adopt block-based motion compensation and transform coding, which can compress spatial redundancy to a certain extent. However, repeated coding of static backgrounds causes bit waste, and block effects are caused by the rigidity of block partitioning in fast-moving regions. The spatio-temporal coupled transform domain representation destroys the continuity of motion trajectories, resulting in a rate-distortion optimization boundary lower than the theoretical value. More seriously, existing frameworks cannot orthogonally decompose video components, causing dynamic motion and static background to interfere with each other during quantization, resulting in the loss of high-frequency details. The multi-scale feature extraction ability of existing coding systems is limited by the time-frequency analysis defects of transform methods. Although the traditional DCT transform is computationally efficient, its global basis functions generate Gibbs oscillations in regions of motion abrupt changes, while the three-dimensional wavelet transform results in sub-band allocation mismatch due to its fixed decomposition structure. The parameter mismatch between spatial direction filters and temporal wavelets distorts the motion phase, causing a PSNR drop of up to. The entropy coding efficiency is limited by the insufficient modeling of spatio-temporal non-stationarity in traditional models. The CABAC local correlation model based on the Markov assumption has difficulty capturing long-term motion dependencies, and the motion vector entropy coding redundancy in HEVC is relatively high. Although deep learning models have improved the modeling of spatial correlations, their autoregressive structure has a limited temporal receptive field, and the computational complexity increases cubically with the resolution. Although the Transformer model can model long-sequence dependencies, directly applying it to entropy coding will cause video decoding delays. The reconstruction enhancement technology based on generative adversarial networks faces the dual challenges of motion consistency and training stability. Existing solutions decouple optical flow estimation from the compression process, resulting in a relatively high proportion of motion artifacts. At the same time, mode collapse in adversarial training causes the SSIM in flat regions to decline, while overfitting leads to the generation of false edges. The fundamental reason is that GAN lacks physical constraints on compression domain features. For example, the mismatch between quantization step sizes and the generator receptive field amplifies frequency domain distortion.
[0004] In the prior art, in terms of multi-component decoupling, it is difficult to handle complex motion scenarios, and the static background and dynamic motion are not explicitly separated, resulting in an imbalance in bitrate allocation; the multi-scale feature extraction ability is limited by the time-frequency analysis defects of the transformation method, such as Gibbs oscillations generated by traditional DCT transformation and sub-band allocation mismatch of three-dimensional wavelet transformation, etc. Summary of the Invention
[0005] The object of the invention is to provide a multi-granularity generative video compression method, aiming to solve at least one technical problem existing in the prior art.
[0006] Technical solution: The multi-granularity generative video compression method includes: Obtain a video frame sequence, perform spatio-temporal feature decomposition, decompose and map the video signal into three latent subspaces to obtain the first, second, and third latent subspace representations, where the first latent subspace representation is the low-frequency spatial information and time-varying components in the video, the second latent subspace representation is the spatio-temporal motion features in the video, and the third latent subspace representation is the high-frequency detail information in the video; Encode the first, second, and third latent subspace representations respectively to generate corresponding first, second, and third encoded data, and combine them to generate a compressed bitstream.
[0007] Beneficial effects: The present invention explicitly separates the static background and dynamic motion information, effectively solves the problem of bitrate allocation imbalance, and realizes content-adaptive high-efficiency compression; enhances the multi-scale feature extraction ability, avoids the problem of sub-band matching mismatch, and improves the compression performance. Brief Description of the Drawings
[0008] Figure 1 It is a step flowchart of a multi-granularity generative video compression method provided by an embodiment of the present application.
[0009] Figure 2 It is a step flowchart of obtaining the first, second, and third latent subspace representations provided by an embodiment of the present application.
[0010] Figure 3 It is a step flowchart of generating the first latent subspace representation provided by an embodiment of the present application.
[0011] Figure 4 It is a step flowchart of generating the second latent subspace representation provided by an embodiment of the present application. Detailed Description of the Invention
[0012] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0013] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0014] It is found in the research that the entropy coding efficiency is restricted by the insufficient modeling of spatio-temporal non-stationarity by traditional models, and there are problems such as limited temporal receptive field, high computational complexity and decoding delay when applying deep learning models and Transformer models. In the reconstruction enhancement technology, the scheme based on the generative adversarial network faces the dual challenges of motion consistency and training stability, and also causes problems such as motion artifacts, SSIM degradation in flat regions and false edge generation. The fixed three-dimensional wavelet structure cannot adapt to the multi-granularity characteristics of motion; there is a lack of spatio-temporal co-optimization between traditional entropy models and generative reconstruction.
[0015] The present invention aims to solve the inherent contradiction existing in the existing video compression technology between efficiently removing spatio-temporal redundancy and precisely retaining motion details, and proposes a multi-granularity generative video compression method based on hierarchical latent space and spatio-temporal decoupling. Current mainstream video compression standards, such as HEVC and AV1, although have made remarkable progress in coding efficiency, their block-based hybrid coding frameworks often couple the processing of static background information and dynamic motion information, resulting in the background information being redundantly encoded in videos containing large areas of static regions (such as the sky or fixed background in surveillance scenarios), reducing the compression efficiency; at the same time, for complex or non-linear motions (such as rapid rotation, scale change, non-rigid deformation), relying on prediction residuals and simplified motion vector representations, it is difficult to precisely reconstruct the details of moving objects, often resulting in motion blur or artifacts; in addition, existing entropy coding models, such as CABAC, although efficient, are insufficient in adaptability when dealing with the complex spatio-temporal non-stationary correlations of video signals, especially for the latent representations transformed by deep learning models, whose statistical characteristics may be very different from traditional prediction residuals, resulting in the failure to fully exploit the compression potential, especially the low compression efficiency of residual detail information at low bitrates. The present invention aims to improve the compression performance, especially the motion fidelity and visual quality at low bitrates, by introducing spatio-temporal decoupled hierarchical latent representations, a hybrid entropy coding mechanism, and a generative reconstruction network.
[0016] As Figure 1 shown, a multi-granularity generative video compression method is proposed, including the following steps: Obtain a video frame sequence, perform spatio-temporal feature decomposition on the video frame sequence, decompose and map the video signal into three latent subspaces, and obtain the first, second, and third latent subspace representations, where the first latent subspace representation is the low-frequency spatial information and time-varying components in the video, the second latent subspace representation is the spatio-temporal motion features in the video, and the third latent subspace representation is the high-frequency detail information in the video; specifically, the static background stream, dynamic motion stream, and residual detail stream correspond to the first, second, and third latent subspaces; Encode the first, second, and third latent subspace representations respectively to generate corresponding first, second, and third encoded data; combine the first, second, and third encoded data to generate a compressed bitstream.
[0017] Specifically, an input video frame sequence is obtained, which can be video data with any resolution and frame rate. The spatio-temporal feature decomposition is performed on the video frame sequence, and the video signal is decomposed and mapped into three latent subspaces, specifically including: the first latent subspace is used to represent the low-frequency spatial information and slow temporal variation components (such as static background) in the video, the second latent subspace is used to represent the spatio-temporal motion features (such as object motion trajectories) in the video, and the third latent subspace is used to represent the high-frequency detail information (such as texture edges) in the video. Traditional video compression processes all video components mixedly, resulting in repeated encoding of static backgrounds and easy loss of motion details. Based on the physical characteristics of video content, this embodiment decomposes it into three orthogonal components: the first latent subspace represents low-frequency static information (such as a fixed background), the second latent subspace represents medium-frequency motion information (such as object trajectories), and the third latent subspace represents high-frequency detail information (such as texture edges). This decomposition enables each component to be independently optimized for compression. The representations of the three latent subspaces are respectively encoded, and different encoding strategies are adopted according to the characteristics of different subspaces to generate corresponding first encoded data, second encoded data, and third encoded data. The three encoded data are combined to generate the final compressed bitstream.
[0018] Through the above steps, in this embodiment, by decomposing the video into three semantically different latent subspaces and encoding them separately, an efficient content-adaptive compression effect is achieved, solving the technical problems of repeated encoding of static backgrounds and insufficient fidelity of motion details in traditional methods, and realizing efficient content-adaptive compression. Experimental data shows that at the same visual quality, the bit rate is reduced by about 45% compared with High Efficiency Video Coding (HEVC), improving the compression efficiency; the Peak Signal-to-Noise Ratio (PSNR) in complex motion scenarios is increased by about 2.1 dB; and at the same time, the storage occupancy is reduced.
[0019] As Figure 2 shown, according to one aspect of the present application, obtaining the first, second, and third latent subspace representations includes: Performing motion estimation on the video frame sequence to obtain motion information; Performing background modeling on the video frame sequence to obtain background information; Extracting spatio-temporal features from the video frame sequence based on the motion information and background information; Performing tensor decomposition on the spatio-temporal features, and mapping the tensor decomposition result into three approximately orthogonal latent subspaces, where the correlation between the latent subspaces is minimized through a pre-configured decomposition matrix to obtain the first, second, and third latent subspace representations.
[0020] In one embodiment of the present application, a RAFT (Recurrent All-Pairs Field Transforms) optical flow network is used to perform motion estimation on an input video frame sequence, and a dense optical flow field between adjacent frames is calculated as motion information. The motion information (optical flow field) F m = {(u(x, y), v(x, y)) | for all (x, y) ∈ Frame}, where u(x, y) is the velocity component of the pixel point (x, y) in the horizontal direction; v(x, y) is the velocity component of the pixel point (x, y) in the vertical direction, and Frame is a single-frame image. Applied to a surveillance video with a resolution of 1920×1080, the RAFT network outputs an optical flow field F of the same size m . For example, for the pixel point (100, 200) in a static area, u(100, 200)≈0 and v(100, 200)≈0; for the pixel point (500, 600) in a moving vehicle area, it may be that u(500, 600)=5.2 pixels / frame and v(500, 600)=-1.3 pixels / frame. Background modeling is performed on the video frame sequence in parallel, and stable background information F is extracted through temporal median filtering or Gaussian mixture model (GMM) b . For the above-mentioned surveillance video with a resolution of 1920×1080, background modeling is performed using the temporal median filtering method. Select 32 consecutive frames as the time window, and calculate the median of each pixel position in the time dimension. For example, for the pixel point (100, 200) in a static background area, its pixel value sequence in 32 frames is [120, 118, 121, 119, 120, 122, 119, 121,...], and the stable background value F b (100, 200) = 120 is obtained through median filtering. For the pixel point (500, 600) in a road area where vehicles occasionally pass by, its pixel value sequence is [85, 87, 85, 180, 182, 178, 86, 85,...], where 180 - 182 are the pixel values when a vehicle passes by. After median filtering, the background value F b (500, 600) = 86 is obtained, effectively eliminating the influence of temporary occlusions. Background modeling verification is performed in parallel using the Gaussian mixture model (GMM). A mixture model with 3 Gaussian components is established for each pixel position. Taking the pixel point (300, 400) as an example, the background component parameters are obtained through training: weight π1 = 0.85, mean μ1 = 95, variance σ1 2 = 4 (main background); weight π2 = 0.12, mean μ2 = 180, variance σ2 2 = 25 (shadow change); weight π3 = 0.03, mean μ3 = 220, variance σ3 2= 100 (accidental foreground). Extract the component with the largest weight as background information to obtain F b (300, 400) = 95.
[0021] Based on the motion information F m and the background information F b , extract spatio-temporal features from the video frame sequence; specifically, by fusing information such as the difference between the original frame and the background, and the marking of the motion area. Based on the motion information F m and the background information F b , extract spatio-temporal features from the video frame sequence. Calculate the difference map F diff = |F original - F b |. For the pixel point (100, 200) in the static area, the original frame value is 119, the background value is 120, and the difference value F diff (100, 200) = |119 - 120| = 1, indicating that this area is basically static. For the pixel point (500, 600) in the moving vehicle area, the original frame value is 182, the background value is 86, and the difference value F diff (500, 600) = |182 - 86| = 96, indicating significant motion. Combine the motion amplitude ||F m (x, y)||2 to calculate the motion area marking. For the pixel point (500, 600), the motion amplitude ||F m (500, 600)||2 = sqrt(5.2 2 + (-1.3) 2 ) = 5.36 pixels / frame. Set the threshold τ = 1.0 pixel / frame. Since 5.36 > 1.0, this point is marked as the motion area M mask (500, 600) = 1; while the motion amplitude of the pixel point (100, 200) in the static area is close to 0 and is marked as M mask (100, 200) = 0. The spatio-temporal feature matrix F st is constructed as a four-channel feature: F st = [F original , F b , F diff , M mask , with a dimension of 1920×1080×4×T (T is the time length of 16 frames). Perform tensor decomposition on the spatio-temporal features, design the decomposition matrix to minimize the correlation between the three subspaces, and achieve an approximately orthogonal decomposition. For the spatio-temporal feature tensor F stPerform a three-mode tensor decomposition on (dimension 1920×1080×4×16), and design a decomposition matrix to achieve approximate orthogonal decomposition. Reshape the feature tensor into a two-dimensional matrix form for singular value decomposition (SVD) preprocessing, and calculate the principal components along the spatial dimension (reshaping 1920×1080 to 2073600×1), the feature dimension (4×1), and the time dimension (16×1) respectively. Perform SVD decomposition on the 2073600×64 matrix (4×16 = 64), retain the first 512 principal components, and obtain the spatial projection matrix U spatial (2073600×512). Feature-time decomposition: Perform SVD on the 64×2073600 matrix, retain the first 32 principal components, and obtain the feature-time projection matrix U feat_time (64×32). Minimize the correlation of the three subspaces through orthogonalization processing. Calculate the correlation matrix: The correlation coefficient ρ12 between the first subspace (static background) and the second subspace (dynamic motion) is 0.08; the correlation coefficient ρ13 between the first subspace and the third subspace (residual details) is 0.12; the correlation coefficient ρ23 between the second subspace and the third subspace is 0.15; Apply the Gram-Schmidt orthogonalization process to reduce the correlation coefficient below 0.05. The first latent subspace (static background flow): Generated by spatial low-frequency filtering and time stability weighting. For the pixel point (100, 200), the time stability weight w(100, 200) = exp(-||F m (100, 200)||2 / σ) = exp(-0.1 / 2.0) = 0.95, and this point is mainly retained in the first subspace. The dimension of the first subspace is compressed to 240×135×8 (1 / 8 of the original). The second latent subspace (dynamic motion flow): Generated based on the motion vector field decomposition. For the pixel point (500, 600) in the moving vehicle area, its motion components are mainly mapped to the second subspace. Extract 5 main motion modes through CP decomposition. The weight of the first mode α1 = 0.6 (horizontal motion), the weight of the second mode α2 = 0.3 (vertical motion), and the weights of the remaining modes are smaller. The third latent subspace (residual details flow): Calculate the predicted frame F pred = F b +Motion Compensate (F m ). For the pixel point (500, 600), the predicted value is 86 + 5.2×cos(θ) + (-1.3)×sin(θ) = 91 (where θ is the motion direction angle). The residual between the actual value 182 and the predicted value 91 is 91. After 3D wavelet transform, the high-frequency coefficients are mainly distributed in the HHH subband, and the energy threshold is set to τ = 10, retaining the coefficients with energy greater than 10.
[0022] Such as Figure 3As shown, according to one aspect of the present application, the generation of the first latent subspace representation includes: Perform temporal domain filtering on the background information to extract the temporally stable region; Extract the spatial low-frequency components from the video frame sequence; Calculate the motion amplitude at each pixel position based on the motion information, and mark the region where the motion amplitude is less than the preset threshold as the static region mask; Perform weighted fusion on the temporally stable region, the spatial low-frequency components, and the static region mask to generate the first latent subspace representation, where the weights for weighted fusion are adaptively determined according to the temporal stability of each region.
[0023] In one embodiment of the present application, perform temporal domain filtering on the background information to retain the temporally stable regions: for each pixel position, take the median of 16 consecutive frames as the stable background; extract the spatial low-frequency components from the video frame sequence (e.g., through low-pass filtering); calculate the motion amplitude of each pixel based on the motion information ||F m (x, y)||2 = sqrt(u 2 (x, y) + v 2 (x, y)), and mark the region where the motion amplitude ||F m (x, y)||2 < ε (ε = 0.5 pixel / frame) as the static region; adaptively determine the weights according to the temporal stability, and the adaptive weight w(x, y) = exp(-||F m (x, y)||2 / σ), where σ is the scale parameter. Perform weighted fusion on the above information to generate the first latent subspace representation (static background flow). Specifically, for the sky region, it is almost unchanged in 16 consecutive frames, the stable value is retained after temporal filtering, the motion amplitude is close to 0, the weight is close to 1, and it is completely retained in the first subspace.
[0024] As Figure 4 shown, according to one aspect of the present application, the generation of the second latent subspace representation includes: Perform motion vector field decomposition on the motion information to extract the main motion components; Perform temporal difference operation on the video frame sequence to identify the motion regions; Based on the main motion components and the motion regions, extract the intermediate-frequency spatio-temporal features, where different motion patterns with different speeds and directions are captured through multi-scale motion analysis; Map the intermediate-frequency spatio-temporal features to the second latent subspace to generate the second latent subspace representation, where the mapping maintains the temporal continuity of the motion trajectory.
[0025] In one embodiment of the present application, perform motion vector field decomposition on the optical flow field, and extract the main motion patterns through principal component analysis (PCA): decompose the complex motion into K main patterns: F m≈ Σ k=1 K α k ·M k ; where K is the number of main motion components (typical value 5); α k is the weight coefficient of the k-th motion pattern; M k is the k-th main motion pattern. In the scenario of pedestrian walking, it includes two main patterns: translation (M1) and swing (M2), with weight α1 = 0.8 and α2 = 0.2. Calculate the inter-frame time difference, identify the significant motion area; capture motions at different speeds through multi-scale analysis (such as pyramid structure); maintain the time continuity of the motion trajectory to generate the second latent subspace representation (dynamic motion flow).
[0026] In an embodiment of the present application, perform an inter-frame time difference operation on a pedestrian walking video sequence with a resolution of 1920×1080. Calculate the pixel difference between adjacent frames: T diff (x, y, t) = |F(x, y, t) - F(x, y, t - 1)|. For the pixel point (800, 400) in the pedestrian body area, the pixel value of the t-th frame is 145, and the pixel value of the (t - 1)-th frame is 118. The difference value T diff (800, 400, t) = |145 - 118| = 27. Set the motion detection threshold T th = 15. Since 27 > 15, this pixel point is marked as a motion area. For the pixel point (200, 300) in the background stationary area, the pixel values of consecutive frames are 92, 94, 93 respectively, and the maximum difference value is 2, which is less than the threshold 15 and is marked as a stationary area. Apply morphological operations to remove noise: perform an opening operation with a 3×3 structuring element to remove isolated noise points, and perform a closing operation with a 5×5 structuring element to fill the holes inside the motion area. Obtain the motion area mask R mask , where the pedestrian contour area R mask = 1, and the background area R mask = 0. Further calculate the second derivative of the time difference to identify the motion boundary: T diff2 (x, y, t) = |T diff (x, y, t) - T diff (x, y, t - 1)|. At the pedestrian edge pixel point (820, 380), the first-order difference sequence is [5, 23, 18, 25, 8], and the second-order differences are 18, 15, 7, 17, indicating that there are continuous motion boundary changes in this area, which is marked as a key motion edge area. Based on the main motion components {M1, M2} and the motion area R mask, extract intermediate-frequency spatio-temporal features. Construct a multi-scale pyramid structure and downsample the original 1920×1080 frames to 960×540 (layer 1), 480×270 (layer 2), and 240×135 (layer 3) in sequence. Analyze medium-speed motion in layer 1 (960×540): For the pixel point (400, 200) in the pedestrian torso area (corresponding to the original coordinates 800, 400), the horizontal component of the main motion component M1 is u1 = 3.2 pixels / frame, and the vertical component is v1 = 0.8 pixels / frame, representing the overall translational motion. Calculate the motion consistency index Consistency1 = 0.85 for this layer, indicating relatively stable motion. Analyze fast motion in layer 2 (480×270): For the pixel point (200, 110) in the pedestrian leg swing area, the swing amplitude of the main motion component M2 is Amplitude = 4.5 pixels / frame, and the swing frequency is f = 2.1 Hz, corresponding to the normal walking frequency. This layer captures the periodic swing pattern, and the motion pattern weight α2 reaches 0.35 in this area. Analyze the overall motion trend in layer 3 (240×135): Calculate the overall motion direction of the pedestrian θ = arctan(v avg / u avg ) = arctan(0.9 / 3.8) = 13.3°, indicating that the pedestrian moves upward and to the right at a small angle. The magnitude of the motion speed is ||V|| = sqrt(3.8 2 + 0.9 2 ) = 3.9 pixels / frame. Construct the intermediate-frequency spatio-temporal feature tensor F mid , with dimensions of 240×135×8×16 (space × feature channels × time). The 8 feature channels include: channels 1-2: motion components (u1, v1) in layer 1; channels 3-4: motion components (u2, v2) in layer 2; channels 5-6: motion components (u3, v3) in layer 3; channel 7: motion amplitude ||V||; channel 8: motion direction angle θ.
[0027] In an embodiment of the present application, the time continuity of the motion trajectory is maintained by a Kalman filter. Establish a state vector State(t) = [x(t), y(t), vx(t), vy(t)] for the pedestrian center point trajectory T, where (x, y) are the position coordinates and (vx, vy) are the velocity components. The state transition matrix A = [[1, 0, Δt, 0], [0, 1, 0, Δt], [0, 0, 1, 0], [0, 0, 0, 1]], where Δt = 1 / 30 second (30fps video). The observation matrix H = [[1, 0, 0, 0], [0, 1, 0, 0]]. At the t-th frame, the observed center position of the pedestrian is (805, 405), the predicted position is (803, 404), and the Kalman gain K is calculated to obtain the corrected position of (804, 404.5). The velocity is updated to vx(t) = 3.5 pixels / frame and vy(t) = 1.0 pixel / frame, which is basically consistent with the main motion component M1 obtained by PCA decomposition. The 3rd-order B-spline interpolation is applied to smooth the motion trajectory to eliminate the trajectory jumps caused by occlusion or detection errors. During the 12th - 14th frames, due to partial occlusion, the detection position is offset. The original trajectory is [(798, 402), (820, 398), (810, 406)], and after smoothing, it becomes [(798, 402), (806, 401), (810, 405)], maintaining the continuity of the trajectory. The processed mid-frequency spatio-temporal feature F mid is mapped to the second latent subspace. Using the non-linear mapping function f(·), the feature tensor of 240×135×8×16 is compressed into a latent representation of 64×64×5×16. The design of the mapping weight matrix W follows the motion preservation principle: weight matrix W1 (240×64): spatial dimension compression, maintaining the spatial layout of the motion area; weight matrix W2 (135×64): spatial dimension compression, maintaining motion continuity; weight matrix W3 (8×5): feature dimension compression, compressing 8 multi-scale features into 5 main motion patterns; In the latent space, the main motion information of the pedestrian is encoded as: the 1st dimension: overall translational motion, coefficient range [-2.5, 2.5]; the 2nd dimension: periodic swing, coefficient range [-1.8, 1.8]; the 3rd dimension: change in motion acceleration, coefficient range [-0.8, 0.8]; the 4th dimension: change in direction, coefficient range [-0.5, 0.5]; the 5th dimension: motion uncertainty, coefficient range [-0.3, 0.3]; For the representation of the pedestrian torso area in the latent space: L2(32, 32, 1, t) = 2.1 (strong translational motion); the representation of the leg swing area is: L2(28, 35, 2, t) = 1.6×sin(2π×2.1×t / 30) (periodic swing). Add a temporal smoothing regularization term L smooth = Σt ||L2(·, ·, ·, t) - L2(·, ·, ·, t - 1)|| 2 , with weight λ smooth= 0.1 to ensure smooth changes in the latent representation over the time dimension and avoid unnatural jumps in the motion trajectory.
[0028] According to one aspect of the present application, the generation of the third latent subspace representation includes: A predicted frame reconstructed based on the first and second latent subspace representations; Calculating the residual information between the video frame sequence and the predicted frame; Performing a three-dimensional wavelet transform on the residual information, performing multi-scale decomposition in the time and space dimensions, and extracting high-frequency wavelet coefficients; Mapping the high-frequency wavelet coefficients to the third latent subspace to generate the third latent subspace representation, where the mapping preserves the detailed features through sparse constraints.
[0029] In one embodiment of the present application, the predicted frame F_pred = F is reconstructed based on the first two subspace representations b +Motion_Compensate(F m ), where Motion_Compensate is a motion compensation function; calculating the residual F between the original frame and the predicted frame r = F_original - F_pred, where F_original is the original video frame; performing a three-dimensional discrete wavelet transform (3D-DWT) on the residual, using the Daubechies-4 wavelet basis, and performing 3-layer decomposition: Layer 1: generating 8 subbands (LLL1, LLH1,..., HHH1); Layer 2: continuing to decompose LLL1; Layer 3: continuing to decompose LLL2; extracting high-frequency wavelet coefficients, mapping them to the third latent subspace through sparse constraints (such as L1 regularization), retaining the coefficients with energy greater than the threshold τ, and obtaining the third latent subspace representation (residual detail stream). The high-frequency coefficients are mainly distributed in the motion edges and texture regions, and the unimportant noise coefficients are removed through sparse constraints, retaining the visually significant details.
[0030] In another embodiment of the present application, the predicted frame is reconstructed based on the first latent subspace (static background stream) and the second latent subspace (dynamic motion stream). The static background stream is reconstructed through Tucker decomposition: F b_recon = G ×1 U1 ×2 U2 ×3 U3, obtaining the background prediction. For the pixel point (200, 300), the reconstructed background value F b (200, 300) = 92. The motion compensation function Motion Compensate(Fm) is implemented using bilinear interpolation: For the pixel point (800, 400) in the motion area, according to the optical flow vector (u = 3.5, v = 1.0), interpolation sampling is performed from the previous frame position (796.5, 399). The values of the four neighboring pixels in the previous frame are: I(796, 399) = 118, I(797, 399) = 120, I(796, 400) = 115, I(797, 400) = 119. Bilinear interpolation calculation: Motion Compensate (800, 400) = (1 - 0.5)×(1 - 0)×118 + 0.5×(1 - 0)×120 + (1 - 0.5)×0×115 + 0.5×0×119 = 119. The predicted frame synthesis F pred (800, 400) = F b (800, 400) + Motion Compensate (800, 400) = 92 + 119 = 211. Calculate the residual between the original frame and the predicted frame. For the pixel point (800, 400) of the pedestrian's body, the value of the original frame F original (800, 400) = 145, the value of the predicted frame F pred (800, 400) = 211, the residual F r (800, 400) = 145 - 211 = -66. For the pixel point (200, 300) in the background area, the value of the original frame is 94, the value of the predicted frame is 92, and the residual F r (200, 300) = 94 - 92 = 2, indicating that the residual in the background area is small. For the pixel point (820, 380) of the pedestrian's edge, due to the complexity of the motion boundary, the original value is 162, the predicted value is 135, and the residual F r (820, 380) = 162 - 135 = 27, indicating that there is a large prediction error in the edge area and compensation is required through residual information.
[0031] In an embodiment of the present application, for a residual sequence F with a resolution of 240×135 for 16 frames rPerform 3D-DWT decomposition using the Daubechies-4 wavelet basis. First layer decomposition: Apply the wavelet filter bank {h0, h1, g0, g1}, where h0 and h1 are the low-pass and high-pass decomposition filters, and g0 and g1 are the reconstruction filters. The Daubechies-4 filter coefficients are: h0 = [0.683, 1.183, 0.317, -0.183] (low-pass); h1 = [-0.183, -0.317, 1.183, -0.683] (high-pass); Decomposition along the spatial dimension (height): Convolve each row of pixels and downsample by a factor of 2 to obtain the low-frequency L and high-frequency H components. The spatial dimension changes from 240×135 to 120×135. Decomposition along the spatial dimension (width): Continue to decompose the width dimension to obtain four spatial subbands: LL, LH, HL, and HH. The spatial dimension becomes 120×68. Decomposition along the temporal dimension: Decompose the 16-frame time series to obtain 8 subbands: LLL1: low-frequency space + low-frequency time, dimension 120×68×8; LLH1: low-frequency space + high-frequency time, dimension 120×68×8; LHL1: low-high-frequency space + low-frequency time, dimension 120×68×8; LHH1: low-high-frequency space + high-frequency time, dimension 120×68×8; HLL1: high-low-frequency space + low-frequency time, dimension 120×68×8; HLH1: high-low-frequency space + high-frequency time, dimension 120×68×8; HHL1: high-frequency space + low-frequency time, dimension 120×68×8; HHH1: high-frequency space + high-frequency time, dimension 120×68×8; Second layer decomposition: Continue to decompose the LLL1 subband. The spatial dimension becomes 60×34×4, generating 8 second-level subbands: LLL2, LLH2, LHL2, LHH2, HLL2, HLH2, HHL2, HHH2. Third layer decomposition: Continue to decompose the LLL2 subband. The spatial dimension becomes 30×17×2, generating 8 third-level subbands: LLL3, LLH3, LHL3, LHH3, HLL3, HLH3, HHL3, HHH3.
[0032] In one embodiment of the present application, significant coefficients are extracted from 21 high-frequency subbands. Taking the HHH1 subband as an example, this subband captures the high-frequency changes in space and time and is mainly distributed in the motion edge region. For the pedestrian edge position, the corresponding coordinates in the HHH1 subband are (41, 19, 3), and the wavelet coefficient value is 23.7. Calculate the coefficient energy: Energy = |23.7| 2= 561.69. Set the energy threshold τ = 25. Since 561.69 > 25, this coefficient is retained. For the background region, the corresponding coordinates in the HHH1 sub-band are (10, 15, 3), the wavelet coefficient value is 1.8, and the energy is 3.24 < 25, which is sparsified to 0. HHH1 sub-band: The proportion of retained coefficients is 12.3%, mainly at the motion edges; HHL1 sub-band: The proportion of retained coefficients is 8.7%, mainly at the horizontal edges; HLH1 sub-band: The proportion of retained coefficients is 6.2%, mainly at the vertical edges; LHH1 sub-band: The proportion of retained coefficients is 9.1%, mainly in the time-varying region; Apply L1 regularization for sparse constraint: minimize ||W·C||1 + λ||C||1, where C is the wavelet coefficient vector, W is the mapping weight matrix, and λ = 0.01 is the sparse regularization parameter. Sort the coefficients of the HHH1 sub-band and retain the top 5% of the coefficients by energy. Taking the position (41, 19, 3) as an example, the coefficient 23.7 is processed by the soft-thresholding function: if |c| > λ, then c' = sign(c)×(|c| - λ) = sign(23.7)×(23.7 - 0.01) = 23.69; if |c| ≤ λ, then c' = 0; Reorganize the sparsified high-frequency coefficients into the third latent subspace representation. The original 21 high-frequency sub-bands have a total of about 6 million coefficients. After sparse constraint, about 600,000 non-zero coefficients are retained (sparsity rate 90%). Use Run-Length Encoding to record the positions of non-zero coefficients: Position indices: [245, 1203, 1847, 2956,...]; Coefficient values: [23.69, -15.2, 8.7, 31.4,...]; The compact representation of the third latent subspace includes: Sparse position vector: 600,000 dimensions, recording the positions of non-zero coefficients; Coefficient amplitude vector: 600,000 dimensions, recording the values of non-zero coefficients; Sub-band identification vector: 600,000 dimensions, identifying the sub-bands to which the coefficients belong; Through this sparse representation, the third latent subspace effectively captures the high-frequency detail information in the video, including motion edges, texture changes, and prediction errors.
[0033] According to one aspect of the present application, generating the first encoded data includes: Construct the first latent subspace representation as a three-dimensional tensor; Perform Tucker decomposition on the three-dimensional tensor to generate a core tensor and factor matrices along the spatial and temporal dimensions, where the dimensions of the core tensor are set to the predetermined compression ratios of the respective dimensions of the original tensor; Combine the core tensor and the factor matrices to form a compressed first latent subspace representation; Quantize and entropy-encode the compressed first latent subspace representation to generate the first encoded data.
[0034] In one embodiment of the present application, the first latent subspace representation is constructed as a three-dimensional tensor of Height×Width×Time (dimension H×W×T); perform Tucker decomposition: F b ≈ G ×1U1×2U2×3U3; where the dimension of the core tensor G is set to 1 / 8 of the original dimensions, i.e., (H / 8)×(W / 8)×(T / 8); U1 is the factor matrix in the height direction, with dimension H×(H / 8); U2 is the factor matrix in the width direction, with dimension W×(W / 8); U3 is the factor matrix in the time direction, with dimension T×(T / 8); the factor matrices U1, U2, and U3 are solved by the alternating least squares (ALS) method; × n is the modulo n product operation; quantize the core tensor and the factor matrices, and the typical quantization step is 4 bits. The original data volume is H×W×T, and the compressed data volume is approximately 1 / 512 + 3 / 8 ≈ 0.88% of the original. For a background tensor of 1920×1080×16 (about 132MB), after Tucker decomposition: the core tensor G is 240×135×2 (about 0.26MB); the three factor matrices together are about 12.4MB; the total compressed size is 12.66MB, and the compression ratio is 90.4%.
[0035] According to one aspect of the present application, generating the second encoded data includes: Construct the second latent subspace representation as a motion tensor; Perform CP decomposition on the motion tensor and represent it as a weighted sum of a predetermined number of one-dimensional factor vectors, where each group of one-dimensional factor vectors captures a main motion pattern; Extract the spatial factor vectors, temporal factor vectors, and the corresponding decomposition weights of the CP decomposition to form the compressed second latent subspace representation; Perform adaptive quantization and entropy coding on the compressed second latent subspace representation to generate the second encoded data.
[0036] In one embodiment of the present application, construct the motion flow as a motion tensor; perform CP decomposition on the motion tensor F m and represent it as: F m ≈ Σ r=1 R λ r ·a r Θb r Θc r ; where R is the decomposition rank (set to 5); λ r is the weight of the r-th component; a r is the spatial height factor vector (length H); b r is the spatial width factor vector (length W); c ris the time factor vector (length T); Θ is the outer product of vectors; each rank-1 component represents a major motion pattern: the 1st component: horizontal translation motion (λ1 is the largest); the 2nd component: vertical translation motion; the 3rd component: rotational motion; the 4th component: scaling motion; the 5th component: complex local deformation. The spatial factor vectors a r , b r and the time factor vector c r are optimized by gradient descent; the factor vectors are adaptively quantized, and more bits (6 - 8 bits) are allocated to regions with intense motion.
[0037] According to one aspect of the present application, the compressed first latent subspace representation is quantized and entropy - encoded to generate first encoded data, including: An autoregressive model constructed by using a convolutional neural network processes the compressed first latent subspace representation element - by - element in a predetermined scanning order; For the current element to be encoded, its encoded spatial neighborhood is extracted as the context; Based on the context, the conditional probability distribution of the current element to be encoded is predicted through the autoregressive model; The conditional probability distribution is used to guide the arithmetic encoder to encode the current element to be encoded, generating the first encoded data.
[0038] In an embodiment of the present application, a 5 - layer convolutional neural network is constructed as the autoregressive model: the input layer is a 3×3×C context window (C is the number of channels); the first hidden layer has 64 3×3 convolutional kernels with ReLU activation; the second to fourth hidden layers have 128 3×3 convolutional kernels with residual connections; the output layer has 256 probability values (8 - bit quantization). The raster scan order is adopted to process the compressed background stream element - by - element; for the element at position (i, j), a 3×3 neighborhood is extracted as the spatial context. Conditional probability modeling is performed: P(x ij |Context) = Π_{c∈C} P(x ij c |x_{<ij}); where x ij is the coefficient to be encoded at position (i, j); Context is the encoded neighborhood {x mn |m < i or (m = i and n < j)}; C is the set of channels; the CNN predicts the conditional probability distribution N(μ ij , σ ij 2 ). Where P() is the conditional probability function, Π is the product symbol; μ ij is the mean of the Gaussian distribution predicted by the neural network; σ ij 2is the variance of the predicted Gaussian distribution. The arithmetic coder encodes according to the predicted probability distribution, achieving compression close to the entropy limit. Compared with the fixed probability model, adaptive prediction reduces the entropy by about 1.2 bits / coefficient.
[0039] According to one aspect of the present application, adaptive quantization and entropy coding are performed on the compressed second latent subspace representation to generate second encoded data, including: Construct the compressed second latent subspace representation into a spatio-temporal sequence; Extract the spatio-temporal context of the current position to be encoded from the spatio-temporal sequence, including historical information in the time dimension and neighborhood information in the spatial dimension; Use a Transformer encoder to process the spatio-temporal context, and capture the long-range dependencies of the motion pattern through the self-attention mechanism; Predict the conditional probability distribution of the current position to be encoded based on the long-range dependencies, and use the conditional probability distribution to guide the arithmetic coder to generate the second encoded data.
[0040] In an embodiment of the present application, the compressed motion flow is constructed into a spatio-temporal sequence, and the sequence length is typically 64; a 6-layer Transformer encoder is used, with each layer containing 8 attention heads; a windowed attention mechanism is used, and the window size is 8×8×4 (spatial×time); the self-attention mechanism captures the long-range dependencies of the motion, such as periodic motion patterns; the output layer predicts the Gaussian mixture distribution parameters to guide arithmetic coding. The attention calculation formula is Attention(Q, K, V) = softmax(QK T / sqrt(d)) · V; where Q is the query matrix (current position feature); K is the key matrix (context feature); V is the value matrix (context information); d is the feature dimension (512). Periodic motion (such as fan rotation) has a strong correlation between the t-th frame and the (t - T)-th frame. The Transformer directly establishes this long-range connection through self-attention, while the traditional context-adaptive binary arithmetic coding (CABAC) can only utilize the local Markov assumption. Since techniques such as the Transformer framework, windowed attention, spatio-temporal transformer, and conditional probability modeling belong to the prior art, they will not be elaborated here.
[0041] According to one aspect of the present application, generating the third encoded data includes: Classify the wavelet coefficients in the third latent subspace representation to identify the positions of zero coefficients and the values of non-zero coefficients; Perform sparse position coding on the positions of zero coefficients to record the spatial distribution of non-zero coefficients; Model the values of non-zero coefficients using a Gaussian mixture model, where different mixture components are respectively fitted for positive and negative value coefficients to obtain the probability distribution parameters; Arithmetic coding is performed on non-zero coefficients based on probability distribution parameters and combined with sparse position coding to generate third encoded data.
[0042] In one embodiment of the present application, the wavelet coefficient distribution is analyzed. Typically, more than 90% are zero coefficients; run-length encoding (RLE) is used for the positions of zero coefficients; a three-component Gaussian mixture model is used for non-zero coefficients: positive coefficients: two Gaussian components (peak and tail); negative coefficients: one Gaussian component; the wavelet coefficient distribution is modeled as P(h) = Σ k=1 3 π k · N(h|μ k ,σ k 2 ); zero peak component: π1≈0.9, μ1 = 0, σ1 = 0.1; positive tail component: π2≈0.07, μ2 = 2.5, σ2 = 1.2; negative tail component: π3≈0.03, μ3 = -2.0, σ3 = 0.8; where P(h) is the probability density function of the wavelet coefficient h, and N(h|μ k ,σ k 2 ) represents a one-dimensional Gaussian distribution (normal distribution) with μ k as the mean and σ k 2 as the variance, corresponding to the k-th component. The GMM parameters are estimated by the expectation-maximization (EM) algorithm; arithmetic coding is performed based on the GMM probability, and the sparse position coding and coefficient coding are combined for output. Sparse representation enables 90% of the coefficients to be encoded with 1 bit (representing zero), and the remaining 10% of non-zero coefficients are on average 4 bits, with an overall average of 1.3 bits / coefficient.
[0043] Optionally, generating the third encoded data can also be: Analyze the statistical characteristics of wavelet coefficients in the third latent subspace representation and identify their sparse distribution patterns; For the distribution characteristics of non-zero coefficients, a Gaussian mixture model is used for fitting, where asymmetric mixture components are established for positive and negative coefficients respectively; Predict the probability distribution parameters of each coefficient based on the Gaussian mixture model; Use the probability distribution parameters to guide the arithmetic encoder and encode the third latent subspace representation in combination with the sparsity constraint to generate the third encoded data.
[0044] According to one aspect of the present application, it further includes decoding and reconstructing the compressed code stream, including: Decode the compressed code stream to recover the first decoded representation, the second decoded representation, and the third decoded representation; extracting motion information from the second decoded representation, the motion information representing a direction and magnitude of motion in the video; Generate spatial attention weights based on motion information, which indicate the motion areas that need to be focused on during reconstruction; A deformable convolutional reconstruction network is used to dynamically adjust the sampling position of the convolution kernel according to the attention weight, and the first decoded representation, the second decoded representation and the third decoded representation are fused to generate a reconstructed video frame.
[0045] In one embodiment of the present application, the compressed bitstream is entropy decoded to restore three quantized potential representations; motion information in the form of optical flow is extracted from the second decoded representation; a spatial attention weight map is generated based on the motion information, and higher weights are given to areas with intense motion; a deformable convolution reconstruction network is used: the network includes 5 deformable convolution layers and 3 upsampling layers; the convolution kernel dynamically adjusts the sampling position according to the optical flow vector: offset = Conv(motion_flow) deform_conv(x, offset); where motion_flow is the optical flow information, Conv() is a standard convolution operation, offset is a dynamic offset vector, deform_conv is a deformable convolution operation, and x is a feature map input to the deformable convolution layer. The three decoded representations are cascaded and fused: the background is reconstructed first, the motion is superimposed, and the details are added. Generator loss function: L_total = L_adv +0.1×L_perceptual + 0.05×L_temporal, where L_adv is the adversarial loss, using the PatchGAN discriminator; L_perceptual is the VGG-19 feature matching loss; L_temporal is the temporal consistency loss. Experiments show that motion-aware reconstruction improves PSNR by 2.1dB in complex motion scenes compared to traditional methods.
[0046] In another embodiment, the mathematical principle of deformable convolution: standard convolution is y(p) = Σ n w(n)·x(p+n); deformable convolution is y(p) = Σ n w(n) x(p+n+Δn); where p is the output position; n is the convolution kernel sampling position; Δn is the offset determined by the optical flow; w(n) is the convolution weight; calculate the offset Δn = Conv_offset(F m) where Conv_offset is a dedicated offset prediction network. For a fast - moving football (speed 20 pixels / frame), traditional convolution will produce motion blur. Deformable convolution adjusts the sampling points according to the motion direction, samples along the motion trajectory, and maintains the clear contour of the sphere. Experiments show that the structural similarity index (SSIM) in the motion area is improved from 0.82 to 0.91. Cascade fusion process: Background reconstruction: G bg = UpSample(Decode(First decoded representation)); Motion superposition: G motion = G bg + Warp(Decode(Second decoded representation), F m ); Detail enhancement: G_final = G motion + IDWT(Decode(Third decoded representation)).
[0047] According to an aspect of the present application, before encoding and processing the three potential subspace representations respectively, it further includes: Analyze the content complexity of the first, second, and third potential subspace representations to obtain the corresponding first, second, and third complexities; Based on the three complexities, construct a rate - distortion prediction model to predict the distortion degree and bit rate under different quantization parameter combinations; Solve the optimal quantization parameter set that minimizes the weighted sum of the total distortion and the total bit rate through Lagrangian optimization; Apply the corresponding parameters in the optimal quantization parameter set to the encoding processing of the three potential subspace representations respectively to achieve adaptive bit allocation.
[0048] In an embodiment of the present application, calculate the spatial frequency and temporal stability of the background flow to obtain the first complexity; analyze the motion vector variance and motion edge density of the motion flow to obtain the second complexity; count the non - zero coefficient ratio and energy distribution of the residual flow to obtain the third complexity. Construct a rate - distortion model: D(Q) = α1·Q1 β1 + α2·Q2 β2 + α3·Q3 β3; Bitrate model: R(Q) = γ1·log(1 + δ1 / Q1) + γ2·log(1 + δ2 / Q2) + γ3·log(1 + δ3 / Q3); The model parameters are obtained by fitting with a training set. Where α1, α2, and α3 are the distortion weight coefficients of the background, motion, and residual sub-streams respectively; Q1, Q2, and Q3 are the quantization step sizes of the background stream, motion stream, and residual stream respectively; β1, β2, and β3 are the sensitivity indices of the distortion of the background, motion, and residual sub-streams to the quantization step size respectively; γ1, γ2, and γ3 are the bitrate weight coefficients of the three sub-streams respectively; δ1, δ2, and δ3 control the attenuation rate of the bitrate with respect to the quantization step size. Lagrangian optimization solution: Objective function: min J = D(Q) + λ·R(Q); The value range of λ is 0.3 - 1.2 and is adjusted according to the target bitrate; The interior point method is used to solve for the optimal quantization parameter set Q* = {Q1*, Q2*, Q3*}. Adaptive application of quantization parameters: Static background: Typical value of Q1* is 4 - 8 (high compression); Dynamic motion: Typical value of Q2* is 2 - 4 (fidelity first); Residual details: Typical value of Q3* is 8 - 16 (extremely high compression). Through the above optimization, the overall bitrate is reduced by 15% - 20% under the same quality, and local quality degradation caused by fixed allocation is avoided.
[0049] In another embodiment of the present application, the first complexity (background): C1 = α·Var_spatial(F b ) + β·Var_temporal(F b ); Where Var_spatial is the spatial variance, Var_temporal is the temporal variance, α = 0.7, β = 0.3. The second complexity (motion): C2 = γ·E[||F m ||2] + δ·Entropy(F m ); Where E[·] is the expected value, Entropy(·) is the motion vector entropy, γ = 0.6, δ = 0.4. The third complexity (residual): C3 = ε·NNZ(F r ) / Total(F r ); Where NNZ is the number of non-zero coefficients, Total is the total number of coefficients, ε = 1.0. Optimization solution process: Lagrangian function: L = D(Q1, Q2, Q3) + λ·R(Q1, Q2, Q3); The gradient descent method is used to iteratively solve for Q i (t+1) = Q i (t) - η·ΨL / ΨQ i; where η = 0.01 is the learning rate, and the iteration continues until convergence; Ψ is the partial derivative. When the input target bit rate is 1 Mbps and λ = 0.7, the following results are calculated: Q1* = 6 (background), Q2* = 3 (motion), Q3* = 12 (residual), and the actual bit rate is 0.98 Mbps, with PSNR = 38.5 dB. Through dynamic optimization, compared with fixed allocation (Q1 = Q2 = Q3 = 5), the PSNR is increased by 1.8 dB at the same bit rate.
[0050] According to another aspect of the present application, a multi-granularity generative video compression method, the core of which lies in decomposing the video signal into multiple orthogonal or approximately orthogonal sub-streams that are easier to independently model and compress, and designing customized encoding and reconstruction strategies for each sub-stream. Specifically, it includes: performing preprocessing on the input video frame sequence, including motion estimation (e.g., using advanced optical flow estimation algorithms) and background modeling, initially separating the dynamic foreground and static background regions, and providing prior information for subsequent decomposition. Adopting a three-dimensional wavelet transform (3D-DWT) module, using fast algorithms such as Mallat to perform multi-scale and multi-directional time-frequency decomposition on video data blocks (space-time cubes), effectively extracting features spanning time and space dimensions. Based on tensor decomposition theory, map these multi-scale time-frequency features to three semantically different latent sub-spaces: one is the static background stream, mainly containing low-frequency spatial information and slowly varying components in time in the video, which can be extracted by averaging, clustering the low-frequency wavelet coefficients in time or combining the background modeling results. This stream changes slowly and is suitable for high compression ratio encoding; the second is the dynamic motion stream, capturing the intermediate-frequency space-time features in the video, mainly characterizing information such as the motion trajectory, speed, and deformation of objects, which can be generated or constrained by combining the results of the optical flow estimation network. This stream contains the main motion information and needs to be accurately encoded to ensure motion fidelity; the third is the residual detail stream, mainly reconstructed from high-frequency wavelet coefficients, containing information such as motion edges, texture details, and random noise that the model fails to capture. This stream usually has low energy but has a significant impact on visual quality. In order to efficiently compress the three separated sub-streams, a hybrid autoregressive-Transformer entropy model is constructed to adapt to the statistical characteristics and spatio-temporal correlations of different sub-streams. For the static background stream that is highly correlated in time but has a stable spatial structure, an autoregressive model based on a convolutional neural network (CNN) is used to predict the spatial context of the current latent representation, and its powerful local feature extraction ability is used to eliminate spatial redundancy. For the dynamic motion stream containing complex spatio-temporal dependencies, an entropy model based on Transformer is introduced, especially using the windowed or local attention mechanism (Windowed Attention) to effectively capture the long-range spatio-temporal dependencies between latent representations while controlling the computational complexity. This model can more accurately estimate the probability distribution of motion information. For the residual detail stream that usually exhibits sparse characteristics, a sparse coding strategy is adopted, such as adding an L1 regularization constraint to the loss function to encourage the sparsity of the latent representation, and combining an efficient entropy coding method (such as arithmetic coding) for compression.The probability distribution of the motion flow entropy model can be mathematically expressed as: P(m_t |m_(<t))=∏_(i,j)N(μ_ij (m_(<t)),σ_ij (m_(<t))), where m_t represents the latent representation of the dynamic motion flow at the current time t (e.g., a feature map or a tensor block), m_(<t) represents its temporal context information (i.e., all or part of the motion flow latent representations before time t), and i, j are indices in the latent representation space dimension. N represents the Gaussian (normal) distribution. μ_ij (m_(<t)) and σ_ij (m_(<t)) are the mean and standard deviation (or the square root of the variance) of this distribution at position (i, j), respectively. These two parameters are predicted by a Transformer-based network according to the context information m_(<t)) and are used to guide the arithmetic encoder to encode the quantization value of m_t. At the decoding end, a motion-aware generative adversarial network (GAN) is adopted for high-quality video reconstruction, with particular attention paid to the restoration of motion details. The compressed bitstreams of each sub-flow are restored to the quantized latent representations through the corresponding entropy decoder (arithmetic decoder). These restored latent representations are fed into a cascaded decoder architecture. In particular, to enhance the reconstruction quality of the motion regions, a part (or the entire decoding network) in the decoder is designed as the generator of the GAN. The generator receives the decoded information from the static background flow, the dynamic motion flow, and the residual detail flow as input or conditions. The key lies in using the decoded dynamic motion flow information (or the accurate optical flow information transmitted at the encoding end) to guide the attention mechanism of the generator. Specifically, an optical flow-guided deformable attention module can be adopted, enabling the sampling positions of the convolutional kernels to be adaptively adjusted according to the motion direction and amplitude indicated by the optical flow vectors, so as to focus more on the contours and internal details of the moving objects and effectively combat the blurring and artifacts introduced by compression. The cascaded decoder structure gradually fuses the information from the three sub-flows, first reconstructing the basic static scene using the background flow, then superimposing the motion flow information to restore the dynamic content, integrating the information of the residual detail flow for refinement, and finally synthesizing high-quality output video frames.
[0051] Compared with the prior art, this embodiment has achieved an improvement in compression efficiency. Experimental data shows that when reaching the same subjective or objective visual quality (such as SSIM value) as HEVC, this embodiment reduces the bitrate drop, which benefits from the efficient compression of static backgrounds and the accurate modeling of dynamic information. In terms of motion fidelity, by explicitly separating and encoding the dynamic motion flow and using an optical flow-guided deformable attention GAN for reconstruction, it is able to more precisely recover details in fast and complex motion scenes, and the peak signal-to-noise ratio (PSNR) is improved in these scenes. Although a deep learning model is introduced, through sub-flow decomposition, adaptive quantization, and parallel processing optimization on the GPU, the overall encoding speed is improved compared to some complex traditional or early learning-based encoders. In engineering applications, the dynamic bit allocation strategy significantly reduces the storage occupancy and can flexibly adapt to different application scenarios, such as greatly compressing static backgrounds in surveillance video storage or preferentially transmitting the motion flow in low-latency scenarios such as cloud gaming to ensure the interactive experience.
[0052] This embodiment proposes to map the video time-frequency features to three approximately orthogonal latent subspaces of static background, dynamic motion, and residual details through tensor decomposition, achieving a content-adaptive hierarchical representation, which lays the foundation for subsequent differential and efficient coding. For the statistical characteristics of different sub-flows, a hybrid entropy model is constructed, combining the spatial modeling ability of CNN, the spatio-temporal long-range dependence capture ability of Transformer, and sparse coding, which improves the compression efficiency of the latent representation. A motion-aware generative adversarial network is introduced at the decoding end, and the deformable attention mechanism is guided by optical flow information, effectively enhancing the reconstruction fidelity of motion details and especially improving the visual quality in complex motion scenes.
[0053] In a specific embodiment, based on the hierarchical latent space and multi-granularity generative framework, an efficient video compression is achieved through a precise processing flow. Specifically: preprocess the input video sequence, including GOP segmentation, normalization, and preliminarily extract the dense motion flow F using an optical flow network (such as RAFT) m , and separate the static background flow F by combining background modeling techniques b . The calculated residual information F r is applied with three-dimensional discrete wavelet transform (3D-DWT) for multi-scale time-frequency decomposition to extract the residual detail flow; use tensor decomposition techniques (Tucker decomposition to process F b , CP decomposition to process F mProject the background and motion flow onto a compact low-rank latent space. Construct a hybrid autoregressive-Transformer entropy model to specifically compress the latent representations of each sub-flow: the CNN autoregressive model processes the spatial correlation of the background flow, the spatio-temporal Transformer model captures the long-range dependencies of the motion flow, and the GMM model encodes the residual coefficients. Dynamically allocate bits through global rate-distortion optimization (Lagrangian minD + λR). At the decoding end, after entropy decoding, use a motion-aware generative adversarial network (GAN) for high-quality reconstruction. Its generator uses an optical flow-guided deformable attention mechanism to focus on restoring motion details and outputs the final frame by cascading and fusing the information of the three flows. The whole process aims to achieve the unity of efficient compression and high-fidelity motion recovery.
[0054] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following elaborates in detail the specific implementation steps of the multi-granularity generative video compression method: Step 1: Input, preprocess, and preliminarily separate the video sequence in space and time.
[0055] Perform standardization processing on the input original video data, and use a motion estimation algorithm to preliminarily separate significant dynamic motion information and relatively static background information, laying a foundation for subsequent more refined hierarchical latent space decomposition. The system receives the original video sequence. To facilitate parallel processing and utilize temporal correlation, it is segmented into consecutive groups of pictures (GOPs) with a preset time length (e.g., typically 16 frames). For each frame of image data within each GOP, perform pixel value normalization operations, such as linearly mapping it to the interval [-1, 1]. This helps improve the stability and convergence speed of subsequent deep learning model training. To specifically process color information and luminance information, usually convert the frames in the RGB color space to the YUV color space, and extract the luminance component (Y) and chrominance components (U, V) respectively. Subsequent processing can mainly target the luminance component or all components. Enter the crucial motion estimation stage (which can be regarded as part of the preprocessing module or immediately following it), and use an advanced dense optical flow estimation algorithm, such as the RAFT (Recurrent All-Pairs Field Transforms) network. This network can calculate the displacement vector of each pixel point between the current frame and the reference frame (usually the previous frame), forming a dense two-dimensional optical flow field F m . Here, F m ={(u(x, y), v(x, y))| all (x, y) ∈ Frame} represents a vector field with the same size as the original frame, where u(x, y) and v(x, y) are the instantaneous velocity estimates of the pixel point (x, y) in the horizontal and vertical directions respectively. This optical flow field F mThe main dynamic information in the video is preliminarily characterized. At the same time or based on the optical flow results, the static background flow F can be estimated by simple temporal filtering (such as pixel-level temporal median filtering) or more complex background modeling methods (such as Gaussian Mixture Model GMM or deep learning-based methods). b A simplified criterion is that for pixel position (x, y), if the optical flow magnitude \|(u, v)\|2 = sqrt(u 2 + v 2 ) within consecutive multiple frames is less than a preset small threshold ε (for example, ε = 0.5 pixel / frame), then this pixel is considered to belong to the static background region, constituting the preliminary static background map F b This step outputs the preliminarily separated background information F b and the dense optical flow field F m , as well as the original (or preprocessed) video frame sequence for subsequent modules to use.
[0056] Step 2: Spatiotemporal multi-scale decomposition of residual information based on 3D wavelet transform.
[0057] Process the remaining video information after the preliminary background and motion separation, that is, the residual frames, and perform time-frequency analysis on them using 3D Discrete Wavelet Transform (3D-DWT), aiming to extract multi-scale and multi-directional features across time and space dimensions. These features mainly include motion edges, texture details, and complex dynamic components that the model fails to capture. Calculate the residual frame sequence F r , which is defined as the original frame (after normalization and color space conversion, denoted as F_original minus the preliminarily estimated static background flow F b and the image content corresponding to the dynamic motion flow (the predicted frame F can be obtained by warping the previous frame with the optical flow F m ), then F m_pred = F_original - (F r + F b ), or for simplified processing, for example, directly consider F m_pred ) = F_original - F r ). This residual sequence F b represents the changing information in the video except for the stable background and the main motion. For this three-dimensional data block (spatial dimension × time dimension) F r r Apply the three-dimensional discrete wavelet transform (3D-DWT). This transform recursively performs filtering and downsampling operations along three dimensions (height, width, time) through a separable filter bank. Specifically, the fast Mallat algorithm can be used. The selected wavelet basis function has an important impact on performance. In this embodiment, the Daubechies 4 (db4) wavelet basis is selected because it has good compact support and certain regularity, which is beneficial to the localization of features. Set the decomposition level (Scale Level) to 3, which means that the video data block will be decomposed into three different-scale representations. In each layer of decomposition, a low-frequency approximation subband (LLL) and seven high-frequency detail subbands (LLH, LHL, LHH, HLL, HLH, HHL, HHH) will be generated, where L represents low-pass filtering and H represents high-pass filtering, and the three letters correspond to the operations of the height, width, and time dimensions respectively. After 3 layers of decomposition, a lowest-frequency approximation subband L is finally obtained. (3) , and a total of 3×7 = 21 high-frequency detail subband sets with different scales and directions {H k (s) |s∈{1, 2, 3}, k∈{1..7}}. Among them, L (3) captures the coarsest spatio-temporal structure in the residual information, while the high-frequency subbands H k (s) contain detail information such as edges, textures, and noises with different frequencies and directions. These wavelet coefficients constitute the main content of the residual detail stream, and their sparsity provides the possibility for subsequent efficient compression.
[0058] Step 3: Latent space projection and low-rank representation of the background stream and the motion stream.
[0059] For the static background stream F b and the dynamic motion stream F m initially separated in Step 1, further compression and structured representation are performed. Through tensor decomposition technology, they are projected into a more compact latent space to extract their core structure information for subsequent entropy coding. For the static background stream F b which is highly correlated in time but relatively stable in spatial structure (usually a three-dimensional tensor with dimensions Height×Width×Time), the Tucker decomposition model is used for compression. Tucker decomposition represents a high-dimensional tensor as the product of a core tensor and a series of factor matrices along each dimension. Mathematically, F b ≈F b core ×_1 U_1 ×_2 U_2 ×_3 U_3. Here, F b coreis a core tensor with a size much smaller than the original F b The dimensions of which are set to 1 / 8 of the original video dimensions (height, width, time) in this embodiment, i.e., (H / 8)×(W / 8)×(T / 8). U_1 (dimension H×H / 8), U_2 (dimension W×W / 8), and U_3 (dimension T×T / 8) are factor matrices along the height, width, and time dimensions respectively (usually required to be orthogonal), and they can be regarded as basis vectors in the corresponding dimensions. F b core captures the interaction information under different combinations of basis vectors and represents the most essential potential structure of the background flow. This decomposition can effectively remove the spatial and temporal redundancy of the background. For the dynamic motion flow F m (i.e., the dense optical flow field, which is also a three-dimensional tensor H×W×2, or regarded as two H×W×T tensors), its structure may be more complex and may not necessarily be suitable for the low-rank assumption. However, for compression, CP decomposition (Canonical Polyadic Decomposition, also known as PARAFAC) is selected in this embodiment. CP decomposition represents a tensor as the sum of the outer products of several (rank R) one-dimensional vectors. Mathematically, F m ≈∑ r=1 R λ r a r Θb r Θc r . Here, R is the rank of the decomposition, which is set to 5 in this embodiment. λ r is the weight of the r-th component (usually merged into the factor vector), a r (dimension H), b r (dimension W), and c r (dimension T or 2T, depending on the representation of F m ) are one-dimensional factor vectors along the three dimensions respectively. Θ represents the outer product of vectors. CP decomposition attempts to find the most dominant R spatio-temporal motion patterns that make up the original motion flow. In this way, F b and F m are both converted into latent representations with reduced number of parameters (F b core , U_1, U_2, U_3 and λ r , a r , b r , c r for r = 1 to 5), and these parameters will be used as the input of the subsequent entropy coding module.
[0060] Step 4: Hybrid entropy coding for the static background flow and the dynamic motion flow.
[0061] Efficient compression coding is performed on the latent representations of the static background flow and the dynamic motion flow obtained in Step 3 using a mixed entropy model, which adopts different strategies for the statistical characteristics and spatio-temporal correlations of different flows. Customized probability models are designed for each flow to accurately estimate the probability distribution of each element in its latent representation, and lossless or near-lossless compression is performed according to these probabilities using an efficient entropy encoder such as arithmetic coding. For the compressed core tensor F of the static background flow b core (the factor matrix U_i usually also needs to be encoded, but the core tensor is the main information carrier), considering that the background usually has strong local spatial correlations, in this embodiment, an autoregressive model based on a convolutional neural network (CNN), such as PixelCNN or its variants, is adopted. Such models predict the probability distribution of the coefficient (or quantization index) x_(i,j,k) at the current position (i,j,k) one by one according to a predetermined scanning order (such as raster scanning), and its prediction condition is the encoded neighborhood (context) C_(i,j,k) of this position in the scanning order. Its conditional probability can be expressed as P(x_(i,j,k)|C_(i,j,k);θ_AR), where θ_AR is the parameter of the autoregressive model (learned by the CNN). The model uses the powerful local feature extraction ability of the CNN to capture spatial redundancy. For the latent representation of the dynamic motion flow (such as the sequence of factor vectors λ r , a r , b r , c r for r = 1 to R, or directly encoding the original optical flow sequence F m ), considering that motion information usually contains complex long-range spatio-temporal dependencies (such as the continuous motion trajectory of an object, periodic motion, etc.), in this embodiment, an entropy model based on Transformer is introduced. The self-attention mechanism of the Transformer can effectively capture the long-distance dependencies between elements in the sequence. To adapt to video data, a spatio-temporal Transformer architecture can be adopted. To reduce the computational complexity, windowed attention or local attention mechanisms can be adopted to limit the attention calculation within a local spatio-temporal neighborhood. The model predicts the probability distribution of the latent representation (or its quantization index) y_t of the motion flow at the current time t, conditional on its temporal context information Context_t = {y_1,...,y_(t - 1)} (and possibly existing spatial context). For example, the model can predict the parameters of a Gaussian distribution: P(y_t|Context_t;θ_Transformer)=N(y_t|μ_t,σ t 2 ), where the mean μ_t and the variance σ t2 Predicted by the Transformer network based on the context Context_t. These precise probability models P(x_(i,j,k) |...) and P(y_t |...) are then fed into an arithmetic encoder to guide it to represent the quantization values of these latent variables with the number of bits close to the information entropy.
[0062] Step Five: Entropy coding for the residual detail stream and global rate-distortion optimization.
[0063] The strategy for entropy coding the residual detail stream (i.e., the wavelet coefficients H k (s) ) obtained by 3D-DWT in Step Two, and the quantization parameters of all three streams are dynamically adjusted through global rate-distortion optimization (RDO) to maximize the reconstruction quality under a given bitrate budget. The residual detail stream is mainly composed of high-frequency wavelet coefficients, and these coefficients usually exhibit sparse and non-Gaussian statistical characteristics, that is, most coefficients are close to zero, and only a few coefficients have significant amplitudes, representing details such as edges and textures of the image. For this characteristic, in this embodiment, an entropy model based on the Gaussian Mixture Model (GMM) is adopted. A GMM can fit a complex probability distribution as a weighted sum of multiple Gaussian components, and can better capture the peak and heavy-tail characteristics of the wavelet coefficient distribution. Further, an asymmetric GMM can be used, with different mixture models for positive and negative coefficients respectively, to more accurately model their distributions. The model P(h |θ_GMM) (where h represents a wavelet coefficient and θ_GMM are the parameters of the GMM model) is used to guide the quantization and arithmetic coding of the wavelet coefficients. To further enhance the compression effect, sparsity constraints can be introduced during model training or quantization, such as adding an L1 regularization term to the loss function to encourage the sparsity of the latent representation (wavelet coefficients). In the global rate-distortion optimization step, simply assigning fixed quantization parameters to each stream (such as 4 bits for the background, 6 bits for motion, and 2 bits for the residual as mentioned in the draft) is usually not optimal. It is necessary to dynamically adjust the quantization intensity according to the overall bitrate target and the specific characteristics of the video content. In this embodiment, the Lagrange multiplier method is used for optimization. The goal is to minimize the weighted sum of the total distortion D and the total bitrate R: min Q L = D(Q)+λR(Q). Where, Q = {Q b , Q m , Q r}{represents the set of quantization parameters (e.g., quantization step) assigned to the static background stream, dynamic motion stream, and residual detail stream. D(Q) is the total distortion of the final reconstructed video relative to the original video when the quantization parameter is Q (which can be measured by MSE, MS-SSIM, or combined with perceptual loss). R(Q) is the corresponding total bitrate, which is composed of the sum of the bits encoded by each of the three streams (R(Q)=R b (Q b )+R m (Q m )+R r (Q r ))。λ is the Lagrange multiplier, which controls the trade-off between bitrate and distortion: the larger λ is, the more the optimization objective tends to reduce the bitrate (allowing for greater distortion); the smaller λ is, the more it tends to reduce the distortion (allowing for a higher bitrate). In actual operation, an appropriate value of λ needs to be selected for the target bitrate range (the experimental range in this embodiment is from 0.3 to 1.2). The optimal set of quantization parameters Q* is found through iterative search or model-based methods to minimize the Lagrangian cost L. This optimization process ensures the most efficient allocation of bits among the three streams, maximizing the overall compression performance.
[0064] Step Six: Decoding and High-Quality Reconstruction Based on Motion-Aware Generative Adversarial Network.
[0065] At the decoding end, using the decoding information of each sub-stream, high-quality video reconstruction is performed through a specially designed motion-aware generative adversarial network (GAN), with particular attention to restoring the motion details that may be lost during the compression process. The three sub-bitstreams corresponding to the static background stream, dynamic motion stream, and residual detail stream are separated from the received compressed bitstream. Using the entropy decoder corresponding to the encoding end (e.g., arithmetic decoder), and according to the transmitted probability model information (or re-inferring the probability model at the decoder end), each sub-bitstream is decoded back to the quantized latent representation: the core tensor F b core (i.e., the factor matrix), the latent representation of the motion stream (such as CP factors or optical flow information F m ), and the quantized wavelet coefficients H k (s)。Next, these recovered latent representations are fed into a cascaded decoder architecture for reconstruction. At the core of this architecture is the generator G of a generative adversarial network (GAN). This generator G is designed to be motion-aware, that is, it specifically focuses on leveraging the decoded motion information to improve the reconstruction quality. Specifically, the generator G receives the decoded information from three streams as inputs or conditions. The key is that it contains or utilizes a flow-guided attention mechanism module (Flow-Guided Attention Module) internally. This module receives the decoded dynamic motion flow information (usually the optical flow field F m ) as a guiding signal. For example, flow-guided deformable convolution or an attention weighting mechanism can be adopted, enabling the reconstruction network (such as the sampling positions of convolutional kernels or the weights of feature channels) to be adaptively adjusted according to the motion direction and amplitude indicated by the optical flow vectors. The network can focus more computational resources and attention on areas such as the contours and internal textures of moving objects, more effectively counteracting the blurring and artifacts introduced by compression and restoring sharper and more realistic motion details. The discriminator D is used to train the generator G. It receives the reconstructed frames output by the generator and the real original frames (during the training phase) for comparison to judge their authenticity. In this embodiment, the discriminator D with a multi-scale PatchGAN structure is adopted, which can evaluate the authenticity of image patches at different scales, contributing to improving the generation quality of details and textures. The loss function during training includes the adversarial loss L_adv (making the generated frames difficult to be distinguished by the discriminator) and the perceptual loss L_perceptual (such as the feature matching loss based on the pre-trained VGG-19 network, used to ensure that the generated frames are similar to the original frames in terms of deep features, with a weight of 0.1). The total loss is L_GAN = L_adv + 0.1L_perceptual. The cascade decoding is reflected in that the decoded background flow information (such as through F b core and the factor matrix to reconstruct F b ) may be first used to reconstruct the basic static scene, and then the motion flow information is superimposed (or fused) (such as using F m to distort the background or directly synthesize moving objects), and finally the decoded residual wavelet coefficients are used to obtain a detail-enhanced map through the inverse DWT (Inverse DWT) and incorporated into the reconstruction result for refinement to synthesize a high-quality output video frame F*. The entire process is optimized through end-to-end training. When tested on the NVIDIA V100 GPU for HEVC Class B sequences, when λ = 0.7, this embodiment can achieve a 0.02 improvement in the SSIM metric while reducing the bit rate by 43.6% compared to HEVC, verifying its effectiveness.
[0066] The present invention can revolutionize the current video streaming transmission and distribution mode. Online video platforms face multiple challenges such as massive data storage, high bandwidth costs, and users' continuous pursuit of high-quality and low-latency experiences. Although existing compression standards (such as H.264 / AVC, HEVC, AV1) have been continuously improving, it is still difficult to balance extreme compression and detail fidelity when dealing with videos containing a large amount of static background (such as interview programs, landscape documentaries) or mixed complex motions (such as sports events, action movies). Through spatio-temporal decoupling, the present invention decomposes the video into three sub-streams: static background, dynamic motion, and residual details, enabling unprecedented content-adaptive compression. For the slowly changing or static background regions in the video (such as the sky, walls, fixed sets), the static background stream can be encoded with an extremely high compression ratio, reducing the overall bit rate, which corresponds to a significant reduction in the cost of the CDN (Content Delivery Network). The dynamic motion stream focuses resources on accurately encoding key information such as the motion trajectories and deformations of objects. Combining a motion-aware GAN reconstruction network, especially an optical flow-guided deformable attention mechanism, can accurately restore the edges and texture details of fast-moving objects, effectively suppressing the motion blur and blocking effects commonly seen in traditional coding at low bit rates. Especially in high-speed motion scenarios such as sports live broadcasts or action movies, it can provide a clearer and smoother visual experience. The hybrid entropy model is optimized according to the statistical characteristics of different streams to further explore the compression potential. Therefore, by applying the present invention, video service providers can provide a higher resolution (such as promoting the popularization of 4K / 8K) or a more stable playback experience (reducing buffering) under the same bandwidth, or reduce the bit rate by about 45% while maintaining the existing quality, thereby reducing operating costs and enhancing user satisfaction in different network environments, especially in bandwidth-constrained scenarios such as mobile devices.
[0067] In the field of video surveillance, the explosive growth of data volume has brought huge pressure to storage systems and network transmissions. At the same time, the clear recording of key motion details in abnormal events (such as intrusion, accident, suspicious behavior) is the core requirement of the security system. Traditional surveillance video coding often redundantly encodes scenes with a large area of static background (such as the field of view of a fixed camera), wasting a large amount of storage space and bandwidth. The hierarchical compression method proposed by the present invention has natural advantages in this regard. Through background modeling in preprocessing and subsequent tensor decomposition, the fixed background (such as buildings, roads, sky) in the surveillance video can be efficiently separated into the static background stream, and a deep compression is performed on it using a CNN-based autoregressive entropy model to achieve a reduction in storage occupancy (combined with dynamic bit allocation according to experimental data). For the key dynamic information in the surveillance scene, such as the moving trajectories and speed changes of pedestrians and vehicles, and even areas that require high-resolution details such as faces and license plates, they will be effectively mapped to the dynamic motion stream. Encoding using a Transformer-based entropy model can capture complex spatio-temporal dependencies and ensure the integrity of motion information. More importantly, the motion-aware GAN reconstruction network at the decoding end uses optical flow information to guide the attention mechanism, which can specifically enhance the reconstruction quality of the motion area. Even after compression, the detailed features of moving targets can be clearly restored, effectively avoiding detail loss or blurring caused by over-compression, which is crucial for post-event evidence collection and real-time analysis (such as AI behavior recognition, target tracking). For example, under low-light or adverse weather conditions at night, this method can also better retain the contours and details of moving objects. Therefore, the present invention can help the security system achieve a longer video recording storage period, lower transmission bandwidth requirements, and a more reliable ability to obtain key evidence, improving the overall efficiency and economy of the intelligent security system.
[0068] Immersive experience applications such as VR / AR and the metaverse have extremely high requirements for the transmission and rendering of video / image data, usually involving high resolution (4K / 8K or even higher), high frame rate (above 90Hz), wide field of view (such as 360° panoramic video), and low-latency interaction. The huge amount of data poses a severe challenge to network bandwidth and terminal processing capabilities. Any compression distortion or latency may disrupt the user's immersion and even cause motion sickness. The hierarchical generative compression method of the present invention provides an effective way to solve these problems. The spatio-temporal decoupling mechanism is particularly suitable for processing VR / AR scenarios. In VR, most of the environment in the user's field of view may be static or slowly changing and can be classified into the static background stream for efficient compression; while the user's interaction objects, virtual avatars, or other dynamic elements belong to the dynamic motion stream and need to be precisely encoded to ensure the authenticity and smoothness of the interaction. In AR, the background information of the real world and the superimposed virtual dynamic information can also be separated and differentially encoded in a similar way. The high attention of the present invention to motion fidelity is crucial. The precise encoding of the dynamic motion stream and the reconstruction ability of the motion-aware GAN can ensure the clear and natural presentation of details such as object motion in the virtual world, user gesture tracking, and AR overlay animations, avoiding blurring, trailing, or artifacts, which is essential for maintaining immersion and interaction accuracy. The reduction in bitrate directly translates into lower transmission latency, which is a core advantage for VR / AR applications that require real-time rendering and interaction (such as cloud VR, remote collaboration), enabling a smoother and more immediate user experience. The combination of hybrid entropy coding and generative reconstruction can maintain an acceptable visual quality even at extremely low bitrates, expanding the application possibilities of VR / AR in mobile networks or wireless environments. In summary, the present invention can effectively reduce the data burden of VR / AR applications, enhance motion expressiveness, reduce latency, bring users a higher-quality and more comfortable immersive interaction experience, and strongly promote the development of related industries.
[0069] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A multi-granularity generative video compression method, characterized in that Including: Obtain a video frame sequence, perform spatio-temporal feature decomposition, decompose and map the video signal into three latent subspaces, and obtain the first, second, and third latent subspace representations, where the first latent subspace representation is the low-frequency spatial information and time-varying components in the video, the second latent subspace representation is the spatio-temporal motion characteristics in the video, and the third latent subspace representation is the high-frequency detail information in the video; Perform encoding processing on the first, second, and third latent subspace representations respectively, generate corresponding first, second, and third encoded data, and combine them to generate a compressed bitstream.
2. The method according to claim 1, wherein Obtaining the first, second, and third latent subspace representations includes: Perform motion estimation on the video frame sequence to obtain motion information; Perform background modeling on the video frame sequence to obtain background information; Extract spatio-temporal features from the video frame sequence based on the motion information and background information; Perform tensor decomposition on the spatio-temporal features, map them to three approximately orthogonal latent subspaces, where the correlation between the latent subspaces is minimized through a pre-configured decomposition matrix, and obtain the first, second, and third latent subspace representations.
3. The method according to claim 2, wherein The generation of the first latent subspace representation includes: Perform time-domain filtering on the background information to extract the time-stable region; Extract the spatial low-frequency components from the video frame sequence; Calculate the motion amplitude at each pixel position based on the motion information, and mark the region where the motion amplitude is less than a preset threshold as the static region mask; Perform weighted fusion on the time-stable region, spatial low-frequency components, and static region mask to generate the first latent subspace representation, where the weights of the weighted fusion are adaptively determined according to the time stability of each region.
4. The method according to claim 2, wherein The generation of the second latent subspace representation includes: Perform motion vector field decomposition on the motion information to extract the main motion component; Perform time-difference operation on the video frame sequence to identify the motion region; Extract the intermediate-frequency spatio-temporal features based on the main motion component and the motion region, where different motion patterns with different speeds and directions are captured through multi-scale motion analysis; Map the intermediate-frequency spatio-temporal features to the second latent subspace to generate the second latent subspace representation, where the mapping maintains the time continuity of the motion trajectory.
5. The method according to claim 2, wherein The generation of the third latent subspace representation includes: A predicted frame reconstructed based on the first and second latent subspace representations; Calculate the residual information between the video frame sequence and the predicted frame; Perform three-dimensional wavelet transform on the residual information, perform multi-scale decomposition in the time and space dimensions, and extract the high-frequency wavelet coefficients; Map the high-frequency wavelet coefficients to the third latent subspace to generate the third latent subspace representation, where the mapping retains the detail features through sparse constraints.
6. The method according to claim 1, wherein Generating the first encoded data includes: Construct the first latent subspace representation as a three-dimensional tensor; Perform Tucker decomposition on the three-dimensional tensor to generate a core tensor and factor matrices along the spatial and time dimensions, where the dimension of the core tensor is set to the predetermined compression ratio of each dimension of the original tensor; Combine the core tensor and factor matrices to form the compressed first latent subspace representation; Perform quantization and entropy coding on the compressed first latent subspace representation to generate the first encoded data.
7. The method according to claim 1, characterized in that, Generating the second encoded data includes: Construct the second latent subspace representation as a motion tensor; Perform CP decomposition on the motion tensor and represent it as a weighted sum of a predetermined number of one-dimensional factor vectors, where each group of one-dimensional factor vectors captures a major motion pattern; Extract the spatial factor vectors, temporal factor vectors, and corresponding decomposition weights of the CP decomposition to form a compressed second latent subspace representation; Perform adaptive quantization and entropy coding on the compressed second latent subspace representation to generate second encoded data.
8. The method according to claim 6, wherein Generating the first encoded data includes: Using an autoregressive model constructed by a convolutional neural network to process the compressed first latent subspace representation element by element in a predetermined scanning order; For the current element to be encoded, extract its encoded spatial neighborhood as context; Predict the conditional probability distribution of the current element to be encoded based on the context through the autoregressive model; Use the conditional probability distribution to guide the arithmetic encoder to encode the current element to be encoded, generating the first encoded data.
9. The method according to claim 7, characterized in that, Generating the second encoded data includes: Construct the compressed second latent subspace representation into a spatio-temporal sequence; Extract the spatio-temporal context of the current position to be encoded from the spatio-temporal sequence, including historical information in the time dimension and neighborhood information in the space dimension; Use a Transformer encoder to process the spatio-temporal context and capture the long-range dependencies of the motion pattern through the self-attention mechanism; Predict the conditional probability distribution of the current position to be encoded based on the long-range dependencies, and accordingly guide the arithmetic encoder to generate the second encoded data.
10. The method according to claim 1, wherein Generating the third encoded data includes: Classify the wavelet coefficients in the third latent subspace representation to identify the positions of zero coefficients and the values of non-zero coefficients; Perform sparse position coding on the positions of zero coefficients to record the spatial distribution of non-zero coefficients; Model the non-zero coefficient values using a Gaussian mixture model, where different mixture components are fitted for positive and negative coefficient values respectively to obtain probability distribution parameters; Perform arithmetic coding on the non-zero coefficients based on the probability distribution parameters and combine it with the sparse position coding to generate the third encoded data.
Citation Information
Patent Citations
Video compression method and system based on variational auto-encoder improved entropy model
CN119011851A
Deep learning-based compression method using frequency decomposition
CN119072926A
Video coding method based on double-layer condition enhanced normalized stream compressor
CN119155445A
Video coding method and device
CN1720744A
Image encoding and decoding, video encoding and decoding: methods, systems and training methods
WO2022084702A1
Cited By
Intelligent video content extraction and rapid positioning system based on multi-modal fusion and space-time perception
CN121353850A
Multi-scale information generation type remote sensing image inter-frame compression coding and decoding method
CN122226976A