Video Compression Method and System Based on Variational Autoencoder Improved Entropy Model

By adopting an entropy model based on variational autoencoder in video compression technology, combining the space-time pyramid structure and continuous coding algorithm, the problems of low encoding efficiency and complex calculation in the existing technology are solved, efficient and adaptable video compression is achieved, and real-time video compression needs are met.

CN119011851BActive Publication Date: 2025-06-17ZHONGKE FANGCUN ZHIWEI (NANJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411481911.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-06-17
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

When existing video compression technology processes video content with complex statistical characteristics, the encoding efficiency is not high, and the encoding steps of the traditional entropy model are complex, making it difficult to meet the real-time video compression requirements.

Method used

Using an improved entropy model based on the variational autoencoder, multi-scale temporal context features are extracted through the spatiotemporal pyramid structure and continuous coding algorithm, combined with super prior and latent prior data, an improved hierarchical conditional variational autoencoder is used to generate latent variables, and context-aware adaptive quantization and dynamic entropy coding optimization based on compressed perception theory.

Benefits of technology

Improve the efficiency and adaptability of video compression, reduce the computational complexity, and make real-time video compression possible, while maintaining high compression efficiency while maximizing the detailed information of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119011851B_ABST
    Figure CN119011851B_ABST
Patent Text Reader

Abstract

The present invention discloses a video compression method and system based on an improved entropy model of a variational autoencoder. This method includes receiving current frame data in a video stream, constructing a spatio-temporal pyramid structure, and generating multi-scale persistence maps; fusing them with preset traditional visual features to obtain temporal context features; extracting hyperprior data and latent prior data, and splicing them to form an input feature set; using an improved hierarchical conditional variational autoencoder to generate multi-level latent variables; calculating probability values under predetermined conditions to generate probability distribution data; performing context-aware adaptive quantization to obtain quantized data; performing dynamic entropy coding optimization to obtain compressed data packets; performing rate-distortion optimization to obtain optimized compressed data packets; and performing perception-guided decoding and reconstruction to obtain a reconstructed video frame with the optimal perception quality. The present invention reduces the computational complexity, making the present invention easier to implement real-time video compression while maintaining high compression efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of video compression, and in particular, to a video compression method and system based on improving the entropy model by a variational autoencoder. Background Art

[0002] Video compression aims to reduce the bit rate of video data for storage and transmission while maintaining video quality as much as possible. Traditional video compression standards, such as H.264 and H.265, rely on hybrid frameworks including steps such as prediction, transformation, quantization, entropy coding, and loop filtering. With the development of deep learning technology, neural network-based video compression methods (NVC) have begun to attract attention, which perform compression by learning the intrinsic representation of data.

[0003] In the prior art, a common method is to use a variational autoencoder (VAE) to improve the entropy model. The advantage of the variational autoencoder is that it can learn the probability distribution of data and perform more efficient coding based on this. In existing video compression, the entropy model usually uses a fixed probability model to generate latent variables and optimize the quantization step, which may lead to low coding efficiency and poor effects, especially when dealing with video content with complex statistical characteristics and unable to adapt to the diversity of video content; at the same time, the entropy model coding step in traditional methods may involve complex calculations, especially when dealing with high-resolution videos; in resource-constrained environments, the prior art cannot meet the requirements of real-time video compression. Summary of the Invention

[0004] The object of the invention is to provide a video compression method and system based on improving the entropy model by a variational autoencoder to solve the above problems existing in the prior art.

[0005] Technical Solution: A video compression method based on improving the entropy model by a variational autoencoder includes the following steps:

[0006] S1. Receive the current frame data in the video stream, and based on the current frame data, construct a spatio-temporal pyramid structure; apply the persistent homology algorithm to the spatio-temporal pyramid structure to generate a multi-scale persistence diagram; fuse the multi-scale persistence diagram with a preset traditional visual feature to obtain an enhanced temporal context feature; based on the current frame data, extract hyperprior data and latent prior data, and splice the temporal context feature, hyperprior data, and latent prior data to form an input feature set;

[0007] S2. Based on the input feature set, use an improved hierarchical conditional variational autoencoder to generate multi-level latent variables; based on the multi-level latent variables, use algebraic geometry methods to estimate the probability distribution parameters; based on the probability distribution parameters, calculate the probability value of each latent variable under a predetermined condition to generate probability distribution data;

[0008] S3. Perform context-aware adaptive quantization based on the compressed sensing theory using multi-level latent variables and probability distribution data to obtain the quantized data;

[0009] S4. Perform dynamic entropy coding optimization based on the quantized data to obtain compressed data packets; perform rate-distortion optimization based on information geometry on the final compressed data packets to obtain optimized compressed data packets;

[0010] S5. Perform perception-guided decoding and reconstruction based on the optimized compressed data packets to obtain the reconstructed video frames with the optimal perception quality.

[0011] A video compression system based on an improved entropy model using a variational autoencoder, comprising:

[0012] At least one processor; and,

[0013] A memory communicatively connected to at least one of the processors; wherein,

[0014] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the video compression method based on the improved entropy model using a variational autoencoder.

[0015] Advantageous effects: Through the adaptive quantization mechanism, the present invention can intelligently adjust the quantization step according to different characteristics of video content, effectively filter noise and retain useful information, thereby improving the compression efficiency; by optimizing the network structure, the computational complexity is reduced, making the present invention easier to implement real-time video compression while maintaining high compression efficiency. Description of the Drawings

[0016] Figure 1 Is the flowchart of the present invention.

[0017] Figure 2 Is the flowchart of step S1 of the present invention.

[0018] Figure 3 Is the flowchart of step S2 of the present invention.

[0019] Figure 4 Is the flowchart of step S3 of the present invention.

[0020] Figure 5 Is the flowchart of step S4 of the present invention.

[0021] Figure 6 Is the flowchart of step S5 of the present invention.

[0022] Figure 7 Is the overall structure diagram of the video compression autoencoder with an improved entropy model according to an embodiment of the present invention. Detailed Implementation Manner

[0023] As Figure 1 shown, the present application proposes a video compression method based on an improved entropy model of a variational autoencoder, including the following steps:

[0024] S1. Receive the current frame data in the video stream, and based on the current frame data, construct a spatio-temporal pyramid structure; apply the persistent homology algorithm to the spatio-temporal pyramid structure to generate a multi-scale persistence diagram; fuse the multi-scale persistence diagram with preset traditional visual features to obtain enhanced temporal context features; based on the current frame data, extract hyperprior data and latent prior data, and splice the temporal context features, hyperprior data, and latent prior data to form an input feature set;

[0025] S2. Based on the input feature set, use an improved hierarchical conditional variational autoencoder to generate multi-level latent variables; based on the multi-level latent variables, use algebraic geometry methods to estimate probability distribution parameters; based on the probability distribution parameters, calculate the probability values of each latent variable under predetermined conditions to generate probability distribution data;

[0026] S3. Based on the multi-level latent variables and probability distribution data, perform context-aware adaptive quantization based on the compressive sensing theory to obtain quantized data;

[0027] S4. Based on the quantized data, perform dynamic entropy coding optimization to obtain compressed data packets; based on the final compressed data packets, perform rate-distortion optimization based on information geometry to obtain optimized compressed data packets;

[0028] S5. Based on the optimized compressed data packets, perform perception-guided decoding and reconstruction to obtain a reconstructed video frame with the optimal perception quality.

[0029] As Figure 2 shown, according to one aspect of the present application, step S1 is further as follows:

[0030] S11. Read the original video data stream from a video input device or a storage medium, extract the video frame data at the current moment to obtain the current frame data; temporarily store the current frame data in a buffer and record its timestamp information; perform preprocessing on the current frame data in the buffer, including denoising, color space conversion, and resolution adjustment, to obtain preprocessed current frame data;

[0031] S12. Based on the preprocessed current frame data, construct a spatio-temporal pyramid structure, and each level of the spatio-temporal pyramid structure contains data with different temporal scales and spatial resolutions; apply the persistent homology algorithm to the data at each level to calculate topological features; summarize the topological features to generate a multi-scale persistence diagram;

[0032] S13. Based on the preprocessed current frame data, use a pre-trained convolutional neural network to extract traditional visual features; adopt a weighted fusion method to combine the multi-scale persistence map and the traditional visual features to obtain enhanced temporal context features;

[0033] S14. Based on the preprocessed current frame data, extract hyperprior data from a pre-trained neural network model; based on the encoded results of pre-stored historical frames, extract latent prior data; perform normalization processing on the hyperprior data and the latent prior data to obtain normalized hyperprior data and latent prior data; concatenate the normalized hyperprior data and latent prior data with the enhanced temporal context features to form a complete input feature set.

[0034] In this embodiment, the hyperprior data contains high-level semantic information of the video content; the latent prior data reflects the temporal correlation of the video sequence.

[0035] In an embodiment of the present application, use the OpenCV library to read the video stream, read one frame each time; convert the frame into the numpy array format, denoted as F t , where t represents the current time step. Apply Gaussian filtering for denoising: F t_denoised = cv2.GaussianBlur(F t , (5, 5), 0); Convert the RGB color space to the YUV color space: F t_yuv = cv2.cvtColor(F t_denoised , cv2.COLOR BGR2YUV ); Adjust the resolution to the standard size (such as 640x480): F t_resized = cv2.resize(F t_yuv , (640, 480)). Where F t represents the original image data of the t-th frame; F t_denoised represents the denoised image data; F t_yuv represents the image data in the YUV color space; F t_resized represents the image data after adjusting the resolution, cv2.GaussianBlur represents the Gaussian blur function, which is used to perform Gaussian filtering on the image to remove noise; cv2.cvtColor represents the color space conversion function, which is used to convert the image from one color space to another color space; cv2.COLOR BGR2YUV represents the color conversion code; cv2.resize represents the image scaling function.

[0036] Construct a spatio-temporal pyramid structure: Define the number of layers L (such as 3 layers), for each layer l (l = 0, 1,..., L - 1), the time window size is Wt = 2 l + 1; The spatial scale is S l = F t_resized.shape / / 2 l ; Store the data of each layer: P l = [downsample(F t_resized , S l ) for t in range(t - W t / / 2, t + W t / / 2 + 1)]. Apply the persistent homology algorithm and use the Persistent Homology Computational Toolkit (ripser) library to calculate the persistent homology of each layer: PH l = ripser.ripser(P l , maxdim = 1)['dgms'], Generate the multi-scale persistent diagram: MPD = [PH l , l ∈ L]; Use the pre-trained 50-layer Residual Network (ResNet50) to extract traditional visual features: VF = ResNet50(weights='imagenet')(F t_resized ); Perform feature fusion: fused feature = concatenate([flatten(MPD), VF]). Where L represents the number of pyramid layers, W t represents the time window size; S l represents the spatial scale of the l-th layer; P l represents the pyramid data of the l-th layer; PH l represents the persistent homology result of the l-th layer; MPD represents the multi-scale persistent diagram; VF represents the traditional visual features; fused feature represents the fused features, F t_resized.shape represents the shape of the resized image; downsample() represents the downsampling function; for t inrange represents the loop statement; ripser.ripser() represents the function in the Ripser library for calculating persistent homology; maxdim represents the maximum dimension for calculating persistent homology; dgms represents the result of persistent homology; weights='imagenet' represents specifying to use the weights pre-trained on the dataset; concatenate() represents the concatenation function; flatten() represents the function for flattening a multi-dimensional array into a one-dimensional array.

[0037] Use the pre-trained 16-layer Visual Geometry Group Network (VGG16) network to extract the hyperprior data: z t = VGG16(weights='imagenet')(F t_resized); Use a pre-trained recurrent neural network (RNN) model to extract latent prior data: y t_1 = RNN(previous frames ); Perform normalization: z t_norm = (z t - mean(z t )) / std(z t ); y t_1_norm = (y t_1 - mean(y t_1 )) / std(y t_1 ); Perform feature concatenation: input features = concatenate([fused feature , z t_norm , y t_1_norm ). Where z t represents the hyperprior data; y t_1 represents the latent prior data; z t_norm represents the normalized hyperprior data; y t_1_norm represents the normalized latent prior data; input features represents the final input feature set, previous frames represents the previous frame data; mean( ) represents the mean function, and std( ) represents the standard deviation function.

[0038] In another embodiment of the present application, read the current frame data from the video stream and store it as a three-dimensional array F t , where t represents the current time step. Apply Gaussian filtering for denoising: Create a 5x5 Gaussian kernel G', with a standard deviation σ = 1; perform two-dimensional convolution on the three-dimensional array F t : F t_denoised = F t * G'. Convert the RGB color space to the YUV color space: Perform the following conversion on each pixel (r, g, b): Y = 0.299R + 0.587G + 0.114B; U = -0.147R - 0.289G + 0.436B; V = 0.615R - 0.515G - 0.100B; Resample the image using the bilinear interpolation algorithm and adjust the resolution to the standard size (such as 640x480). Where F t represents the original image data of the t-th frame; F t_denoisedDenote the denoised image data; Y, U, and V represent the components of the YUV color space respectively, and r, g, and b represent the lowercase variables of the red, green, and blue components in the RGB color space; R, G, and B represent the uppercase variables of the red, green, and blue components in the RGB color space.

[0039] Construct a spatio-temporal pyramid structure: Define the number of layers L (e.g., 3 layers). For each layer l (l = 0, 1, ..., L - 1), the time window size is W t = 2 l + 1; The spatial scale is S l = F t_resized.shape / / 2 l ; Use average pooling with a stride of 2 l Downsample each layer. Apply the persistent homology algorithm: Construct a distance matrix D for the data of each layer, calculate the simplicial complex at different thresholds ε, and calculate the persistent homology group H k (ε), k = 0, 1; Generate a multi-scale persistence diagram: Record the birth and death times of each homology class and store the results in MPD. Extract traditional visual features: Implement a simplified convolutional neural network including 3 convolutional layers, each followed by ReLU activation and max pooling; 2 fully connected layers. Concatenate MPD and the traditional visual features into a vector.

[0040] Construct a simplified VGG network: 5 convolutional blocks, each containing 2 - 3 convolutional layers and a max pooling layer, using pre-trained weights. Extract latent prior data: Implement a simple RNN, an LSTM layer, and a fully connected layer, using the features of the previous N frames as input to predict the features of the current frame. Calculate the mean μ and standard deviation σ, and normalize the data: x norm = (x - μ) / σ; Concatenate all feature vectors into a large vector, where x norm represents the normalized data.

[0041] In this embodiment, by introducing a spatio-temporal pyramid structure and a persistent homology algorithm, the extraction efficiency and quality of temporal context features are improved. The spatio-temporal pyramid structure can capture video features at different temporal scales and spatial resolutions, enabling the subsequent encoding process to better adapt to the spatio-temporal changes of video content. The application of the persistent homology algorithm enables the system to extract topological features of video data, which are highly robust to changes in motion patterns and scene structures in the video. By fusing topological features with traditional visual features, a more comprehensive and stable temporal context representation is formed. This fused feature can not only more accurately reflect the temporal continuity of video content but also maintain stability when the video scene undergoes drastic changes. The introduction of the hyperprior and the latent prior provides additional prior knowledge for the subsequent encoding process, helping to improve the encoding efficiency. Especially when dealing with complex video sequences, this prior information can help the system better predict and encode the correlation between video frames. In addition, by normalizing and splicing features from different sources, a unified and information-rich input feature set is formed, providing high-quality input data for the subsequent variational autoencoder. The fusion of such multi-source information not only improves the expressive power of the features but also enhances the system's adaptability to different types of video content. This embodiment lays a solid foundation for the video compression process through multi-scale and multi-modal feature extraction and fusion, effectively improving the efficiency and quality of subsequent encoding.

[0042] As Figure 3 shown, according to one aspect of the present application, step S2 is further as follows:

[0043] S21. Input the input feature set into the encoder network to generate an initial latent variable; obtain the conditional information of the current frame data, and combine the conditional information with the initial latent variable to construct a multi-level latent variable generation network model; wherein the conditional information includes the scene type and the motion complexity;

[0044] S22. Based on the conditional information and the initial latent variable, use an adaptive prior network to dynamically adjust the prior distribution parameters of each layer in the multi-level latent variable generation network model to obtain an adjusted multi-level latent variable generation network model; use the adjusted multi-level latent variable generation network model to generate latent variables at different abstraction levels layer by layer; combine the latent variables of all levels to form multi-level latent variables;

[0045] S23. Map the multi-level latent variables onto a predefined algebraic variety to construct a geometric structure representation of the probability distribution; obtain the encoding results of the most recent N frames as the observation data from the pre-stored data cache, and based on the observation data and the geometric structure representation, use algebraic statistical methods to estimate the parameters on the algebraic variety; adopt the semi-algebraic set theory to constrain the parameter space on the algebraic variety using the preset constraints to obtain the constrained parameters; where N is a natural number greater than 0.

[0046] S24. Based on the constrained parameters, construct a conditional probability model of the latent variables, calculate the probability values of each latent variable under the predefined conditions, and generate probability distribution data.

[0047] In this embodiment, an improved hierarchical conditional variational autoencoder is used for encoding. Specifically, the input feature set is received and input into the encoder network to generate the initial latent variables. The conditional information such as the scene type and motion complexity of the current frame is obtained from the scene analysis module, and these conditional information are combined with the initial latent variables and input into the multi-level latent variable generation network. This network generates latent variables at different abstraction levels layer by layer, and the output of each layer is used as the input of the next layer. At the same time, an adaptive prior network is used to dynamically adjust the prior distribution parameters of each layer according to the current input. The latent variables at all levels are combined to form the final multi-level latent representation.

[0048] Map the multi-level latent representation onto a predefined algebraic variety to construct a geometric structure representation of the probability distribution. Obtain the encoding results of the most recent N frames from the data cache as the observation data, and use algebraic statistical methods to estimate the parameters on the algebraic variety based on these observation data. Apply the semi-algebraic set theory to constrain the parameter space using the preset constraints to ensure the rationality of the estimation results. Use the estimated parameters to construct a conditional probability model of the latent variables. Calculate the probability values of each latent variable under the given conditions, generate the probability distribution data, and transmit it together with the multi-level latent representation to the following steps for adaptive quantization.

[0049] In an embodiment of the present application, an improved hierarchical conditional variational autoencoder (HCVAE) model structure is constructed: The encoder is input features -> Dense layers -> μ, σ; Sampling layer: ε ~ N(0, I), z = μ + σ * ε; Decoder: z -> Dense layers -> reconstructed features . Conditional information processing: The scene type is scene type = classify scene (F t_resized ) ; The motion complexity is motion complexity = estimatemotion (F t , F t_1 )。Generate multi - level latent variables: for l in range(num layers ): μ l , σ l = encoder l (z l1 , scene type , motion complexity ); z l = sample(μ l , σ l ). Adaptive prior network: p z = prior network (scene type , motion complexity ). Train HCVAE: loss = reconstruction loss + KL divergence (q(z|x), p z ); optimize(loss). Where μ and σ represent the mean and standard deviation of the latent space; ε represents random noise; z represents the sampled latent variable; scene type represents the scene type; motion complexity represents the motion complexity; z l represents the latent variable of the l - th layer; p z represents the adaptive prior distribution, Dense layers represents the dense layer or fully - connected layer; reconstructed features represents the reconstructed feature; classify scene represents the scene classification; estimate motion represents the motion estimation, which refers to the function for calculating the motion information between adjacent frames; num layers represents the number of layers; encoder l ( ) represents the encoder function, which is the encoder of the l - th layer and is used to generate the mean and standard deviation of the latent variable; scene type represents the scene type; motion complexity represents the motion complexity; sample( ) represents the sampling function; prior network represents the prior network; reconstruction loss represents the reconstruction loss; KL divergenceKL divergence, which represents the difference between two probability distributions and is used to measure the difference between the model distribution and the prior distribution; q(z|x) is the posterior distribution, which represents the conditional probability distribution of the latent variable z given the input x; optimize(loss) represents optimizing the loss.

[0050] Define the algebraic variety: V = {(x, y, z) ∈ R³ | f(x, y, z) = 0}, where f is a polynomial function representing the geometric structure of the probability distribution. Apply algebraic statistical methods: Use maximum likelihood estimation (MLE) to solve for the parameter: θ MLE = argmax θ Σ log p(x i |θ); Apply semi-algebraic set theory to constrain the parameter space: S = {θ ∈ R d | g1(θ) ≥ 0,..., g k (θ) ≥ 0}, where g i is a polynomial inequality. Construct the conditional probability model: p(z | x, θ) = exp(f(x, z, θ)) / Σ z exp(f(x, z, θ)); Calculate the conditional probability of each latent variable: for z i in z, p i = p(z i |x, θ MLE ). Where V represents the algebraic variety of the probability distribution; f represents the polynomial function defining the algebraic variety; θ represents the model parameter; θ MLE represents the parameter of the maximum likelihood estimation; S represents the semi-algebraic set of the parameter space; g i represents the polynomial inequality defining the semi-algebraic set; p(z | x, θ) represents the conditional probability of z given x and the parameter θ.

[0051] In another embodiment of the present application, the HCVAE model structure is implemented: the encoder is a multi-layer perceptron that outputs μ and σ; the sampling layer is z = μ + σ * ε, where ε ~ N(0, I); the decoder is a multi-layer perceptron that reconstructs the input. Conditional information processing: The scene type is to implement a simple scene classifier (such as a decision tree); the motion complexity is to calculate the optical flow of adjacent frames and take the average of its modulus. Multi-level latent variable generation: Implement multi-level encoders and decoders, and add conditional information at each level. The adaptive prior network specifically implements a small neural network to generate prior distribution parameters according to conditional information. When training the HCVAE, the reconstruction loss is implemented: mean squared error; the KL divergence is implemented: DKL = 0.5 * Σ(μ² + σ - log(σ²) - 1); the total loss = reconstruction loss + β * KL divergence; stochastic gradient descent is used for optimization. Where μ and σ represent the mean and standard deviation of the latent space, ε represents random noise; z represents the sampled latent variable; DKL represents the KL divergence; β represents the weight of the KL divergence.

[0052] Define an algebraic variety: Select a quadric surface as the simplified model: ax² + by² + cz² + dxy + exz + fyz + gx + hy + iz + j = 0. Apply algebraic statistical methods: Collect data points (x i , y i , z i ), construct a least squares problem: min Σ(ax i ² + by i ² + cz i ² + dx i y i + ex i z i + fy i z i + gx i + hy i + iz i+ j)²; Solve the linear equations using SVD to obtain the parameters (a, b, c, d, e, f, g, h, i, j). Apply the semi - algebraic set theory to constrain the parameter space: Define inequality constraints: such as a > 0, b > 0, c > 0 (to ensure the surface is an ellipsoid). Construct a conditional probability model: Use the distance from a point to the surface as the probability metric: p(z|x, y) ∝ exp(-k * distance((x, y, z), surface)), where k is a scaling factor. Calculate the conditional probability of each latent variable: For each point, calculate its distance to the surface and use the above formula to calculate the conditional probability. Here, a, b, c, d, e, f, g, h, i, j represent the parameters of the algebraic variety, k represents the scaling factor in probability calculation, p(z|x, y) represents the conditional probability of the latent variable z given the input x and y; distance( ) represents the distance function; surface represents the surface.

[0053] In this embodiment, by combining the improved hierarchical conditional variational auto - encoder (HCVAE) and algebraic geometry methods, the encoding efficiency and accuracy of the entropy model are improved. The multi - level structure of HCVAE enables the system to capture the features of video data at different abstraction levels, thus representing the complexity of video content more comprehensively. The introduction of conditional information, such as scene type and motion complexity, enables the encoding process to be dynamically adjusted according to the characteristics of video content, improving the adaptability of encoding. The application of the adaptive prior network further enhances the system's ability to model different types of video data, making the generated latent representation more compact and information - rich. The introduction of algebraic geometry methods provides a theoretical basis for probability distribution estimation. By mapping the probability distribution to an algebraic variety, the system can more accurately capture the internal structure of the data. It not only improves the accuracy of probability estimation but also enhances the model's ability to express complex data distributions. The application of semi - algebraic set theory ensures the rationality of parameter estimation, effectively avoiding overfitting problems and improving the generalization ability of the model. This probability model optimization based on algebraic geometry is particularly suitable for dealing with the common non - linear and non - Gaussian distributions in video data and can more accurately capture the complex dependencies between video frames. In addition, by constructing a conditional probability model, the system can more precisely predict the probability distribution of each latent variable, providing more accurate prior information for subsequent entropy coding. This improvement is particularly effective for processing highly dynamic and complex video sequences, capable of maximizing the retention of video detail information while maintaining a high compression ratio. This embodiment improves the encoding efficiency and accuracy of the entropy model through advanced probability modeling and optimization techniques, laying a foundation for realizing efficient video compression.

[0054] As Figure 4 shown, according to one aspect of the present application, step S3 is further:

[0055] S31. Use a pre-trained sparse dictionary to project the multi-level latent variables into the sparse domain to obtain a sparse domain representation; select a quantization matrix from a preset quantization matrix library based on the content features of the current frame data; use a context analysis module to dynamically adjust the parameters of the quantization matrix based on the probability distribution data to obtain adjusted quantization parameters.

[0056] S32. Based on the sparse domain representation and the adjusted quantization parameters, perform a quantization operation to obtain quantized discrete values, that is, quantized data.

[0057] According to one aspect of the present application, it further includes: S33. Based on the quantized discrete values, use a convex optimization algorithm to reconstruct the original data; calculate the perceptual quality difference between the reconstructed original data and the multi-level latent variables, and based on the perceptual quality difference, further fine-tune the adjusted quantization parameters to obtain fine-tuned quantization parameters.

[0058] According to another aspect of the present application, in step S31, using the context analysis module to dynamically adjust the parameters of the quantization matrix based on the probability distribution data to obtain the adjusted quantization parameters is further as follows:

[0059] S311. Based on the multi-level latent variables, use a convolutional neural network to extract local features, including edges and textures; use a long short-term memory network to extract global features, including motion trajectories and scene changes.

[0060] S312. Based on the local features, global features, and probability distribution data, use an attention mechanism for weighted fusion to obtain a comprehensive feature vector.

[0061] S313. Based on the comprehensive feature vector, use a reinforcement learning algorithm to dynamically adjust the parameters of the quantization matrix to obtain adjusted quantization parameters.

[0062] In this embodiment, a pre-trained sparse dictionary is used to project the multi-level latent representation into the sparse domain. Based on the content features of the current frame, a quantization matrix that satisfies the principle of restricted synchronization is selected from the quantization matrix library. The context analysis module is applied to dynamically adjust the quantization parameters by considering the local and global video content features and combining the probability distribution data. The quantization operation is performed to convert the sparse domain data into discrete values. Convex optimization algorithms such as LASSO are applied to reconstruct the original data from the quantized discrete values. The perceptual quality difference between the reconstructed data and the original multi-level latent representation is calculated, and the quantization parameters are fine-tuned according to this difference to balance the bit rate and visual quality. The quantized data and the adjusted quantization parameters are passed to the following steps for entropy coding.

[0063] In an embodiment of the present application, a sparse representation dictionary is constructed: use the K-SVD algorithm to learn the dictionary D, and perform sparse coding on the latent variable z: α = argminα ||z - Dα||² 2 + λ||α||₁. Construct the RIP quantization matrix: Generate a random Gaussian matrix Φ, orthogonalize Q = orth(Φ), scale A = sqrt(m / n) * Q, where m < n. Perform context analysis: Extract local features: local features =extract local _ features (z); Extract global features: global features =extract global _ features (z); Fuse features: context = concatenate([local features , global features ). Dynamically adjust the quantization parameter: Construct a neural network f q : q = f q (context); Quantization operation: z q = round(z / q) * q. Reconstruct using the LASSO algorithm: z hat = argmin z ||Az - y||² 2 + λ||z||₁; Calculate the perceptual quality: SSIM = structural similarity (z, z hat ); Update the quantization parameter: q = update quantization (q, SSIM). Where D represents the sparse representation dictionary; α represents the sparse coefficient; Φ represents the random Gaussian matrix; Q represents the orthogonalized matrix; A represents the RIP quantization matrix; z represents the original latent variable; z q represents the quantized latent variable; q represents the quantization step size; z hat represents the reconstructed latent variable; SSIM represents the structural similarity index, extract local _ features represents extracting local features; extract global _ features represents extracting global features, round( ) represents the rounding function.

[0064] In another embodiment of the present application, a sparse representation dictionary is constructed to implement the K-SVD algorithm: initialize the dictionary D, iterative optimization: use the orthogonal matching pursuit (OMP) algorithm to update the sparse coding step; for each atom, use SVD to update the dictionary update step. Construct a RIP quantization matrix: generate a random Gaussian matrix Φ, implement the Gram-Schmidt orthogonalization process, and normalize the matrix columns. Perform context analysis: local features are: calculate the mean, variance, and gradient of the local area; global features are: calculate the color histogram and edge direction histogram of the entire frame. Dynamically adjust the quantization parameters to implement a simple feedforward neural network, with context features as input; output as the quantization step size. Reconstruct and optimize: use the proximal gradient descent method to solve the optimization problem; implement the calculation of the mean, variance, and covariance, calculate the similarity according to the SSIM formula, and adjust the quantization step size according to SSIM.

[0065] This embodiment improves the efficiency and adaptability of the quantization process by introducing a context-aware adaptive quantization (CS-CAAQ) method based on compressed sensing theory. The application of compressed sensing theory enables the system to represent data in a sparse domain, reducing the amount of information that needs to be encoded. It is particularly suitable for processing natural video data because video signals usually exhibit sparse characteristics in certain transform domains (such as wavelet domain or DCT domain). By constructing a quantization matrix that satisfies restricted synchronization (RIP), the system can reduce the sampling rate while ensuring the reconstruction quality, thereby improving the compression efficiency. The introduction of the context-aware mechanism enables the quantization process to be dynamically adjusted according to the local and global characteristics of the video content, which is particularly important for processing complex and changing video scenes. For areas with intense motion, the system can adopt a more detailed quantization strategy to retain more details; while for static background areas, a coarser quantization can be used to improve the compression rate. The application of convex optimization algorithms such as LASSO ensures high-quality reconstruction from quantized discrete values ​​to original data, effectively reducing quantization errors. This reconstruction process is particularly suitable for processing texture details and edge information in videos, and can maintain good visual quality at high compression rates. The introduction of perceptual quality indicators enables the system to achieve a better balance between bit rate and visual quality. This optimization based on the characteristics of the human visual system is particularly suitable for video compression applications because it can maximize compression efficiency while ensuring subjective visual quality. In addition, the dynamic adjustment mechanism of the adaptive quantization parameters enables the system to quickly respond to changes in video content, such as scene switching or sudden changes in lighting conditions, thereby maintaining stable compression performance throughout the video sequence. This embodiment combines compressed sensing theory and context-aware technology to achieve an efficient and adaptive quantization process, improve the efficiency and quality of video compression, and especially performs well when processing complex and dynamically changing video content.

[0066] like Figure 5As shown, according to one aspect of the present application, step S4 is further as follows:

[0067] S41. Obtain the statistical characteristics of the most recent N coded frames from the pre-stored coding statistics cache, and dynamically update the pre-configured symbol occurrence probability model based on the statistical characteristics and the quantized data to obtain an updated symbol occurrence probability model;

[0068] S42. Based on the updated symbol occurrence probability model, assign an optimal coding length to each symbol; based on the coding length and the quantized data, use the context adaptive arithmetic coding algorithm to calculate the coded data; based on the coded data, construct and optimize the variable length coding strategy to obtain optimized coded data; organize the optimized coded data into a bitstream in a predefined format, including necessary header information and metadata, to form a final compressed data packet;

[0069] S43. Based on the final compressed data packet, construct a statistical manifold, calculate the Fisher information matrix; based on the Fisher information matrix, construct an iterative optimization algorithm based on the natural gradient, evaluate the rate-distortion performance under the current parameter configuration of the iterative optimization algorithm until a preset number of iterations is reached; based on the evaluation result, output the optimal parameter configuration to obtain an optimized compressed data packet.

[0070] This embodiment receives the quantized data and the adjusted quantization parameters. Obtain the statistical characteristics of the most recent N coded frames from the coding statistics cache, and dynamically update the symbol occurrence probability model in combination with the data of the current frame. According to the updated probability model, assign an optimal coding length to each symbol. Considering the local spatial and temporal correlations, apply the context adaptive arithmetic coding algorithm to further improve the compression efficiency. Based on the data characteristics of the current frame, adjust the Huffman coding table in real time to optimize the variable length coding strategy. Organize the coded data into a bitstream in a predefined format, including necessary header information and metadata, to form a final compressed data packet, and transfer it to the following steps for global optimization, and at the same time store it in the coding result cache for subsequent frames to use.

[0071] Receive compressed data packets and collect intermediate data and adjustable parameters of each module from the parameter caches of the respective modules. Construct a statistical manifold and map the rate-distortion problem onto this manifold. Calculate the Fisher information matrix to quantify the importance of each direction in the parameter space. Construct an iterative optimization algorithm based on the natural gradient to search for the optimal parameter configuration on the statistical manifold. In each iteration, update the parameters of each module and re-execute steps S1 to S4 to evaluate the rate-distortion performance under the current parameter configuration. Adjust the search direction and step size according to the evaluation results. When the preset number of iterations is reached or the performance improvement is less than the threshold, output the optimal parameter configuration, update the parameter caches of each module, and pass the optimized compressed data packet to the following steps for decoding and reconstruction.

[0072] In one embodiment of the present application, the symbol probability model is dynamically updated: maintain a sliding window W that contains the statistical information of the most recent N encoded frames; for each symbol s, calculate its occurrence frequency in W: P(s) = count(s) / Σcount(s i ); use an adaptive hybrid model: P new (s) = λP(s) + (1 - λ)P prior (s), use the Shannon coding principle for optimal coding length allocation: L(s) = -log2(P new (s)). Perform context-adaptive arithmetic coding: define the context C(s) as the pixels around s, calculate the conditional probability: P(s|C(s)) = P(s, C(s)) / P(C(s)); use an arithmetic encoder to encode symbol s according to P(s|C(s)). Dynamically update the Huffman coding table: construct a Huffman tree according to P new (s) and update the Huffman coding table periodically (e.g., every 100 frames). Organize the bitstream, including a frame header that contains metadata such as frame type and quantization parameters; encoded data that contains the bitstream of Huffman coding or arithmetic coding; and a frame tail for synchronization and error detection. Where W represents the sliding window; P(s) represents the empirical probability of symbol s; P prior (s) represents the prior probability of symbol s; P new (s) represents the updated symbol probability; λ represents the mixing weight; L(s) represents the optimal coding length of symbol s; C(s) represents the context of symbol s; P(s|C(s)) represents the conditional probability of symbol s given context C(s), and count( ) represents the counting function.

[0073] Construct a statistical manifold: define the parameter space Θ = {θ|θ ∈ R d}; Define the probability distribution family P = {p(x|θ) | θ ∈ Θ}; Statistical manifold S = (Θ, g), where g is the Fisher information matrix; Calculate the Fisher information matrix: g ij (θ) = E x [-Ψ² / Ψθ i Ψθ j log p(x|θ)]; Construct the natural gradient optimization algorithm: θ new = θ old - η g -1 θ old ▽Lθ old , where L(θ) is the rate distortion function and η is the learning rate. Iterative optimization process, when not converged: 1. Execute steps S1 to S4 using the current parameter θ; 2. Calculate the rate distortion performance: RD = λR + D; where R is the bit rate, D is the distortion degree, and λ is the Lagrange multiplier; 3. Calculate the natural gradient: ▽ nat L = g -1 (θ) ▽L(θ); 4. Update the parameter: θ = θ - η ▽ nat L; 5. Check the convergence condition. Output the optimal parameter configuration. Where g ij represents the i, j-th element of the Fisher information matrix; θ represents the model parameter; ▽L(θ) represents the gradient of the rate distortion function; ▽ nat L represents the natural gradient of the rate distortion function; RD represents the rate distortion performance metric, and Ψ represents the partial derivative symbol.

[0074] In another embodiment of the present application, dynamically update the symbol probability model: Maintain a counter array to record the occurrence times of each symbol; Regularly update the probability: P(s) = count(s) / total count ; Implement an adaptive hybrid model: P new (s) = λP(s) + (1 - λ)P prior (s). Optimal coding length allocation, implement the calculation of -log2(P new (s)). Perform context - adaptive arithmetic coding: Implement the context model, use the surrounding pixel values as the context; Implement the conditional probability calculation; Implement arithmetic coding: Maintain an interval [low, high); Subdivide the interval according to the symbol probability and output the binary number that can uniquely determine the interval. Dynamically update the Huffman coding table: Implement the Huffman tree construction algorithm and regularly reconstruct the Huffman tree. Organize the bitstream: Construct the frame header format, including the necessary metadata; Organize the encoded data into a byte stream. Where P(s) represents the probability of symbol s; λ represents the mixing weight; P new (s) represents the updated symbol probability; [low, high) represents the current interval of arithmetic coding.

[0075] Construct a statistical manifold: Select key parameters as the optimization objects and define the parameter space; Implement probability distribution calculation: such as Gaussian distribution, Laplace distribution, etc. Calculate the Fisher information matrix: Implement the calculation of the second derivative of the log-likelihood; Use the Monte Carlo method to estimate and calculate the expectation. Construct a natural gradient optimization algorithm to implement the inverse operation of the Fisher information matrix (for low dimensions, the inverse can be directly calculated, and for high dimensions, the conjugate gradient method can be used for approximation); Implement parameter update: θ new = θ old - η g -1 θ old ▽Lθ old . Iterate the optimization process to implement the calculation of the rate-distortion function: The bit rate R is to calculate the number of bits after encoding; The distortion D is to calculate the MSE or the perceptual quality metric; Use the finite difference method to implement the gradient calculation; Implement parameter update and convergence check. Output the optimal parameter configuration.

[0076] In this embodiment, by introducing the Dynamic Entropy Coding Optimization (DECO) method, the efficiency and adaptability of bitstream coding are improved. The application of the dynamic symbol probability model enables the system to adapt to the statistical property changes of video content in real time, which is particularly effective for processing long-sequence videos or videos with variable content. By analyzing the statistical properties of the most recent N encoded frames, the system can accurately capture the local temporal correlation of video data, thereby assigning the optimal coding length to each symbol. This adaptive coding strategy not only improves the compression efficiency but also enhances the system's response ability to video content changes. The introduction of context-adaptive arithmetic coding further utilizes the local spatial and temporal correlation of video data, improving the coding efficiency. It is particularly suitable for processing texture regions and motion sequences in videos, and can effectively capture the conditional probability relationship between pixel values, thereby achieving a more compact representation. The real-time adjusted Huffman coding table provides another layer of adaptability for the system, enabling the coding process to quickly respond to sudden changes in video content, such as scene transitions or lighting changes. This dynamic coding table strategy not only improves the coding efficiency but also enhances the system's adaptability to different types of video content. In addition, by organizing the encoded data into a predefined bitstream format, including necessary header information and metadata, the system ensures the integrity and decodability of the compressed data. This structured data organization not only facilitates subsequent transmission and storage but also provides the necessary context information for the decoding end, helping to improve the accuracy and efficiency of decoding. This embodiment realizes an efficient and flexible bitstream generation process through a multi-level adaptive coding strategy. It can effectively handle the spatio-temporal correlation and content diversity of video data, while ensuring a high compression ratio, maximizing the retention of the information content of the video. Especially for long-time sequences or videos with complex and variable content, this dynamic optimization method can continuously maintain high coding performance, providing key support for achieving high-quality, low-bitrate video compression.

[0077] By introducing a rate-distortion optimization method based on information geometry, global optimization of the entire video compression process is achieved, improving the overall performance of the system. The rate-distortion problem is modeled as an optimization problem on a statistical manifold, enabling the system to understand and optimize the compression process from a higher level of abstraction. By constructing a statistical manifold, the system can capture the geometric structure in the parameter space, which is crucial for understanding the impact of parameter changes on compression performance. The application of the Fisher information matrix further quantifies the importance of each direction in the parameter space, enabling the optimization process to more effectively allocate computational resources and focus on optimizing the parameters that have the greatest impact on performance. The construction of an iterative optimization algorithm based on the natural gradient enables the system to perform efficient parameter search on the statistical manifold. It is particularly suitable for dealing with high-dimensional parameter spaces, can quickly converge to the optimal solution, and avoids the local optimum problems that may be encountered by traditional gradient descent methods. By re-executing steps S1 to S4 and evaluating the rate-distortion performance in each iteration, the system can comprehensively consider the mutual influence between various modules and achieve true end-to-end optimization. This global optimization strategy is particularly effective for dealing with complex video compression tasks and can find the best parameter configuration under different compression rate requirements. The mechanism for dynamically adjusting the search direction and step size further enhances the adaptability of the optimization process, enabling the system to quickly respond to changes in video content, such as scene changes or changes in motion complexity. This adaptive optimization is particularly suitable for dealing with long video sequences or video streams with variable content. By setting reasonable termination conditions, such as a preset number of iterations or a performance improvement threshold, the system can achieve a good balance between optimization effect and computational overhead. This is particularly important for real-time or near-real-time video compression applications and can achieve the best compression performance with limited computational resources. Through the global optimization strategy in this embodiment, collaborative optimization of the entire video compression system is achieved, improving the adaptability and performance of the system under different video contents and compression requirements. It can effectively handle complex non-linear relationships in video compression and provides strong support for achieving efficient and high-quality video compression.

[0078] As Figure 6 shown, according to one aspect of the present application, step S5 is further:

[0079] S51. Based on the header information and metadata in the optimized compression data packet, parse the encoding parameters, and reverse-execute steps S4 to S2 to obtain the preliminarily reconstructed video frame data; use a pre-trained perceptual quality evaluation module to evaluate the perceptual quality of the preliminarily reconstructed video frame data to obtain an evaluation result;

[0080] S52. Based on the evaluation results, adaptively select and apply post-processing techniques from a preset post-processing technology library, including deblocking filtering and detail enhancement, to process the preliminarily reconstructed video frame data to obtain processed frames; input the processed frames into a pre-trained adversarial network to improve the visual quality and obtain enhanced frames;

[0081] S53. Compare the enhanced frames with the current frame data in step S1, calculate the perceptual similarity; based on the perceptual similarity, adjust the post-processing parameters; based on the adjusted post-processing parameters, output the reconstructed video frame with the optimal perceptual quality.

[0082] In this embodiment, the optimized compressed data packet is received. According to the header information and metadata in the data packet, the encoding parameters are parsed, and the processes of steps S4 to S2 are executed reversely to obtain the preliminarily reconstructed video frame data. A pre-trained perceptual quality evaluation module is used to simulate the human eye visual system to evaluate the perceptual quality of the reconstructed frames. According to the evaluation results, post-processing techniques such as deblocking filtering and detail enhancement are adaptively selected and applied from the post-processing technology library. The processed frames are input into a pre-trained adversarial network to further improve the visual quality. The enhanced frames are compared with the original input frames, the perceptual similarity is calculated, and the post-processing parameters are adjusted according to the similarity. The reconstructed video frame with the optimal perceptual quality is output, stored in the reconstructed frame buffer for subsequent processing, and at the same time passed to the video output module for display or storage.

[0083] In an embodiment of the present application, for preliminary reconstruction: the processes of steps S4 to S2 are executed reversely to obtain the preliminary reconstruction frame F r . For perceptual quality evaluation: use the pre-trained perceptual quality evaluation network Q, q = Q(F r ), q ∈ [0, 1], where 1 represents the highest quality. For adaptive post-processing: Deblocking: F d = deblock(F r, strength = f(q)); Detail enhancement: F e = enhance(F d , factor = g(q)), where f and g are adaptive functions based on q. Adversarial network enhancement: Generator G: F g = G(F e ); Discriminator D: authenticity = D(F g ); Optimization objective: min G max D E[log(D(F original )) + log(1 - D(G(F e)))]。Perform perceptual similarity calculation and parameter adjustment: Calculate SSIM: s = SSIM(F original , F g ); Adjust the post-processing parameters: strength new = strength * (1 + α(1 - s)) factor new = factor * (1 + β(1 - s)), where α and β are adjustment coefficients. Select the frame with the highest perceptual quality as the final output: F final = argmax F Q(F) for F in {F r , F d , F e , F g}. Where q represents the perceptual quality score, F d represents the frame after deblocking, F e represents the frame after detail enhancement, F g represents the frame enhanced by the adversarial network, F original represents the original input frame, s represents the structural similarity index (SSIM), F final represents the final output frame. The perceptual quality assessment network Q can use a pre-trained VGG19 network and be fine-tuned to adapt to the video quality assessment task. The deblocking function deblock() can be implemented as: def deblock(frame, strength): return cv2.fastNlMeansDenoisingColored(frame, None, strength, strength, 7, 21); The detail enhancement function enhance() can be implemented as: def enhance(frame, factor): kernel = np.array([[-1, -1, -1], [-1, 9, -1], [-1, -1, -1]]) return cv2.filter2D(frame, -1, kernel * factor); The adversarial network G and D can adopt an architecture similar to SRGAN, but need to be adjusted for the video frame reconstruction task. The SSIM calculation can use the skimage library: from skimage.metrics import structural similarity as ssim s = ssim(F original ,F g , multichannel=True); The parameter adjustment function can be implemented as: def adjust params(strength, factor, ssim, alpha = 0.1, beta = 0.2): strength new = strength * (1 + alpha * (1 - ssim)) factor new = factor * (1 + beta * (1 - ssim)) return strength new , factor new 。

[0084] In another embodiment of the present application, preliminary reconstruction is performed: the previous steps are executed in reverse to obtain a preliminary reconstruction frame. Perceptual quality assessment is carried out: a simplified version of the perceptual quality assessment network is implemented, using a pre-trained VGG feature extractor, and several fully connected layers are added for quality scoring. Adaptive post-processing is carried out: a deblocking effect filter is implemented, and a mean filter with an adaptive window is constructed; detail enhancement is implemented: an adaptive sharpening filter is constructed. Adversarial network enhancement is carried out: a simplified version of the generator G is implemented: using a U-Net structure with skip connections; a simplified version of the discriminator D is implemented: using a PatchGAN structure, and an adversarial training process is implemented. Perceptual similarity calculation and parameter adjustment are carried out: SSIM calculation is implemented: local mean, variance, and covariance are calculated; similarity is calculated according to the SSIM formula; a parameter adjustment function is implemented: the processing intensity is linearly adjusted according to SSIM. All candidate frames are scored using the perceptual quality assessment network, and the frame with the highest score is selected as the final output.

[0085] According to one aspect of the present application, a pre-trained perceptual quality assessment module is used to evaluate the perceptual quality of the preliminarily reconstructed video frame data, and the evaluation result is specifically as follows: The preliminarily reconstructed video frame data is normalized to ensure that the input data meets the requirements of the pre-trained model. Normalization can include resizing the image, normalizing pixel values, etc. The high-level features of the video frame are extracted using the feature extraction layer of the pre-trained model. These features can include edges, textures, color distributions, etc. The pre-processed features are input into the pre-trained perceptual quality assessment module. This module is usually a deep neural network that has been trained on a large amount of labeled data. The model outputs a perceptual quality score, which is usually a scalar value representing the visual quality of the video frame. The score can be based on multiple metrics, such as structural similarity (SSIM), peak signal-to-noise ratio (PSNR), etc. The video frame is processed at multiple scales to extract features at different scales. A pyramid structure or a multi-scale convolutional neural network (MSCNN) can be used to achieve this. The perceptual quality scores at different scales are weighted and averaged to obtain a comprehensive perceptual quality score. This step can improve the robustness and accuracy of the evaluation. Detect and process possible abnormal scores, such as overly high or low scores. Statistical methods or rule-based detection algorithms can be used. The perceptual quality scores of consecutive frames are smoothed to reduce the fluctuations in the scores. Methods such as moving average or Kalman filtering can be used. The final perceptual quality score is stored in a specified data structure for subsequent processing and analysis. If necessary, the evaluation result can be visually displayed, such as plotting a quality score curve or generating a quality report.

[0086] According to one aspect of the present application, the enhanced frame is compared with the current frame data in step S1 to calculate the perceptual similarity, which is specifically as follows: Features are extracted from the enhanced frame and the current frame data. Deep learning models such as convolutional neural networks (CNNs) can be used to extract high-level visual features. The extracted features are matched, and metrics such as Euclidean distance and cosine similarity are used to calculate the similarity between the features. Based on the result of the feature matching, the overall perceptual similarity is calculated. A weighted average method can be used to combine the similarities of different features to obtain a comprehensive perceptual similarity score. The calculated perceptual similarity score is compared with a preset threshold to evaluate whether the similarity is within an acceptable range. If the similarity is low, the post-processing parameters may need to be adjusted. According to the result of the similarity evaluation, the post-processing parameters are adjusted. Optimization algorithms such as gradient descent can be used to gradually adjust the parameters to improve the perceptual similarity.

[0087] In this embodiment, by introducing the Perceptual-Guided Reconstruction Enhancement (PGRE) technology, the quality of video decoding and reconstruction is improved. By reversely executing the encoding steps, the system can accurately parse the encoding parameters and data in the bitstream, providing a reliable basis for the subsequent reconstruction process. The introduction of the pre-trained perceptual quality assessment module simulates the characteristics of the human visual system and can accurately evaluate the perceptual quality of the reconstructed frames. This assessment method based on human perception is particularly suitable for video applications because it can capture the visual features that the human eye is most sensitive to, such as edge sharpness, texture details, and motion coherence. The adaptive post-processing technology selection mechanism based on the assessment results enables the system to adopt the most appropriate recovery strategy for different types of video content and compression artifacts. For example, for regions containing a large amount of texture, the system may select a stronger deblocking filter; while for regions containing clear edges, detail enhancement technology may be applied. This adaptive processing strategy ensures that while improving the overall visual quality, important image details are not overly smoothed. The application of the pre-trained adversarial network further enhances the visual quality of the reconstructed video. The adversarial learning mechanism can generate high-frequency details closer to the original video. Especially when dealing with highly compressed videos, it can effectively recover the texture and detail information lost during the compression process. While maintaining the overall structure of the video, it can generate more realistic and natural visual effects. By comparing the enhanced frames with the original input frames and calculating the perceptual similarity, the system can objectively evaluate the reconstruction quality and dynamically adjust the post-processing parameters accordingly. This closed-loop feedback mechanism ensures the stability and consistency of the reconstruction process. Especially when dealing with long video sequences, it can maintain a stable reconstruction quality. In addition, this quality assessment and parameter adjustment mechanism based on the original frames can effectively prevent artifacts caused by overprocessing, such as oversharpening or unnatural textures. The strategy of storing the reconstructed video frames in a cache for subsequent processing provides additional support for processing video sequences with temporal correlation. This cache mechanism enables the system to utilize the information of the previous and subsequent frames to improve the reconstruction quality of the current frame. Especially when dealing with motion sequences or scene transitions, it can provide a more coherent and natural visual effect. This embodiment realizes a high-quality video reconstruction process by combining advanced perceptual quality assessment, adaptive post-processing, and adversarial learning technologies. It can not only effectively recover the visual information lost during the compression process but also optimize the reconstruction results according to the characteristics of the human visual system, ensuring the high perceptual quality of the reconstructed video. It is particularly suitable for video applications that require high visual quality, such as high-definition video streaming, professional video production, etc., and can still maintain excellent visual effects at high compression rates.

[0088] According to one aspect of the present application, a video compression system that improves the entropy model based on a variational autoencoder includes:

[0089] At least one processor; and,

[0090] A memory communicatively connected to at least one of the processors; wherein,

[0091] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the video compression method based on the variational autoencoder improved entropy model described in any one of the above embodiments.

[0092] In another embodiment of the present application, a quantization scaling mechanism is learned through a variational autoencoder to modulate the latent features of the current frame to support a wider quality range. And a new spatial-channel intelligent quantization mechanism is constructed to adaptively determine the quantization step according to the cross-entropy and KL divergence mixed loss function, so as to improve the compression efficiency and retain more useful information. At the same time, a lightweight framework is proposed, which achieves a better quality and complexity trade-off in video compression compared with previous state-of-the-art methods.

[0093] As Figure 7 shown, the basic model of this embodiment is divided into two parts, the upper part of the upper structure diagram is the basic structure of general video compression encoding and decoding, and the lower part named the entropy model is the main improved part of this patent, the entropy model improved based on the variational autoencoder. The neural network parts of both parts are replaced with corresponding residual networks to reduce irrelevant connections between neurons and achieve lightweight improvement of the model.

[0094] (1), Inputs of Entropy Model: The inputs of the entropy model include the hyperprior z', t the temporal context prior C, t and the latent prior y', t-1 These inputs provide rich context information to help the model better predict the probability distribution.

[0095] (2), Outputs of Entropy Model: In addition to the probability distribution parameters, the entropy model also outputs quantization steps, and these steps are determined at different granularity levels, including global, channel level, and spatial-channel level.

[0096] (3), Encoding and Decoding Process: In the encoding process, the current frame is converted into a latent representation, and then quantized and encoded into a bitstream. In the decoding process, the bitstream is decoded and inverse quantized to reconstruct a high-quality frame.

[0097] According to the design of video compression, the neural compression module consists of a variational autoencoder, a quantization process, and a prior model. First, the temporal feature C t is encoded into a compact latent variable e t through a feature encoder: t e t = Encoder(C

[0098] To achieve robustness to noise, quantization is applied to the latent variable e t . Most existing quantization solutions in video compression only use a fixed quantization step size. In fact, content features vary greatly spatially. A fixed quantization step size cannot handle various complex contents well. For example, a fixed small quantization step size cannot effectively remove noise information, while a fixed large quantization step size will result in large information loss (i.e., inherent quantization noise). Therefore, an adaptive quantization mechanism is proposed, where the quantization step size is learnable. Currently, the latest method for learnable quantization step size is achieved through simple addition, subtraction, multiplication, division, and rounding, and finally calculating the cross-entropy loss. However, since the purpose of the entropy model is to find a latent variable to guide the learning of the quantization step size, this simple operation cannot learn the latent variable well. So, a variational autoencoder (VAE) with a residual neural network module is used here. According to the input latent variable e t and the probability distribution (μ t , σ t ) processed by the parameters, it automatically learns the noise-resistant latent variable e' t . At the same time, to achieve noise robustness, it is necessary to appropriately learn the data distribution and the quantization step size. In practice, the data distribution is unknown, so a prior model is used to estimate it, and then a joint loss is used to guide the learning of the data distribution and the quantization step size. The loss formula is as follows:

[0099] Loss = Loss CE + Loss KL = -log2p(e' t ) + ∑ x p(e' t ) log2(q(e' t ) / p(e' t ));

[0100] where p(e' t ) is the estimated probability mass function of the noise-resistant latent variable e' t , following a Laplace distribution. The prior model composed of a neural network is used to estimate the distribution parameters. The joint loss guides the compression module to learn the appropriate data distribution and quantization step size, and then achieves robustness to noise. In the framework of this model, the quantization step size q scIt is learnable in terms of spatial-channel intelligence and can adapt to regions with different content features. Finally, the processed quantization step size q sc is sent back to the video compression and decoding structure. After synchronous processing, the video compression operation is finally achieved.

[0101] This embodiment supports a wider PSNR range, exceeding the average range of the prior art, through a learning-based quantization scaling mechanism; by learning quantization scaling with a variational autoencoder, the present invention can dynamically adjust the quantization step size according to different video contents, thereby achieving better compression effects at different quality levels; the spatial-channel intelligent quantization mechanism allows the model to apply different quantization strategies to different regions of the video, which can reduce the bit rate while maintaining the image quality; by optimizing the network structure, the amount of computation and the number of model parameters are reduced, making the model more lightweight and easy to deploy in practical applications, while maintaining high compression performance; it not only performs excellently on synthetic datasets but also achieves good performance in real-world video compression tasks.

[0102] The present invention has broad application prospects. The following are some known and potential technical / product application fields and their application methods: By using the video compression technology of the present invention, streaming media platforms can transmit high-quality video content at a lower data rate, thereby reducing bandwidth costs and improving the user experience. The security monitoring field can adopt the technology of the present invention to efficiently store and transmit video data while maintaining high-definition requirements for details for post-event analysis and evidence preservation. In a mobile network environment, using the video compression technology of the present invention can optimize the quality and smoothness of video calls, especially under network conditions with limited or unstable bandwidth. Video editing software can integrate the technology of the present invention to support more efficient video file processing and storage, speed up the editing process and reduce hardware resource consumption. Enterprises and cloud service providers can utilize the video compression technology of the present invention to provide a more cost-effective solution for users' data storage and reduce storage costs. By optimizing the compression of video content through the present invention, video transmission and distribution networks (CDNs) can more efficiently cache and distribute video data, improving the overall performance and reliability of the system. Virtual reality (VR) and augmented reality (AR) applications can integrate the video compression technology of the present invention to support higher-quality image transmission and rendering, enhancing the user's immersive experience. In intelligent transportation monitoring and management, it can be used to optimize the compression and analysis of video data captured by traffic cameras, improving processing efficiency. Independent video producers and content creators can utilize the present invention to produce and share high-quality video works while controlling the file size for easy dissemination on social platforms. Educational institutions can adopt the video compression technology of the present invention to provide high-quality video teaching resources and optimize the performance of online teaching platforms. In the medical field, the present invention can be used to efficiently store and transmit high-resolution medical imaging data such as MRI and CT scans, facilitating remote diagnosis and medical cooperation. Cloud gaming platforms can utilize the video compression technology of the present invention to provide high-quality game streaming services, reducing latency and improving the player experience. In satellite imagery and geographic information systems (GIS), when processing and transmitting high-resolution satellite imagery data, the present invention can reduce the required storage space and transmission bandwidth. Smart devices and Internet of Things (IoT) devices can integrate the video compression technology of the present invention to optimize the transmission of video data, extend the battery life of the devices and improve network efficiency. These application fields demonstrate the diversity and flexibility of the present invention and are expected to have a significant impact on the field of video processing and transmission. With further development and optimization of the technology, its application scope may be further expanded.

[0103] By combining variational autoencoders with advanced mathematical and computer science methods, this video compression method based on improving the entropy model with variational autoencoders enhances the performance and efficiency of video compression in multiple aspects. In the feature extraction stage, the application of the spatio-temporal pyramid structure and persistent homology algorithm enables the system to capture the multi-scale spatio-temporal features and topological structure of video data, which enhances the ability to represent complex video content. This high-quality feature extraction lays a solid foundation for the subsequent encoding process, enabling the system to more accurately capture the temporal correlation and spatial structure in the video, thereby maximizing the retention of detail information in the video while maintaining a high compression ratio. In the entropy model encoding stage, the combination of the improved hierarchical conditional variational autoencoder (HCVAE) and algebraic geometry methods improves the accuracy and flexibility of probability distribution estimation. The multi-level structure of HCVAE enables the system to capture video features at different levels of abstraction, while algebraic geometry methods provide a powerful tool for dealing with complex non-linear data distributions. This combination not only improves the encoding efficiency but also enhances the system's adaptability to different types of video content, especially performing well in processing highly dynamic and complex video sequences. The context-aware adaptive quantization (CS-CAAQ) method based on compressive sensing theory introduced in the adaptive quantization stage improves the efficiency and adaptability of the quantization process. It can dynamically adjust the quantization strategy according to the local and global features of the video content, maximizing the compression efficiency while ensuring visual quality. Especially when dealing with complex and changing video scenes, this adaptive quantization strategy can improve the compression performance. The dynamic entropy coding optimization (DECO) method in the bitstream coding stage further improves the encoding efficiency, especially performing well in processing long-sequence videos or videos with changing content. By dynamically adjusting the symbol probability model and encoding strategy, the system can continuously maintain high encoding performance and effectively cope with changes in video content. The rate-distortion optimization method based on information geometry introduced in the global optimization stage realizes the end-to-end optimization of the entire compression process. This global optimization strategy can coordinate the mutual influence between various modules, find the optimal parameter configuration under different compression ratio requirements, and improve the overall performance and adaptability of the system. In the decoding and reconstruction stage, the application of the perception-guided reconstruction enhancement (PGRE) technique improves the visual quality of the reconstructed video. By combining perceptual quality assessment, adaptive post-processing, and adversarial learning techniques, the system can effectively recover the visual information lost during the compression process and optimize the reconstruction results according to the human visual characteristics. While maintaining a high compression ratio, it ensures the high perceptual quality of the reconstructed video, which is especially suitable for video applications requiring high visual quality.

[0104] The video compression method based on the improved entropy model of variational autoencoder of the present invention realizes an overall improvement in video compression performance. It not only improves the compression efficiency and reconstruction quality, but also enhances the adaptability of the system to different types of video content. It is especially suitable for processing complex video content with high resolution and high frame rate, and can achieve a high compression ratio while maintaining high visual quality, providing strong technical support for application fields such as video streaming, video surveillance, and telemedicine that require efficient compression and high-quality reconstruction. In addition, the adaptability and scalability of this method enable it to adapt to the development of future video technologies, such as emerging video formats like 8K resolution and high dynamic range (HDR), laying a foundation for the development of the next generation of video compression technology.

[0105] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.

Claims

1. A video compression method based on a variational autoencoder improved entropy model, characterized in that: The steps include: S1. Receive the current frame data in the video stream, and construct a spatiotemporal pyramid structure based on the current frame data; apply the continuous homology algorithm to the spatiotemporal pyramid structure to generate a multi-scale persistence map; fuse the multi-scale persistence map with the preset traditional visual features to obtain enhanced temporal context features; extract super prior data and potential prior data based on the current frame data, and splice the temporal context features, super prior data and potential prior data to form an input feature set; S2, based on the input feature set, an improved hierarchical conditional variational autoencoder is used to generate multi-level latent variables; Based on multi-level latent variables, the probability distribution parameters are estimated using algebraic geometry methods; based on the probability distribution parameters, the probability value of each latent variable under predetermined conditions is calculated to generate probability distribution data; S3, based on multi-level latent variables and probability distribution data, context-aware adaptive quantization based on compressed sensing theory is performed to obtain quantized data; S4, based on the quantized data, dynamic entropy coding optimization is performed to obtain a compressed data packet; Based on the final compressed data packet, performing rate-distortion optimization based on information geometry to obtain an optimized compressed data packet; S5. Based on the optimized compressed data packet, a perception-guided decoding reconstruction is performed to obtain a reconstructed video frame with optimal perception quality; Step S1 is further as follows: S11, reading the original video data stream from the video input device or storage medium, extracting the video frame data at the current moment, and obtaining the current frame data; temporarily storing the current frame data in a buffer, and recording its timestamp information; preprocessing the current frame data in the buffer, including denoising, color space conversion, and resolution adjustment, to obtain the preprocessed current frame data; S12, constructing a spatiotemporal pyramid structure based on the preprocessed current frame data, wherein each level of the spatiotemporal pyramid structure contains data of different time scales and spatial resolutions; Apply the persistence homology algorithm to the data at each level to calculate the topological features; summarize the topological features to generate a multi-scale persistence graph; S13, based on the preprocessed current frame data, using the pre-trained convolutional neural network, extracting traditional visual features; using a weighted fusion method, combining the multi-scale persistence map and the traditional visual features to obtain enhanced temporal context features; S14, extracting super-prior data from a pre-trained neural network model based on the pre-processed current frame data; extracting potential prior data based on the encoding results of the pre-stored historical frames; standardizing the super-prior data and the potential prior data to obtain standardized super-prior data and potential prior data; splicing the standardized super-prior data and the potential prior data with the enhanced temporal context features to form a complete input feature set; Step S2 is further as follows: S21, inputting the input feature set into the encoder network to generate initial latent variables; obtaining conditional information of the current frame data, combining the conditional information with the initial latent variables, and constructing a multi-level latent variable generation network model; The condition information includes scene type and motion complexity; S22. Based on the conditional information and the initial latent variables, an adaptive prior network is used to dynamically adjust the prior distribution parameters of each layer in the multi-level latent variable generation network model to obtain an adjusted multi-level latent variable generation network model; using the adjusted multi-level latent variable generation network model, latent variables of different abstraction levels are generated layer by layer; and latent variables of all levels are combined to form a multi-level latent variable; S23, map multi-level latent variables to predefined algebraic clusters to construct a geometric structure representation of probability distribution; Obtain the encoding results of the most recent N frames from the pre-stored data cache as observation data, and estimate the parameters on the algebraic cluster using algebraic statistical methods based on the observation data and geometric structure representation; use semi-algebraic set theory to constrain the space of parameters on the algebraic cluster using pre-set constraints to obtain the constrained parameters; wherein N is a natural number greater than 0; S24, constructing a conditional probability model of latent variables based on the constrained parameters, calculating the probability value of each latent variable under predetermined conditions, and generating probability distribution data; Step S3 is further as follows: S31, using a pre-trained sparse dictionary, projecting the multi-level latent variables into a sparse domain to obtain a sparse domain representation; based on the content features of the current frame data, selecting a quantization matrix from a preset quantization matrix library; Based on the probability distribution data, a context analysis module is used to dynamically adjust the parameters of the quantization matrix to obtain the adjusted quantization parameters; S32, performing a quantization operation based on the sparse domain representation and the adjusted quantization parameter to obtain a quantized discrete value, that is, quantized data; S33, based on the quantized discrete values, using a convex optimization algorithm to reconstruct the original data; Calculating the perceived quality difference between the reconstructed original data and the multi-level latent variables, and further fine-tuning the adjusted quantization parameters based on the perceived quality difference to obtain the fine-tuned quantization parameters; Step S4 is further as follows: S41, obtaining statistical characteristics of the most recent N coded frames from a pre-stored coding statistics cache, and dynamically updating a pre-configured symbol occurrence probability model based on the statistical characteristics and quantized data to obtain an updated symbol occurrence probability model; S42, assigning an optimal coding length to each symbol based on the updated symbol occurrence probability model; Based on the coding length and the quantized data, a context adaptive arithmetic coding algorithm is used to calculate the coded data; Based on the encoded data, a variable-length encoding strategy is constructed and optimized to obtain optimized encoded data; Organize the optimized coded data into a bit stream according to a predefined format, including header information and metadata, to form a final compressed data packet; S43, based on the final compressed data packet, construct a statistical manifold and calculate the Fisher information matrix; Based on the Fisher information matrix, an iterative optimization algorithm based on natural gradient is constructed. Based on the current parameters of the iterative optimization algorithm, the rate-distortion performance under the current parameter configuration is evaluated until the preset number of iterations is reached. Based on the evaluation results, the optimal parameter configuration is output to obtain the optimized compressed data packet. Step S5 is further as follows: S51, based on the header information and metadata in the optimized compressed data packet, parsing the encoding parameters, and reversely executing steps S4 to S2 to obtain preliminarily reconstructed video frame data; using a pre-trained perceptual quality assessment module, assessing the perceptual quality of the preliminarily reconstructed video frame data to obtain an assessment result; S52, based on the evaluation result, adaptively selecting and applying a post-processing technology from a preset post-processing technology library, including deblocking filtering and detail enhancement, to process the preliminarily reconstructed video frame data to obtain a processed frame; Input the processed frames into the pre-trained adversarial network to improve the visual quality and obtain enhanced frames; S53, comparing the enhanced frame with the current frame data in step S1, calculating the perceptual similarity; adjusting the post-processing parameters based on the perceptual similarity; Based on the adjusted post-processing parameters, the reconstructed video frames with optimal perceptual quality are output.

2. The video compression method based on the variational autoencoder improved entropy model according to claim 1, characterized in that: In step S31, based on the probability distribution data, the context analysis module is used to dynamically adjust the parameters of the quantization matrix, and the adjusted quantization parameters are further obtained as follows: S311. Based on multi-level latent variables, use convolutional neural networks to extract local features, including edges and textures; use long short-term memory networks to extract global features, including motion trajectories and scene changes; S312, based on local features, global features and probability distribution data, use the attention mechanism for weighted fusion to obtain a comprehensive feature vector; S313. Based on the comprehensive feature vector, a reinforcement learning algorithm is used to dynamically adjust the parameters of the quantization matrix to obtain adjusted quantization parameters.

3. A video compression system based on a variational autoencoder improved entropy model, characterized in that: include: at least one processor; as well as, a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the video compression method based on the variational autoencoder improved entropy model as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Method and system for image processing

    CA3152644A1

  • Image encoding and decoding, video encoding and decoding: methods, systems, and training methods

    CN116584098A