Learning video compression method and system based on spatial correlation and layering

By using a learning video compression method based on spatial correlation and layering, dynamically adjusting the encoding order, and utilizing motion vectors and temporal context, the problem of low encoding efficiency in existing video compression methods is solved, achieving more efficient video compression and better image quality.

CN120640015APending Publication Date: 2025-09-12HOHAI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510739439.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing video compression methods cannot dynamically adjust the encoding order according to the autocorrelation of video content during entropy coding, resulting in a decrease in coding efficiency and failing to fully utilize the correlation between channels, limiting the effective use of spatiotemporal information.

Method used

Through a learning video compression method based on spatial correlation and layering, using motion vector and temporal context extraction technology, combined with a spatial correlation prior entropy model and layered temporal attention, the encoding order is dynamically adjusted, the entropy coding process is optimized, and the coding efficiency is improved.

Benefits of technology

The utilization rate of spatial and temporal information in entropy coding is improved, the image quality of reconstructed frames and the compression ratio of video frames are improved, and the limitations of existing methods in the utilization of spatial and temporal information are overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640015A_ABST
    Figure CN120640015A_ABST
Patent Text Reader

Abstract

The invention discloses a learning video compression method and system based on spatial correlation and layering, and relates to the technical field of video coding and decoding. The invention discloses a learning video compression method and system based on spatial correlation and layering, and the method comprises the following steps: selecting a current frame and a reference frame in a given video sequence, and sequentially carrying out the preprocessing, motion vector compression processing, layering time attention refined time context compression processing and frame reconstruction processing, and obtaining a reconstruction frame of the current frame; a loss function is constructed, a guided entropy model is coded, and an evaluation index model is trained; and performing the operation processing on the current frame to be compressed and the reference frame, compressing the video frame by the training model, and outputting the compressed and reconstructed current frame. According to the learning video compression method based on spatial correlation and layering, the problem of low utilization rate of spatial-temporal information of entropy coding is solved, and the image quality of reconstructed frames and the compression ratio of video frames are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video coding and decoding, and in particular to a learning video compression method and system based on spatial correlation and layering. Background Art

[0002] With the growing demand for high-quality video content in modern media applications such as streaming services and video conferencing, video compression technology has become a key enabling technology. Significant progress has been made in the field of video compression over the past few decades, with traditional standards such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC achieving high visual quality while reducing data rates. These methods rely on techniques such as transform coding, motion compensation, quantization, and entropy coding to achieve compression, but their coding strategies are typically fixed and lack flexibility, making end-to-end rate-distortion optimization difficult. This limits the potential of deep learning models for modern video compression.

[0003] In recent years, learning-based video compression methods have gradually emerged, driving a paradigm shift in the field. These methods achieve higher compression efficiency and superior reconstruction quality through deep learning models. However, despite extensive exploration of entropy coding techniques in image compression (such as autoregressive entropy coding, super-prior entropy coding, checkerboard entropy coding, channel autoregressive entropy coding, and hybrid probabilistic model entropy coding), most methods in video compression still rely on existing entropy models used in image compression and fail to fully utilize the inherent spatiotemporal information of video. For example, existing methods (such as DVC, DVCPro, and DCVC) directly utilize autoregressive entropy coding in image compression and use super-priors to provide auxiliary information. To address the parallelization issues of autoregressive models, DCVC-HEM proposed a spatiotemporal hybrid entropy coding scheme, replacing serial autoregressive entropy coding with a more efficient two-step checkerboard method. Subsequently, DCVC-DC was further optimized to a four-step checkerboard method using a quadtree entropy coding strategy.

[0004] However, these checkerboard-based entropy coding schemes typically use fixed coding patterns and are unable to dynamically adjust the coding order based on the autocorrelation of the video content. When sections with strong autocorrelation are preferentially encoded, sections with weaker autocorrelation may not fully utilize prior information, resulting in reduced coding efficiency. Furthermore, existing methods fail to fully capture inter-channel correlations during channel-level segmentation, limiting the effective use of spatiotemporal information. Summary of the Invention

[0005] The purpose of the present invention is to provide a learning-based image compression method and system based on hybrid local and non-local correlation, so as to solve the problem that the checkerboard-based entropy coding scheme cannot dynamically adjust the coding order according to the autocorrelation of the video content, and when the parts with stronger autocorrelation are preferentially encoded, the parts with weaker autocorrelation may not be able to fully utilize the prior information, resulting in a decrease in coding efficiency. The existing methods fail to fully capture the correlation between channels in the channel-level segmentation processing, which limits the effective use of spatiotemporal information.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] A video compression method based on spatial correlation and layering learning includes the following steps:

[0008] Step 1: Select the current frame and reference frame in a given video sequence for preprocessing to obtain the motion vector between the current frame and the reference frame and the potential representation of the current frame;

[0009] Step 2: Process the motion vector between the current frame and the reference frame in step 1 and the reference motion vector in the buffer to obtain a reconstructed motion vector;

[0010] Step 3: Perform temporal context extraction on the motion vector reconstructed in step 2 and the reference features in the buffer to obtain the temporal context of the current frame;

[0011] Step 4: Perform context encoding and decoding and frame reconstruction on the potential representation of the current frame in step 1, the temporal context of the current frame in step 3, and the reference potential representation in the buffer to obtain the reconstructed features and reconstructed frame of the current frame;

[0012] Step 5: Using a spatial correlation prior entropy model to predict the probability distribution of the latent representation to be encoded in step 1 and the motion vector to be encoded in step 2, to obtain a probability distribution, a reconstructed latent representation, and a reconstructed motion vector;

[0013] Step 6: Construct a loss function to calculate the error between the reconstructed frame of the current frame and the current frame, and use the probability distribution of step 5 to calculate the bits required to reconstruct the current frame. Train the evaluation index model under the PSNR metric and the MS-SSIM metric to obtain a trained model.

[0014] In step 7, the operations of steps 1 to 5 are performed on the current frame to be compressed and the reference frame, and the model trained in step 6 is input to compress the video frame, and the compressed and reconstructed current frame is output.

[0015] Furthermore, the specific content of step 3 is:

[0016] Extract the temporal context from the reconstructed motion vector and the reference features in the buffer to obtain a rough temporal context;

[0017] The coarse temporal context is processed by averaging, convolution and activation function to obtain the channel attention of the hierarchical temporal context;

[0018] The channel attention of the coarse temporal context and the hierarchical temporal context are multiplied to obtain the final temporal context.

[0019] Furthermore, the specific content of step 5 is:

[0020] Extract spatial correlation between the latent representation and motion vector to be encoded to obtain spatial correlation prior;

[0021] Perform contrast processing on the spatial correlation prior to obtain the prior mask;

[0022] Multiply the latent representation and motion vector to be encoded by the mask to obtain latent representations and motion vectors with different spatial correlations;

[0023] The super-prior side information, the reference potential representation (the reference motion vector if the coded motion vector is used), and the temporal context (none if the coded motion vector is used) are fused a priori to obtain the prior mean and variance.

[0024] Encoding the prior mean and variance as well as the potential representations and motion vectors of different spatial correlations to obtain quantized potential representations and motion vectors as well as new prior mean and variance;

[0025] The encoded latent representations and motion vectors of different spatial correlations as well as the prior means and variances are added together to obtain the final probability distribution and the reconstructed latent representation and reconstructed motion vector.

[0026] Furthermore, in step 6, according to the rate-distortion optimization function, different evaluation index models are trained using loss functions under the PSNR metric standard and the MS-SSIM standard. The formula is as follows:

[0027]

[0028] Where λ is a Lagrange multiplier, which is the coordination of distortion D and rate R, and d(·) represents the current frame x t and the current reconstructed frame The distortion of the PSNR standard is d(·) represents the mean square error; under the MS-SSIM standard, d(·) represents 1-MSSSIM, R mv , R y represent the number of bits per pixel required to encode motion vectors and latent representation, respectively.

[0029] The present invention provides the following technical solutions:

[0030] A video compression system based on spatial correlation and layering learning, comprising:

[0031] A preprocessing module is configured to select a current frame and a reference frame in a given video sequence, and obtain a motion vector between the current frame and the reference frame and a potential representation of the current frame; wherein the reference frame is a reconstructed frame of the previous frame of the current frame;

[0032] a motion compression module, configured to obtain a reconstructed motion vector based on the motion vector between the previous frame and the reference frame output by the motion estimation module;

[0033] A temporal context extraction module, configured to obtain a temporal context of a current frame based on the reconstructed motion vector output by the motion compression module and a reference feature in a buffer; wherein the reference feature is a reconstructed feature of a frame previous to the current frame;

[0034] A context compression module is used to obtain a reconstructed potential representation based on the potential representation of the current frame output by the preprocessing module, the temporal context of the current frame output by the temporal context extraction module, and the reference potential representation in the buffer; wherein the reference potential representation is the reconstructed potential representation of the previous frame of the current frame.

[0035] a frame reconstruction module, configured to obtain a reconstructed frame of a current frame and a reconstructed feature of the current frame based on the reconstructed potential representation output by the context compression module;

[0036] a spatial correlation prior entropy model module, configured to obtain a probability distribution and a quantized latent representation and a quantized motion vector based on the motion vector output by the motion compression module and the latent representation output by the context compression module;

[0037] The model training module is used to construct a loss function, calculate the error between the reconstructed frame of the current frame and the current frame, and calculate the bits required to reconstruct the current frame using the probability distribution output by the spatial correlation prior entropy model module; using the constructed loss function, the evaluation index model is trained under the PSNR metric standard and the MS-SSIM metric standard respectively to obtain a trained model.

[0038] The beneficial effects of the present invention are as follows: on the one hand, this application effectively improves the utilization of spatial information during entropy coding based on spatial correlation priors, thereby improving the image quality of reconstructed frames; on the other hand, it effectively extracts the utilization of temporal information in the temporal context using layered temporal attention, further improving the compression ratio of video frames. This overcomes the limitations of existing video compression methods in utilizing spatiotemporal information during entropy coding, effectively improving the efficiency and quality of video compression.

[0039] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a flow chart showing an embodiment of the present invention.

[0041] Figure 2 FIG. 4 is a flowchart of spatial correlation a priori entropy coding according to an embodiment of the present invention.

[0042] Figure 3 Flowchart of hierarchical temporal attention according to one embodiment of the present invention. DETAILED DESCRIPTION

[0043] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0044] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0045] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0046] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0047] See Figure 1-3A video compression method based on spatial correlation and layered learning as shown in a preferred embodiment of the present application includes the following steps:

[0048] Step 1: For a given video sequence}x 1, x 2,…, x n-1, x n} for preprocessing, the present invention uses the spynet optical flow network (opticalflow net) to extract the current frame x t and reference frame Motion vector information v between two adjacent frames t . Use the context encoder to extract the output frame x t The high-dimensional latent representation y t .

[0049] Step 2: The motion vector v between the current frame and the reference frame in step 1 is t Input the motion encoder, then input it into the super prior module to obtain the side information of the motion vector Utilizing side information of motion vectors and the reference motion vector in the buffer As a priori information, the spatial correlation prior entropy model is used to input the motion vector v t Then, the result of entropy coding is input into the motion decoder to obtain the reconstructed motion vector

[0050] Step 3: reconstruct the motion vector from step 2 and reference features in the buffer Perform temporal context extraction to obtain a rough temporal context of the current frame The formula is as follows:

[0051]

[0052] Where T represents temporal context extraction;

[0053] like Figure 3 As shown, the rough time context As input, it is input to the hierarchical temporal attention module. First, the rough temporal context of its non-channel part is calculated. The mean of The formula is as follows:

[0054]

[0055] Among them, mean means calculating the mean of its non-channel part.

[0056] Next, Input a 3×3 convolution layer and a sigmoid activation function layer to obtain the channel attention map in the time context. The formula is as follows:

[0057]

[0058] Finally, the attention map is combined with the coarse temporal context Multiply to get a temporal context rich in channel information The formula is as follows:

[0059]

[0060] in, Represents multiplication.

[0061] The present invention sets different context quality requirements for each level and adopts a hierarchical channel attention mechanism. Specifically, the corresponding network model is dynamically selected according to the index of the video frame. The index of the video frame is frameindex, and the index calculation method is as follows:

[0062] index=frameindex%4

[0063] Therefore, the channel attention map can be expressed in more detail using the following formula:

[0064]

[0065] Among them, the network weight parameters of different layers are different.

[0066] Step 4: The potential representation y of the current frame in step 1 t , Step 3 Time context of the current frame and the reference latent representation in the buffer Perform context encoding and decoding. In the time context of the current frame and the reference latent representation in the buffer With the help of t is reconstructed as the latent representation of the current frame

[0067] Next, the latent representation of the current frame is reconstructed The input frame reconstruction module processes and obtains the reconstruction features of the current frame and reconstructed frames

[0068] Step 5: The potential representation to be encoded in step 1 and the motion vector to be encoded in step 2 The spatial correlation prior entropy model is used to predict its probability distribution. According to the conditional entropy principle, the entropy value H is defined as follows:

[0069]

[0070] Among them, params is the prior parameters used in the encoding process. The higher the correlation, the lower the entropy value, and the fewer bits required for encoding.

[0071] In order to give priority to encoding positions with less reference information during the encoding process, this paper designs a mask based on spatial correlation to optimize the encoding order. First, the target y is captured by a 3×3 convolution operation. t The spatial correlation of , and normalized, to get y 空间相关 , the formula is as follows:

[0072] y 空间相关 =normalize(conv(y t ))

[0073] According to y 空间相关 The spatial correlation of the representation is divided into four parts (denoted as masks i = 1, 2, ..., n, where n = 4) and sorted from low to high correlation;

[0074] Finally, the mask is compared with y t Multiply them together to generate the corresponding checkerboard pattern. By integrating the prior information With the coded information as a reference, checkerboard entropy coding is performed step by step, and finally all coding results are summed up to get the final The formula is as follows:

[0075]

[0076] in, Represents multiplication, in encoding motion vector When params includes hyper-prior boundary information and the reconstructed motion vector of the previous frame In encoding latent representation When params includes hyper-prior boundary information Reconstructed latent representation of the previous frame and time context

[0077] Finally, we get the probability distribution (μ t ,σ t ) and reconstruct the latent representation and reconstructed motion vectors

[0078] Step 6: Construct a loss function to calculate the error between the reconstructed frame of the current frame and the current frame, and use the probability distribution of step 5 to calculate the bits required to reconstruct the current frame. Under the PSNR metric and MS-SSIM standard, use the loss function to train different evaluation index models. The formula is as follows:

[0079]

[0080] Where λ is a Lagrange multiplier, which is the coordination of distortion D and rate R, and d(·) represents the current frame x t and the current reconstructed frame In the PSNR standard, d(·) represents the mean square error, and λ is set to {85, 170, 380, 840}; in the MS-SSIM standard, d(·) represents 1-MSSSIM, and λ is set to {8, 16, 32, 64}, R mv , R y represent the number of bits per pixel required to encode motion vectors and latent representation, respectively.

[0081] The present invention utilizes gradient descent to optimize hyperparameters, employing, including but not limited to, the Adam optimizer for parameter updating and learning. During network training, the present invention uses the Vimeo-90k dataset as a training set, which contains 89,800 video clips, each containing 7 consecutive frames. The training process includes 35 rounds, with a batch size of 4 for each round. The learning rate is set to 1e-4 for the first 26 rounds, 5e-5 for rounds 26 to 28, 5e-6 for rounds 28 to 30, and 1e-6 for rounds 30 to 32. Finally, a cascaded training strategy is used for the final 32 to 35 rounds to reduce error propagation, with a learning rate set to 5e-6.

[0082] Step 7: When t=1, use the image compression model to compress; when t≥2, the current frame x to be compressed t and reference frame Perform steps 1 to 5, and input the trained model in step 6 to compress the video frame, and output the compressed and reconstructed current frame. As the next frame x t+1 reference frame.

[0083] The present invention proposes a learning video compression system based on spatial correlation and layering, comprising:

[0084] A preprocessing module is configured to select a current frame and a reference frame in a given video sequence, and obtain a motion vector between the current frame and the reference frame and a potential representation of the current frame; wherein the reference frame is a reconstructed frame of the previous frame of the current frame;

[0085] a motion compression module, configured to obtain a reconstructed motion vector based on the motion vector between the previous frame and the reference frame output by the motion estimation module;

[0086] A temporal context extraction module, configured to obtain a temporal context of a current frame based on the reconstructed motion vector output by the motion compression module and a reference feature in a buffer; wherein the reference feature is a reconstructed feature of a frame previous to the current frame;

[0087] A context compression module is used to obtain a reconstructed potential representation based on the potential representation of the current frame output by the preprocessing module, the temporal context of the current frame output by the temporal context extraction module, and the reference potential representation in the buffer; wherein the reference potential representation is the reconstructed potential representation of the previous frame of the current frame.

[0088] a frame reconstruction module, configured to obtain a reconstructed frame of a current frame and a reconstructed feature of the current frame based on the reconstructed potential representation output by the context compression module;

[0089] a spatial correlation prior entropy model module, configured to obtain a probability distribution and a quantized latent representation and a quantized motion vector based on the motion vector output by the motion compression module and the latent representation output by the context compression module;

[0090] The model training module is used to construct a loss function, calculate the error between the reconstructed frame of the current frame and the current frame, and calculate the bits required to reconstruct the current frame using the probability distribution output by the spatial correlation prior entropy model module; using the constructed loss function, the evaluation index model is trained under the PSNR metric standard and the MS-SSIM metric standard respectively to obtain a trained model.

[0091] In summary, the present invention proposes a learning video compression system based on spatial correlation and layering, which overcomes the limitations of existing video compression methods in utilizing spatiotemporal information during entropy coding. The method includes the following steps: selecting a current frame and a reference frame in a given video sequence for preprocessing to obtain a motion vector between the current frame and the reference frame and a potential representation of the current frame; processing the motion vector between the current frame and the reference frame and the reference motion vector in the buffer to obtain a reconstructed motion vector; performing temporal context extraction processing on the reconstructed motion vector and the reference features in the buffer to obtain a temporal context of the current frame; performing context encoding and decoding on the potential representation of the current frame, the temporal context of the current frame and the reference potential representation in the buffer; Frame reconstruction processing is performed to obtain the reconstructed features and reconstructed frame of the current frame; the spatial correlation prior entropy model is used to predict the probability distribution of the potential representation to be encoded in step 1 and the motion vector to be encoded in step 2, and the probability distribution and reconstructed potential representation and reconstructed motion vector are obtained; a loss function is constructed to calculate the error between the reconstructed frame of the current frame and the current frame and the bits required to calculate the reconstructed frame of the current frame using the probability distribution of step 5, and the evaluation index model is trained under the PSNR metric and the MS-SSIM metric respectively to obtain a trained model; the above operations are performed on the current frame to be compressed and the reference frame, and the trained model is input to compress the video frame, and the compressed and reconstructed current frame is output. The present invention proposes a learning video compression method based on spatial correlation and layering. On the one hand, based on the spatial correlation prior, it effectively improves the utilization rate of spatial information during entropy coding and improves the image quality of the reconstructed frame; on the other hand, it uses layered temporal attention to effectively extract the utilization rate of temporal information in the temporal context, further improving the compression ratio of the video frame.

[0092] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0093] The above-described embodiments merely illustrate the implementation methods of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.

Claims

1. A video compression method based on spatial correlation and layering learning, characterized in that The steps include: Step 1: Select the current frame and reference frame in a given video sequence for preprocessing to obtain the motion vector between the current frame and the reference frame and the potential representation of the current frame; Step 2: Process the motion vector between the current frame and the reference frame in step 1 and the reference motion vector in the buffer to obtain a reconstructed motion vector; Step 3: Perform temporal context extraction on the motion vector reconstructed in step 2 and the reference features in the buffer to obtain the temporal context of the current frame; Step 4: Perform context encoding and decoding and frame reconstruction on the potential representation of the current frame in step 1, the temporal context of the current frame in step 3, and the reference potential representation in the buffer to obtain the reconstructed features and reconstructed frame of the current frame; Step 5: Using a spatial correlation prior entropy model to predict the probability distribution of the latent representation to be encoded in step 1 and the motion vector to be encoded in step 2, to obtain a probability distribution, a reconstructed latent representation, and a reconstructed motion vector; Step 6: Construct a loss function to calculate the error between the reconstructed frame of the current frame and the current frame, and use the probability distribution of step 5 to calculate the bits required to reconstruct the current frame. Train the evaluation index model under the PSNR metric and the MS-SSIM metric to obtain a trained model. In step 7, the operations of steps 1 to 5 are performed on the current frame to be compressed and the reference frame, and the model trained in step 6 is input to compress the video frame, and the compressed and reconstructed current frame is output.

2. A video compression method based on spatial correlation and layering learning as claimed in claim 1, characterized in that: The specific content of step 3 is: Extract the temporal context from the reconstructed motion vector and the reference features in the buffer to obtain a rough temporal context; The coarse temporal context is processed by averaging, convolution and activation function to obtain the channel attention of the hierarchical temporal context; The channel attentions of the coarse temporal context and the hierarchical temporal context are multiplied to obtain the final temporal context.

3. A video compression method based on spatial correlation and layering learning as claimed in claim 1, characterized in that: The specific content of step 5 is: Extract spatial correlation between the latent representation and motion vector to be encoded to obtain spatial correlation prior; Perform contrast processing on the spatial correlation prior to obtain the prior mask; Multiply the latent representation and motion vector to be encoded by the mask to obtain latent representations and motion vectors with different spatial correlations; The super-prior side information, the reference potential representation (the reference motion vector if the coded motion vector is used), and the temporal context (none if the coded motion vector is used) are fused a priori to obtain the prior mean and variance. Encoding the prior mean and variance as well as the potential representations and motion vectors of different spatial correlations to obtain quantized potential representations and motion vectors as well as new prior mean and variance; The encoded latent representations and motion vectors of different spatial correlations as well as the prior means and variances are added together to obtain the final probability distribution and the reconstructed latent representation and motion vector.

4. A video compression method based on spatial correlation and layering learning as claimed in claim 1, characterized in that: In step 6, according to the rate-distortion optimization function, different evaluation index models are trained using loss functions under the PSNR metric standard and the MS-SSIM standard. The formula is as follows: Where λ is a Lagrange multiplier, which is the coordination of distortion D and rate R, and d(·) represents the current frame x t and the current reconstructed frame The distortion of the PSNR standard is d(·) represents the mean square error; under the MS-SSIM standard, d(·) represents 1-MSSSIM, R mv , R y represent the number of bits per pixel required to encode motion vectors and latent representation, respectively.

5. A video compression system based on spatial correlation and layering learning, characterized in that include: A preprocessing module is configured to select a current frame and a reference frame in a given video sequence, and obtain a motion vector between the current frame and the reference frame and a potential representation of the current frame; wherein the reference frame is a reconstructed frame of the previous frame of the current frame; a motion compression module, configured to obtain a reconstructed motion vector based on the motion vector between the previous frame and the reference frame output by the motion estimation module; A temporal context extraction module, configured to obtain a temporal context of a current frame based on the reconstructed motion vector output by the motion compression module and a reference feature in a buffer; wherein the reference feature is a reconstructed feature of a frame previous to the current frame; A context compression module is used to obtain a reconstructed potential representation based on the potential representation of the current frame output by the preprocessing module, the temporal context of the current frame output by the temporal context extraction module, and the reference potential representation in the buffer; wherein the reference potential representation is the reconstructed potential representation of the previous frame of the current frame. a frame reconstruction module, configured to obtain a reconstructed frame of a current frame and a reconstructed feature of the current frame based on the reconstructed potential representation output by the context compression module; a spatial correlation prior entropy model module, configured to obtain a probability distribution and a quantized latent representation and a quantized motion vector based on the motion vector output by the motion compression module and the latent representation output by the context compression module; The model training module is used to construct a loss function, calculate the error between the reconstructed frame of the current frame and the current frame, and calculate the bits required to reconstruct the current frame using the probability distribution output by the spatial correlation prior entropy model module; using the constructed loss function, the evaluation index model is trained under the PSNR metric standard and the MS-SSIM metric standard respectively to obtain a trained model.