A context-guided converter entropy coding video compression device and method

Through the context-guided Transformer conditional entropy coding video compression device, the model algorithm of the motion network and the context network is used to map the video frames. Combined with the teacher-student network architecture, the problem of difficult balance between computational efficiency and compression rate in the existing technology is solved, and efficient video compression and decoding quality improvement are achieved.

CN120343275BActive Publication Date: 2025-10-03NINGBO ORIENTAL UNIV OF TECH (TEMPORARY NAME)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510828647.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing conditional entropy coding methods face the problem of balancing computational efficiency and compression rate in temporal context modeling and spatial context dependency modeling, making it difficult to optimize both computational cost and compression efficiency at the same time.

Method used

A context-guided Transformer conditional entropy coding video compression device is adopted. The current video frame and the historical reconstructed frame are mapped into the temporal context and the current frame latent variable through the model algorithm of the motion network and the context network. The context-guided Transformer conditional entropy coding model is used for reconstruction. The teacher-student network architecture is combined to improve the training consistency and decoding efficiency of the model.

Benefits of technology

It significantly reduces the computational cost, retains sufficient time-dependent information, optimizes the utilization efficiency of spatial context, improves the decoding quality of video frames, and achieves a good balance between computational efficiency and compression rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343275B_ABST
    Figure CN120343275B_ABST
Patent Text Reader

Abstract

The present invention relates to a context-guided transformer entropy coding video compression device and method. By configuring a video compression device, a motion network and context network model algorithm are used to map an input current video frame and historical reconstructed frames into a temporal context and a current frame latent variable. By configuring a context allocation device, a context-guided Transformer conditional entropy coding model configured thereon can be invoked to obtain a reconstructed video frame based on the temporal context and the current frame latent variable. The combination of the two effectively compresses the temporal context, significantly reducing computational costs while retaining sufficient temporal dependency information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer models, and in particular to a context-guided converter entropy coding video compression device and method. Background Art

[0002] Video compression plays a key role in the efficient transmission and storage of digital video content. Traditional video compression methods rely on complex hand-crafted structures and prior knowledge, which limits their flexibility and adaptability. In recent years, deep learning-driven video compression methods have developed rapidly, especially neural video compression methods based on conditional entropy coding (hereinafter referred to as conditional entropy coding methods), which have demonstrated excellent compression performance.

[0003] Currently, the mainstream technical solutions for conditional entropy coding methods include temporal context modeling and spatial context modeling. However, existing conditional entropy coding methods still face several key challenges, mainly including the computational cost of temporal context modeling and the limitation of lacking spatial context dependency modeling.

[0004] Temporal context modeling methods include fixed-window-based temporal context modeling methods and Markov-based temporal context simplification methods. Fixed-window-based temporal context modeling methods reduce computational costs by limiting the window length of the temporal context (such as the first two frames). This method can reduce computational overhead, but it will lead to insufficient utilization of temporal information and difficulty in capturing long-term temporal dependencies, especially when processing fast-motion or complex dynamic scenes. The temporal context simplification method based on the Markov assumption simplifies video frame processing into a first-order Markov process, that is, it only relies on the information of the most recent frame for entropy estimation. Although this method can reduce computational costs, its ability to model long-range temporal dependencies is weak, resulting in reduced compression efficiency.

[0005] Spatial context modeling methods include autoregressive spatial context modeling, fixed-pattern spatial context selection, and minimum entropy greedy strategies. Autoregressive spatial context modeling methods typically employ an autoregressive strategy, gradually inferring undecoded regions based on decoded content. However, this approach imposes fixed constraints on the direction of information flow, limiting the scope of spatial dependencies. Furthermore, the step-by-step inference used during decoding results in high computational overhead, impacting real-time performance. Fixed-pattern spatial context selection methods use a fixed pattern (such as a checkerboard decoding strategy) to select the decoding order to achieve partially parallel decoding. However, this approach still relies on manually specified context positions, lacks adaptability, and cannot guarantee the selection of optimal information regions. Minimum entropy greedy strategies, based on the minimum entropy strategy, prioritize decoding regions with the lowest entropy to reduce overall bit overhead. However, due to their local greedy strategy, they fail to fully consider dependencies between spatial contexts, leading to suboptimal global decisions. Furthermore, during training, the video compression model learns based on random masks, while during inference, the optimal decoding path must be actively selected. This mismatch between training and inference can affect model generalization.

[0006] Therefore, the conditional entropy coding method in the existing technology still has deficiencies in the information utilization of temporal context modeling and spatial context dependency modeling, and it is difficult to achieve a good balance between computational efficiency and compression rate. Summary of the Invention

[0007] The technical problem to be solved by the present invention is how to overcome the technical deficiency of existing conditional entropy coding methods, which have difficulty in achieving a good balance between computational efficiency and compression rate. To overcome the above-mentioned disadvantages of the prior art, the present invention provides a context-guided transformer entropy coding video compression device and method, which specifically include a context-guided transformer entropy coding video compression device and a context-guided transformer entropy coding video compression method.

[0008] The present invention provides a context-guided converter entropy coding video compression device, comprising:

[0009] a memory configured to store historical reconstructed frames and a temporal context;

[0010] a video compression device, electrically connected to the memory, and configured to map a current video frame input to the video compression device and the historical reconstructed frame into the temporal context and current frame latent variables through a motion network and a context network model algorithm;

[0011] The context allocation device is electrically connected to the video compression device and is configured to obtain a reconstructed video frame according to the temporal context and the current frame latent variable by calling the context-guided Transformer conditional entropy coding model contained therein.

[0012] The context-guided transformer entropy coding video compression device disclosed in the present invention, by setting a video compression device, adopts a model algorithm of a motion network and a context network to map the input current video frame and the historical reconstructed frame into a temporal context and a current frame latent variable, and the obtained result has multi-scale and multi-type characteristics. By setting a context allocation device, the context-guided Transformer conditional entropy coding model contained therein can be called to obtain a reconstructed video frame based on the temporal context and the current frame latent variable. The combination of the two can effectively compress the temporal context, significantly reduce the computational cost, and retain sufficient time-dependent information. Moreover, the context-guided Transformer conditional entropy coding model can also achieve modeling of the importance and determinism of spatial context, optimize the utilization efficiency of spatial context, effectively improve the decoding quality of video frames, and thus overcome the technical defects of existing conditional entropy coding methods, which are difficult to achieve a good balance between computational efficiency and compression rate.

[0013] In one possible implementation, the memory is divided into a frame buffer area and a latent variable buffer area, both the frame buffer area and the latent variable buffer area are provided with input pins and output pins, the frame buffer area is used to store the historical reconstructed frames, and the latent variable buffer area is used to store the time context; thereby, the historical reconstructed frames and the time context can be stored separately to avoid information crosstalk.

[0014] In one possible embodiment, the video compression device includes a motion network, a context network, a frame encoder, an entropy encoder and a frame decoder, the input end of the motion network is electrically connected to the output pin of the frame buffer area, the input end of the context network is electrically connected to the output end of the motion network and the output end of the frame decoder at the same time, the input end of the frame encoder is electrically connected to the output end of the context network, the input end of the entropy encoder is electrically connected to the output end of the frame encoder, the input end of the frame decoder is electrically connected to the output end of the context network and the output end of the entropy encoder at the same time, the input pin of the frame buffer area is electrically connected to the output end of the frame decoder, the input pin of the latent variable buffer area is electrically connected to the output end of the context network, and the context allocation device is electrically connected to the output pin of the latent variable buffer area and the output end of the frame encoder at the same time.

[0015] In a possible implementation, the video compression device is configured to perform the following steps:

[0016] A1: calling the motion network to receive the input current video frame, and mapping the input current video frame and the historical reconstructed frame into a motion vector through the motion network model algorithm;

[0017] A2: calling the context network to map the motion vector and the latent variable output by the frame decoder into the temporal context through a context network model algorithm;

[0018] A3: calling the frame encoder to execute a frame encoding algorithm to encode the input current video frame and the temporal context into a current frame latent variable and a reconstructed frame latent variable;

[0019] A4: calling the entropy encoder to distribute and output the reconstructed frame latent variables according to a probability mass function;

[0020] A5: Call the frame decoder to decode the reconstructed frame latent variable assigned and output by the entropy encoder in a frame decoding manner according to the time context to obtain a reconstructed frame, and transfer the reconstructed frame as the historical reconstructed frame to the frame buffer.

[0021] The video compression device with the above technical features can effectively compress the temporal context, significantly reduce the computational cost, while retaining sufficient temporal dependency information, and further optimize the utilization efficiency of the spatial context, effectively improving the decoding quality of the video frame.

[0022] In one possible embodiment, the context-guided Transformer conditional entropy coding model includes a time context resampler, a teacher network and a student network, the input end of the time context resampler is electrically connected to the output pin of the latent variable buffer and the output end of the frame encoder at the same time, the input end of the teacher network is electrically connected to the output end of the time context resampler, the output end of the entropy encoder is electrically connected to the output end of the teacher network, the input end of the student network is electrically connected to the output end of the time context resampler, and the output end of the entropy encoder is electrically connected to the output end of the student network.

[0023] In one possible implementation, the context-guided Transformer conditional entropy coding model is configured to perform the following steps:

[0024] B1: calling the temporal context resampler to execute a window cross attention mechanism to extract and fuse features of the temporal context and the current frame latent variable to obtain a temporal context representation;

[0025] B2: Determine whether the current training phase is in the context-guided Transformer conditional entropy coding model.

[0026] If yes, proceed to the next step;

[0027] If not, proceed to step B4;

[0028] B3: calling the teacher network to obtain and output the reconstructed video frame using the temporal context representation, and obtaining the estimated parameters of the probability mass function using the reconstructed video frame through a parameter mapping layer algorithm;

[0029] B4: Calling the student network to obtain the reconstructed video frame and the estimated parameters of the probability mass function by simulating the operation mode of the teacher network.

[0030] In a possible implementation, step B3 includes the following steps:

[0031] B31: calling the teacher network to execute the window attention mechanism of Swin Transformer to generate an attention map representing the importance of spatial position and an entropy map representing the certainty of spatial position using the temporal context representation;

[0032] B32: calling the teacher network to normalize the attention map and the entropy map, and generating a spatial dependency score matrix through a weighted summation calculation formula;

[0033] B33: calling the teacher network to use the top-k algorithm to sort the positions of the elements in the spatial dependency score matrix and determine the k positions with the most context value;

[0034] B34: calling the teacher network to decode the k positions with the most context value to determine the spatial context, obtain and output the reconstructed video frame;

[0035] B35: Calling the teacher network to execute a parameter mapping layer algorithm to obtain estimated parameters of the probability mass function using the reconstructed video frame.

[0036] The context-guided Transformer conditional entropy coding model, equipped with these technical features, utilizes a teacher-student network architecture, ensuring that the model's training process is highly consistent with the inference process during production. This effectively improves the model's generalization performance and adaptability to practical decoding tasks. Furthermore, by using only the student network during inference, computing resource requirements are reduced, further enhancing decoding efficiency in practical application scenarios.

[0037] In a possible implementation, the teacher network and the student network both include a Swin Transformer encoder, a Swin Transformer decoder, and a parameter mapping layer model, the input end of the Swin Transformer encoder is electrically connected to the output end of the time context resampler, the input end of the Swin Transformer decoder is electrically connected to the output end of the Swin Transformer encoder, the input end of the parameter mapping layer model is electrically connected to the output end of the Swin Transformer decoder, and the input end of the entropy encoder is electrically connected to the output end of the parameter mapping layer model;

[0038] The Swin Transformer encoder is configured to utilize the temporal context representation to generate an attention map representing the importance of spatial positions and an entropy map representing the certainty of spatial positions through the windowed attention mechanism of the Swin Transformer; the attention map and the entropy map are then normalized, and a spatial dependency score matrix is ​​generated through a weighted sum calculation formula;

[0039] The Swin Transformer decoder is configured to use a top-k algorithm to sort the positions of elements in the spatial dependency score matrix to determine the k positions with the most context value; then decode the k positions with the most context value to determine the spatial context, obtain and output the reconstructed video frame;

[0040] The parameter mapping layer model is configured to obtain estimated parameters of the probability mass function using the reconstructed video frame through a parameter mapping layer algorithm.

[0041] The teacher network and student network with the above technical features can not only realize parameter sharing, but also further reduce the demand for computing resources and improve the decoding efficiency in actual application scenarios.

[0042] Another technical solution of the present invention is to provide a context-guided converter entropy coding video compression method, comprising the following steps:

[0043] S1: Mapping a current video frame input to the video compression device and a historical reconstructed frame into a temporal context and a current frame latent variable through a video compression device;

[0044] S2: Obtain a reconstructed video frame according to the temporal context and the current frame latent variable through a context allocating device.

[0045] The method disclosed in this application first uses a video compression device to map the input current video frame and historical reconstructed frames into a temporal context and current frame latent variables, resulting in a result with multi-scale and multi-type features. A context allocation device then uses the temporal context and current frame latent variables to obtain a reconstructed video frame. The combination of these two methods effectively compresses the temporal context, significantly reducing computational costs while retaining sufficient temporal dependency information, effectively improving the decoding quality of the video frames. This overcomes the technical drawback of existing conditional entropy coding methods, which struggle to achieve a good balance between computational efficiency and compression rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a schematic structural diagram of a context-guided transformer entropy coding video compression device disclosed in an embodiment of the present application;

[0047] Figure 2 A vector diagram corresponding to the video compression device disclosed in the embodiment of this application;

[0048] Figure 3 A vector diagram corresponding to the context-guided Transformer conditional entropy coding model disclosed in the embodiments of this application;

[0049] Figure 4 This is a graph showing the experimental curve results of the MCL-JCV dataset disclosed in the examples of this application;

[0050] Figure 5 This is a graph showing the test curve results of the UVG data set disclosed in the examples of this application;

[0051] Figure 6 This is a graph showing the experimental curve results on the HEVC-B dataset disclosed in the embodiments of this application. DETAILED DESCRIPTION

[0052] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of the present application and are not intended to limit the scope of protection of the embodiments of the present application. Those skilled in the art may adjust them as needed to suit specific application scenarios.

[0053] In the embodiments of the present application, unless otherwise clearly specified and limited, the electrical connection between the first feature and the second feature means that there is transmission of electrical signals between the first feature and the second feature, that is, there is an electrical relationship, and the way to achieve the transmission of electrical signals may be electrical connection of wires, radio connection, electrical connection of electromagnetic media (such as semiconductors), communication achieved by channels, etc.

[0054] In the embodiments of the present application, unless otherwise expressly specified or limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, a first feature being "above," "above," and "above" a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being "below," "below," and "below" a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.

[0055] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0056] See also Figures 1 to 6 The embodiment of the present application discloses a context-guided converter entropy coding video compression device, the structural diagram of the video compression device is shown in FIG. Figure 1 As shown, the video compression device includes a memory, a video compression device and a context allocation device, wherein the video compression device is electrically connected to the memory, and the context allocation device is electrically connected to the video compression device.

[0057] See also Figure 1 and Figure 2 In the video compression device, the memory is configured to store historical reconstructed frames and time context. In this embodiment, the memory is divided into a frame buffer area and a latent variable buffer area. Both the frame buffer area and the latent variable buffer area are provided with input pins and output pins. The frame buffer area is used to store historical reconstructed frames, and the latent variable buffer area is used to store time context, thereby avoiding information crosstalk.

[0058] See also Figure 1 and Figure 2 In the video compression device, the video compression device is configured to map the current video frame and the historical reconstructed frame input to the video compression device into a temporal context and a current frame latent variable through a model algorithm of a motion network and a context network. Figure 2 As shown, in this embodiment, the video compression device includes a motion network, a context network, a frame encoder, an entropy encoder and a frame decoder. The input end of the motion network is electrically connected to the output pin of the frame buffer area, the input end of the context network is electrically connected to the output end of the motion network and the output end of the frame decoder at the same time, the input end of the frame encoder is electrically connected to the output end of the context network, the input end of the entropy encoder is electrically connected to the output end of the frame encoder, the input end of the frame decoder is electrically connected to the output end of the context network and the output end of the entropy encoder at the same time, the input pin of the frame buffer area is electrically connected to the output end of the frame decoder, the input pin of the latent variable buffer area is electrically connected to the output end of the context network, and the context allocation device is electrically connected to the output pin of the latent variable buffer area and the output end of the frame encoder at the same time.

[0059] Please continue to see Figure 2 ,In the video compression device, the motion network is set to receive the current video frame as input (denoted by Instead of the current video frame, the current video frame is abbreviated as the current frame), and the motion network model algorithm is used to compare the input current video frame with the historical reconstructed frame (in the figure with Instead of the historical reconstructed frame) is mapped to a motion vector (in the figure The context network is configured to map the motion vector and the latent variable output by the frame decoder into a temporal context through a context network model algorithm. The frame encoder is configured to execute a frame encoding algorithm to encode the input current video frame and the temporal context into a current frame latent variable and a reconstructed frame latent variable (denoted by Replace the current frame hidden variable with The entropy encoder is configured to allocate and output the reconstructed frame latent variables according to the probability mass function. The frame decoder is configured to decode the reconstructed frame latent variables allocated and output by the entropy encoder in a frame decoding manner according to the temporal context to obtain a reconstructed frame, and transmit the reconstructed frame as a historical reconstructed frame to the frame buffer.

[0060] In this embodiment, the video compression device is configured to perform the following steps: A1: calling the motion network to receive the input current video frame, and mapping the input current video frame and the historical reconstructed frame into a motion vector through the motion network model algorithm; A2: calling the context network to map the motion vector and the latent variable output by the frame decoder into a time context through the context network model algorithm; A3: calling the frame encoder to execute the frame encoding algorithm to encode the input current video frame and the time context into the current frame latent variable and the reconstructed frame latent variable; A4: calling the entropy encoder to allocate and output the reconstructed frame latent variable according to the probability mass function; A5: calling the frame decoder to decode the reconstructed frame latent variable allocated and output by the entropy encoder in a frame decoding manner according to the time context, obtain the reconstructed frame, and transmit the reconstructed frame as the historical reconstructed frame to the frame buffer.

[0061] See also Figure 2 and Figure 3 In the video compression device, the context allocation device is configured to obtain a reconstructed video frame according to the temporal context and the current frame latent variable by calling the context-guided Transformer conditional entropy coding model contained therein. Figure 3As shown, the context-guided Transformer conditional entropy coding model includes a time context resampler, a teacher network and a student network. The input end of the time context resampler is electrically connected to the output pin of the latent variable buffer and the output end of the frame encoder at the same time, the input end of the teacher network is electrically connected to the output end of the time context resampler, the output end of the entropy encoder is electrically connected to the output end of the teacher network, the input end of the student network is electrically connected to the output end of the time context resampler, and the output end of the entropy encoder is electrically connected to the output end of the student network.

[0062] See also Figure 3 In the context-guided Transformer conditional entropy coding model, the temporal context resampler is set to extract and fuse the temporal context and the current frame latent variables through the window cross attention mechanism to obtain the temporal context representation. The teacher network is set to generate an attention map representing the importance of spatial positions and an entropy map representing the certainty of spatial positions by using the window attention mechanism of Swin Transformer when training the context-guided Transformer conditional entropy coding model; then the attention map and the entropy map are normalized, and the spatial dependency score matrix is ​​generated by weighted summation calculation; then the top-k algorithm is used to sort the positions of the elements in the spatial dependency score matrix to determine the k positions with the most context value; then the k positions with the most context value are decoded to determine the spatial context, obtain and output the reconstructed video frame; finally, the parameter mapping layer algorithm is used to use the reconstructed video frame to obtain the estimated parameters of the probability mass function, and the estimated parameters are the mean and standard deviation The student network is set to obtain the estimated parameters of the reconstructed video frame and probability mass function, i.e. the mean of the probability mass function, by simulating the operation of the teacher network when the context-guided Transformer conditional entropy coding model is officially running. and standard deviation . Figure 3 The “learnable tokens” in can be either the current frame latent variables or the reconstructed tokens during training, and the current frame latent variables during runtime.

[0063] Please continue to see Figure 3In this embodiment, the teacher network and the student network both include a Swin Transformer encoder, a Swin Transformer decoder and a parameter mapping layer model (referred to as the parameter mapping layer in the figure). The input end of the Swin Transformer encoder is electrically connected to the output end of the time context resampler, the input end of the Swin Transformer decoder is electrically connected to the output end of the Swin Transformer encoder, the input end of the parameter mapping layer model is electrically connected to the output end of the Swin Transformer decoder, and the input end of the entropy encoder is also electrically connected to the output end of the parameter mapping layer model. Figure 3 The paper presents a scenario where the teacher network and the student network share a Swin Transformer encoder, which can reduce computational dissipation. However, the teacher network and the student network cannot share a Swin Transformer decoder. The Swin Transformer encoder is configured to utilize the temporal context representation through the Swin Transformer's windowed attention mechanism to generate an attention map representing the importance of spatial positions and an entropy map representing the certainty of spatial positions. The attention map and entropy map are then normalized and a weighted summation formula is used to generate a spatial dependency score matrix. The Swin Transformer decoder is configured to use a top-k algorithm to sort the positions of the elements in the spatial dependency score matrix and determine the k most contextually valuable positions. The k most contextually valuable positions are then decoded to determine the spatial context, obtaining and outputting a reconstructed video frame. The parameter mapping layer model is configured to use the reconstructed video frame to obtain the estimated parameters of the probability mass function through the parameter mapping layer algorithm.

[0064] In this embodiment, the weighted sum calculation formula is as follows:

[0065] ,

[0066] Where,

[0067] represents the spatial dependency scoring matrix;

[0068] Represents the normalized attention map;

[0069] represents the normalized entropy map;

[0070] Represents the weighting coefficient.

[0071] In this embodiment, the context-guided Transformer conditional entropy coding model is configured to perform the following steps: B1: calling the temporal context resampler to execute the window cross-attention mechanism to perform feature extraction and feature fusion on the temporal context and the current frame latent variables to obtain a temporal context representation; B2: determining whether the current training phase of the context-guided Transformer conditional entropy coding model is in progress, and if so, executing the next step; if not, executing step B4; B3: calling the teacher network to obtain and output a reconstructed video frame using the temporal context representation, and obtaining estimated parameters of a probability mass function using the reconstructed video frame through a parameter mapping layer algorithm; B4: calling the student network to obtain the reconstructed video frame and the estimated parameters of the probability mass function by simulating the operation of the teacher network.

[0072] In this embodiment, step B3 includes the following steps; B31: calling the teacher network to execute the window attention mechanism of Swin Transformer to use the temporal context representation to generate an attention map representing the importance of spatial positions and an entropy map representing the certainty of spatial positions; B32: calling the teacher network to normalize the attention map and the entropy map, and generate a spatial dependency score matrix through a weighted summation calculation formula; B33: calling the teacher network to use the top-k algorithm to sort the positions of the elements in the spatial dependency score matrix, and determine the k positions with the most contextual value; B34: calling the teacher network to decode the k positions with the most contextual value to determine the spatial context, obtain and output the reconstructed video frame; B35: calling the teacher network to execute the parameter mapping layer algorithm to obtain the estimated parameters of the probability mass function using the reconstructed video frame.

[0073] The following will further disclose a method for using the context-guided converter entropy coding video compression device in this embodiment. The method includes the following steps: S1: using a video compression device to map the current video frame of the input video compression device and the historical reconstructed frame into a time context and a current frame latent variable; S2: using a context allocation device to obtain a reconstructed video frame based on the time context and the current frame latent variable.

[0074] The following is a detailed description of the technical effects of the context-guided transformer entropy coding video compression device in this embodiment. This embodiment conducts a performance comparison experiment on the video compression task of the MCL-JCV dataset, UVG ​​dataset, and HEVC-B dataset using the video compression device (abbreviated as Ours in the figure), the DMC method, the VTM method, and the HM method. In the experiment, DMC is used as the frame encoder, and the peak signal-to-noise ratio is finally calculated. The comparative experimental results are shown in Figure 2. Figure 4 、 Figure 5 and Figure 6 As shown. Among them, Figure 4The test curve results of the MCL-JCV data set are shown in Figure 2. Figure 5 The test curve results of the UVG data set are shown in Figure 2. Figure 6 These are the test curve results for the HEVC-B dataset. From these figures, it can be seen that the video compression device has a good peak signal-to-noise ratio, and thus its performance is relatively excellent.

[0075] The context-guided transformer entropy coding video compression device disclosed in this embodiment utilizes a video compression device and a motion network and context network model algorithm to map the input current video frame and historical reconstructed frames into temporal context and current frame latent variables. The resulting result has multi-scale and multi-type features. By providing a context allocation device, the context-guided Transformer conditional entropy coding model contained therein can be called to obtain a reconstructed video frame based on the temporal context and current frame latent variables. The combination of the video compression device and the context allocation device can effectively compress the temporal context, significantly reducing computational cost while retaining sufficient temporal dependency information. Furthermore, the context-guided Transformer conditional entropy coding model can model the importance and determinism of spatial context, optimize the utilization efficiency of spatial context, and effectively improve the decoding quality of video frames. This overcomes the technical drawback of existing conditional entropy coding methods, which struggle to achieve a good balance between computational efficiency and compression rate.

[0076] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present application.

[0077] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are mutually inconsistent.

[0078] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A context-guided transformer entropy coding video compression device, characterized in that include: a memory configured to store historical reconstructed frames and a temporal context; a video compression device, electrically connected to the memory, and configured to map a current video frame input to the video compression device and the historical reconstructed frame into the temporal context and current frame latent variables through a motion network and a context network model algorithm; A context allocation device, electrically connected to the video compression device, configured to obtain a reconstructed video frame according to the temporal context and the current frame latent variable by calling a context-guided Transformer conditional entropy coding model contained therein; Wherein, the context-guided Transformer conditional entropy coding model includes a time context resampler, a teacher network and a student network, the input end of the time context resampler is electrically connected to the output pin of the latent variable buffer and the output end of the frame encoder at the same time, the input end of the teacher network is electrically connected to the output end of the time context resampler, the output end of the entropy encoder is electrically connected to the output end of the teacher network, the input end of the student network is electrically connected to the output end of the time context resampler, and the output end of the entropy encoder is electrically connected to the output end of the student network; The context-guided Transformer conditional entropy coding model is configured to perform the following steps: B1: calling the temporal context resampler to execute a window cross attention mechanism to extract and fuse features of the temporal context and the current frame latent variable to obtain a temporal context representation; B2: Determine whether the current training phase is in the context-guided Transformer conditional entropy coding model. If yes, proceed to the next step; If not, proceed to step B4; B3: calling the teacher network to obtain and output the reconstructed video frame using the temporal context representation, and obtaining estimated parameters of a probability mass function using the reconstructed video frame through a parameter mapping layer algorithm; B4: calling the student network to obtain the reconstructed video frame and the estimated parameters of the probability mass function by simulating the operation mode of the teacher network; The step B3 includes the following steps: B31: calling the teacher network to execute the window attention mechanism of Swin Transformer to generate an attention map representing the importance of spatial position and an entropy map representing the certainty of spatial position using the temporal context representation; B32: calling the teacher network to normalize the attention map and the entropy map, and generating a spatial dependency score matrix through a weighted summation calculation formula; B33: calling the teacher network to use the top-k algorithm to sort the positions of the elements in the spatial dependency score matrix and determine the k positions with the most context value; B34: calling the teacher network to decode the k positions with the most context value to determine the spatial context, obtain and output the reconstructed video frame; B35: Calling the teacher network to execute a parameter mapping layer algorithm to obtain estimated parameters of the probability mass function using the reconstructed video frame.

2. The context-guided transformer entropy coding video compression apparatus according to claim 1, wherein The memory is divided into a frame buffer area and a latent variable buffer area. Both the frame buffer area and the latent variable buffer area are provided with input pins and output pins. The frame buffer area is used to store the historical reconstructed frame, and the latent variable buffer area is used to store the time context.

3. The context-guided transformer entropy coding video compression apparatus according to claim 2, wherein The video compression device includes a motion network, a context network, a frame encoder, an entropy encoder and a frame decoder. The input end of the motion network is electrically connected to the output pin of the frame buffer area, the input end of the context network is electrically connected to the output end of the motion network and the output end of the frame decoder at the same time, the input end of the frame encoder is electrically connected to the output end of the context network, the input end of the entropy encoder is electrically connected to the output end of the frame encoder, the input end of the frame decoder is electrically connected to the output end of the context network and the output end of the entropy encoder at the same time, the input pin of the frame buffer area is electrically connected to the output end of the frame decoder, the input pin of the latent variable buffer area is electrically connected to the output end of the context network, and the context allocation device is electrically connected to the output pin of the latent variable buffer area and the output end of the frame encoder at the same time.

4. The context-guided transformer entropy coding video compression apparatus according to claim 3, wherein The video compression device is configured to perform the following steps: A1: calling the motion network to receive the input current video frame, and mapping the input current video frame and the historical reconstructed frame into a motion vector through the motion network model algorithm; A2: calling the context network to map the motion vector and the latent variable output by the frame decoder into the temporal context through a context network model algorithm; A3: calling the frame encoder to execute a frame encoding algorithm to encode the input current video frame and the temporal context into a current frame latent variable and a reconstructed frame latent variable; A4: calling the entropy encoder to distribute and output the reconstructed frame latent variables according to a probability mass function; A5: Call the frame decoder to decode the reconstructed frame latent variable assigned and output by the entropy encoder in a frame decoding manner according to the time context to obtain a reconstructed frame, and transfer the reconstructed frame as the historical reconstructed frame to the frame buffer.

5. The context-guided transformer entropy coding video compression apparatus according to claim 4, wherein The teacher network and the student network both include a Swin Transformer encoder, a Swin Transformer decoder and a parameter mapping layer model, the input end of the Swin Transformer encoder is electrically connected to the output end of the temporal context resampler, the input end of the Swin Transformer decoder is electrically connected to the output end of the Swin Transformer encoder, the input end of the parameter mapping layer model is electrically connected to the output end of the Swin Transformer decoder, and the input end of the entropy encoder is electrically connected to the output end of the parameter mapping layer model; The Swin Transformer encoder is configured to utilize the temporal context representation to generate an attention map representing the importance of spatial positions and an entropy map representing the certainty of spatial positions through the windowed attention mechanism of the Swin Transformer; the attention map and the entropy map are then normalized, and a spatial dependency score matrix is ​​generated through a weighted sum calculation formula; The Swin Transformer decoder is configured to use a top-k algorithm to sort the positions of elements in the spatial dependency score matrix to determine the k positions with the most context value; then decode the k positions with the most context value to determine the spatial context to obtain the reconstructed video frame; The parameter mapping layer model is configured to obtain estimated parameters of the probability mass function using the reconstructed video frame through a parameter mapping layer algorithm.

6. A context-guided transformer entropy coding video compression method, characterized in that A context-guided transformer entropy coding video compression device according to any one of claims 1 to 5, comprising the following steps: S1: Mapping a current video frame input to the video compression device and a historical reconstructed frame into a temporal context and a current frame latent variable through a video compression device; S2: Obtain a reconstructed video frame according to the temporal context and the current frame latent variable through a context allocating device.

Citation Information

Patent Citations

  • Video compression method and device based on conditional scale spatial stream

    CN116347081A