Video compression and high-fidelity reconstruction system driven by deep learning
Through the video compression method of large language model and cross-modal attention mechanism, the problems of semantic information missing and poor time continuity in the prior art are solved, and efficient video compression and high-fidelity reconstruction are achieved.
Patent Information
- Application Number
- CN202510403633.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-09-05
AI Technical Summary
The existing video compression methods do not fully explore semantic content, lack of cross-modal fusion, limited dynamic optimization, and poor time continuity, resulting in compression redundancy, loss of reconstruction details and motion jitter.
A large language model is used to analyze text descriptions, combine cross-modal attention mechanisms and symmetric hierarchical coding networks, and generate high-fidelity reconstruction videos through spatial and temporal encoders respectively, and dynamically jointly trained.
Maintain high fidelity and time fluency at low bit rates, significantly improve compression efficiency and reconstruct semantic consistency, and reduce information loss.
Smart Images

Figure CN120602680A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video compression and reconstruction, and specifically relates to a deep learning-driven video compression and high-fidelity reconstruction system. Background Art
[0002] With the explosive growth of video data, efficient compression and high-quality reconstruction technologies have become key to addressing storage and transmission bottlenecks. Traditional video compression methods rely on inter-frame prediction and transform coding, but suffer from the following problems: Lack of semantic information: Existing methods do not fully exploit the semantic content of the video, resulting in incomplete redundancy elimination during compression and severe loss of detail during reconstruction; Insufficient cross-modal fusion: The correlation between text descriptions and visual features is not effectively utilized, making it difficult to achieve semantically driven compression and reconstruction; Dynamic optimization limitations: Static loss weight distribution during training cannot balance the latent space distribution and pixel-level reconstruction requirements, affecting model convergence and performance; Poor temporal continuity: Temporal compression easily destroys inter-frame motion coherence, resulting in jitter and blur in the reconstructed video.
[0003] This paper proposes a video compression and reconstruction method based on semantic enhancement by combining the semantic parsing capability of the large language model (LLM) with the cross-modal attention mechanism, which solves the above problems. It can still maintain high fidelity and temporal fluency at low bit rates and is suitable for ultra-high-definition video and AR / VR scenarios. Summary of the Invention
[0004] The present invention discloses a high-quality video compression and reconstruction method based on semantic decomposition of a large language model and a symmetric layered coding network. Through cross-modal attention injection, dynamic joint training and dimensional decoupling optimization, the compression efficiency and reconstruction fidelity are significantly improved.
[0005] Technical solution: A deep learning-driven video compression and high-fidelity reconstruction system, characterized by comprising the following steps:
[0006] S1. Construct a dual-path architecture consisting of a spatial encoder, a temporal encoder, and a spatiotemporal fusion module: the spatial encoder uses a three-level residual convolution block to perform one-eighth spatial compression in the height and width dimensions. Each level contains a three-dimensional convolution kernel, a batch normalization layer, a ReLU activation function, and a jump connection, and outputs spatial latent features with 256 channels. The temporal encoder uses a three-level temporal convolution block to perform one-eighth compression in the temporal dimension. Each level contains a one-dimensional temporal convolution kernel with a group normalization layer and a GELU activation function, and outputs temporal latent features with 128 channels. The spatial latent features and temporal latent features are input into a four-layer Transformer module equipped with eight parallel attention heads, and global dependency modeling is performed by injecting position encoding of spatiotemporal coordinate information to generate fused latent features. In the cross-attention module of the spatial encoder, the pre-trained CLIP model is used to extract visual semantic features. In the cross-attention module of the temporal encoder, the pre-trained T5 model is used to extract text semantic features. The input text description is separated by either GPT-4 or Deepseek3 models, and the parameters of the CLIP and T5 models are fixed and do not participate in training.
[0007] S2. Based on the fused latent features generated in S1, the input text is parsed using a structured hint template containing an object noun phrase defining the main part and a directional verb combined with an adverb phrase defining the motion part. The original text description is separated using either the GPT-4 or Deepseek3 model to generate a spatial entity embedding vector and a motion description embedding vector, respectively. In the spatial encoding path, the spatial entity embedding vector and the visual features are subjected to a multi-head cross-attention calculation using the visual features as the query vector and the text vector as the key vector and value vector. In the temporal encoding path, a gating coefficient is generated by a one-dimensional convolutional layer activated by a sigmoid function, and the motion description embedding vector and the visual features are weightedly fused according to the gating coefficient, where the text vector weight is the gating coefficient value and the visual feature weight is 1 minus the gating coefficient value. The one-dimensional convolutional layer covers contextual information of 5 time steps, and the generated gating coefficient is constrained to an interval using a sigmoid function.
[0008] S3. Based on the features after cross-modal fusion in S2, the KL divergence loss with a weight coefficient of 0.7 is used to optimize the potential space distribution in the first twenty cycles of training, and the initial learning rate is set to 3×10 -4 The KL divergence loss weight decreases linearly from the initial 0.7 to 0.3, and the reconstruction loss weight increases from 0.3 to 0.7 according to the cosine function. The weight conversion is completed within 20 training cycles. The initial learning rate decays by 10% every 5 training cycles, with a minimum limit of 1×10 -5When the training process advances to the 20th epoch, multi-level supervision is activated: a peak signal-to-noise ratio loss with a weight of 0.4 is applied to the output of the spatial decoder at the quarter-temporal resolution stage; the optical flow consistency loss is calculated based on the RAFT optical flow field for the output of the temporal decoder at a quarter-temporal resolution; at the final output stage, a composite loss function is constructed by combining the structural similarity index, the perceptual loss based on the third layer features of the third convolutional block of the VGG-16 network, and the adversarial loss of the PatchGAN discriminator, with the weights of the three losses being 0.5, 0.3, and 0.2, respectively; distributed training is performed using 8 NVIDIA A800 graphics cards, and training is terminated early if the PSNR and SSIM indicators of the validation set do not improve for three consecutive epochs;
[0009] S4. Using the model parameters optimized after training in S3, the feature tensors of the temporal decoder with a quarter of the temporal resolution are dimensionally expanded and spliced during the quarter-downsampling stage of the spatial decoder, and cross-modal fusion is achieved through 3D convolution kernels. A deformable convolutional network with three levels of 3D convolutional layers is used to predict pixel offsets based on the motion features output by the temporal decoder. Eightfold upsampling is performed through sub-pixel convolutional layers to generate an output consistent with the original video resolution. When the input video resolution is 312×600, the frame rate is 25fps, and the duration is 9 seconds, the peak signal-to-noise ratio (PSNR) of the reconstructed video is ≥32dB, the structural similarity (SSIM) is ≥0.95, and the frame-by-frame texture detail error rate is ≤3%.
[0010] Preferably, the S1 specifically includes the following steps:
[0011] S1-1. Residual block structure optimization: In the residual convolution block of the spatial encoder, the skip connection achieves channel matching through a 1×1 convolution layer. Each residual block performs batch normalization, ReLU activation function operation, and 3×3×3 convolution operation in sequence.
[0012] S1-2. Enhanced temporal feature extraction: In the temporal convolution block of the temporal encoder, the 1×3×3 convolutional layer preserves local motion features between consecutive frames in the temporal dimension, and the group normalization layer divides the feature channels into 8 groups for independent normalization.
[0013] S1-3. Spatiotemporal Position Code Generation: The final spatiotemporal position code is used as learnable spatiotemporal prior information and is input into the Transformer global dependency modeling module described in claim S2. Through coordinate-aware attention calculation of key-value pairs, cross-spatiotemporal semantic association of video sequences is achieved, completing the technical chain from local feature encoding to global relationship reasoning.
[0014] Preferably, the step S2 specifically includes the following steps:
[0015] S2-1. Structured template design: The structured prompt template is limited to the main part as an object noun phrase containing color and shape attributes, the motion part is an action description combined with direction and speed adverbs, and the object name is selected from the 80 categories of entities predefined in the COCO dataset;
[0016] S2-2. Cross-modal attention calculation: In the multi-head cross-attention mechanism, eight parallel computing heads each process 32-dimensional features. The attention weights are calculated by dividing the dot product of the query vector and the key vector by the square root of the key vector dimension, and then applying the softmax function to calculate the final weighted aggregate value vector.
[0017] S2-3. Dynamic Gating Fusion Mechanism: The gating coefficients are generated using a one-dimensional convolutional layer with 128 convolution kernels. The convolution kernels cover contextual information for five time steps. The output is constrained to the range of 0 to 1 using a sigmoid function, achieving soft fusion of text and visual features.
[0018] S2-4. Text separation constraint: Use either GPT-4 or Deepseek3 to ensure that the semantics of the subject information and motion information do not overlap;
[0019] Preferably, the step S3 specifically includes the following steps:
[0020] S3-1. Dynamic weight scheduling strategy: The KL divergence loss weight is linearly decreased from the initial 0.7 to 0.3, and the reconstruction loss weight is increased from 0.3 to 0.7 according to the cosine function. Both weight conversions are completed within 20 training cycles.
[0021] S3-2. Learning rate step decay: initial learning rate 3×10 -4 The decay rate is 10% every 5 training cycles, with a minimum limit of 1×10 -5 , ensuring the stability of parameter fine-tuning in the later stages of training;
[0022] S3-3. Multi-scale supervision mechanism: At the quarter-resolution stage, the spatial decoder outputs the peak signal-to-noise ratio loss, and the temporal decoder outputs the optical flow field extracted by the L1 norm matching RAFT algorithm;
[0023] Preferably, the S4 specifically includes the following steps:
[0024] S4-1. Cross-path feature fusion: The quarter-resolution features (256 channels) of the spatial decoder and the features (128 channels) of the temporal decoder are concatenated into 384 channels, compressed to 256 channels via a 3×3×3 convolution, and then processed using the LeakyReLU activation function.
[0025] S4-2. Deformable Convolution Configuration: The offset prediction network uses three levels of 3D convolutional layers to gradually reduce the dimensionality, ultimately outputting a 3×3×2 offset field to guide the dynamic sampling of the convolution kernel in the spatial and temporal dimensions.
[0026] S4-3. Sub-pixel reconstruction optimization: Sub-pixel upsampling is performed using a 4×4 convolution kernel. The convolution parameters are initialized according to the He normal distribution, and the output is normalized to the range [-1, 1] using a tanh function to match the video pixel standard.
[0027] S4-4. When the input video description is "Woman using smartphone in cafe", the reconstructed video must maintain the original resolution, frame rate, and duration, with a texture detail error rate ≤ 3% and a motion coherence error rate ≤ 2%.
[0028] Compared with the prior art, the advantages of the present invention are:
[0029] (1) Semantic-driven multimodal compression: parse text descriptions through a large language model, separate spatial entities and temporal motion information, combine with pre-trained models to generate semantic vectors, and achieve cross-modal attention injection, significantly improving compression efficiency and reconstructing semantic consistency;
[0030] (2) Symmetric hierarchical coding network: The spatial encoder and temporal encoder compress the spatial and temporal dimensions respectively, dynamically align semantic and visual features through the cross-attention module, retain key edge and motion information, and reduce information loss;
[0031] (3) Dynamic joint training strategy: In the early stage, the KL divergence loss is used to optimize the latent space distribution, and the pixel-level reconstruction loss is gradually increased in the middle and late stages, with the final weight reaching 60%, balancing the compression representation and reconstruction quality. The model converges faster and is more robust. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is the overall flow chart of the deep learning-driven video compression and high-fidelity reconstruction system proposed in the present invention;
[0033] Figure 2 Schematic block diagram of the spatial encoder and spatial decoder of the deep learning-driven video compression and high-fidelity reconstruction system proposed in the present invention;
[0034] Figure 3 Schematic block diagram of the temporal encoder and temporal decoder of the deep learning-driven video compression and high-fidelity reconstruction system proposed in this invention. DETAILED DESCRIPTION
[0035] Example
[0036] A deep learning-driven video compression and high-fidelity reconstruction system, characterized by comprising the following steps:
[0037] S1. Construct a dual-path architecture consisting of a spatial encoder, a temporal encoder, and a spatiotemporal fusion module: the spatial encoder uses a three-level residual convolution block to perform one-eighth spatial compression in the height and width dimensions. Each level contains a three-dimensional convolution kernel, a batch normalization layer, a ReLU activation function, and a jump connection, and outputs spatial latent features with 256 channels. The temporal encoder uses a three-level temporal convolution block to perform one-eighth compression in the temporal dimension. Each level contains a one-dimensional temporal convolution kernel with a group normalization layer and a GELU activation function, and outputs temporal latent features with 128 channels. The spatial latent features and temporal latent features are input into a four-layer Transformer module equipped with eight parallel attention heads, and global dependency modeling is performed by injecting position encoding of spatiotemporal coordinate information to generate fused latent features. In the cross-attention module of the spatial encoder, the pre-trained CLIP model is used to extract visual semantic features. In the cross-attention module of the temporal encoder, the pre-trained T5 model is used to extract text semantic features. The input text description is separated by either GPT-4 or Deepseek3 models, and the parameters of the CLIP and T5 models are fixed and do not participate in training.
[0038] The S1 specifically includes the following steps:
[0039] S1-1. Residual block structure optimization: In the residual convolution block of the spatial encoder, the skip connection achieves channel matching through a 1×1 convolution layer. Each residual block performs batch normalization, ReLU activation function operation, and 3×3×3 convolution operation in sequence.
[0040] S1-2. Enhanced temporal feature extraction: In the temporal convolution block of the temporal encoder, the 1×3×3 convolutional layer preserves local motion features between consecutive frames in the temporal dimension, and the group normalization layer divides the feature channels into 8 groups for independent normalization.
[0041] S1-3. Spatiotemporal Position Code Generation: The final spatiotemporal position code is used as learnable spatiotemporal prior information and is input into the Transformer global dependency modeling module described in claim S2. Through coordinate-aware attention calculation of key-value pairs, cross-spatiotemporal semantic association of video sequences is achieved, completing the technical chain loop from local feature encoding to global relationship reasoning:
[0042] S1-3-1. Generate a learnable embedding vector for each spatiotemporal position in the video, where is the time dimension index and is the spatial height and width index;
[0043] S1-3-2. The spatiotemporal position encoding is generated by the following formula:
[0044] Among them,,, are the sinusoidal position encoding functions of time, height and width dimensions respectively;
[0045] S1-3-3. After splicing the spatiotemporal position code with the spatial latent features and temporal latent features described in claim S1, the code is input into the Transformer module described in claim S2;
[0046] S1-3-4. In the coordinate-aware attention computation of key-value pairs, the key vector is generated by linearly mapping text features and spatiotemporal position encodings, and the value vector is generated by linearly mapping visual features and spatiotemporal position encodings. A multi-head attention mechanism is used to achieve cross-spatial semantic association.
[0047] S2. Based on the fused latent features generated in S1, the input text is parsed using a structured hint template containing an object noun phrase defining the main part and a directional verb combined with an adverb phrase defining the motion part. The original text description is separated using either the GPT-4 or Deepseek3 model to generate a spatial entity embedding vector and a motion description embedding vector, respectively. In the spatial encoding path, the spatial entity embedding vector and the visual features are subjected to a multi-head cross-attention calculation using the visual features as the query vector and the text vector as the key vector and value vector. In the temporal encoding path, a gating coefficient is generated by a one-dimensional convolutional layer activated by a sigmoid function, and the motion description embedding vector and the visual features are weightedly fused according to the gating coefficient, where the text vector weight is the gating coefficient value and the visual feature weight is 1 minus the gating coefficient value. The one-dimensional convolutional layer covers contextual information of 5 time steps, and the generated gating coefficient is constrained to an interval using a sigmoid function.
[0048] The S2 specifically includes the following steps:
[0049] S2-1. Structured template design: The structured prompt template is limited to the main part as an object noun phrase containing color and shape attributes, the motion part is an action description combined with direction and speed adverbs, and the object name is selected from the 80 categories of entities predefined in the COCO dataset;
[0050] S2-2. Cross-modal attention calculation: In the multi-head cross-attention mechanism, each of the eight parallel computing heads processes 32-dimensional features. The attention weights are calculated by dividing the dot product of the query vector and the key vector by the square root of the key vector dimension, and then applying the softmax function to calculate the final weighted aggregate value vector. Specifically, it includes the following:
[0051] In the multi-head cross-attention mechanism, eight parallel computing heads each process 32-dimensional features.
[0052] S2-2-2. The query vector is generated by linear mapping the visual features, the key vector is generated by linear mapping the text features (spatial entity embedding vector, motion description embedding vector), and the value vector is generated by linear mapping the visual features.
[0053] S2-2-3. The attention weight is calculated as follows:
[0054] Among them, is the dimension of the key vector; finally, the cross-modal fusion feature is generated by weighted aggregation value vector
[0055] S2-3. Dynamic Gating Fusion Mechanism: The gating coefficients are generated using a one-dimensional convolutional layer with 128 convolution kernels. The convolution kernels cover contextual information for five time steps. The output is constrained to the range of 0 to 1 using a sigmoid function, achieving soft fusion of text and visual features.
[0056] S2-4. Text separation constraint: Use either GPT-4 or Deepseek3 to ensure that the semantics of the subject information and motion information do not overlap;
[0057] S3. Based on the features after cross-modal fusion in S2, the KL divergence loss with a weight coefficient of 0.7 is used to optimize the potential space distribution in the first twenty cycles of training, and the initial learning rate is set to 3×10 -4 The KL divergence loss weight decreases linearly from the initial 0.7 to 0.3, and the reconstruction loss weight increases from 0.3 to 0.7 according to the cosine function. The weight conversion is completed within 20 training cycles. The initial learning rate decays by 10% every 5 training cycles, with a minimum limit of 1×10 -5 When the training process advances to the 20th epoch, multi-level supervision is activated: a peak signal-to-noise ratio loss with a weight of 0.4 is applied to the output of the spatial decoder at the quarter-temporal resolution stage; the optical flow consistency loss is calculated based on the RAFT optical flow field for the output of the temporal decoder at a quarter-temporal resolution; at the final output stage, a composite loss function is constructed by combining the structural similarity index, the perceptual loss based on the third layer features of the third convolutional block of the VGG-16 network, and the adversarial loss of the PatchGAN discriminator, with the weights of the three losses being 0.5, 0.3, and 0.2, respectively; distributed training is performed using 8 NVIDIA A800 graphics cards, and training is terminated early if the PSNR and SSIM indicators of the validation set do not improve for three consecutive epochs;
[0058] The S3 specifically includes the following steps:
[0059] S3-1. Dynamic weight scheduling strategy: The KL divergence loss weight is linearly decreased from the initial 0.7 to 0.3, and the reconstruction loss weight is increased from 0.3 to 0.7 according to the cosine function. Both weight conversions are completed within 20 training cycles.
[0060] S3-2. Learning rate step decay: initial learning rate 3×10 -4 The decay rate is 10% every 5 training cycles, with a minimum limit of 1×10 -5 , ensuring the stability of parameter fine-tuning in the later stages of training;
[0061] S3-3. Multi-scale supervision mechanism: At the quarter-resolution stage, the spatial decoder outputs the calculated peak signal-to-noise ratio loss, and the temporal decoder outputs the optical flow field extracted by the L1 norm matching RAFT algorithm. Specifically, it includes the following:
[0062] S3-3-1. Calculate the peak signal-to-noise ratio loss PSNRLoss for the output of the spatial decoder quarter downsampling stage:
[0063] Among them, is the maximum value of the pixel (usually 255), is the reconstructed frame, and is the real frame;
[0064] S3-3-2. For the output of the temporal decoder with a quarter temporal resolution, the optical flow field of adjacent frames is extracted using the RAFT algorithm, and the optical flow consistency loss is calculated:
[0065] Among them, is the optical flow field predicted by the model;
[0066] S4. Using the model parameters optimized after training in S3, the feature tensors of the temporal decoder with a quarter of the temporal resolution are dimensionally expanded and spliced during the quarter-downsampling stage of the spatial decoder, and cross-modal fusion is achieved through 3D convolution kernels. A deformable convolutional network with three levels of 3D convolutional layers is used to predict pixel offsets based on the motion features output by the temporal decoder. Eightfold upsampling is performed through sub-pixel convolutional layers to generate an output consistent with the original video resolution. When the input video resolution is 312×600, the frame rate is 25fps, and the duration is 9 seconds, the peak signal-to-noise ratio (PSNR) of the reconstructed video is ≥32dB, the structural similarity (SSIM) is ≥0.95, and the frame-by-frame texture detail error rate is ≤3%.
[0067] The S4 specifically includes the following steps:
[0068] S4-1. Cross-path feature fusion: The quarter-resolution features (256 channels) of the spatial decoder and the features (128 channels) of the temporal decoder are concatenated into 384 channels, compressed to 256 channels via a 3×3×3 convolution, and then processed using the LeakyReLU activation function.
[0069] S4-2. Deformable Convolution Configuration: The offset prediction network uses three levels of 3D convolutional layers to gradually reduce the dimensionality, ultimately outputting a 3×3×2 offset field to guide the dynamic sampling of the convolution kernel in the spatial and temporal dimensions. Specifically, it includes the following:
[0070] Supplementary deformable convolution implementation details:
[0071] S4-2-1. Offset prediction network structure:
[0072] First layer: 3×3×3 convolution, output 128 channels;
[0073] Second layer: dilated convolution (dilation=2), output 64 channels;
[0074] The third layer: separable convolution, output 3×3×2 offset field;
[0075] S4-2-2. Dynamic sampling uses a bilinear interpolation kernel:
[0076] Offset field regularization constraint: L2 norm ≤ 1.5;
[0077] S4-3. Sub-pixel reconstruction optimization: Sub-pixel upsampling is performed using a 4×4 convolution kernel. The convolution parameters are initialized according to the He normal distribution, and the output is normalized to the range [-1, 1] using the tanh function to match the video pixel standard. Specifically, it includes the following:
[0078] S4-3-1. Sub-pixel upsampling is achieved through a phase shift operation. The specific process is as follows:
[0079] Reorganize the channel dimension of the input feature map and expand the number of channels from (which is the spatial upsampling multiple);
[0080] Rearrange the reorganized feature map according to the spatial block size to generate a high-resolution feature map;
[0081] The final refinement is performed by a 4×4 convolution kernel, the convolution parameters are initialized according to the He normal distribution, and the output is normalized to the range by the tanh function;
[0082] S4-3-2. Temporal upsampling is achieved through cubic linear interpolation to ensure that the reconstructed video frame rate is consistent with the original input;
[0083] S4-4. When the input video description is "Woman using smartphone in cafe", the reconstructed video must maintain the original resolution, frame rate, and duration, with a texture detail error rate ≤ 3% and a motion coherence error rate ≤ 2%.
[0084] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. Deep learning-driven video compression and high-fidelity reconstruction system, characterized by: The following steps are involved: S1. Construct a dual-path architecture consisting of a spatial encoder, a temporal encoder, and a spatiotemporal fusion module. The spatial encoder uses a three-level residual convolution block to perform one-eighth spatial compression in the height and width dimensions. Each level contains a 3D convolution kernel, a batch normalization layer, a ReLU activation function, and a skip connection, outputting a spatial latent feature with 256 channels. The temporal encoder uses a three-level temporal convolution block to perform one-eighth compression of the temporal dimension. Each level contains a one-dimensional temporal convolution kernel with a group normalization layer and a GELU activation function, and outputs a temporal latent feature with 128 channels. The spatial latent features and temporal latent features are input into a four-layer Transformer module equipped with eight parallel attention heads. Global dependency modeling is performed by injecting position encoding of spatiotemporal coordinate information to generate fused latent features. In the cross-attention module of the spatial encoder, the pre-trained CLIP model is used to extract visual semantic features. In the cross-attention module of the temporal encoder, the pre-trained T5 model is used to extract text semantic features. The input text description is separated by using either the GPT-4 or Deepseek3 model. The parameters of the CLIP and T5 models are fixed and do not participate in training. S2. Based on the fused latent features generated in S1, the input text is parsed using a structured hint template that includes an object noun phrase to define the main part and a directional verb combined with an adverb phrase to define the motion part. The original text description is separated using either GPT-4 or Deepseek3 models to generate spatial entity embedding vectors and motion description embedding vectors, respectively. In the spatial encoding path, the spatial entity embedding vector and the visual features are subjected to a multi-head cross-attention calculation, with the visual features as the query vector and the text vector as the key vector and value vector. In the temporal encoding path, a one-dimensional convolutional layer activated by a sigmoid function generates a gating coefficient, which is used to perform a weighted fusion of the motion description embedding vector and the visual features, where the text vector weight is the gating coefficient value and the visual feature weight is 1 minus the gating coefficient value. The one-dimensional convolutional layer covers context information of 5 time steps, and the generated gating coefficient is constrained to an interval using a sigmoid function. S3. Based on the features after cross-modal fusion in S2, the KL divergence loss with a weight coefficient of 0.7 is used to optimize the potential space distribution in the first twenty cycles of training. The initial learning rate is set to: 3×10 -4 The KL divergence loss weight decreases linearly from the initial 0.7 to 0.3, and the reconstruction loss weight increases from 0.3 to 0.7 according to the cosine function. The weight conversion is completed within 20 training cycles. The initial learning rate decays by 10% every 5 training cycles, with a minimum limit of 1×10 -5 ; When the training process advances to the twentieth cycle, multi-level supervision is started: a peak signal-to-noise ratio loss with a weight of 0.4 is applied to the output of the spatial decoder at the quarter downsampling stage; the optical flow consistency loss is calculated based on the RAFT optical flow field for the output of the temporal decoder at a quarter time resolution; At the final output stage, a composite loss function is constructed by combining the structural similarity index, the perceptual loss based on the third layer features of the third convolutional block of the VGG-16 network, and the adversarial loss of the PatchGAN discriminator. The weights of these three losses are 0.5, 0.3, and 0.2, respectively. Distributed training is performed using 8 NVIDIA A800 graphics cards. If the PSNR and SSIM indicators of the validation set do not improve for three consecutive cycles, training is terminated early. S4. Using the model parameters optimized after training in S3, the feature tensors of the temporal decoder with a quarter temporal resolution are dimensionally expanded and concatenated during the quarter-downsampling stage of the spatial decoder, achieving cross-modal fusion using a 3D convolution kernel. A deformable convolutional network with three levels of 3D convolutional layers is used to predict pixel offsets based on the motion features output by the temporal decoder. Eightfold upsampling is performed through the sub-pixel convolution layer to generate an output consistent with the original video resolution; when the input video resolution is 312×600, the frame rate is 25fps, and the length is 9 seconds, the peak signal-to-noise ratio (PSNR) of the reconstructed video is ≥32dB, the structural similarity (SSIM) is ≥0.95, and the frame-by-frame texture detail error rate is ≤3%.
2. The deep learning-driven video compression and high-fidelity reconstruction system according to claim 1, characterized in that The S1 specifically includes the following steps: S1-1. Residual block structure optimization: In the residual convolution block of the spatial encoder, the skip connection achieves channel matching through a 1×1 convolution layer. Each residual block performs batch normalization, ReLU activation function operation, and 3×3×3 convolution operation in sequence. S1-2. Enhanced temporal feature extraction: In the temporal convolution block of the temporal encoder, the 1×3×3 convolutional layer preserves local motion features between consecutive frames in the temporal dimension, and the group normalization layer divides the feature channels into 8 groups for independent normalization. S1-3. Generation of spatiotemporal position coding: The final spatiotemporal position coding is input into the Transformer global dependency modeling module described in claim S2 as learnable spatiotemporal prior information. Through the coordinate-aware attention calculation of key-value pairs, the cross-spatiotemporal semantic association of the video sequence is realized, completing the technical chain loop from local feature coding to global relationship reasoning.
3. The deep learning driven video compression and high-fidelity reconstruction system according to claim 1, characterized in that The S2 specifically includes the following steps: S2-1. Structured template design: The structured prompt template is limited to the main part as an object noun phrase containing color and shape attributes, the motion part is an action description combined with direction and speed adverbs, and the object name is selected from the 80 categories of entities predefined in the COCO dataset; S2-2. Cross-modal attention calculation: In the multi-head cross-attention mechanism, eight parallel computing heads each process 32-dimensional features. The attention weights are calculated by dividing the dot product of the query vector and the key vector by the square root of the key vector dimension, and then applying the softmax function to calculate the final weighted aggregate value vector. S2-3. Dynamic Gating Fusion Mechanism: The gating coefficients are generated using a one-dimensional convolutional layer with 128 convolution kernels. The convolution kernels cover contextual information for five time steps. The output is constrained to the range of 0 to 1 using a sigmoid function, achieving soft fusion of text and visual features. S2-4. Text separation constraint: Use either GPT-4 or Deepseek3 to ensure that the semantics of the subject information and motion information do not overlap.
4. The deep learning driven video compression and high-fidelity reconstruction system according to claim 1, characterized in that The S3 specifically includes the following steps: S3-1. Dynamic weight scheduling strategy: The KL divergence loss weight is linearly decreased from the initial 0.7 to 0.3, and the reconstruction loss weight is increased from 0.3 to 0.7 according to the cosine function. Both weight conversions are completed within 20 training cycles. S3-2. Learning rate step decay: initial learning rate 3×10 -4 The decay rate is 10% every 5 training cycles, with a minimum limit of 1×10 -5 , ensuring the stability of parameter fine-tuning in the later stages of training; S3-3. Multi-scale supervision mechanism: At the quarter-resolution stage, the spatial decoder output calculates the peak signal-to-noise ratio loss, and the temporal decoder outputs the optical flow field extracted by the L1 norm matching RAFT algorithm.
5. The deep learning driven video compression and high-fidelity reconstruction system according to claim 1, characterized in that The S4 specifically includes the following steps: S4-1. Cross-path feature fusion: The quarter-resolution features (256 channels) of the spatial decoder and the features (128 channels) of the temporal decoder are concatenated into 384 channels, compressed to 256 channels via a 3×3×3 convolution, and then processed using the LeakyReLU activation function. S4-2. Deformable Convolution Configuration: The offset prediction network uses three levels of 3D convolutional layers to gradually reduce the dimensionality, ultimately outputting a 3×3×2 offset field to guide the dynamic sampling of the convolution kernel in the spatial and temporal dimensions. S4-3. Sub-pixel reconstruction optimization: Sub-pixel upsampling is performed using a 4×4 convolution kernel. The convolution parameters are initialized according to the He normal distribution, and the output is normalized to the range [-1, 1] using a tanh function to match the video pixel standard. S4-4. When the input video description is "Woman using smartphone in cafe", the reconstructed video must maintain the original resolution, frame rate, and duration, with a texture detail error rate ≤ 3% and a motion coherence error rate ≤ 2%.