A multimodal approach to real-time high-fidelity video transmission based on semantic streaming

Through a multimodal method based on semantic streams, the potential video representation is extracted and multimodal fusion is carried out, which solves the problem of insufficient video compression efficiency and semantic utilization in the prior art, and realizes efficient real-time high-fidelity video transmission.

CN119728994BActive Publication Date: 2025-06-06GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510213570.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-06
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing video compression technology has significant limitations in long-term time-dependent modeling, perceptual optimization, generative model stability and semantic utilization, and it is difficult to meet the needs of real-time high-fidelity video transmission.

Method used

Using a multimodal method based on semantic streams, the potential representation of video is extracted through a space-time compressor, the semantic translator transforms visual features into text features, the Transformer fusion model performs multimodal fusion, the codebook model performs quantization processing, and the reconstructed video sequence is generated and reconstructed through the video control network.

Benefits of technology

It significantly improves video compression efficiency, realizes priority transmission of key semantic information, maintains efficient compression and perceived correlation under bandwidth constraints, and ensures semantic consistency and temporal coherence of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119728994B_ABST
    Figure CN119728994B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention relates to the field of video transmission technology, and specifically discloses a multimodal method for real-time high-fidelity video transmission based on semantic streams. The embodiment of the present invention receives a multi-frame video sequence, extracts spatial and temporal correlations through a spatiotemporal compressor, and outputs a potential representation; maps the potential representation to a semantic space through a semantic translator, and gradually transforms visual features and text features; multimodally fuses the potential representation and text features through a preset Transformer fusion model, and outputs a fused representation; quantizes the fused representation into a quantized representation through a preset codebook model; and processes the quantized representation and text features through a video control network to generate a reconstructed video sequence. The compression efficiency can be significantly improved, and priority transmission of key semantic information can be achieved, thereby maintaining efficient compression and perceptual correlation under bandwidth-constrained conditions, and ensuring the semantic consistency and temporal coherence of the video content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video transmission, and in particular relates to a multimodal method for real-time high-fidelity video transmission based on semantic streams. Background Art

[0002] Currently, video compression technology mainly includes traditional methods and deep learning-based methods.

[0003] Traditional methods (such as H.264 / AVC, HEVC, etc.) reduce the amount of data by compressing single frames or adjacent frames. Although they can effectively utilize spatial redundancy and short-term temporal correlation, it is difficult to model long-term temporal dependence, which limits the compression efficiency. In addition, existing methods use pixel fidelity as the optimization goal and do not fully consider the perceptual characteristics of the human visual system (HVS). They perform poorly when bandwidth is limited or network conditions change dynamically.

[0004] Deep learning-based methods use neural networks to extract high-level semantic features, significantly improving compression efficiency, optimizing perceptual indicators through adaptive modeling of spatial and temporal redundancy, and more in line with human visual preferences. However, generative models (such as GAN and diffusion models) have stability issues, which may lead to inconsistent reconstruction, especially in terms of high-frequency details and temporal coherence. In addition, existing methods do not make sufficient use of key semantic features of video content and fail to prioritize important information, limiting compression efficiency and dynamic network adaptability.

[0005] Therefore, although existing technical solutions have made some progress in compression performance, they still have significant limitations in long-term temporal dependency modeling, perceptual optimization, generative model stability, and semantic utilization, and need further improvement to meet the needs of real-time video transmission scenarios. Summary of the invention

[0006] The purpose of the embodiments of the present invention is to provide a multimodal method for real-time high-fidelity video transmission based on semantic streams, aiming to solve the problems raised in the background technology.

[0007] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0008] A multimodal method for real-time high-fidelity video transmission based on semantic stream, the method specifically comprises the following steps:

[0009] Receive a multi-frame video sequence, extract spatial and temporal correlations through a spatiotemporal compressor, and output a latent representation;

[0010] Mapping the latent representation to a semantic space through a semantic translator, gradually transforming visual features and text features;

[0011] The latent representation and the text feature are multimodally fused through a preset Transformer fusion model, and a fused representation is output;

[0012] quantizing the fused representation into a quantized representation through a preset codebook model;

[0013] The quantized representation and the text features are processed through a video control network to generate a reconstructed video sequence.

[0014] As a further limitation of the technical solution of the embodiment of the present invention, the expression of the multi-frame video sequence is:

[0015] ;

[0016] in, is a multi-frame video sequence, Indicates the frame number;

[0017] The expression of the potential representation is:

[0018] ;

[0019] in, For potential representation, are the parameters of the space-time compressor;

[0020] The expression of the visual feature is:

[0021] ;

[0022] in, For visual features, is the first parameter of the semantic translator;

[0023] The expression of the text feature is:

[0024] ;

[0025] in, is the text feature, It is the second parameter of the semantic translator;

[0026] The expression of the fusion expression is:

[0027] ;

[0028] in, For fusion representation, are the parameters of the Transformer fusion model;

[0029] The expression of the quantitative representation is:

[0030] ;

[0031] in, To quantify, is the codebook parameter;

[0032] The expression of the reconstructed video sequence is:

[0033] ;

[0034] in, represents the reconstructed video sequence, Parameters of the video control network.

[0035] As a further limitation of the technical solution of the embodiment of the present invention, the receiving of a multi-frame video sequence, extracting spatial and temporal correlations through a spatiotemporal compressor, and outputting a potential representation specifically include the following steps:

[0036] Receiving a multi-frame video sequence;

[0037] Capturing the spatial structure and motion pattern between adjacent frames of the multi-frame video sequence through a 3D convolutional layer, and outputting joint spatiotemporal features;

[0038] Based on the joint spatiotemporal features, encoding the short-term and long-term dependencies in the multi-frame video sequence through a recurrent layer, and outputting encoded data;

[0039] Through the downsampling layer, cross-row convolution is applied to compress the encoded data into a latent representation.

[0040] As a further limitation of the technical solution of the embodiment of the present invention, the expression of the joint spatiotemporal feature is:

[0041] ;

[0042] in, is the joint spatiotemporal feature, are the parameters of the convolutional layer;

[0043] The expression of the coded data is:

[0044] ;

[0045] in, represents the encoded data, are the parameters of the recurrent network;

[0046] The expression of the potential representation is:

[0047] ;

[0048] in, are the parameters of the downsampling layer.

[0049] As a further limitation of the technical solution of the embodiment of the present invention, mapping the potential representation to the semantic space by the semantic translator and gradually transforming the visual features and the text features specifically includes the following steps:

[0050] extracting high-level semantic features from the latent representation through a visual encoder to generate a visual feature map;

[0051] Dynamically associating the visual feature map with a pre-trained language embedding obtained from a preset large language model through a cross-attention layer, establishing a correspondence between visual elements and their language counterparts, and calculating alignment weights to obtain a highly aligned feature set;

[0052] Text features are generated by using a text decoder, fine-tuned to capture video-specific semantics based on the highly aligned feature set.

[0053] As a further limitation of the technical solution of the embodiment of the present invention, the expression of the visual feature mapping is:

[0054] ;

[0055] in, is the visual feature map, are the parameters of the visual encoder;

[0056] The expression of the highly aligned feature set is:

[0057] ;

[0058] in, represents a highly aligned feature set, for pre-trained language embeddings, is the parameter across the attention layer;

[0059] The expression of the text feature is:

[0060] ;

[0061] in, Parameters for the text decoder.

[0062] As a further limitation of the technical solution of the embodiment of the present invention, the method of performing multimodal fusion of the latent representation and the text feature through a preset Transformer fusion model and outputting the fused representation specifically includes the following steps:

[0063] Embedding the text features into a shared feature space of a preset Transformer fusion model through a text embedding layer to generate a first dense vector;

[0064] Embedding the latent representation into a shared feature space of a preset Transformer fusion model through a visual embedding layer to generate a second dense vector;

[0065] A cross attention layer is used to perform fusion and alignment processing of semantic information and visual information on the first dense vector and the second dense vector, and a fusion representation is output.

[0066] As a further limitation of the technical solution of the embodiment of the present invention, the expression of the first dense vector is:

[0067] ;

[0068] in, represents the first dense vector, Parameters for the text embedding layer;

[0069] The expression of the second dense vector is:

[0070] ;

[0071] in, represents the second dense vector, are the parameters of the visual embedding layer;

[0072] The expression of the fusion expression is:

[0073] ;

[0074] in, are the parameters of the Transformer layer.

[0075] As a further limitation of the technical solution of the embodiment of the present invention, the processing of the quantized representation and the text features through the video control network to generate the reconstructed video sequence specifically includes the following steps:

[0076] De-noising the quantized representation and the text feature by using a preset denoising diffusion probability model to obtain a denoised representation;

[0077] By using a cross-attention layer and a preset semantic alignment control mechanism, the quantitative representation and the text feature are dynamically aligned at each time step to obtain an aligned expression;

[0078] Through the upsampling module, convolution and super-resolution techniques are applied to refine spatial details and produce high-resolution frames to generate reconstructed video sequences.

[0079] As a further limitation of the technical solution of the embodiment of the present invention, the expression of the denoising representation is:

[0080] ;

[0081] in, For denoising, represents the current time step, are the parameters of the denoising diffusion probability model;

[0082] The expression of the alignment expression is:

[0083] ;

[0084] in, To align the expression, is the parameter of the cross-attention layer.

[0085] Compared with the prior art, the present invention has the following beneficial effects:

[0086] The embodiment of the present invention is capable of processing a framework of multi-frame spatiotemporal compression, making full use of spatiotemporal redundancy across frames, and significantly improving compression efficiency. At the same time, it combines a multimodal integration strategy of visual and textual semantics, and realizes priority transmission of key semantic information through the encoding method of semantic sketches and descriptive texts, thereby maintaining efficient compression and perceptual relevance under bandwidth-constrained conditions. In addition, in order to solve the stability problem of the generation model in restoring fidelity, a diffusion-based video control network is adopted to balance compression efficiency and reconstruction quality through a two-stage training process to ensure semantic consistency and temporal coherence of video content. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention.

[0088] Figure 1 A flow chart of a multimodal method for real-time high-fidelity video transmission based on semantic streams provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0089] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0090] It is understandable that in the prior art, video compression technology mainly includes traditional methods and deep learning-based methods, among which: traditional methods (such as H.264 / AVC, HEVC, etc.) reduce the amount of data by compressing single frames or adjacent frames. Although they can effectively utilize spatial redundancy and short-term temporal correlation, they are difficult to model long-term temporal dependence, which limits the compression efficiency. In addition, the existing methods use pixel fidelity as the optimization goal, do not fully consider the perceptual characteristics of the human visual system (HVS), and perform poorly when bandwidth is limited or network conditions change dynamically; deep learning-based methods use neural networks to extract high-level semantic features, significantly improve compression efficiency, and optimize perceptual indicators by adaptively modeling spatial and temporal redundancy, which is more in line with human visual preferences. However, generative models (such as GAN and diffusion models) have stability issues, which may lead to inconsistent reconstruction, especially in terms of high-frequency details and temporal coherence. In addition, the existing methods do not make sufficient use of the key semantic features of video content and cannot prioritize important information, which limits compression efficiency and dynamic network adaptability.

[0091] To solve the above problems, the embodiment of the present invention receives a multi-frame video sequence, extracts spatial and temporal correlations through a spatiotemporal compressor, and outputs a potential representation; maps the potential representation to a semantic space through a semantic translator, and gradually transforms visual features and text features; multimodally fuses the potential representation and text features through a preset Transformer fusion model, and outputs a fused representation; quantizes the fused representation into a quantized representation through a preset codebook model; and processes the quantized representation and text features through a video control network to generate a reconstructed video sequence. The compression efficiency can be significantly improved, and priority transmission of key semantic information can be achieved, thereby maintaining efficient compression and perceptual correlation under bandwidth-constrained conditions, and ensuring semantic consistency and temporal coherence of video content.

[0092] in, Figure 1 A flow chart of a multimodal method for real-time high-fidelity video transmission based on semantic streams provided by an embodiment of the present invention is shown.

[0093] Specifically, a multimodal method for real-time high-fidelity video transmission based on semantic streams comprises the following steps:

[0094] Step S101, receiving a multi-frame video sequence, extracting spatial and temporal correlations through a spatiotemporal compressor, and outputting a latent representation.

[0095] In the embodiment of the present invention, a multi-frame video sequence is received, and the spatial structure and motion pattern between adjacent frames of the multi-frame video sequence are captured through a 3D convolution layer, and a joint spatiotemporal feature is output. Then, based on the joint spatiotemporal feature, the short-term and long-term dependencies in the multi-frame video sequence are encoded through a recurrent layer, and the encoded data is output. Then, the cross-row convolution is applied through a downsampling layer to compress the encoded data into a potential representation. Specifically, the expression of the multi-frame video sequence is:

[0096] ;

[0097] in, is a multi-frame video sequence, Indicates the frame number;

[0098] The overall expression of the potential representation is:

[0099] ;

[0100] in, For potential representation, are the parameters of the space-time compressor;

[0101] The expression of joint spatiotemporal features is:

[0102] ;

[0103] in, is the joint spatiotemporal feature, are the parameters of the convolutional layer;

[0104] The expression for encoding the data is:

[0105] ;

[0106] in, represents the encoded data, are the parameters of the recurrent network;

[0107] The detailed expression of the potential representation is:

[0108] ;

[0109] in, are the parameters of the downsampling layer.

[0110] Understandably, the spatiotemporal compressor is a key component of the semantic flow framework for processing multi-frame video sequences and extracting spatial and temporal correlations. Its goal is to transform high-dimensional video data into a compact latent representation that retains the necessary semantic information while discarding redundant details. The compressor integrates a hybrid architecture that combines convolutional layers and recurrent networks to ensure seamless representation of spatiotemporal dynamics.

[0111] It can be understood that the 3D convolutional layer captures both the spatial structure and motion patterns between adjacent frames, allowing the network to identify dynamic changes, such as object movement or transformation, while maintaining the structural integrity of static regions. To further capture temporal dependencies, the output of the 3D convolutional layer is passed through a recurrent layer, such as a GRU or LSTM unit. These layers are particularly effective for sequence data because they maintain a memory state that reflects the relationship between frames.

[0112] It can be understood that the processing of recurrent layers ensures that the latent representation reflects dynamic interactions across frames, such as object tracking and motion flow, while maintaining a compact feature set.

[0113] It can be understood that the downsampling layer reduces the dimensionality of the time-aware features. By applying cross-row convolutions, this module compresses the data into a compact latent representation, which not only minimizes the memory and computation requirements of subsequent modules, but also ensures Keep it semantically rich.

[0114] It can be understood that the combination of 3D convolutions, recurrent networks, and downsampling creates a powerful pipeline for spatiotemporal-pore feature extraction, optimizes the latent representation to preserve key information in multi-frame video sequences, making it suitable for integration with the downstream semantic processing components of the framework, and ensures that the compressor effectively balances compression efficiency and semantic fidelity, providing a solid foundation for the semantic flow architecture.

[0115] Step S102, mapping the latent representation to a semantic space through a semantic translator, and gradually transforming visual features and text features.

[0116] In the embodiment of the present invention, a visual encoder is used to extract high-level semantic features from a potential representation to generate a visual feature map. Then, through a cross-attention layer, the visual feature map is dynamically associated with a pre-trained language embedding obtained from a preset large language model to establish a correspondence between visual elements and language counterparts, and an alignment weight is calculated to obtain a highly aligned feature set. Then, by using a text decoder, the semantics specific to the video are fine-tuned to capture the highly aligned feature set to generate text features. Specifically, the expression of the visual feature is:

[0117] ;

[0118] in, For visual features, is the first parameter of the semantic translator;

[0119] The overall expression of text features is:

[0120] ;

[0121] in, is the text feature, It is the second parameter of the semantic translator;

[0122] The expression of visual feature mapping is:

[0123] ;

[0124] in, is the visual feature map, are the parameters of the visual encoder;

[0125] The expression of highly aligned feature set is:

[0126] ;

[0127] in, represents a highly aligned feature set, for pre-trained language embeddings, is the parameter across the attention layer;

[0128] The detailed expression of text features is:

[0129] ;

[0130] in, Parameters for the text decoder.

[0131] It can be understood that the semantic translator is a key component in the semantic flow framework to bridge the gap between visual semantics and language semantics. Its role is to convert the latent representation obtained from the spatiotemporal compressor into meaningful text features, encapsulating the core semantic content of the video. By adopting an encoder-decoder architecture with an enhanced attention mechanism, the semantic translator ensures that basic visual details are effectively translated into language-based representations.

[0132] It can be understood that the visual encoder extracts high-level semantic features from the latent representation by utilizing a lightweight convolutional neural network (CNN), which is designed to operate on a compressed latent space without adding significant computational overhead. The lightweight convolutional neural network (CNN) is able to apply multiple convolutional layers to gradually refine spatial features and abstract features to generate visual feature maps that are able to encapsulate the key semantic structures and patterns embedded in video sequences while maintaining the compactness of the latent representation.

[0133] It can be understood that cross-attention layers dynamically associate visual feature maps with pre-trained language embeddings obtained from a Large Language Model (LLM), establishing a correspondence between visual elements (such as objects, motion patterns, or scene compositions) and their language counterparts.

[0134] Understandably, in order to align these visual semantics with the language information, the semantic translator employs a cross-attention mechanism that computes alignment weights that prioritize the most semantically relevant visual features while conditioning them according to the language embeddings ELLM:.

[0135] Understandably, the semantic translator uses a text decoder that leverages pre-trained LLM weights, specifically fine-tuned to capture video-specific semantics, ensuring that the generated text is both contextually accurate and semantically rich. The decoding process produces a sequence of words that encapsulates the core meaning of the visual input.

[0136] Understandably, the semantic translator ensures that the complex and multifaceted semantics of video content are summarized into an interpretable and compact text format by effectively combining visual encoding, cross-attention alignment, and text decoding. This text-based semantic representation plays a key role in guiding the video control network, ensuring that the reconstructed video remains faithful to its original semantic intent while maintaining high fidelity and temporal consistency.

[0137] Step S103, performing multimodal fusion on the latent representation and the text feature through a preset Transformer fusion model, and outputting a fused representation.

[0138] In the embodiment of the present invention, the text features are embedded into the shared feature space of the preset Transformer fusion model through the text embedding layer to generate a first dense vector, and then the potential representation is embedded into the shared feature space of the preset Transformer fusion model through the visual embedding layer to generate a second dense vector. The cross attention layer is used to perform fusion alignment processing of semantic information and visual information on the first dense vector and the second dense vector, and the fusion representation is output. Specifically, the overall expression of the fusion representation is:

[0139] ;

[0140] in, For fusion representation, are the parameters of the Transformer fusion model;

[0141] The expression of the first dense vector is:

[0142] ;

[0143] in, represents the first dense vector, Parameters for the text embedding layer;

[0144] The expression of the second dense vector is:

[0145] ;

[0146] in, represents the second dense vector, are the parameters of the visual embedding layer;

[0147] The detailed expression of the fusion representation is:

[0148] ;

[0149] in, are the parameters of the Transformer layer.

[0150] Understandably, the Transformer fusion module plays a key role in integrating the textual features generated by the semantic translator and the latent representation produced by the spatiotemporal compressor. This integration ensures that the semantic and visual information are seamlessly combined into a unified representation, enabling downstream components to reconstruct high-quality and semantically accurate videos.

[0151] It can be understood that textual and visual features are processed through a series of transformation layers, allowing each modality to refine its internal representation, such that for textual features, semantic dependencies between words or tokens are preserved, while for visual features, spatial and temporal correlations are preserved.

[0152] It is understandable that in order to achieve inter-modal fusion, a cross-attention layer is adopted to align semantic information with visual information, allowing textual features to guide visual features in representing context, while visual features provide well-founded references for semantics. The fused representation encapsulates the semantic consistency from textual features and the spatiotemporal context from visual features. This representation is specifically optimized for downstream tasks, ensuring that the generated videos remain faithful to the semantic intent while maintaining high visual fidelity. By effectively aligning textual guidance with visual context, the Transformer fusion module significantly improves the quality and consistency of the reconstructed video. The design of the Transformer fusion module helps to address the challenges of multimodal integration. By leveraging the attention mechanism and shared embedding space, it ensures that the complementary information of text and vision is fully utilized, paving the way for high-quality video reconstruction in the semantic flow framework.

[0153] Step S104: quantize the fused representation into a quantized representation using a preset codebook model.

[0154] In the embodiment of the present invention, the fusion representation is quantized through a preset codebook model, ensuring efficient compression while maintaining semantic fidelity. Specifically, the expression of the quantized representation is:

[0155] ;

[0156] in, To quantify, is the codebook parameter.

[0157] Step S105: Process the quantized representation and the text features through a video control network to generate a reconstructed video sequence.

[0158] In the embodiment of the present invention, the quantized representation and the text features are denoised by a preset denoising diffusion probability model to obtain a denoised representation, and then the cross-attention layer is used to dynamically align the quantized representation and the text features at each time step by adopting a preset semantic alignment control mechanism to obtain an aligned expression, and then the upsampling module is used to apply convolution and super-resolution technology to refine the spatial details and generate high-resolution frames to generate a reconstructed video sequence. Specifically, the overall expression of the reconstructed video sequence is:

[0159] ;

[0160] in, represents the reconstructed video sequence, Parameters of the video control network;

[0161] The expression of denoising is:

[0162] ;

[0163] in, For denoising, represents the current time step, are the parameters of the denoising diffusion probability model;

[0164] The expression for the alignment expression is:

[0165] ;

[0166] in, To align the expression, is the parameter of the cross-attention layer.

[0167] It can be understood that the Video Control Network, as the generative backbone of the semantic flow framework, is tasked with reconstructing high-quality video sequences under the guidance of semantics and latent representations. It adopts a diffusion-based generation process and enhances semantic conditioning to ensure visual fidelity and adherence to the original semantic intent. Its architecture integrates a denoising diffusion probability model (DDPM), a semantic alignment control mechanism, and an upsampling network for high-resolution output. At its core, the diffusion backbone operates as a denoising process, gradually refining the noisy latent input into coherent video frames, by iteratively generating videos, starting from a noise distribution and combining semantic and visual guidance at each step.

[0168] It can be understood that during the denoising process, by iteratively reducing noise over multiple time steps, the video frames generated by the diffusion backbone are temporally consistent and semantically consistent with the original content.

[0169] It can be understood that the use of a preset semantic alignment control mechanism to dynamically align the quantized representation and text features at each time step can ensure that the generated frames accurately reflect the semantic intent of the input, such as object identity, motion, or scene context. This not only improves the model's dependence on semantic cues, but also stabilizes the generation process and ensures consistency of output across frames.

[0170] It is understandable that the upsampling module aims to improve the spatial resolution of the generated frames. While the initial output of the diffusion backbone operates in a lower-dimensional latent space, the upsampling module applies convolution and super-resolution techniques to refine the spatial details and produce high-resolution frames, ensuring that the final reconstructed video not only retains semantic accuracy but also achieves a visually pleasing quality suitable for practical applications.

[0171] As can be understood, by combining the denoising diffusion model, semantic control mechanism and upsampling module, Video Control Net achieves a strong synergy between generation flexibility and semantic accuracy. This integration enables the module to effectively process a variety of video content and generate outputs that are both visually high-fidelity and semantically coherent. Therefore, Video Control Net is the cornerstone of the semantic streaming framework, which enables robust video reconstruction even under challenging conditions such as high compression ratios or transmission quality degradation.

[0172] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0173] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0174] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0175] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

[0176] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A multimodal method for real-time high-fidelity video transmission based on semantic streaming, characterized in that: The method specifically comprises the following steps: Receive a multi-frame video sequence, extract spatial and temporal correlations through a spatiotemporal compressor, and output a latent representation; Mapping the latent representation to a semantic space through a semantic translator, gradually transforming visual features and text features; The latent representation and the text feature are multimodally fused through a preset Transformer fusion model, and a fused representation is output; quantizing the fused representation into a quantized representation through a preset codebook model; Processing the quantized representation and the text features through a video control network to generate a reconstructed video sequence; The expression of the multi-frame video sequence is: ; in, is a multi-frame video sequence, Indicates the frame number; The expression of the potential representation is: ; in, For potential representation, are the parameters of the space-time compressor; The expression of the visual feature is: ; in, For visual features, is the first parameter of the semantic translator; The expression of the text feature is: ; in, is the text feature, It is the second parameter of the semantic translator; The expression of the fusion expression is: ; in, For fusion representation, are the parameters of the Transformer fusion model; The expression of the quantitative representation is: ; in, To quantify, is the codebook parameter; The expression of the reconstructed video sequence is: ; in, represents the reconstructed video sequence, Parameters of the video control network; The step of mapping the potential representation to the semantic space by the semantic translator and gradually transforming the visual features and the text features specifically includes the following steps: extracting high-level semantic features from the latent representation through a visual encoder to generate a visual feature map; Dynamically associating the visual feature map with a pre-trained language embedding obtained from a preset large language model through a cross-attention layer, establishing a correspondence between visual elements and their language counterparts, and calculating alignment weights to obtain a highly aligned feature set; Generate text features by using a text decoder, based on the highly aligned feature set, fine-tuned to capture video-specific semantics; The expression of the visual feature map is: ; in, is the visual feature map, are the parameters of the visual encoder; The expression of the highly aligned feature set is: ; in, represents a highly aligned feature set, for pre-trained language embeddings, is the parameter across the attention layer; The expression of the text feature is: ; in, Parameters for the text decoder; The step of processing the quantized representation and the text features through the video control network to generate a reconstructed video sequence specifically comprises the following steps: De-noising the quantized representation and the text feature by using a preset denoising diffusion probability model to obtain a denoised representation; By using a cross-attention layer and a preset semantic alignment control mechanism, the quantitative representation and the text feature are dynamically aligned at each time step to obtain an aligned expression; Through the upsampling module, convolution and super-resolution techniques are applied to refine spatial details and produce high-resolution frames to generate reconstructed video sequences; The expression of the denoising expression is: ; in, For denoising, represents the current time step, are the parameters of the denoising diffusion probability model; The expression of the alignment expression is: ; in, To align the expression, is the parameter of the cross-attention layer.

2. The multimodal method for real-time high-fidelity video transmission based on semantic stream according to claim 1, characterized in that: The receiving of a multi-frame video sequence, extracting spatial and temporal correlations through a spatiotemporal compressor, and outputting a potential representation specifically comprises the following steps: Receiving a multi-frame video sequence; Capturing the spatial structure and motion pattern between adjacent frames of the multi-frame video sequence through a 3D convolutional layer, and outputting joint spatiotemporal features; Based on the joint spatiotemporal features, encoding the short-term and long-term dependencies in the multi-frame video sequence through a recurrent layer, and outputting encoded data; Through the downsampling layer, cross-row convolution is applied to compress the encoded data into a latent representation.

3. The multimodal method for real-time high-fidelity video transmission based on semantic stream according to claim 2, characterized in that: The expression of the joint spatiotemporal feature is: ; in, is the joint spatiotemporal feature, are the parameters of the convolutional layer; The expression of the coded data is: ; in, represents the encoded data, are the parameters of the recurrent network; The expression of the potential representation is: ; in, are the parameters of the downsampling layer.

4. The multimodal method for real-time high-fidelity video transmission based on semantic stream according to claim 1, characterized in that: The method of performing multimodal fusion of the latent representation and the text feature through a preset Transformer fusion model and outputting the fused representation specifically includes the following steps: Embedding the text features into a shared feature space of a preset Transformer fusion model through a text embedding layer to generate a first dense vector; Embedding the latent representation into a shared feature space of a preset Transformer fusion model through a visual embedding layer to generate a second dense vector; A cross attention layer is used to perform fusion and alignment processing of semantic information and visual information on the first dense vector and the second dense vector, and a fusion representation is output.

5. The multimodal method for real-time high-fidelity video transmission based on semantic stream according to claim 4, characterized in that: The expression of the first dense vector is: ; in, represents the first dense vector, Parameters for the text embedding layer; The expression of the second dense vector is: ; in, represents the second dense vector, are the parameters of the visual embedding layer; The expression of the fusion expression is: ; in, are the parameters of the Transformer layer.

Citation Information

Patent Citations

  • Video semantic representation method and system based on multi-mode fusion mechanism and medium

    CN109472232A

  • Low-bandwidth crowd scene security monitoring method and system based on semantic coding and decoding

    CN116708725A