A video semantic communication system and apparatus

CN122554441APending Publication Date: 2026-08-11TSINGHUA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

该类方法在稳定信道条件下具有较好应用效果,但在短码传输、强时变信道、严格时延约束等场景下,码率、可靠性与时延之间往往难以兼顾

Benefits of technology

本申请提供了一种视频语义通信系统,包括:在发送端部署以下多个模块,得到理想接收分支:嵌入有可学习的特征提取向量的待训练特征提取器、嵌入有可学习的上下文提取向量的待训练上下文提取器、待训练超先验编码器、待训练超先验解码器、待训练空间先验模块、嵌入有可学习的生成向量的待训练生成器;利用样本视频帧对所述理想接收分支进行训练,得到训练完毕的理想接收分支、训练完毕的特征提取向量、训练完毕的上下文提取向量、训练完毕的生成向量;所述训练完毕的特征提取向量用于对训练完毕的特征提取器所提取的语义特征进行调制,所述训练完毕的上下文提取向量用于对训练完毕的上下文提取器所提取的上下文进行调制,所述训练完毕的生成向量用于对训练完毕的生成器生成的重建特征进行调制;基于所述训练完毕的理想接收分支,将待训练特征编码器部署在所述发送端,将待训练特征解码器部署在发送端缓冲区更新分支和接收端,得到待训练语义通信系统;利用所述样本视频帧对所述待训练特征编码器和所述待训练特征解码器进行参数更新,得到第一视频语义通信系统。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554441A_ABST
    Figure CN122554441A_ABST
Patent Text Reader

Abstract

This application proposes a video semantic communication system and apparatus, relating to the field of semantic communication technology. The system is obtained by first training an ideal receiving branch and then updating the parameters of the feature encoder and feature decoder, which improves training stability and convergence efficiency. Furthermore, by modulating semantic features, contextual information, and reconstructed features, the representation capability of key semantic information in the video is enhanced, improving the perceptual quality of the reconstructed video. In addition, through advanced prior and spatial prior modeling, bitrate utilization efficiency can be improved; and through collaborative decoding at the transmitter and receiver, error accumulation caused by reference drift can be reduced, thereby improving the transmission robustness and temporal stability of the video under bandwidth-constrained and channel fluctuation conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of semantic communication technology, and in particular to a video semantic communication system and apparatus. Background Technology

[0002] With the development of 4K / 8K ultra-high-definition video live streaming, mobile short videos, and immersive interactive services, video transmission services have placed comprehensive demands on communication systems, requiring high resolution, low latency, and stable reconstruction quality. In communication scenarios such as satellite links and cellular networks, factors such as bandwidth limitations, channel fluctuations, and limited end-side computing power are prevalent, posing significant challenges to the real-time transmission of ultra-high-definition video.

[0003] Existing video transmission systems typically employ a source-channel separation architecture, where video source coding is performed first, followed by channel coding and modulation transmission. The source coding side often uses standards such as H.264 / AVC (Advanced Video Coding), H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Coding), while the channel coding side often uses error correction coding methods such as LDPC (Low-Density Parity-Check) and Polar Code. While these methods perform well under stable channel conditions, they often struggle to balance bit rate, reliability, and latency in scenarios involving short bitrate transmission, highly time-varying channels, and strict latency constraints.

[0004] Furthermore, traditional methods often use pixel-level metrics such as PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index) as optimization targets. However, these metrics do not always align with subjective perception and are prone to issues such as texture blurring, edge distortion, and temporal flicker. In recent years, while deep learning-based intelligent compression and semantic communication methods have made some progress, they still have shortcomings in areas such as model lightweighting, multi-rate adaptation, control of side information overhead, and temporal stability.

[0005] Therefore, this application proposes a novel video semantic communication system. Summary of the Invention

[0006] In view of the above problems, embodiments of this application provide a video semantic communication system and apparatus to overcome or at least partially solve the above problems.

[0007] A first aspect of this application provides a video semantic communication system, comprising: The following modules are deployed at the transmitting end to obtain the ideal receiving branch: a trainable feature extractor embedded with learnable feature extraction vectors, a trainable context extractor embedded with learnable context extraction vectors, a trainable super-prior encoder, a trainable super-prior decoder, a trainable spatial prior module, and a trainable generator embedded with learnable generation vectors. The ideal receiving branch is trained using sample video frames to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector. The trained feature extraction vector is used to modulate the semantic features extracted by the trained feature extractor, the trained context extraction vector is used to modulate the context extracted by the trained context extractor, and the trained generation vector is used to modulate the reconstructed features generated by the trained generator. Based on the ideal receiving branch that has been trained, the feature encoder to be trained is deployed at the sending end, and the feature decoder to be trained is deployed at the sending end buffer update branch and the receiving end to obtain the semantic communication system to be trained. The parameters of the feature encoder and the feature decoder to be trained are updated using the sample video frames to obtain the first video semantic communication system.

[0008] A second aspect of this application provides a video semantic communication device, the device comprising: The first deployment module is used to deploy the following modules at the sending end to obtain the ideal receiving branch: a feature extractor to be trained with embedded learnable feature extraction vectors, a context extractor to be trained with embedded learnable context extraction vectors, a super prior encoder to be trained, a super prior decoder to be trained, a spatial prior module to be trained, and a generator to be trained with embedded learnable generation vectors. The training module is used to train the ideal receiving branch using sample video frames to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector. The trained feature extraction vector is used to modulate the semantic features extracted by the trained feature extractor, the trained context extraction vector is used to modulate the context extracted by the trained context extractor, and the trained generation vector is used to modulate the reconstructed features generated by the trained generator. The second deployment module is used to deploy the feature encoder to be trained at the sending end and the feature decoder to be trained at the sending end buffer update branch and the receiving end based on the ideal receiving branch that has been trained, so as to obtain the semantic communication system to be trained. The parameter update module is used to update the parameters of the feature encoder and the feature decoder to be trained using the sample video frames, so as to obtain the first video semantic communication system.

[0009] The beneficial effects of this application are: This application provides a video semantic communication system, comprising: deploying multiple modules at the transmitting end to obtain an ideal receiving branch: a feature extractor to be trained embedded with learnable feature extraction vectors, a context extractor to be trained embedded with learnable context extraction vectors, a pre-prior encoder to be trained, a pre-prior decoder to be trained, a spatial prior module to be trained, and a generator to be trained embedded with learnable generation vectors; training the ideal receiving branch using sample video frames to obtain a trained ideal receiving branch, a trained feature extraction vector, a trained context extraction vector, and a trained generation vector; the trained feature extraction vector is used... The semantic features extracted by the trained feature extractor are modulated. The trained context extraction vector is used to modulate the context extracted by the trained context extractor, and the trained generation vector is used to modulate the reconstructed features generated by the trained generator. Based on the trained ideal receiving branch, the feature encoder to be trained is deployed at the transmitting end, and the feature decoder to be trained is deployed at the transmitting end buffer update branch and the receiving end to obtain the semantic communication system to be trained. The parameters of the feature encoder and the feature decoder to be trained are updated using the sample video frames to obtain the first video semantic communication system.

[0010] The video semantic communication system provided in this application improves training stability and convergence efficiency by first training an ideal receiving branch and then updating the parameters of the feature encoder and feature decoder. Furthermore, by modulating semantic features, contextual information, and reconstructed features, the system enhances the representation of key semantic information in the video, thereby improving the perceptual quality of the reconstructed video. In addition, the use of advanced prior and spatial prior modeling improves bitrate utilization efficiency; and collaborative decoding between the transmitter and receiver reduces error accumulation caused by reference drift, thus enhancing the transmission robustness and temporal stability of the video under bandwidth-constrained and channel fluctuation conditions. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating the steps of a training method for a video semantic communication system provided in an embodiment of this application. Figure 2 This is a schematic diagram of an ideal receiving branch architecture provided in an embodiment of this application; Figure 3 This is an architecture diagram of a sender buffer update branch provided in an embodiment of this application; Figure 4 This is a schematic diagram of the overall architecture of a video semantic communication system provided in an embodiment of this application; Figure 5 This is a schematic diagram of a training device for a video semantic communication system provided in an embodiment of this application. Detailed Implementation

[0013] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.

[0014] In a first aspect, this application provides a video semantic communication system, such as... Figure 1 The diagram shows a training method for the video semantic communication system provided in this application, specifically including: S101, the following modules are deployed at the transmitting end to obtain the ideal receiving branch: a feature extractor to be trained with embedded learnable feature extraction vectors, a context extractor to be trained with embedded learnable context extraction vectors, a super prior encoder to be trained, a super prior decoder to be trained, a spatial prior module to be trained, and a generator to be trained with embedded learnable generation vectors. S102, the ideal receiving branch is trained using sample video frames to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector; the trained feature extraction vector is used to modulate the semantic features extracted by the trained feature extractor, the trained context extraction vector is used to modulate the context extracted by the trained context extractor, and the trained generation vector is used to modulate the reconstructed features generated by the trained generator. S103, based on the ideal receiving branch that has been trained, the feature encoder to be trained is deployed at the sending end, and the feature decoder to be trained is deployed at the sending end buffer update branch and the receiving end to obtain the semantic communication system to be trained. S104, the parameters of the feature encoder and the feature decoder to be trained are updated using the sample video frames to obtain the first video semantic communication system.

[0015] In S101, at the transmitting end, a trainable feature extractor embedded with learnable feature extraction vectors, a trainable context extractor embedded with learnable context extraction vectors, a trainable super-prior encoder, a trainable super-prior decoder, a trainable spatial prior module, and a trainable generator embedded with learnable generation vectors are deployed respectively. The trainable feature extractor is used to extract semantic latent features from the current input video frame under temporal context conditions. The trainable context extractor is used to extract temporal context information from the reference buffer maintained at the transmitting end to characterize the prior constraints of historical reference frames or historical features on the reconstruction of the current frame. The trainable super-prior encoder and the trainable super-prior decoder are used to generate and recover side information to enhance the modeling ability of discrete semantic feature probability distributions. The trainable spatial prior module is used to output the conditional distribution parameters of each dimension of semantic features under the combined effect of temporal context and side information. The trainable generator is used to generate reconstructed video frames based on the recovered features and synchronously update intermediate reference features.

[0016] In step S102, the ideal receiving branch is trained using sample video frames, resulting in the trained ideal receiving branch and corresponding trained feature extraction vector, context extraction vector, and generation vector. In this step, only the parameters of the modules related to the ideal receiving branch are updated, while the parameters related to the feature encoding path are temporarily frozen, allowing the system to first obtain relatively stable video generation and reference buffer update capabilities. During specific training, the quantization process can initially use additive uniform noise approximation to improve training stability and accelerate convergence, and then switch to a pass-through estimator to reduce inconsistencies between the training and inference phases; perceptual constraints can also be gradually introduced after the quantization process has stabilized. The feature extraction vector obtained through this step is used to modulate the semantic features extracted by the feature extractor; the trained context extraction vector is used to modulate the temporal context output by the context extractor; and the trained generation vector is used to modulate the reconstructed features output by the generator.

[0017] In S103, after completing the training of the ideal receiving branch, the feature encoder to be trained is further deployed at the transmitting end, and the feature decoder to be trained is deployed at the transmitting end buffer update branch and the receiving end, respectively, to obtain the semantic communication system to be trained. The feature encoder is mainly used to map the quantized semantic features into a symbol sequence suitable for channel transmission; the feature decoder is used to recover the received symbol sequence into discrete features. The reason this application deploys the feature decoder simultaneously at the transmitting end buffer update branch and the receiving end is that this application not only focuses on video reconstruction at the receiving end but also emphasizes the consistency of the reference states between the transmitting and receiving ends. Through its local ideal receiving branch and its corresponding feature recovery and generation process, the transmitting end can synchronously update the transmitting end reference buffer, ensuring that the reference state maintained internally by the transmitting end is as consistent as possible with that of the receiving end, thereby suppressing the reference drift accumulation problem during long-sequence transmission.

[0018] In S104, the parameters of the feature encoder and decoder to be trained are updated using sample video frames to obtain the first video semantic communication system. In this step, since the ideal receiving branch already possesses relatively stable feature recovery and reference update capabilities, the training focus can be placed on the feature mapping and recovery process between the transmitter and receiver, i.e., optimizing the feature encoder and decoder to achieve more robust semantic feature transmission and recovery under given side information, temporal context, and channel disturbances. During training, an explicit wireless channel can be incorporated into the training loop, allowing the encoder and decoder to perform joint optimization under conditions of channel fluctuations, noise, or fading, thereby improving the system's transmission robustness. Furthermore, in subsequent training stages, the frozen state of more modules can be unfrozen, performing end-to-end joint optimization, and pruning the rate embedding and corresponding mapping layers used less frequently during training to reduce inference complexity and side information overhead.

[0019] This application constructs and pre-trains an ideal receiving branch to first obtain stable feature recovery, video generation, and reference update capabilities. Then, based on this, it trains the feature encoder and feature decoder to undertake the actual transmission tasks, thereby achieving phased training of the video semantic communication system. This training method can improve training stability and convergence efficiency, and helps enhance the system's transmission robustness, reconstruction perception quality, and long-term temporal stability under bandwidth-constrained and channel fluctuation scenarios.

[0020] In one embodiment, training the ideal receiving branch using sample video frames to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector includes: The context extractor to be trained extracts the context of the sending reference features in the sending decoding buffer to obtain the sending super-prior context and the sending feature context. Using the feature extractor to be trained, semantic features are obtained by extracting features from sample video frames under the feature context of the sending end. Based on the semantic features and the sending end super-prior context, the ideal receiving semantic features are obtained by sequentially using the super-prior encoder to be trained, the super-prior decoder to be trained, and the spatial prior module to be trained. The ideal reconstructed video frame is obtained by using the ideal receiving semantic features under the condition of the sending end feature context through the generator to be trained. Based at least on the difference between the sample video frame and the ideal reconstructed video frame, the parameters of each module in the ideal receiving branch and each learnable vector are updated to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector.

[0021] In this embodiment, the architecture diagram of the ideal receiving branch is as follows: Figure 2 As shown, the transmitter first reads reference features from the transmitter's decoding buffer and then extracts context from these features using a context extractor to obtain the transmitter's hyper-prior context and transmitter's feature context. The hyper-prior context provides a temporal reference for subsequent hyper-prior side information modeling, while the transmitter's feature context provides temporal correlation information for semantic feature extraction of the current video frame and subsequent video generation. Since the context information directly originates from historical reference features in the transmitter's buffer, this step can establish an implicit temporal correlation between the current frame and historical frames without transmitting explicit motion information, providing a priori basis for subsequent semantic residual extraction and stable reconstruction.

[0022] Furthermore, after obtaining the sender's prior context and feature context, the feature extractor to be trained extracts features from the sample video frames under the conditions of the sender's feature context, thus obtaining semantic features. Here, feature extraction does not involve isolated encoding of the current video frame, but rather extracts the semantic latent representation of the current frame relative to the historical reference state under temporal context constraints. Therefore, the obtained semantic features can more effectively represent the information that needs to be supplemented or reconstructed in the current frame. To further achieve single-model multi-rate adaptation, a learnable feature extraction vector is also introduced inside the feature extractor to modulate the features of intermediate channels, enabling different channels to form a response capability adapted to bitrate control and quantization control during training, thereby enhancing the expressive efficiency of semantic features.

[0023] Furthermore, after obtaining the semantic features and the transmitter's prior context, the semantic features are sequentially recovered through a trainable prior encoder, a trainable prior decoder, and a trainable spatial prior module to obtain the ideal received semantic features. Specifically, the prior encoder generates side information representations based on the semantic features, the prior decoder recovers the side information representations and outputs the basic quantization scale and initial statistical parameters, and the spatial prior module, under the combined effect of the transmitter's prior context and the recovered side information, outputs the conditional distribution parameters for each dimension of the semantic features. Based on the aforementioned basic quantization scale, statistical parameters, and conditional distribution parameters, the semantic features can be modulated, quantized, and recovered, thereby forming the ideal received semantic features corresponding to the recovered results at the receiver locally at the transmitter.

[0024] After obtaining the ideal received semantic features, the generator to be trained reconstructs video frames using these features within the context of the sending end's features, resulting in ideally reconstructed video frames. Simultaneously, the generator can also synchronously generate updated intermediate features to update the reference state in the sending end's decoding buffer, ensuring that the reference features used by the sending end to extract context in subsequent time steps remain consistent with the currently trained reconstruction results.

[0025] After generating the ideal reconstructed video frame, the parameters and learnable vectors of each module in the ideal receiving branch are updated based on the differences between the sample video frame and the ideal reconstructed video frame. This results in the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector. The trained feature extraction vector can be used for channel-level modulation of the semantic features extracted by the feature extractor; the trained context extraction vector can be used for channel-level modulation of the context information extracted by the context extractor; and the trained generation vector can be used for channel-level modulation of the reconstructed features generated by the generator.

[0026] In one embodiment, based on the trained ideal receiving branch, the feature encoder to be trained is deployed at the transmitting end, and the feature decoder to be trained is deployed at the transmitting end buffer update branch and the receiving end to obtain the semantic communication system to be trained, including: At the transmitting end, the trained context extractor, the trained feature extractor, the trained super prior encoder, the trained super prior decoder, the first spatial prior module, and the feature encoder to be trained are connected in series in sequence. At the output of the feature encoder to be trained, the feature decoder to be trained, the second spatial prior module, and the trained generator are connected in series to form the sending buffer update branch; the first spatial prior module and the second spatial prior module are both the trained spatial prior modules. At the receiving end, the feature decoder to be trained, the trained super-prior decoder, the trained spatial prior module, the trained generator, and the trained context extractor are deployed.

[0027] In this embodiment, reference Figure 3The diagram shown illustrates the architecture of the sender buffer update branch. At the sender, the following components are sequentially connected: a trained context extractor, a trained feature extractor, a trained super-prior encoder, a trained super-prior decoder, a first spatial prior module, and a feature encoder to be trained. This deployment corresponds to the forward processing chain at the sender. The context extractor extracts temporal context from reference features in the sender buffer. The feature extractor extracts semantic latent features of the current video frame under the temporal context. The super-prior encoder and super-prior decoder generate and recover side information. The first spatial prior module combines the temporal context and side information to output the conditional distribution parameters required for feature recovery. The feature encoder to be trained further maps the features to be transmitted into a channel symbol sequence.

[0028] At the output of the feature encoder to be trained, the feature decoder to be trained, the second spatial prior module, and the trained generator are sequentially connected in series to form the sender buffer update branch. This sender buffer update branch does not undertake the actual external transmission function, but instead recovers the symbol-corresponding features already output by the feature encoder to be trained locally at the sender end, and further generates local reconstruction results and updated intermediate features, which are then used to update the sender buffer. In this way, the sender end can simulate the recovery and generation process of the receiver end locally, so that the reference state on which the sender end subsequently extracts the time context is as consistent as possible with that of the receiver end, thereby suppressing the accumulation of reference drift during long sequence transmission.

[0029] At the receiving end, a feature decoder to be trained, a pre-prior decoder already trained, a spatial prior module already trained, a generator already trained, and a context extractor already trained are deployed. This deployment corresponds to the actual recovery link at the receiving end. Specifically, the feature decoder to be trained is used to recover the received channel symbols into discrete features; the pre-prior decoder is used to recover side information and generate relevant statistical parameters; the spatial prior module is used to output the conditional distribution parameters of each dimension at the receiving end; the generator is used to generate reconstructed video frames based on the recovered features; and the context extractor is used to extract temporal context from the receiving end buffer for use in subsequent frame reconstruction.

[0030] This embodiment sets up a buffer update branch at the sending end and deploys the feature decoder to be trained simultaneously on both the sending end's buffer update branch and the receiving end. This allows the sending end to locally simulate the feature recovery and video generation process at the receiving end, updating the reference state in the sending end's buffer in a timely manner. This approach helps ensure the consistency of reference information between the sending and receiving ends, reducing error accumulation caused by reference drift. Furthermore, it provides more accurate temporal context for subsequent video frames, thereby improving the temporal stability and reconstruction quality of the video semantic communication system.

[0031] In one embodiment, the step of updating the parameters of the feature encoder and the feature decoder to be trained using the sample video frames to obtain a first video semantic communication system includes at least: At the transmitting end, the trained context extractor extracts the context of the transmitting end reference features in the transmitting end decoding buffer to obtain the transmitting end super-prior context and the transmitting end feature context. At the transmitting end, the trained feature extractor extracts features from the sample video frames under the feature context of the transmitting end to obtain semantic features; At the transmitting end, based on the semantic features, the trained super-prior encoder, the trained super-prior decoder, and the first spatial prior module are used sequentially to obtain the symbol length factor; wherein, at the transmitting end, the trained super-prior decoder is used to perform super-prior decoding under the conditions of the super-prior context at the transmitting end. At the transmitting end, using the feature encoder to be trained, under the condition of the transmitting end semantic context in the transmitting end decoding buffer, rate matching and symbol mapping are performed according to the symbol length factor under an unrestricted rate set to obtain a semantic symbol sequence; The semantic symbol sequence flows into the sending buffer update branch to obtain the sending-end updated and reconstructed video frame; At least based on the difference between the sample video frame and the video frame reconstructed by the sending end, the parameters of the feature encoder to be trained deployed at the sending end are updated, and the parameters of the feature decoder to be trained deployed in the sending end buffer update branch and the receiving end are updated.

[0032] In this embodiment, at the sending end, the context extractor extracts the sending end reference features in the sending end decoding buffer to obtain the sending end prior context and the sending end feature context. Semantic features are obtained by extracting features from sample video frames under the feature context of the sending end using a feature extractor. Furthermore, based on semantic features, the symbol length factor is obtained by sequentially processing through a super-prior encoder, a super-prior decoder, and a first spatial prior module. The super-prior decoder performs super-prior decoding under the conditions of the super-prior context at the transmitting end to generate parameters related to rate matching and symbol mapping.

[0033] In this embodiment, after obtaining the symbol length factor, the transmitter uses the feature encoder to be trained to perform rate matching and symbol mapping under the transmitter semantic context in the transmitter decoding buffer, based on the symbol length factor and an unrestricted rate set, to obtain a semantic symbol sequence. The semantic symbol sequence then flows into the transmitter buffer update branch, where feature recovery, video generation, and buffer update are completed locally at the transmitter to obtain the transmitter-updated reconstructed video frame. Finally, based at least on the difference between the sample video frame and the video frame reconstructed by the sender, the parameters of the feature encoder to be trained in the sender, as well as the parameters of the feature decoder to be trained deployed in the update branch of the sender buffer and in the receiver are updated, thereby obtaining the first video semantic communication system.

[0034] This embodiment further trains the feature encoder and feature decoder jointly after the ideal receiver branch is pre-trained, enabling the system to complete symbol length prediction, rate matching and symbol mapping under the reference context at the transmitter, thereby enhancing the adaptability of the single model to different symbol length outputs.

[0035] In one embodiment, the step of updating the parameters of the feature encoder and the feature decoder to be trained using the sample video frames to obtain the first video semantic communication system further includes: At the transmitting end, based on the semantic features, the trained super-prior encoder and quantizer are used sequentially to obtain the quantized super-prior latent vector, which is then converted into super-prior side information and sent to the receiving end. At the transmitting end, the symbol length factor is converted into symbol length factor side information and sent to the receiving end; The sending end sends the semantic symbol sequence to the receiving end; At the receiving end, the trained context extractor extracts the context of the receiving end reference features in the receiving end decoding buffer to obtain the receiving end prior context and the receiving end feature context. Using the trained modules and the feature decoder to be trained in the receiving end, the reconstructed video frame is obtained; wherein, in the receiving end, the feature decoder to be trained performs rate matching and symbol mapping on the received semantic symbol sequence under the unrestricted rate set, based on the receiving end semantic context in the receiving end decoding buffer and the symbol length factor corresponding to the received symbol length factor side information; in the receiving end, the trained super-prior decoder performs super-prior decoding under the receiving end super-prior context; in the receiving end, the feature decoder to be trained performs video frame reconstruction under the receiving end feature context. In addition to updating the reconstructed video frame based on the difference between the sample video frame and the reconstructed video frame at the sending end, the parameters of the feature encoder to be trained deployed at the sending end are also updated based on the difference between the sample video frame and the reconstructed video frame at the receiving end. Furthermore, the parameters of the feature decoder to be trained deployed at the sending end and the buffer update branch at the sending end and the receiving end are also updated.

[0036] In this embodiment, at the transmitting end, based on semantic features, the quantized super-prior latent vector is obtained by sequentially passing it through a trained super-prior encoder and quantizer, and further converted into super-prior side information before being sent to the receiving end. The transmitting end also converts the symbol length factor obtained from the first spatial prior module into symbol length factor side information and sends this side information to the receiving end. Finally, the transmitting end sends the semantic symbol sequence obtained after rate matching and symbol mapping based on the symbol length factor to the receiving end.

[0037] At the receiving end, the trained context extractor extracts the context of the receiving end reference features in the receiving end decoding buffer, obtaining the receiving end hyper-prior context and the receiving end feature context. Based on the received hyper-prior information, symbol length factor information, and semantic symbol sequence, the receiving end uses the trained modules and the feature decoder to be trained to generate the reconstructed video frame. Specifically, the feature decoder to be trained, under the receiving end semantic context in the receiving end decoding buffer, performs rate matching and symbol mapping on the received semantic symbol sequence according to the symbol length factor corresponding to the received symbol length factor information, under an unrestricted rate set; the trained hyper-prior decoder performs hyper-prior decoding under the receiving end hyper-prior context to recover modulation, quantization, and statistical parameter information; the trained spatial prior module further outputs the conditional distribution parameters required for discrete feature recovery; based on this, the feature decoder and generator, combined with the receiving end feature context, complete the video frame reconstruction.

[0038] Furthermore, during the parameter update phase, in addition to updating the parameters of the feature encoder deployed at the transmitter and the feature decoder deployed in the transmitter's buffer update branch based on the differences between the sample video frames and the reconstructed video frames at the transmitter, the parameters of the feature encoder deployed at the transmitter and the feature decoder deployed at the receiver are also updated based on the differences between the sample video frames and the reconstructed video frames at the receiver. This embodiment simultaneously utilizes the local reconstruction results at the transmitter and the actual reconstruction results at the receiver as supervision criteria, enabling the feature encoder and feature decoder to adapt not only to the reference alignment requirements in the transmitter's buffer update branch but also to the feature recovery requirements in the receiver's actual recovery link.

[0039] This embodiment, based on the established local recovery link at the sending end, further incorporates the actual recovery link at the receiving end into the training process. The local branch at the sending end can be used to update the sending end reference state, providing a more accurate temporal context for subsequent moments; the actual reconstruction branch at the receiving end can be used to enable the model to better adapt to the recovery process under real transmission scenarios. By simultaneously utilizing both types of reconstruction results for parameter updates, the first video semantic communication system can balance reference consistency, transmission stability, and video reconstruction quality during subsequent operation.

[0040] In one embodiment, the code rate loss is obtained based on the semantic symbol sequence, the prior information, and the symbol length factor information. In addition to updating the reconstructed video frame based on the difference between the sample video frame and the reconstructed video frame at the transmitting end, and the difference between the sample video frame and the reconstructed video frame at the receiving end, the parameters of the feature encoder to be trained deployed at the transmitting end are also updated based on the bitrate loss, as are the parameters of the feature decoder to be trained deployed in the updating branch of the transmitting end buffer and the receiving end.

[0041] Based on the above embodiments, the code rate loss can also be obtained based on the semantic symbol sequence, prior side information, and symbol length factor side information. In this embodiment, the end-to-end optimization objective adopts a joint objective of "code rate loss + reconstruction consistency loss + perceptual quality constraint loss", where the code rate loss includes the spatial unit symbol budget and side information entropy term. Therefore, the code rate loss can be calculated jointly based on the semantic symbol sequence, prior side information, and symbol length factor side information to reflect the comprehensive cost of the current training sample in terms of both symbol transmission and side information transmission.

[0042] In one embodiment, at the sending end, the semantic symbol sequence flows into the feature decoder to be trained in the sending end buffer update branch, and feature decoding is performed under the conditions of the sending end semantic context in the sending end decoding buffer to obtain a new sending end semantic context and store it in the sending end decoding buffer. The generator trained in the sender buffer update branch performs video frame reconstruction under the conditions of the sender feature context. In addition to outputting the sender-updated and reconstructed video frame, it also outputs new sender reference features and stores the new sender reference features in the sender decoding buffer. The feature decoder to be trained in the receiving end performs feature decoding under the condition of the receiving end semantic context in the receiving end decoding buffer, obtains a new receiving end semantic context, and stores it in the receiving end decoding buffer. The generator trained in the receiver reconstructs video frames under the conditions of the receiver feature context. In addition to outputting the updated and reconstructed video frames, it also outputs new receiver reference features and stores the new receiver reference features in the receiver decoding buffer.

[0043] In this embodiment, the feature decoder to be trained performs feature decoding on the semantic symbol sequence under the conditions of the sending end semantic context in the sending end decoding buffer, obtains a new sending end semantic context, and stores the new sending end semantic context in the sending end decoding buffer. After the sending end completes the generation of the semantic symbol sequence at the current time, it can synchronously update the semantic context state in its own buffer to provide a new temporal reference for semantic feature extraction, rate matching, and feature recovery at the next time step.

[0044] The trained generator in the sender buffer update branch performs video frame reconstruction under the sender feature context. Besides outputting the updated and reconstructed video frame, it also outputs new sender reference features and stores these new features in the sender decoding buffer. In other words, the sender buffer update branch not only handles local reconstruction but also reference feature updates, ensuring that the reference features used by the sender to subsequently extract context remain consistent with the current local reconstruction result.

[0045] The feature decoder in the receiver performs feature decoding under the receiver semantic context in the receiver decoding buffer, obtaining a new receiver semantic context, and stores the new receiver semantic context in the receiver decoding buffer. Then, the trained generator in the receiver performs video frame reconstruction under the receiver feature context. In addition to outputting the reconstructed video frame updated by the receiver, it also outputs new receiver reference features, which are stored in the receiver decoding buffer. Thus, while completing the reconstruction of the video frame at the current moment, the receiver can simultaneously update the semantic context and reference features in its own buffer, thereby providing a new temporal reference for the reconstruction of the video frame at the next moment.

[0046] In one embodiment, after obtaining the first video semantic communication system, the method further includes: With the goal of pruning the unrestricted rate set, the first semantic communication system is fine-tuned using sample video frames until the unrestricted rate set is pruned to contain K rate levels. At the transmitting end, the trained feature decoder, the second spatial prior module, and the trained generator, which are sequentially connected, are removed from the output of the trained feature encoder at the transmitting end to obtain the second video semantic communication system.

[0047] In this embodiment, after obtaining the first video semantic communication system, with the goal of pruning the unrestricted rate set, the first video semantic communication system is fine-tuned using sample video frames until the unrestricted rate set is pruned to contain K rate levels.

[0048] Specifically, in the aforementioned training process, the unrestricted rate set is used to cover a finer-grained bitrate space, enabling the system to learn feature mapping and recovery capabilities under different symbol lengths over a wider rate range. After the model reaches the preset convergence condition, the rate levels used less frequently during training can be gradually pruned, retaining only K representative rate levels. On this basis, fine-tuning can continue using sample video frames, so that the system can still maintain good video reconstruction and transmission performance even with a reduced number of rate levels.

[0049] After completing the above fine-tuning and pruning the unrestricted rate set to include K rate levels, the trained feature decoder, the second spatial prior module, and the trained generator, which are sequentially connected, can be removed from the output of the trained feature encoder at the transmitting end to obtain the second video semantic communication system.

[0050] In this embodiment, the sender-side local recovery link, set up during the aforementioned training phase to achieve sender-side reference alignment and local buffer updates, can be removed when forming a system for actual deployment. This is because this link is primarily used to reproduce the receiver's feature recovery and reference update process during training to ensure consistency between the sender and receiver reference states. After completing the aforementioned training and fine-tuning, this link has fulfilled its training assistance role and does not need to be retained during actual transmission. Therefore, the sender-side structure can be further simplified, reducing the computational overhead during system inference, making the resulting second video semantic communication system more suitable for practical deployment.

[0051] This embodiment, based on the first video semantic communication system, prunes and fine-tunes the unrestricted rate set, retaining only K representative rate levels. This helps to reduce rate search complexity, decrease redundant rate embeddings and corresponding mapping layers, and reduce side information overhead while maintaining multi-rate adaptive capabilities. Furthermore, by removing the local recovery link used only for training in the transmitter, the transmitter structure can be simplified, reducing computational complexity and overhead during actual deployment.

[0052] In one embodiment, a schematic diagram of the overall architecture of a video semantic communication system is provided, such as... Figure 4As shown: At the transmitting end, the transmitting end reference features are first read from the transmitting end decoding buffer, and the transmitting end super-prior context and transmitting end feature context are obtained through the context extractor. Then, semantic features are extracted based on the sample video frames. Subsequently, on the one hand, super-prior side information is generated based on the semantic features through super-prior encoding, quantization, arithmetic encoding, etc., and symbol length factor side information is obtained by combining super-prior decoding, spatial prior and rate matching. On the other hand, the semantic features are mapped into a semantic symbol sequence through the feature encoder. All three are sent to the receiving end through the wireless channel. At the same time, the semantic symbol sequence also flows into the transmitting end buffer update branch, and the transmitting end updated and reconstructed video frame is obtained through the feature decoder and generator, and the transmitting end semantic context and transmitting end reference features are updated. After receiving the semantic symbol sequence, prior information, and symbol length factor information, the receiver first extracts the receiver's prior context and feature context from the receiver's decoding buffer. Then, it completes the video frame reconstruction through the prior decoder, spatial prior module, feature decoder, and generator to obtain the reconstructed video frame. At the same time, it updates the receiver's semantic context and receiver's reference features to provide temporal reference for the reconstruction of subsequent video frames.

[0053] Based on the same inventive concept, a second aspect of the embodiments of this application provides a video semantic communication device, such as... Figure 5 As shown, the device includes: The first deployment module 201 is used to deploy the following multiple modules at the sending end to obtain an ideal receiving branch: a feature extractor to be trained with embedded learnable feature extraction vectors, a context extractor to be trained with embedded learnable context extraction vectors, a super prior encoder to be trained, a super prior decoder to be trained, a spatial prior module to be trained, and a generator to be trained with embedded learnable generation vectors. The training module 202 is used to train the ideal receiving branch using sample video frames to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector. The trained feature extraction vector is used to modulate the semantic features extracted by the trained feature extractor, the trained context extraction vector is used to modulate the context extracted by the trained context extractor, and the trained generation vector is used to modulate the reconstructed features generated by the trained generator. The second deployment module 203 is used to deploy the feature encoder to be trained at the sending end and the feature decoder to be trained at the sending end buffer update branch and the receiving end based on the ideal receiving branch that has been trained, so as to obtain the semantic communication system to be trained. The parameter update module 204 is used to update the parameters of the feature encoder and the feature decoder to be trained using the sample video frames to obtain the first video semantic communication system.

[0054] Optionally, the step of training the ideal receiving branch using sample video frames to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector, wherein the training module 202 includes: The first context extraction submodule is used to extract the context of the sending end reference features in the sending end decoding buffer through the context extractor to be trained, so as to obtain the sending end super-prior context and the sending end feature context. The first feature extraction submodule is used to extract features from sample video frames under the feature context of the sending end through the feature extractor to be trained, so as to obtain semantic features. The first determining submodule is used to obtain ideal received semantic features based on the semantic features and the sending end super-prior context, by sequentially using the super-prior encoder to be trained, the super-prior decoder to be trained, and the spatial prior module to be trained. The first video frame reconstruction submodule is used to reconstruct video frames using the ideal received semantic features under the condition of the sending end feature context through the generator to be trained, so as to obtain an ideal reconstructed video frame. The vector update submodule is used to update the parameters of each module in the ideal receiving branch and each learnable vector based at least on the difference between the sample video frame and the ideal reconstructed video frame, so as to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector.

[0055] Optionally, based on the ideal receiving branch after training, the feature encoder to be trained is deployed at the sending end, and the feature decoder to be trained is deployed at the sending end buffer update branch and the receiving end to obtain the semantic communication system to be trained. The second deployment module 203 includes: The first concatenated submodule is used to sequentially concatenate the trained context extractor, the trained feature extractor, the trained super-prior encoder, the trained super-prior decoder, the first spatial prior module, and the feature encoder to be trained at the transmitting end. The second concatenated submodule is used to sequentially connect the feature decoder to be trained, the second spatial prior module, and the trained generator at the output of the feature encoder to be trained, to form the sending buffer update branch; the first spatial prior module and the second spatial prior module are both the trained spatial prior modules. The first deployment submodule is used to deploy the feature decoder to be trained, the trained super-prior decoder, the trained spatial prior module, the trained generator, and the trained context extractor at the receiving end.

[0056] Optionally, the parameter update module 204, which updates the parameters of the feature encoder and the feature decoder to be trained using the sample video frames to obtain the first video semantic communication system, includes at least: The second context extraction submodule is used at the sending end to extract the context of the sending end reference features in the sending end decoding buffer through the trained context extractor, so as to obtain the sending end super-prior context and the sending end feature context. The second feature extraction submodule is used at the sending end to extract features from sample video frames under the feature context of the sending end through the trained feature extractor to obtain semantic features. The second determining submodule is used at the sending end to obtain the symbol length factor by sequentially using the trained super-prior encoder, the trained super-prior decoder, and the first spatial prior module based on the semantic features; wherein, at the sending end, the trained super-prior decoder is used to perform super-prior decoding under the conditions of the super-prior context at the sending end. The rate matching and symbol mapping submodule is used at the transmitting end to perform rate matching and symbol mapping under an unrestricted rate set, based on the symbol length factor, using the feature encoder to be trained and the transmitting end semantic context in the transmitting end decoding buffer, to obtain a semantic symbol sequence. The third determining submodule is used to receive the semantic symbol sequence into the sending buffer update branch to obtain the sending updated and reconstructed video frame; The first parameter update submodule is used to update the parameters of the feature encoder to be trained deployed at the transmitting end based at least on the difference between the sample video frame and the updated and reconstructed video frame at the transmitting end, and to update the parameters of the feature decoder to be trained deployed in the buffer update branch at the transmitting end and the receiving end.

[0057] Optionally, the parameter update module 204, which updates the parameters of the feature encoder and the feature decoder to be trained using the sample video frames to obtain the first video semantic communication system, further includes: The first conversion and transmission submodule is used at the transmitting end to obtain the quantized super-prior latent vector by sequentially using the trained super-prior encoder and quantizer based on the semantic features, convert it into super-prior side information, and send it to the receiving end. The second conversion and transmission submodule is used at the transmitting end to convert the symbol length factor into symbol length factor side information and send it to the receiving end; A sending submodule is used by the sending end to send the semantic symbol sequence to the receiving end; The third context extraction submodule is used at the receiving end to extract the context of the receiving end reference features in the receiving end decoding buffer through the trained context extractor, so as to obtain the receiving end prior context and the receiving end feature context. The third determining submodule is used to obtain the reconstructed video frame at the receiving end using the trained modules and the feature decoder to be trained at the receiving end; wherein, at the receiving end, the feature decoder to be trained performs rate matching and symbol mapping on the received semantic symbol sequence under the unrestricted rate set, based on the symbol length factor corresponding to the received symbol length factor side information, under the receiving end semantic context in the decoding buffer at the receiving end; at the receiving end, the trained super-prior decoder performs super-prior decoding under the super-prior context at the receiving end; at the receiving end, the feature decoder to be trained performs video frame reconstruction under the feature context at the receiving end. The second parameter update submodule is used to update the parameters of the feature encoder to be trained deployed at the sending end, in addition to updating the parameters of the sample video frame and the reconstructed video frame at the sending end based on the difference between the sample video frame and the reconstructed video frame at the receiving end, and to update the parameters of the feature decoder to be trained deployed at the sending end buffer update branch and the receiving end.

[0058] Optionally, the system further includes: The fourth determining submodule is used to obtain the code rate loss based on the semantic symbol sequence, the prior information, and the symbol length factor information. The third parameter update submodule is used to update the parameters of the feature encoder to be trained deployed at the sending end based on the difference between the sample video frame and the reconstructed video frame at the sending end, and the difference between the sample video frame and the reconstructed video frame at the receiving end, as well as the parameters of the feature decoder to be trained deployed at the receiving end based on the bitrate loss.

[0059] Optionally, the system further includes: The first feature decoding submodule is used at the sending end, where the semantic symbol sequence flows into the sending end buffer update branch of the feature decoder to be trained, and performs feature decoding under the conditions of the sending end semantic context in the sending end decoding buffer to obtain a new sending end semantic context and store it in the sending end decoding buffer. The second video frame reconstruction submodule is used by the generator that has been trained in the sending end buffer update branch to reconstruct video frames under the conditions of the sending end feature context. In addition to outputting the sending end updated and reconstructed video frames, it also outputs new sending end reference features and stores the new sending end reference features into the sending end decoding buffer. The second feature decoding submodule is used for the feature decoder to be trained in the receiving end to perform feature decoding under the condition of the receiving end semantic context in the receiving end decoding buffer, to obtain a new receiving end semantic context and store it in the receiving end decoding buffer. The third video frame reconstruction submodule is used by the generator trained in the receiving end to reconstruct video frames under the condition of the receiving end feature context. In addition to outputting the updated and reconstructed video frames of the receiving end, it also outputs new receiving end reference features and stores the new receiving end reference features in the receiving end decoding buffer.

[0060] Optionally, the system further includes: The fine-tuning submodule is used to fine-tune the first semantic communication system using sample video frames with the goal of pruning the unrestricted rate set until the unrestricted rate set is pruned to contain K rate levels. A removal submodule is used to remove the sequentially connected trained feature decoder, second spatial prior module, and trained generator from the output of the trained feature encoder at the transmitting end, thereby obtaining a second video semantic communication system.

[0061] Each embodiment in this specification focuses on the differences from other embodiments. For the same or similar parts between the embodiments, please refer to each other.

[0062] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0063] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0064] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0066] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0067] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0068] The video semantic communication system and apparatus provided above have been described in detail. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A video semantic communication system, characterized by, include: The following modules are deployed at the transmitting end to obtain the ideal receiving branch: a trainable feature extractor embedded with learnable feature extraction vectors, a trainable context extractor embedded with learnable context extraction vectors, a trainable super-prior encoder, a trainable super-prior decoder, a trainable spatial prior module, and a trainable generator embedded with learnable generation vectors. The ideal receiving branch is trained using sample video frames to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector. The trained feature extraction vector is used to modulate the semantic features extracted by the trained feature extractor, the trained context extraction vector is used to modulate the context extracted by the trained context extractor, and the trained generation vector is used to modulate the reconstructed features generated by the trained generator. Based on the ideal receiving branch that has been trained, the feature encoder to be trained is deployed at the sending end, and the feature decoder to be trained is deployed at the sending end buffer update branch and the receiving end to obtain the semantic communication system to be trained. The parameters of the feature encoder and the feature decoder to be trained are updated using the sample video frames to obtain the first video semantic communication system.

2. The video semantic communication system of claim 1, wherein, The step of training the ideal receiving branch using sample video frames to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector includes: The context extractor to be trained extracts the context of the sending reference features in the sending decoding buffer to obtain the sending super-prior context and the sending feature context. Using the feature extractor to be trained, semantic features are obtained by extracting features from sample video frames under the feature context of the sending end. Based on the semantic features and the sending end super-prior context, the ideal receiving semantic features are obtained by sequentially using the super-prior encoder to be trained, the super-prior decoder to be trained, and the spatial prior module to be trained. The ideal reconstructed video frame is obtained by using the ideal receiving semantic features under the condition of the sending end feature context through the generator to be trained. Based at least on the difference between the sample video frame and the ideal reconstructed video frame, the parameters of each module in the ideal receiving branch and each learnable vector are updated to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector.

3. The video semantic communication system of claim 1, wherein, Based on the ideal receiving branch that has been trained, the feature encoder to be trained is deployed at the sending end, and the feature decoder to be trained is deployed at the sending end buffer update branch and the receiving end, thus obtaining the semantic communication system to be trained, including: At the transmitting end, the trained context extractor, the trained feature extractor, the trained super prior encoder, the trained super prior decoder, the first spatial prior module, and the feature encoder to be trained are connected in series in sequence. At the output of the feature encoder to be trained, the feature decoder to be trained, the second spatial prior module, and the trained generator are connected in series to form the sending buffer update branch; the first spatial prior module and the second spatial prior module are both the trained spatial prior modules. At the receiving end, the feature decoder to be trained, the trained super-prior decoder, the trained spatial prior module, the trained generator, and the trained context extractor are deployed.

4. The video semantic communication system of claim 3, wherein, The method of updating the parameters of the feature encoder and the feature decoder to be trained using the sample video frames to obtain the first video semantic communication system includes at least the following: At the transmitting end, the trained context extractor extracts the context of the transmitting end reference features in the transmitting end decoding buffer to obtain the transmitting end super-prior context and the transmitting end feature context. At the transmitting end, the trained feature extractor extracts features from the sample video frames under the feature context of the transmitting end to obtain semantic features; At the transmitting end, based on the semantic features, the trained super-prior encoder, the trained super-prior decoder, and the first spatial prior module are used sequentially to obtain the symbol length factor; wherein, at the transmitting end, the trained super-prior decoder is used to perform super-prior decoding under the conditions of the super-prior context at the transmitting end. At the transmitting end, using the feature encoder to be trained, under the condition of the transmitting end semantic context in the transmitting end decoding buffer, rate matching and symbol mapping are performed according to the symbol length factor under an unrestricted rate set to obtain a semantic symbol sequence; The semantic symbol sequence flows into the sending buffer update branch to obtain the sending-end updated and reconstructed video frame; At least based on the difference between the sample video frame and the video frame reconstructed by the sending end, the parameters of the feature encoder to be trained deployed at the sending end are updated, and the parameters of the feature decoder to be trained deployed in the sending end buffer update branch and the receiving end are updated.

5. The video semantic communication system of claim 4, wherein, The method of updating the parameters of the feature encoder and the feature decoder to be trained using the sample video frames to obtain the first video semantic communication system further includes: At the transmitting end, based on the semantic features, the trained super-prior encoder and quantizer are used sequentially to obtain the quantized super-prior latent vector, which is then converted into super-prior side information and sent to the receiving end. At the transmitting end, the symbol length factor is converted into symbol length factor side information and sent to the receiving end; The sending end sends the semantic symbol sequence to the receiving end; At the receiving end, the trained context extractor extracts the context of the receiving end reference features in the receiving end decoding buffer to obtain the receiving end prior context and the receiving end feature context. Using the trained modules and the feature decoder to be trained in the receiving end, the reconstructed video frame is obtained; wherein, in the receiving end, the feature decoder to be trained performs rate matching and symbol mapping on the received semantic symbol sequence under the unrestricted rate set, based on the receiving end semantic context in the receiving end decoding buffer and the symbol length factor corresponding to the received symbol length factor side information; in the receiving end, the trained super-prior decoder performs super-prior decoding under the receiving end super-prior context; in the receiving end, the feature decoder to be trained performs video frame reconstruction under the receiving end feature context. In addition to updating the reconstructed video frame based on the difference between the sample video frame and the reconstructed video frame at the sending end, the parameters of the feature encoder to be trained deployed at the sending end are also updated based on the difference between the sample video frame and the reconstructed video frame at the receiving end. Furthermore, the parameters of the feature decoder to be trained deployed at the sending end and the buffer update branch at the sending end and the receiving end are also updated.

6. The video semantic communication system according to claim 5, characterized in that, The code rate loss is obtained based on the semantic symbol sequence, the prior information, and the symbol length factor information. In addition to updating the reconstructed video frame based on the difference between the sample video frame and the reconstructed video frame at the transmitting end, and the difference between the sample video frame and the reconstructed video frame at the receiving end, the parameters of the feature encoder to be trained deployed at the transmitting end are also updated based on the bitrate loss, as are the parameters of the feature decoder to be trained deployed in the updating branch of the transmitting end buffer and the receiving end.

7. The video semantic communication system according to claim 5, characterized in that, At the sending end, the semantic symbol sequence flows into the feature decoder to be trained in the sending end buffer update branch. Under the conditions of the sending end semantic context in the sending end decoding buffer, feature decoding is performed to obtain a new sending end semantic context and store it in the sending end decoding buffer. The generator trained in the sender buffer update branch performs video frame reconstruction under the conditions of the sender feature context. In addition to outputting the sender-updated and reconstructed video frame, it also outputs new sender reference features and stores the new sender reference features in the sender decoding buffer.

8. The video semantic communication system according to claim 7, characterized in that, The feature decoder to be trained in the receiving end performs feature decoding under the condition of the receiving end semantic context in the receiving end decoding buffer, obtains a new receiving end semantic context, and stores it in the receiving end decoding buffer. The generator trained in the receiver reconstructs video frames under the conditions of the receiver feature context. In addition to outputting the updated and reconstructed video frames, it also outputs new receiver reference features and stores the new receiver reference features in the receiver decoding buffer.

9. The video semantic communication system of claim 5, wherein, After obtaining the first video semantic communication system With the goal of pruning the unrestricted rate set, the first semantic communication system is fine-tuned using sample video frames until the unrestricted rate set is pruned to contain K rate levels. At the transmitting end, the trained feature decoder, the second spatial prior module, and the trained generator, which are sequentially connected, are removed from the output of the trained feature encoder at the transmitting end to obtain the second video semantic communication system.

10. A video semantic communication apparatus, characterized by, The device includes: The first deployment module is used to deploy the following modules at the sending end to obtain the ideal receiving branch: a feature extractor to be trained with embedded learnable feature extraction vectors, a context extractor to be trained with embedded learnable context extraction vectors, a super prior encoder to be trained, a super prior decoder to be trained, a spatial prior module to be trained, and a generator to be trained with embedded learnable generation vectors. The training module is used to train the ideal receiving branch using sample video frames to obtain the trained ideal receiving branch, the trained feature extraction vector, the trained context extraction vector, and the trained generation vector. The trained feature extraction vector is used to modulate the semantic features extracted by the trained feature extractor, the trained context extraction vector is used to modulate the context extracted by the trained context extractor, and the trained generation vector is used to modulate the reconstructed features generated by the trained generator. The second deployment module is used to deploy the feature encoder to be trained at the sending end and the feature decoder to be trained at the sending end buffer update branch and the receiving end based on the ideal receiving branch that has been trained, so as to obtain the semantic communication system to be trained. The parameter update module is used to update the parameters of the feature encoder and the feature decoder to be trained using the sample video frames, so as to obtain the first video semantic communication system.