Semantic video transmission method and apparatus, electronic device, and storage medium

CN122845804APending Publication Date: 2026-09-29BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610912574.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]然而,上述技术方案存在以下技术问题:现有技术多从单一环节分别优化,任务意图与紧凑语义参数之间的转化关系不明确,导致语义参数与任务目标难以对应验证;此外,关键帧调度、稀疏监督等各子阶段缺乏统一约束,导致时序一致性差、评测标准碎片化,端到端质量难以客观衡量

Benefits of technology

[0012]本公开提供了语义视频传输方法、装置、设备以及存储介质,本申请通过获取原始视频帧并进行预处理得到规范帧序列,为后续意图掩码生成与低秩反演提供统一像素坐标系的数据基础。接着,获取与视频处理任务对应的任务文本,根据该任务文本生成与规范帧序列空间对齐的帧级意图掩码,通过任务文本与视频内容的空间对齐,确保"关注哪里、完成何种任务"的语义信息能够准确映射到像素级区域,有效减少因意图区域定位偏差导致的语义传输失真。进一步地,从规范帧序列中提取关键帧并记录关键帧索引,针对每个关键帧初始化低秩提示参数,将低秩提示参数输入扩散生成模型生成重建结果,根据重建结果与关键帧在帧级意图掩码加权下的差异迭代优化低秩提示参数,使低秩提示参数在任务相关区域获得更高优化权重,突破了均匀优化导致任务语义被背景信息稀释的瓶颈,使有限参数资源优先保留任务关键视觉语义。最终将各关键帧优化后的低秩提示参数、帧级意图掩码及关键帧索引封装为紧凑语义参数包输出至接收端,实现了以低秩提示参数为语义载体、以意图掩码为区域引导、以关键帧索引为时序锚定的紧凑传输格式,显著降低传输码率与存储占用的同时,使接收端能够基于同一解析约定完成关键帧重建与中间时刻插值补全,在算力受限条件下仍能获得与规范帧序列时序对齐的重建视频。综上,本申请通过帧级意图掩码引导实现任务语义优先保留、关键帧稀疏反演实现传输参数显著压缩,以及紧凑语义参数包实现收发两端解析约定统一,在有限码率与算力约束下实现了任务意图、关注区域与像素坐标的端到端一致表达,降低了接收端整体算力与时延开销,提升了重建视频在几何与时间参数上的可复现性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845804A_ABST
    Figure CN122845804A_ABST
Patent Text Reader

Abstract

The present disclosure provides a semantic video transmission method and device, electronic equipment and storage medium, and relates to the technical field of semantic communication. The method comprises: obtaining an original video frame and preprocessing to obtain a standard frame sequence; obtaining a task text corresponding to a video processing task, and generating a frame-level intent mask spatially aligned with the standard frame sequence according to the task text; extracting a key frame from the standard frame sequence and recording a key frame index, initializing a low-rank prompt parameter for each key frame, inputting the low-rank prompt parameter into a diffusion generation model to generate a reconstruction result, iteratively optimizing the low-rank prompt parameter according to the difference between the reconstruction result and the key frame under the frame-level intent mask weighting, and obtaining an optimized low-rank prompt parameter; encapsulating the optimized low-rank prompt parameter of each key frame, the frame-level intent mask and the key frame index into a compact semantic parameter package, and outputting to a receiving end. The present disclosure improves the reproducibility of the reconstructed video in geometric and temporal parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of communication technology, specifically to the field of semantic communication technology, and in particular to semantic video transmission methods, apparatus, electronic devices, and storage media. Background Technology

[0002] With the continuous growth of video and streaming media traffic, the constraints of wireless access and terminal-side computing power and power consumption are becoming increasingly prominent. Traditional transmission methods based on dense pixel-based bitstreams struggle to maintain acceptable subjective quality under low bitrate conditions. Generative semantic video communication has gained attention as a solution. Existing technologies have proposed frameworks that combine multimodal semantic decomposition, channel adaptation, and pre-trained diffusion receiver synthesis. Other works combine low-rank cue streams with diffusion decoding to achieve cue streaming, or elevate video coding to the semantic layer and employ decoupled diffusion multi-frame compensation at the receiver.

[0003] However, the aforementioned technical solutions suffer from the following technical problems: existing technologies mostly optimize each individual stage separately, resulting in an unclear transformation relationship between task intent and compact semantic parameters, making it difficult to verify the correspondence between semantic parameters and task objectives; furthermore, the lack of unified constraints in various sub-stages such as keyframe scheduling and sparse supervision leads to poor temporal consistency, fragmented evaluation standards, and difficulty in objectively measuring end-to-end quality. Both lack a comprehensive approach to task intent, supervision signals, and unified evaluation constraints, resulting in insufficient closure and reproducibility of the end-to-end semantic video reconstruction chain. Summary of the Invention

[0004] This disclosure provides a semantic video transmission method, apparatus, electronic device, and storage medium.

[0005] According to one aspect of this disclosure, a semantic video transmission method is provided, executed by a sending end, the method comprising: The original video frames are acquired and preprocessed to obtain a standardized frame sequence. Obtain the task text corresponding to the video processing task, and generate a frame-level intent mask that is spatially aligned with the standard frame sequence based on the task text; Keyframes are extracted from the canonical frame sequence and keyframe indices are recorded. Low-rank cue parameters are initialized for each keyframe. The low-rank cue parameters are input into the diffusion generation model to generate reconstruction results. Based on the difference between the reconstruction results and the keyframes under the frame-level intent mask weighting, the low-rank cue parameters are iteratively optimized to obtain optimized low-rank cue parameters. The optimized low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index are encapsulated into a compact semantic parameter package and output to the receiving end. The compact semantic parameter package is used at the receiving end to reconstruct a video frame sequence that is time-aligned with the canonical frame sequence.

[0006] According to another aspect of this disclosure, a semantic video transmission method is provided, executed by a receiving end, the method comprising: The system receives a compact semantic parameter packet from a sending end. The sending end is used to acquire and preprocess raw video frames to obtain a canonical frame sequence. It acquires task text corresponding to the video processing task and generates a frame-level intent mask spatially aligned with the canonical frame sequence based on the task text. It extracts keyframes from the canonical frame sequence and records the keyframe indices. For each keyframe, it initializes low-rank cue parameters, inputs the low-rank cue parameters into a diffusion generation model to generate a reconstruction result, and iteratively optimizes the low-rank cue parameters based on the difference between the reconstruction result and the keyframe under the weighting of the frame-level intent mask to obtain optimized low-rank cue parameters. The optimized low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe indices are encapsulated into a compact semantic parameter packet. Parse the compact semantic parameter packet to obtain the low-rank cue parameters of each key frame, the frame-level intent mask, and the key frame index; Based on the low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index, video reconstruction is performed to obtain a reconstructed video frame sequence that is time-aligned with the normalized frame sequence.

[0007] According to a third aspect of this disclosure, a semantic video transmission apparatus is provided, comprising: The first acquisition module is used to acquire raw video frames and preprocess them to obtain a standardized frame sequence. The second acquisition module is used to acquire the task text corresponding to the video processing task and generate a frame-level intent mask that is aligned with the space of the standard frame sequence based on the task text. The optimization module is used to extract keyframes from the canonical frame sequence and record keyframe indices, initialize low-rank cue parameters for each keyframe, input the low-rank cue parameters into the diffusion generation model to generate reconstruction results, and iteratively optimize the low-rank cue parameters based on the difference between the reconstruction results and the keyframes under the frame-level intent mask weighting to obtain optimized low-rank cue parameters. The output module is used to encapsulate the optimized low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index into a compact semantic parameter package and output it to the receiving end. The compact semantic parameter package is used to reconstruct a video frame sequence that is time-aligned with the canonical frame sequence at the receiving end.

[0008] According to a fourth aspect of this disclosure, a semantic video transmission apparatus is provided, comprising: A receiving module is used to receive a compact semantic parameter packet from a sending end. The sending end is used to acquire raw video frames and preprocess them to obtain a canonical frame sequence; acquire task text corresponding to the video processing task, and generate a frame-level intent mask spatially aligned with the canonical frame sequence based on the task text; extract keyframes from the canonical frame sequence and record keyframe indices; initialize low-rank cue parameters for each keyframe; input the low-rank cue parameters into a diffusion generation model to generate a reconstruction result; iteratively optimize the low-rank cue parameters based on the difference between the reconstruction result and the keyframe under the weighting of the frame-level intent mask to obtain optimized low-rank cue parameters; and encapsulate the optimized low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe indices into a compact semantic parameter packet. The parsing module is used to parse the compact semantic parameter package to obtain the low-rank cue parameters of each key frame, the frame-level intent mask, and the key frame index. The reconstruction module is used to perform video reconstruction based on the low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index, to obtain a reconstructed video frame sequence that is time-aligned with the normalized frame sequence.

[0009] According to a fifth aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described in any of the above technical solutions.

[0010] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any one of the methods described above.

[0011] According to a seventh aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in any one of the above technical solutions.

[0012] This disclosure provides a semantic video transmission method, apparatus, device, and storage medium. This application obtains a standardized frame sequence by acquiring and preprocessing raw video frames, providing a unified pixel coordinate system as the data foundation for subsequent intent mask generation and low-rank inversion. Next, task text corresponding to the video processing task is acquired, and a frame-level intent mask spatially aligned with the standardized frame sequence is generated based on this task text. This spatial alignment of the task text and video content ensures that the semantic information of "where to focus and what task to complete" is accurately mapped to pixel-level regions, effectively reducing semantic transmission distortion caused by intent region positioning deviations. Furthermore, keyframes are extracted from the standardized frame sequence, and keyframe indices are recorded. Low-rank cue parameters are initialized for each keyframe, and these parameters are input into a diffusion generation model to generate reconstruction results. The low-rank cue parameters are iteratively optimized based on the difference between the reconstruction results and the keyframes under frame-level intent mask weighting, giving them higher optimization weights in task-related regions. This overcomes the bottleneck of uniform optimization causing task semantics to be diluted by background information, allowing limited parameter resources to prioritize the preservation of key visual semantics of the task. Finally, the optimized low-rank cue parameters, frame-level intent masks, and keyframe indices of each keyframe are encapsulated into a compact semantic parameter packet and output to the receiving end. This achieves a compact transmission format with low-rank cue parameters as the semantic carrier, intent masks as region guidance, and keyframe indices as temporal anchors. This significantly reduces the transmission bitrate and storage footprint, while enabling the receiving end to complete keyframe reconstruction and intermediate time-interpolation based on the same parsing convention. Even under computational constraints, it can still obtain reconstructed video that is temporally aligned with the standard frame sequence. In summary, this application achieves priority preservation of task semantics through frame-level intent mask guidance, significant compression of transmission parameters through keyframe sparse inversion, and unified parsing conventions at both the sending and receiving ends through compact semantic parameter packets. Under limited bitrate and computational constraints, it achieves end-to-end consistent expression of task intent, region of interest, and pixel coordinates, reducing the overall computational cost and latency at the receiving end and improving the reproducibility of the reconstructed video in terms of geometric and temporal parameters.

[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0014] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic diagram of the steps of a semantic video transmission method in one embodiment of the present disclosure; Figure 2 This is a schematic diagram of the steps of a semantic video transmission method in another embodiment of this disclosure; Figure 3This is a schematic diagram of the overall process of a semantic video transmission method in one embodiment of the present disclosure; Figure 4 This is a schematic block diagram of the semantic video transmission device in the embodiments of this disclosure; Figure 5 This is a schematic block diagram of the semantic video transmission device in the embodiments of this disclosure; Figure 6 This is a block diagram of an electronic device used to implement the semantic video transmission method of the embodiments of this disclosure. Detailed Implementation

[0015] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0016] This disclosure provides a semantic video transmission method, which is applied at the sending end; see [link to relevant documentation]. Figure 1 As shown, it includes: Step S101: Obtain the original video frames and perform preprocessing to obtain a standardized frame sequence.

[0017] Specifically, a raw video frame refers to a single frame of image data directly acquired or decoded from a video source without any preprocessing or compression. It contains complete pixel information (such as pixel values ​​of each channel in the RGB or YUV color space), original resolution, and frame rate timing information, serving as the initial input to the video processing flow. After acquiring the raw video frames, preprocessing is performed. For example, sampling rate conversion, aspect ratio-preserving scaling, and edge padding can be performed frame by frame to match the size of each frame with the network input size used for subsequent intent mask generation and latent space coding. At the same time, affine transformation parameters and padding region masks are recorded as geometric metadata. The processed frames are then arranged in temporal order to obtain a standardized frame sequence with pixel coordinates consistent with those in subsequent processing stages.

[0018] Step S102: Obtain the task text corresponding to the video processing task, and generate a frame-level intent mask that is aligned with the standard frame sequence space based on the task text.

[0019] Specifically, "task text corresponding to the video processing task" refers to natural language text describing the video's area of ​​interest and task objectives, used to clearly define the semantic instructions of "where to focus and what task to complete." In this application, the task text originates from explicit user input (such as user questions in intelligent agent video question answering), pre-set procedures of the business system (such as equipment tag numbers and operating procedures in industrial inspection), or task summaries (such as key point descriptions in field records). Generating a frame-level intent mask spatially aligned with the standard frame sequence based on the task text means: transforming the natural language task text describing the video's area of ​​interest and task objectives into a binary spatial mask that corresponds one-to-one with the pixel positions of each frame in the standard frame sequence through a multimodal intent understanding and segmentation model.

[0020] The specific implementation process of this scheme includes: encoding the task text into text features, dividing the canonical frame into image blocks and encoding them into image features, fusing the two through a visual language understanding network to obtain a segmentation feature vector indicating the target region, and simultaneously generating two-dimensional image features that preserve intra-frame positional relationships through a segmentation network. After matching the segmentation feature vector with the two-dimensional image features, a continuous mask response map is generated. This map is then restored to the canonical frame size and thresholded to obtain a frame-level intent mask, ensuring that the marked regions in the mask precisely correspond to the pixel positions pointed to by the task text. This process is sequentially performed on each canonical frame to obtain a mask sequence arranged in temporal order. This scheme achieves a precise mapping of task intent from the linguistic space to the pixel space, providing a basis for locating task-related regions for subsequent low-rank inversion optimization, enabling the optimization process to prioritize the preservation of visual semantics related to the task objective.

[0021] Step S103: Extract keyframes from the canonical frame sequence and record the keyframe indices. Initialize low-rank cue parameters for each keyframe. Input the low-rank cue parameters into the diffusion generation model to generate reconstruction results. Based on the difference between the reconstruction results and the keyframes under frame-level intent mask weighting, iteratively optimize the low-rank cue parameters to obtain the optimized low-rank cue parameters.

[0022] Specifically, a "keyframe" refers to a representative frame extracted from the canonical frame sequence at preset intervals, used to carry the main visual content and semantic information of the video. A "keyframe index" refers to the position number of each keyframe in the canonical frame sequence, used for subsequent temporal localization and alignment with the receiver's reconstruction. A "low-rank cue parameter" refers to a semantic condition parameter, composed of the product of two low-rank matrices, with a parameter count much smaller than that of dense pixel data, used to drive the diffusion generation model to reconstruct the keyframe. A "diffusion generation model" refers to a generative neural network based on a diffusion probability model, capable of progressively denoising noise to generate an image matching the conditional input. A "reconstruction result" refers to the predicted frame corresponding to the keyframe generated by the diffusion generation model based on the low-rank cue parameters. "Frame-level intent mask weighted difference" refers to using a frame-level intent mask to increase the error weight of task-related regions when calculating the pixel difference between the reconstructed result and the keyframe, making the error in the task region contribute more to the overall loss.

[0023] The specific implementation process of this scheme includes: extracting representative frames as keyframes from the standardized frame sequence at preset intervals, and recording the position number of each keyframe in the standardized frame sequence. For each keyframe, an initial optimizable low-rank cue parameter is set. This parameter is composed of the product of two low-rank matrices, and its parameter size is much smaller than that of dense pixel data. The low-rank cue parameter is decomposed into two low-rank matrices and multiplied to obtain a condition matrix. This condition matrix is ​​input into a diffusion generation model, which generates a reconstruction result corresponding to the keyframe. The pixel differences between the reconstruction result and the keyframe are calculated. Frame-level intent masks are used to increase the difference weight of the masked region, making the error in the task-related region contribute more to the overall loss, resulting in a weighted difference. Based on this weighted difference, the low-rank cue parameter is adjusted backward using gradient descent. This process of generating reconstruction results, calculating weighted differences, and adjusting parameters backward is repeated until the loss is below a threshold or the maximum number of iterations is reached. Ultimately, the low-rank cue parameter prioritizes the retention of semantic information in the task-related region, resulting in the optimized low-rank cue parameter. This scheme reduces the scale of inversion computation by sparse extraction of keyframes, achieves controllable compression of parameters by low-rank matrix factorization, guides the priority preservation of task semantics by intention mask weighting, and ensures optimization stability by iterative convergence, thus achieving efficient generation of task alignment semantic parameters under the constraint of limited computing power.

[0024] Step S104: The optimized low-rank cue parameters, frame-level intent mask and keyframe index of each keyframe are encapsulated into a compact semantic parameter package and output to the receiving end. The compact semantic parameter package is used to reconstruct the video frame sequence that is time-aligned with the canonical frame sequence at the receiving end.

[0025] Specifically, a "compact semantic parameter packet" refers to a structured data transmission unit formed by the sender after quantizing, compressing, and encapsulating low-rank cue parameters, frame-level intent masks, and keyframe indexes using unified fields. Its data volume is much smaller than the original video pixel stream. "Temporal alignment" means that the temporal order of the video reconstructed by the receiver is completely consistent with the standardized frame sequence of the sender.

[0026] The specific implementation process of this scheme includes: First, octet-specific point quantization and serialization are performed on the optimized low-rank cue parameters and associated scale information of each keyframe to obtain a cue parameter file; lossless compression is then performed on the frame-level intent mask sequence to obtain compressed mask data. Second, the cue parameter file, compressed mask data, keyframe index, and metadata required for parsing are written into a data structure according to a unified field convention to obtain a compact semantic parameter packet. Finally, the compact semantic parameter packet is output to the receiving end. After parsing, the receiving end reconstructs the keyframes based on the low-rank cue parameters, locates the timing based on the keyframe index, maintains the semantic consistency of the task region based on the frame-level intent mask, and completes the intermediate frames through interpolation to obtain a reconstructed video frame sequence aligned with the timing of the standard frame sequence. In this way, efficient transmission of low-rank semantic parameters is achieved through quantization compression and unified encapsulation, and the closed-loop parsing convention ensures the consistency of the reconstruction results at both the sending and receiving ends.

[0027] This disclosure provides a semantic video transmission method, apparatus, device, and storage medium. This application obtains a standardized frame sequence by acquiring and preprocessing raw video frames, providing a unified pixel coordinate system as the data foundation for subsequent intent mask generation and low-rank inversion. Next, task text corresponding to the video processing task is acquired, and a frame-level intent mask spatially aligned with the standardized frame sequence is generated based on this task text. This spatial alignment of the task text and video content ensures that the semantic information of "where to focus and what task to complete" is accurately mapped to pixel-level regions, effectively reducing semantic transmission distortion caused by intent region positioning deviations. Furthermore, keyframes are extracted from the standardized frame sequence, and keyframe indices are recorded. Low-rank cue parameters are initialized for each keyframe, and these parameters are input into a diffusion generation model to generate reconstruction results. The low-rank cue parameters are iteratively optimized based on the difference between the reconstruction results and the keyframes under frame-level intent mask weighting, giving them higher optimization weights in task-related regions. This overcomes the bottleneck of uniform optimization causing task semantics to be diluted by background information, allowing limited parameter resources to prioritize the preservation of key visual semantics of the task. Finally, the optimized low-rank cue parameters, frame-level intent masks, and keyframe indices of each keyframe are encapsulated into a compact semantic parameter packet and output to the receiving end. This achieves a compact transmission format with low-rank cue parameters as the semantic carrier, intent masks as region guidance, and keyframe indices as temporal anchors. This significantly reduces the transmission bitrate and storage footprint, while enabling the receiving end to complete keyframe reconstruction and intermediate time-interpolation based on the same parsing convention. Even under computational constraints, it can still obtain reconstructed video that is temporally aligned with the standard frame sequence. In summary, this application achieves priority preservation of task semantics through frame-level intent mask guidance, significant compression of transmission parameters through keyframe sparse inversion, and unified parsing conventions at both the sending and receiving ends through compact semantic parameter packets. Under limited bitrate and computational constraints, it achieves end-to-end consistent expression of task intent, region of interest, and pixel coordinates, reducing the overall computational cost and latency at the receiving end and improving the reproducibility of the reconstructed video in terms of geometric and temporal parameters.

[0028] In some optional embodiments, the original video frames are acquired and preprocessed to obtain a standardized frame sequence, including: The original video is sampled frame by frame to obtain sampled frames that match the target resolution; The sampled frame is scaled while preserving its aspect ratio to obtain a scaled frame; Edge padding is applied to the scaled frame to obtain a normalized frame; Arrange the standard frames in chronological order to obtain the standard frame sequence.

[0029] Specifically, "sample rate conversion" refers to resampling each frame of the original video according to the target frame rate, adjusting the temporal density of the frames to ensure the frame rate matches the requirements of subsequent processing networks. "Target resolution" refers to the frame rate specification of the network input used for subsequent intent mask generation and latent space coding. "Aspect ratio-preserving scaling" refers to scaling the sampled frames to a preset size while maintaining the original width and height ratio of the image, preventing image stretching or compression distortion. "Scaled frames" are intermediate frames that have been scaled proportionally, maintaining the original spatial relationships but not perfectly matching the network input in size. "Edge padding" refers to adding pixel values ​​(such as black or mirrored pixels) to the edges of scaled frames to ensure the overall size of the frame perfectly matches the size of subsequent network input. "Canonical frames" are standard frames that, after sample rate conversion, aspect ratio-preserving scaling, and edge padding, have the same size as the network input and retain the original temporal relationships. "Canonical frame sequence" refers to organizing canonical frames into a continuous sequence according to the original temporal order, ensuring the preservation of video temporal continuity.

[0030] The specific implementation process of this scheme includes: First, the sampling rate of the original video is converted frame by frame to obtain sampling frames that match the target resolution. For example, the original 60 frames per second video is downsampled to 30 frames per second to make the frame rate consistent with the input frame rate of the subsequent diffusion generation model, avoiding timing distortion or computational redundancy caused by frame rate mismatch. Second, the sampling frames are scaled while maintaining the aspect ratio to obtain scaled frames. For example, a 1920×1080 sampling frame is scaled to 512×288 with an aspect ratio of 16:9 to ensure that circular objects in the image remain circular and square objects remain square, preventing target shape distortion from affecting the accurate positioning of the subsequent intent mask. Then, the scaled frames are edge-padded to obtain normalized frames. For example, a 512×288 scaled frame is padded with 112-pixel black borders above and below, resulting in a 512×512 canonical frame. This ensures the frame's size perfectly matches the input size of the subsequent image encoder (512×512). Simultaneously, the affine transformation parameters (scaling ratio 0.267, padding offset 112) and the padding region mask are recorded as geometric metadata. Finally, the canonical frames are arranged chronologically to obtain a canonical frame sequence. For instance, performing this process sequentially on a 10-second original video (600 frames) yields 300 512×512 canonical frames, which are then organized into a canonical frame sequence according to the original chronological order. This provides a unified pixel coordinate system for subsequent keyframe extraction, low-rank inversion optimization, and receiver-side temporal reconstruction. Thus, frame rate consistency is ensured through sampling rate conversion, target spatial relationships are maintained through aspect ratio-preserving scaling, network input size matching is achieved through edge padding, and video continuity is preserved through temporal arrangement, providing a geometric normalization foundation for end-to-end semantic video transmission links.

[0031] In this way, by performing sampling rate conversion frame by frame on the original video, sampling frames matching the target resolution are obtained, ensuring that the frame rate is consistent with the requirements of subsequent processing networks and avoiding temporal distortion caused by frame rate mismatch. The sampling frames are then scaled while maintaining aspect ratio to obtain scaled frames, adapting to the network input size while preserving the original image proportions, preventing image distortion from disrupting semantic spatial relationships. Edge padding is applied to the scaled frames to obtain normalized frames, ensuring that the frame size perfectly matches the subsequent network input. Simultaneously, affine transformation parameters and the padding region mask are recorded as geometric metadata, providing a basis for cross-stage alignment of the intent mask and pixel coordinates. Finally, the normalized frames are arranged temporally to obtain a normalized frame sequence, ensuring that the video temporal relationships are preserved. In summary, the above preprocessing steps ensure that the video frames, subsequent intent mask generation, low-rank inversion, and receiver reconstruction are in a unified pixel coordinate system, reducing cross-stage semantic bias, improving the spatial alignment accuracy of the intent mask and reconstruction results, and providing a geometric basis for the closure and reproducibility of the end-to-end semantic video transmission link.

[0032] In some optional embodiments, generating a frame-level intent mask spatially aligned with the canonical frame sequence based on the task text includes: For each specification frame, the task text is encoded into text features, and the current specification frame is divided into image blocks and encoded into image features; The text features and image features are fused to obtain a segmentation feature vector. The current canonical frame is then encoded with position preservation to obtain two-dimensional image features. The segmentation feature vector is matched with the two-dimensional image features to generate a continuous mask response map; The continuous mask response map is restored to the current standard frame size, and then the frame-level intent mask is obtained after thresholding. Arrange the frame-level intent masks corresponding to each standard frame in time sequence to obtain a frame-level intent mask sequence that is spatially aligned with the standard frame sequence.

[0033] Specifically, "text features" refers to converting the natural language task text describing the video's regions of interest and task objectives into a high-dimensional vector representation using a text encoder, enabling the model to understand the text's semantics. "Image patches" refers to dividing the current canonical frame into several local regions (e.g., 16×16 pixel patches) based on spatial location; each image patch contains local visual information. "Image features" refers to converting each image patch into a high-dimensional vector representation using an image encoder, preserving the local visual content. "Segmentation feature vectors" are feature vectors specifically used to indicate the target segmentation region, generated by fusing text and image features, recording the semantic information of the target region in the current frame. "Position-preserving encoding" refers to processing the canonical frame through the image encoder of the segmentation network, generating a two-dimensional feature map that preserves the original spatial relationships of each region within the frame, rather than compressing it into a one-dimensional sequence. "Two-dimensional image features" refers to the two-dimensional planar feature distribution recording the appearance and position of each region in the frame, used in conjunction with the segmentation feature vectors to achieve pixel-level localization. "Continuous mask response map" refers to the continuous probability distribution map of each position within the frame belonging to the target region, output by the segmentation decoder after matching and calculating the segmentation feature vectors with the two-dimensional image features; higher values ​​indicate a higher probability that the position belongs to the target region. "Thresholding" refers to converting a continuous response map into a binary mask according to a set threshold, distinguishing the target region from the background region. "Frame-level intent mask" refers to a binary mask that corresponds one-to-one with the pixel position of a single canonical frame, marking the target region pointed to by the task text. "Frame-level intent mask sequence" refers to a set of frame masks arranged in temporal order, perfectly aligned with the canonical frame sequence in both time and space.

[0034] The specific implementation process of this scheme includes: First, for each canonical frame, the task text is encoded into text features, and the current canonical frame is divided into image blocks and encoded into image features. For example, the task text "object being operated by a person" is converted into a 512-dimensional text feature vector using the CLIP text encoder, and the 512×512 canonical frame is divided into 32×32 16×16 pixel image blocks, which are then processed by the ViT image encoder to obtain a 1024-dimensional image feature vector. Second, the text features and image features are fused to obtain a segmentation feature vector, and the current canonical frame is simultaneously encoded with position preservation to obtain two-dimensional image features. For example, the text features and image features are fused with cross-attention in a visual language understanding network to generate a segmentation feature vector indicating "operated object"; at the same time, the canonical frame is processed by the SAM image encoder to obtain a 64×64×256 two-dimensional image feature map, preserving the original spatial layout of each region. Then, the segmentation feature vector is matched with the two-dimensional image features to generate a continuous mask response map. For example, the similarity between the segmentation feature vector and the features at each location in the 2D image feature map is calculated to obtain a 64×64 continuous response value distribution, with high response values ​​corresponding to the "operation object" region. Next, the continuous mask response map is restored to the current canonical frame size and thresholded to obtain a frame-level intent mask. For example, the 64×64 response map is upsampled to 512×512 and binarized with a threshold of 0.5 to obtain a frame-level intent mask marking the pixel positions of the "operation object". Finally, the frame-level intent masks corresponding to each canonical frame are arranged in temporal order to obtain a frame-level intent mask sequence spatially aligned with the canonical frame sequence. For example, the above process is performed sequentially on 100 canonical frames to obtain 100 masks arranged in temporal order, with each mask perfectly matching the pixel coordinates of its corresponding canonical frame. Thus, text-image cross-modal fusion achieves accurate mapping of task intent to pixel space, position-preserving encoding ensures spatial alignment accuracy, and threshold binarization yields clear region boundaries, providing a reliable basis for task region localization for subsequent low-rank inversion optimization.

[0035] In this way, by encoding the task text into text features for each canonical frame and dividing the current canonical frame into image blocks and encoding them into image features, the text semantics and visual content are converted into vector representations that the model can process, laying the foundation for cross-modal fusion. Secondly, by fusing text features and image features to obtain segmentation feature vectors, and performing position-preserving encoding on the current canonical frame to obtain two-dimensional image features, the model can locate target regions in the video by combining textual intent, while preserving the spatial relationship between regions within the frame, avoiding misalignment between target location and pixel coordinates. Further, by matching the segmentation feature vectors with the two-dimensional image features to generate a continuous mask response map, restoring the continuous mask response map to the current canonical frame size and thresholding it to obtain a frame-level intent mask, the marked regions in the mask correspond one-to-one with the pixel positions pointed to by the task text, ensuring that the semantic information of "where to focus" can be accurately mapped to the pixel-level space. Finally, by arranging the frame-level intent masks corresponding to each canonical frame in temporal order, a frame-level intent mask sequence aligned with the canonical frame sequence space is obtained, ensuring that the task intent is continuous and consistent in the video temporal sequence. In summary, the above-mentioned technical means achieve accurate mapping and cross-modal alignment of task text semantics to video pixel space, enabling subsequent low-rank inversion optimization to prioritize the preservation of task-related semantics based on accurate regional positioning, avoiding semantic transmission deviations caused by spatial mismatch between intent mask and video frames, and improving the regional positioning accuracy and task consistency of the end-to-end semantic video reconstruction link.

[0036] In some optional embodiments, keyframes are extracted from the canonical frame sequence and their indices are recorded. Low-rank cue parameters are initialized for each keyframe. The low-rank cue parameters are input into a diffusion generation model to generate a reconstruction result. Based on the difference between the reconstruction result and the keyframes under frame-level intent mask weighting, the low-rank cue parameters are iteratively optimized to obtain optimized low-rank cue parameters, including: Keyframes are selected from the standard frame sequence at preset intervals, and their indices are recorded. For each keyframe, low-rank cue parameters are initialized, and the low-rank cue parameters are decomposed into a first low-rank matrix and a second low-rank matrix. The product of the first low-rank matrix and the second low-rank matrix is ​​used as a conditional matrix and input into the diffusion generation model to generate the reconstruction result. The pixel differences between the reconstruction result and the keyframe are calculated, and the difference weights of the corresponding regions of the mask are increased by using frame-level intent masks to obtain weighted differences; The low-rank cue parameters are adjusted in reverse based on the weighted differences. The generation, difference calculation and inverse adjustment are performed iteratively until the convergence condition is met, and the optimized low-rank cue parameters are obtained.

[0037] Specifically, the "first low-rank matrix" and the "second low-rank matrix" refer to the two sub-matrices after the low-rank cue parameter decomposition, and their product constitutes the complete condition matrix. The "diffusion generative model" refers to a generative neural network based on a diffusion probability model, capable of progressively denoising noise to generate an image matching the conditional input. The "reconstruction result" refers to the predicted frame corresponding to the keyframe generated by the diffusion generative model based on the condition matrix. "Pixel difference" refers to the point-by-point deviation of the reconstruction result from the keyframe in pixel values, typically measured using L1 or L2 distance. "Weighted difference" refers to the comprehensive loss value after increasing the weight of task-related regions using frame-level intent masks based on the pixel difference. The "convergence condition" refers to at least one of the following: loss below a threshold, reaching the maximum number of iterations, or the change in adjacent iterations satisfying the early stopping condition.

[0038] The specific implementation process of this scheme includes: First, selecting keyframes from the canonical frame sequence at preset intervals and recording their indices. For example, selecting one frame every 10 frames from a canonical frame sequence of 30 frames per second as a keyframe and recording its position number as 0, 10, 20, etc. Second, initializing low-rank cue parameters for each keyframe, decomposing the low-rank cue parameters into a first low-rank matrix and a second low-rank matrix, and inputting their product as a condition matrix into the diffusion generation model to generate the reconstruction result. For example, initializing a first low-rank matrix of size 77×r and a second low-rank matrix of size r×1024, multiplying them to obtain a 77×1024 condition matrix, inputting it into a diffusion generation model with frozen weights, and generating a 256×256 pixel reconstructed frame. Then, calculating the pixel difference between the reconstructed result and the keyframe, and using a frame-level intent mask to increase the difference weight of the corresponding region of the mask to obtain a weighted difference. For example, the L2 distance between the reconstructed frame and the canonical keyframe is calculated. The "personnel operation area" marked by the mask is assigned a weight of 2.0, and the background area is assigned a weight of 0.5. The weighted sum is then used to obtain the overall loss. Next, the low-rank cue parameters are adjusted in reverse based on the weighted difference. This process of generation, difference calculation, and inverse adjustment is iteratively executed until the convergence condition is met, resulting in the optimized low-rank cue parameters. For example, the AdamW optimizer with a learning rate of 1e-4, combined with gradient pruning, is used. The process stops after 500 iterations or when the loss is below 0.01, ultimately ensuring that the details of the personnel operation area in the reconstructed frame are consistent with those in the keyframe. Thus, sparse keyframe extraction reduces the computational scale, low-rank matrix factorization achieves controllable parameter compression, intention mask weighting guides the priority preservation of task semantics, and iterative convergence ensures optimization stability. This achieves efficient generation of task alignment semantic parameters under limited computational constraints.

[0039] In this way, by selecting keyframes from the standardized frame sequence at preset intervals and recording their indices, a sparse sampling strategy is used instead of full-frame processing, concentrating the inversion calculation on representative frames and reducing the overall optimization complexity and computational overhead. By initializing low-rank cue parameters for each keyframe, these parameters are decomposed into a first low-rank matrix and a second low-rank matrix. Their product is used as a conditional matrix input to the diffusion generation model to generate the reconstruction result, making the number of parameters to be transmitted approximately linear with rank, significantly compressing transmission and storage compared to dense pixel payloads. Furthermore, by calculating the pixel differences between the reconstruction result and the keyframes, a frame-level intent mask is used to increase the difference weight of the corresponding region, resulting in a weighted difference. This allows task-related regions to receive higher attention during optimization, guiding the low-rank cue parameters to prioritize retaining intent semantics and avoiding excessive parameter resources from non-task regions. Finally, by adjusting the low-rank cue parameters in reverse according to the weighted difference, the generation, difference calculation, and inverse adjustment are iteratively executed until the convergence condition is met, resulting in optimized low-rank cue parameters. This achieves a balance between task region fidelity and overall structural stability in the optimization process. In summary, the above-mentioned techniques reduce the computational scale through sparse keyframe extraction, achieve controllable compression of parameters through low-rank matrix factorization, guide semantic priority preservation through intention mask weighting, and ensure optimization stability through iterative convergence. Under the constraints of limited bit rate and computing power, these techniques achieve efficient generation of low-rank semantic parameters and task-aligned transmission.

[0040] In some optional embodiments, the optimized low-rank cue parameters, frame-level intent masks, and keyframe indices for each keyframe are encapsulated into a compact semantic parameter packet and output to the receiving end, including: The low-rank cue parameters and associated scale information of each keyframe are subjected to octet point quantization and serialization to obtain the cue parameter file; Lossless compression is performed on the frame-level intent mask sequence to obtain compressed mask data; The prompt parameter file, compressed mask data, keyframe index and metadata required for parsing are written into the data structure according to a unified field convention to obtain a compact semantic parameter package; Output the compact semantic parameter packet to the receiving end.

[0041] Specifically, "octet-specific-point quantization" refers to the discretization process of converting low-rank hint parameters represented by floating-point numbers into 8-bit integer representations. "Auxiliary scale information" includes the quantization scale (scaling factor) and zero point (offset), used to restore the original numerical range during dequantization. "Serialization" refers to arranging the quantized integer matrix into a continuous data stream for easy storage and transmission. "Lossless compression" refers to compressing binary mask data using lossless coding algorithms (such as run-length encoding, entropy encoding, etc.) to ensure no loss of mask accuracy after decompression. "Unified field conventions" refer to the pre-agreed data structure format between the sender and receiver, specifying the order of components, field lengths, and verification methods. "Metadata required for parsing" includes parameters such as frame rate, keyframe interval, and rank used for reconstruction and interpolation calculations at the receiver.

[0042] The specific implementation process of this scheme includes: First, performing octet-specific quantization and serialization on the optimized low-rank cue parameters and associated scale information of each keyframe to obtain a cue parameter file. For example, the floating-point low-rank matrix is ​​scaled by the quantization scale and a zero-point offset is added to map it to integers in the range of 0 to 255, and then arranged into a byte sequence in row-major order. Second, lossless compression is performed on the frame-level intent mask sequence to obtain compressed mask data. For example, run-length encoding is used on the binarized mask sequence to compress consecutive 0 or 1 sequences into count values, reducing the mask transmission volume. Then, the cue parameter file, compressed mask data, keyframe index, and metadata required for parsing are written into a data structure according to a unified field convention to obtain a compact semantic parameter packet. For example, the data is organized in the order of "packet header (including version and check fields) → keyframe index array → compressed mask data → cue parameter file → parsing metadata (frame rate, keyframe interval, rank)". Finally, the compact semantic parameter packet is output to the receiving end through the communication link. The receiving end parses each component according to the same field convention, completing dequantization, decompression, and conditional assembly. Thus, through quantization compression and unified encapsulation, efficient transmission of low-rank semantic parameters and a closed loop of data parsing between the sending and receiving ends are achieved.

[0043] In this way, by performing octet-specific point quantization and serialization on the low-rank cue parameters and associated scale information of each keyframe, a cue parameter file is obtained. Converting the floating-point low-rank matrix to an integer representation significantly reduces the number of bits after parameter quantization, thus reducing transmission and storage overhead. Lossless compression of the frame-level intent mask sequence yields compressed mask data, compressing the data volume while maintaining the mask space precision, avoiding excessive bandwidth consumption during mask transmission. By writing the cue parameter file, compressed mask data, keyframe index, and metadata required for parsing into a data structure according to a unified field convention, a compact semantic parameter packet is obtained. This allows the sender and receiver to perform data encapsulation and decapsulation based on the same parsing convention, avoiding parsing failures caused by interface mismatch. By outputting the compact semantic parameter packet to the receiver, the receiver can perform dequantization, conditional assembly, and reconstructed timing location based on a unified data structure. In summary, the above-mentioned technical means achieve efficient compression of low-rank cue parameters through octet point quantization, maintain the spatial alignment accuracy of the mask through lossless compression, and achieve data parsing closure at both ends through unified field conventions. While reducing transmission bitrate and storage usage, it improves the closure and reproducibility of the end-to-end semantic video reconstruction link.

[0044] This disclosure provides a semantic video transmission method, which is applied at the receiving end. See [link to relevant documentation]. Figure 2 As shown, it includes: Step S201: Receive a compact semantic parameter packet from the sending end; the sending end is used to acquire and preprocess the original video frames to obtain a canonical frame sequence; acquire the task text corresponding to the video processing task, and generate a frame-level intent mask spatially aligned with the canonical frame sequence based on the task text; extract keyframes from the canonical frame sequence and record the keyframe indices; initialize low-rank cue parameters for each keyframe; input the low-rank cue parameters into a diffusion generation model to generate reconstruction results; iteratively optimize the low-rank cue parameters based on the difference between the reconstruction results and the keyframes under frame-level intent mask weighting to obtain optimized low-rank cue parameters; encapsulate the optimized low-rank cue parameters, frame-level intent mask, and keyframe indices of each keyframe into a compact semantic parameter packet.

[0045] Specifically, a compact semantic parameter packet is received from the sending end. Since the compact semantic parameter packet has been described in detail in the previous embodiments, it will not be repeated here.

[0046] Step S202: Parse the compact semantic parameter packet to obtain the low-rank cue parameters, frame-level intent mask and keyframe index of each keyframe.

[0047] Specifically, the "compact semantic parameter package" refers to a structured data transmission unit formed by the sending end after quantizing, compressing, and encapsulating data such as optimized low-rank cue parameters, frame-level intent masks, and keyframe indices for each keyframe, using a unified field format. "Parsing" refers to the process by which the receiving end extracts the components from the compact semantic parameter package according to a pre-agreed unified field format with the sending end. The "low-rank cue parameters" are semantic condition parameters formed by multiplying two low-rank matrices after octet-specific quantization, used to drive the diffusion generation model to reconstruct keyframes. The "frame-level intent mask" is a lossless compressed binary mask sequence aligned with the canonical frame sequence space, used to mark the target region pointed to by the task text. The "keyframe index" is the position number of each keyframe in the canonical frame sequence, used for temporal positioning and interpolation completion during reconstruction at the receiving end.

[0048] The specific implementation process of this scheme includes: First, the receiving end reads the header information of the compact semantic parameter packet according to the unified field convention, verifies the version and check fields, and confirms the data integrity. Second, it reads the keyframe index array from the compact semantic parameter packet to obtain the position number of each keyframe in the canonical frame sequence, for example, extracting the index sequence [0,10,20,30] to clarify the distribution of keyframes in the original time sequence. Then, it reads the compressed mask data from the compact semantic parameter packet and decompresses it to restore it into a frame-level intent mask sequence aligned with the canonical frame sequence space. For example, it restores the run-length encoded compressed mask data into a 300-frame 512×512 binarized mask, with each mask marking the pixel position of the "personnel operation area". Next, the cue parameter file and associated quantization parameters are read from the compact semantic parameter package. Based on the quantization scale and zero point, the eight-bit integers are dequantized into floating-point matrices to obtain the low-rank cue parameters for each keyframe. For example, 77×r and r×1024 integer matrices are read and dequantized into floating-point matrices using a scale of 2.5 and a zero point of 128. Finally, the dequantized low-rank cue parameters, the decompressed frame-level intent mask, and the keyframe index are assembled using a unified field to obtain the parsing result, providing complete input for subsequent conditional assembly, latent space reconstruction, and interpolation completion. Thus, unified field parsing achieves a closed-loop data decapsulation between the transmitting and receiving ends, dequantization restores the original numerical range of the low-rank semantic parameters, and decompression maintains the spatial alignment accuracy of the mask, providing a parsing basis for the consistency between the receiver's reconstruction and the transmitter's optimization results.

[0049] Step S203: Perform video reconstruction based on the low-rank cue parameters, frame-level intent mask and keyframe index of each keyframe to obtain a reconstructed video frame sequence that is time-aligned with the normalized frame sequence.

[0050] Specifically, "reconstructing a video frame sequence" refers to a complete video formed by sequentially splicing pixel-domain keyframes and interpolated intermediate time frames, which is consistent with the time order of the standard frame sequence.

[0051] The specific implementation process of this scheme includes: First, the dequantized low-rank cue parameters are decomposed into a first low-rank matrix and a second low-rank matrix. The product of the two is used as the condition matrix and input into the diffusion generation model. Simultaneously, the initialized latent variables are read, and the initialized latent variables, the condition matrix, and the frame-level intent mask are input into the diffusion generation model. Latent space reconstruction is performed on the keyframes corresponding to the keyframe indices to obtain pixel-domain keyframes. For example, multiplying the 77×r and r×1024 floating-point matrices yields a 77×1024 condition matrix, which, along with 512×512 initialized latent variables and the intent mask of the corresponding frame, is input into the diffusion generation model. After 50 iterations of denoising starting from the noise, a 512×512 pixel-domain keyframe is generated. Second, a learning-based video interpolation network is used to perform recursive binary interpolation on two adjacent pixel-domain keyframes to obtain the intermediate time frame. For example, frame 5 is generated from frames 0 and 10. Then, frames 2 and 3 are generated from frames 0 and 5, and frames 7 and 8 are generated from frames 5 and 10, repeating this process until all frames 1 through 9 are complete. Next, the pixel-domain keyframes and intermediate timeframes are sorted and concatenated according to the timestamps corresponding to the keyframe indices, resulting in a reconstructed video frame sequence aligned with the normalized frame sequence. For example, keyframes 0, 10, 20, etc., and interpolated intermediate frames 1-9, 11-19, etc., are sorted and concatenated according to timestamps 0, 1, 2, etc., outputting a 300-frame 512×512 reconstructed video. Its temporal order is completely consistent with the normalized frame sequence at the sending end, and its resolution and color space match the geometric normalization convention. Thus, effective recovery of semantic information was achieved through conditional assembly of low-rank parameters, priority preservation of task semantics was achieved through latent space reconstruction and intent masking, efficient completion of non-key frames was achieved through recursive binary interpolation, and consistency of geometric and temporal parameters between the transmitting and receiving ends was achieved through temporal stitching, thereby completing end-to-end semantic video reconstruction under limited computing power.

[0052] In the embodiments of this disclosure, by receiving a compact semantic parameter packet from the sending end, the receiving end does not need to process the dense pixel data of the original video frames and can directly reconstruct based on the compressed semantic parameters, reducing the data processing pressure on the receiving end. By parsing the compact semantic parameter packet to obtain the low-rank cue parameters, frame-level intent masks, and keyframe indices of each keyframe, the receiving end can recover the optimization results and regional guidance information of the sending end based on the same parsing convention, avoiding reconstruction failures caused by interface mismatch. By performing video reconstruction based on the low-rank cue parameters, frame-level intent masks, and keyframe indices of each keyframe, a reconstructed video frame sequence aligned with the time sequence of the standard frame sequence is obtained. This allows the receiving end to reconstruct keyframes based on the low-rank parameter-driven diffusion generation model, locate the time sequence based on the keyframe indices, maintain the semantic consistency of the task region based on the frame-level intent mask, and complete intermediate time frames through interpolation, thus obtaining a complete duration video even under computationally limited conditions. In summary, this application achieves efficient transmission of semantic parameters and task-aligned reconstruction through the collaboration of low-rank inversion at the transmitting end and reconstruction at the receiving end. Under the constraints of limited bit rate and computing power, it reduces the overall computing power and latency overhead at the receiving end and improves the reproducibility of the reconstructed video in terms of geometric and temporal parameters.

[0053] In some optional embodiments, the compact semantic parameter packet is parsed to obtain the low-rank cue parameters, frame-level intent mask, and keyframe index for each keyframe, including: Read the quantized low-rank hint parameters and their corresponding quantization scale and zeros from the compact semantic parameter package, and dequantize the low-rank hint parameters into a floating-point matrix based on the quantization scale and zeros; Read the compressed frame-level intent mask sequence from the compact semantic parameter packet, decompress the frame-level intent mask sequence, and restore it to a frame-level intent mask that is spatially aligned with the canonical frame sequence. Read the keyframe index and parsing metadata required from the compact semantic parameter package. The parsing metadata includes at least one of frame rate, keyframe interval and rank. The dequantized low-rank cue parameters, the decompressed frame-level intent mask, and the keyframe index are assembled using a unified field to obtain the parsing result.

[0054] Specifically, "quantized low-rank cue parameters" refers to the integer matrix representation obtained by the sender after performing octet-specific-point quantization on the optimized floating-point low-rank matrix. "Quantization scale" refers to the scaling factor used to scale the original floating-point values ​​during quantization, determining the quantization precision. "Zero point" refers to the offset used to deflect the original floating-point values ​​during quantization, allowing negative numbers to be mapped to the non-negative integer range; "Dequantization" refers to the process by which the receiver, based on the quantization scale and zero point, restores the integer matrix to a floating-point matrix close to the original floating-point values. "Compressed frame-level intent mask sequence" refers to the data obtained by the sender after lossless encoding and compression of the binarized mask sequence, with a smaller volume than the original mask. "Decompression" refers to the process by which the receiver recovers the original mask sequence using a lossless decoding algorithm. "Metadata required for parsing" refers to the configuration parameters required by the receiver to complete reconstruction and interpolation calculations, including at least one of the following: frame rate (frames per second), keyframe interval (number of frames between adjacent keyframes), and rank (rank value of the low-rank matrix). "Unified Fields" refers to the pre-agreed data structure format between the sending and receiving ends, specifying the order of the components, field lengths, and verification methods. "Assembly" refers to the process of organizing the dequantized and decompressed components into structured data according to the unified fields. "Parsing Result" refers to the complete data set obtained by the receiving end after decapsulation, dequantization, decompression, and assembly, which can be directly used for subsequent reconstruction.

[0055] The specific implementation process of this scheme includes: First, reading the quantized low-rank cue parameters and their corresponding quantization scale and zeros from the compact semantic parameter package, and then dequantizing the low-rank cue parameters into a floating-point matrix based on the quantization scale and zeros. For example, reading a 77×r integer matrix A, an r×1024 integer matrix B, and their corresponding quantization scale of 2.5 and zeros of 128 from the parameter package, and restoring the integer matrix to a floating-point matrix according to the formula "floating-point value = (integer value - zeros) × scale", yields low-rank cue parameters close to the original optimized values. Second, reading the compressed frame-level intent mask sequence from the compact semantic parameter package, and decompressing the frame-level intent mask sequence to restore it to a frame-level intent mask spatially aligned with the canonical frame sequence. For example, reading the run-length encoded compressed mask data, and decoding it to restore it to a 300-frame 512×512 binary mask, where each mask marker region corresponds perfectly to the pixel coordinates of the canonical frame sequence. Then, the keyframe index and metadata required for parsing are read from the compact semantic parameter package. This metadata includes at least one of frame rate, keyframe interval, and rank. For example, the keyframe index sequence [0,10,20,30], frame rate 30fps, keyframe interval 10 frames, and rank value 64 are extracted to provide a basis for subsequent temporal localization and model configuration. Finally, the dequantized low-rank cue parameters, the decompressed frame-level intent mask, and the keyframe index are assembled using a unified field to obtain the parsing result. Thus, the original numerical range of the low-rank semantic parameters is restored through precise dequantization of the quantization scale; the spatial alignment accuracy of the mask is maintained through lossless decompression; and a closed-loop data parsing mechanism is achieved between the transmitting and receiving ends through unified field assembly, providing a complete parsing foundation for the consistency between the reconstruction results at the receiving end and the optimization results at the sending end.

[0056] In this way, by reading the quantized low-rank cue parameters and their corresponding quantization scale and zeros from the compact semantic parameter package, and dequantizing the low-rank cue parameters into a floating-point matrix based on the quantization scale and zeros, the receiver can recover the original numerical range of the low-rank cue parameters based on the quantization parameters agreed upon by the sender, ensuring that the diffusion generation model obtains usable conditional input. By reading the compressed frame-level intent mask sequence from the compact semantic parameter package and decompressing it, a frame-level intent mask aligned with the normal frame sequence space is restored, ensuring that the task area positioning information maintains lossless accuracy after transmission and avoiding task semantic offset during reconstruction due to mask distortion. By reading the keyframe index and metadata required for parsing from the compact semantic parameter package, including at least one of frame rate, keyframe interval, and rank, the receiver obtains keyframe temporal positioning and model configuration parameters, providing necessary basis for subsequent reconstruction and interpolation calculations. By assembling the dequantized low-rank cue parameters, the decompressed frame-level intent mask, and the keyframe index according to a unified field, the parsing result is obtained, enabling both the sender and receiver to complete data decapsulation based on the same field agreement, avoiding parsing failure caused by interface mismatch. In summary, this application achieves efficient transmission and accurate reconstruction of low-rank cue parameters through inverse quantization recovery using octet point quantization, maintains the spatial alignment accuracy of the mask through lossless compression decompression, and realizes a closed-loop data parsing between the transmitting and receiving ends through unified field assembly. While reducing the transmission bitrate, it ensures the consistency between the reconstruction conditions at the receiving end and the optimization results at the sending end, thereby improving the reliability and reproducibility of the end-to-end semantic video reconstruction link.

[0057] In some optional embodiments, video reconstruction is performed based on the low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index to obtain a reconstructed video frame sequence that is temporally aligned with the canonical frame sequence, including: The dequantized low-rank cue parameters are decomposed into a first low-rank matrix and a second low-rank matrix. The product of the first low-rank matrix and the second low-rank matrix is ​​used as the conditional matrix input to the diffusion generation model. Read the initial latent variables, input the initial latent variables, condition matrix and frame-level intent mask into the diffusion generation model, and perform latent space reconstruction on the keyframes corresponding to the keyframe index to obtain pixel domain keyframes. A recursive binary interpolation method is used to perform interpolation on keyframes in two adjacent pixel domains to obtain intermediate time frames. The pixel-domain keyframes and intermediate timeframes are sorted and concatenated according to the timestamps corresponding to the keyframe indices to obtain a reconstructed video frame sequence that is time-aligned with the normalized frame sequence.

[0058] Specifically, "dequantized low-rank cue parameters" refers to the floating-point matrix after quantization and zero-point restoration, used to drive the diffusion generation model to reconstruct keyframes. "First low-rank matrix" and "second low-rank matrix" refer to the two sub-matrices after decomposing the low-rank cue parameters; their product constitutes the complete condition matrix. "Condition matrix" refers to the cross-attention conditions input to the diffusion generation model, controlling the direction and semantics of the generated content. "Diffusion generation model" refers to a generative neural network based on a diffusion probability model, capable of progressively denoising noise to generate an image matching the conditional input. "Initialized latent variables" refers to the initial latent representation corresponding to the first frame or scene reset position, serving as the starting point of the diffusion generation process. "Latent space reconstruction" refers to the process of iteratively denoising from the initialized latent variables along the reverse diffusion step to the latent variables, and then converting it into pixel-domain keyframes via a decoder. "Pixel-domain keyframes" refer to the visualized keyframes in pixel space after reconstruction by the diffusion generation model. "Learned video interpolation network" refers to a trained and optimized neural network capable of inferring the motion state and image content at intermediate moments based on two adjacent frames. "Recursive binary interpolation" refers to an interpolation strategy that continuously divides the time interval of adjacent keyframes into two parts, generating midpoint time frames layer by layer until the entire interval is filled. "Intermediate time frames" refer to non-keyframes between adjacent keyframes that are completed through interpolation. "Reconstructed video frame sequence" refers to a complete video formed by sequentially splicing pixel-domain keyframes and interpolated intermediate time frames, which is consistent with the time order of the standard frame sequence.

[0059] The specific implementation process of this scheme includes: First, the dequantized low-rank cue parameters are decomposed into a first low-rank matrix and a second low-rank matrix, and their product is used as the condition matrix input to the diffusion generation model. For example, a 77×r floating-point matrix is ​​used as the first low-rank matrix, and an r×1024 floating-point matrix is ​​used as the second low-rank matrix. Their product is used to obtain a 77×1024 condition matrix, which is then written into the cross-attention condition input of the diffusion generation model. Second, the initialized latent variables are read, and the initialized latent variables, the condition matrix, and the frame-level intent mask are input into the diffusion generation model. Latent space reconstruction is performed on the keyframes corresponding to the keyframe indices to obtain pixel-domain keyframes. For example, 512×512 initialized latent variables, along with the condition matrix and the corresponding frame's intent mask, are input into the diffusion generation model. Iterative 50-step reverse denoising is performed starting from Gaussian noise, and then the model is converted into 512×512 pixel-domain keyframes by a VAE decoder. The intent mask guides the model to focus on task-related regions in the conditional path. Then, a learned video interpolation network is used to recursively perform binary interpolation on adjacent pixel-domain keyframes to obtain intermediate timeframes. For example, frames 0 and 10 are input into the learned video interpolation network to generate frame 5; then frames 2 and 3 are generated using frames 0 and 5, and frames 7 and 8 are generated using frames 5 and 10; this process is repeated until all frames 1 to 9 are completed, making the motion completion coarse to fine. Finally, the pixel-domain keyframes and intermediate timeframes are sorted and concatenated according to the timestamps corresponding to the keyframe indices to obtain a reconstructed video frame sequence that is time-aligned with the normalized frame sequence. For example, keyframes 0, 10, 20, etc., and the interpolated intermediate frames 1-9, 11-19, etc., are sorted and concatenated according to timestamps 0, 1, 2, etc., to output a 300-frame 512×512 reconstructed video, whose time order is completely consistent with the normalized frame sequence at the sending end, and whose resolution and color space match the geometric normalization convention. Thus, effective recovery of semantic information was achieved through low-rank parameter decomposition and conditional assembly; priority preservation of task semantics was achieved through latent space reconstruction and intent masking guidance; efficient completion of non-key frames was achieved through recursive binary interpolation; and consistency of geometric and temporal parameters between the transmitting and receiving ends was achieved through temporal stitching. End-to-end semantic video reconstruction was completed under the condition of limited computing power.

[0060] In this way, by decomposing the dequantized low-rank cue parameters into a first low-rank matrix and a second low-rank matrix, and using their product as the condition matrix input to the diffusion generation model, the low-rank cue parameters are restored from a compact matrix product form to a complete conditional input usable by the diffusion generation model, achieving an effective conversion from low-rank semantic parameters to generation conditions. By reading the initialized latent variables, the initialized latent variables, the condition matrix, and the frame-level intent mask are input into the diffusion generation model to reconstruct the latent space of the keyframes corresponding to the keyframe indices, resulting in pixel-domain keyframes. The receiver then iteratively denoises and generates keyframes from the latent space along the backward diffusion step based on the initial state and semantic conditions provided by the sender. The frame-level intent mask guides the model to focus on task-related regions in the conditional path, ensuring that the reconstructed keyframes prioritize the preservation of task semantics. By using a learned video interpolation network to perform recursive binary interpolation on adjacent pixel-domain keyframes, intermediate time frames are obtained. This eliminates the need for the receiver to perform diffusion generation on non-keyframes; intermediate time frames can be completed solely through learned interpolation, significantly reducing the receiver's computational power and latency overhead. By sorting and concatenating pixel-domain keyframes and intermediate timeframes according to the timestamps corresponding to the keyframe indices, a reconstructed video frame sequence aligned temporally with the normalized frame sequence is obtained. This ensures that the temporal order of the reconstructed frames is completely consistent with the normalized frame sequence at the transmitting end, and that the output resolution and color space match the geometric normalization convention. In summary, this application achieves effective recovery of semantic conditions through low-rank parameter decomposition and conditional assembly, prioritizes the preservation of task semantics through latent space reconstruction and intent masking guidance, efficiently completes non-keyframes through recursive binary interpolation, and ensures consistency of geometric and temporal parameters between the transmitting and receiving ends through temporal concatenation. Under computationally limited conditions, this improves the receiving end's ability to output full-duration video and the reproducibility of the reconstruction results.

[0061] In some optional embodiments, a learned video interpolation network is used to perform recursive binary interpolation on keyframes in adjacent pixel domains to obtain intermediate timeframes, including: Input two adjacent pixel domain keyframes into a learning video interpolation network to generate the midpoint time frame of the time interval corresponding to the two keyframes. The time interval is divided into a left sub-interval and a right sub-interval, with the midpoint time frame as the boundary; The two adjacent frames on the left are input into the learning video interpolation network to generate the midpoint time frame of the left sub-interval; The two adjacent frames on the right are input into the learning video interpolation network to generate the midpoint time frame of the right sub-interval; Repeat the division and interpolation process for each sub-interval until the entire interval between two keyframes is filled, thus obtaining all intermediate time frames.

[0062] Specifically, a "learning-based video interpolation network" refers to a trained and optimized neural network capable of inferring the motion state and image content at an intermediate moment based on the pixel content of two adjacent frames, generating a transition frame between the two frames. A "midpoint frame" refers to the interpolation frame corresponding to the midpoint of the time interval between two adjacent keyframes, generated by the learning-based video interpolation network based on motion estimation and pixel synthesis of the two keyframes. "Left sub-interval" and "right sub-interval" refer to two sub-intervals to be further interpolated, divided from the original keyframe time interval by the midpoint frame. "Recursive binary interpolation" refers to an interpolation strategy that repeatedly performs midpoint interpolation and interval division for each sub-interval, refining it layer by layer until the entire interval is filled.

[0063] The specific implementation process of this scheme includes: First, inputting two adjacent pixel domain keyframes into a learning-based video interpolation network to generate the midpoint time frame of the corresponding time interval between the two keyframes. For example, inputting frame 0 (timestamp 0) and frame 10 (timestamp 10) into the learning-based video interpolation network, which calculates the motion vector between the two frames through optical flow estimation and generates frame 5 (timestamp 5) based on a pixel synthesis mechanism to fill the midpoint time between the two keyframes. Second, dividing the time interval into a left sub-interval and a right sub-interval using the midpoint time frame as the boundary. For example, using frame 5 as the boundary, the original interval [0,10] is divided into a left sub-interval [0,5] and a right sub-interval [5,10]. Then, inputting two adjacent frames on the left into the learning-based video interpolation network to generate the midpoint time frame of the left sub-interval; and inputting two adjacent frames on the right into the learning-based video interpolation network to generate the midpoint time frame of the right sub-interval. For example, inputting frames 0 and 5 into a learning video interpolation network generates frame 2 (timestamp 2), the midpoint of the left sub-interval; inputting frames 5 and 10 into the same network generates frame 7 (timestamp 7), the midpoint of the right sub-interval. Then, the division and interpolation process is repeated for each sub-interval until the entire interval between two keyframes is filled, resulting in all intermediate timeframes. For example, using frame 2 as a boundary, [0,2] and [2,5] are further binary interpolated; using frame 7 as a boundary, [5,7] and [7,10] are further binary interpolated. This process is repeated until frames 1, 3, 4, 6, 8, and 9 are all filled in, ultimately obtaining the complete sequence of frames 0 to 10. Therefore, a learning-based video interpolation network was used to achieve pixel-level transition frame generation based on motion estimation. By using a recursive binary search strategy, global interpolation was decomposed into multi-level processing of local refinement. Each time, only the local motion relationship between two adjacent frames needs to be processed, avoiding motion estimation distortion caused by one-time interpolation across the entire range. While reducing the computational complexity of a single step, temporal coherence was maintained, and motion completion was gradually refined from coarse to fine, adapting to complex motion changes.

[0064] In this way, by inputting keyframes from two adjacent pixel domains into a learning-based video interpolation network, midpoint frames of the corresponding time intervals between the two keyframes are generated. This eliminates the need for the receiver to perform computationally intensive diffusion generation on non-keyframes; instead, the motion state and image content at intermediate moments are inferred solely through the interpolation network, significantly reducing the computational overhead of single-frame generation. By dividing the time interval into left and right sub-intervals with the midpoint frames as the boundary, and inputting the two adjacent frames on the left and right sides into the learning-based video interpolation network to generate midpoint frames for each sub-interval, the interpolation process is progressively decomposed from global coarse-grained to local fine-grained. Each time, only the local motion relationship between adjacent frames needs to be processed, avoiding motion estimation distortion caused by one-time interpolation across the entire interval. By repeating the division and interpolation process for each sub-interval until the entire interval between two keyframes is filled, all intermediate frames are obtained. This recursively refines motion completion from coarse to fine, maintaining temporal continuity while adapting to complex motion changes. In summary, this application replaces full-frame diffusion reconstruction with a recursive binary interpolation strategy, focusing the receiver's computation on keyframe generation and learning-based interpolation. This improves the ability to output full-length video under limited computing power, while also improving the temporal coherence of motion regions, ensuring that the reconstruction results are consistent with the sending end's agreement in terms of geometric and temporal parameters.

[0065] Correspondingly, see Figure 3 , Figure 3This is a schematic diagram of the overall flow of a semantic video transmission method according to one embodiment of this disclosure. The sending end first acquires the task intent and the original video. The task intent describes the video's region of interest and task objective using natural language text, while the original video is the source video data to be transmitted. The original video undergoes geometric normalization processing, i.e., frame-by-frame sampling rate conversion, aspect ratio-preserving scaling, and edge padding to match the size of each frame with the subsequent network input. Simultaneously, affine transformation parameters and the padding region mask are recorded as geometric metadata, resulting in a normalized frame sequence. The task intent and the normalized frame sequence are input together into a frame-level intent mask generation module. The task text is encoded into text features, and the normalized frames are encoded into image features. These are fused by a visual language understanding network to obtain a segmentation feature vector. This vector is matched with two-dimensional image features that preserve positional relationships to generate a continuous mask response map. After size restoration and thresholding, a frame-level intent mask sequence spatially aligned with the normalized frame sequence is obtained. The obtained frame-level intent mask, along with geometric metadata, task text, frame rate, keyframe interval, and other parsing parameters, is written into a data structure in the mask and metadata organization module according to a unified field convention, forming a parsing convention shared by the sending and receiving ends. The semantic cue inversion and optimization module then extracts keyframes from the canonical frame sequence and records their indices. For each keyframe, low-rank cue parameters are initialized and decomposed into a first low-rank matrix and a second low-rank matrix. Their product is used as a conditional matrix input to the diffusion generation model to generate the reconstruction result. The low-rank cue parameters are iteratively optimized based on the difference between the reconstruction result and the keyframes under frame-level intent mask weighting until convergence is achieved. Finally, in the cue parameter quantization and transmission module, the optimized low-rank cue parameters and associated scale information are subjected to octet-specific point quantization and serialization. The frame-level intent mask sequence is losslessly compressed, and the quantized cue parameters, compressed mask, keyframe index, and parsed metadata are encapsulated into a compact semantic parameter packet and output to the receiving end.

[0066] The receiving end first receives the compact semantic parameter packet through the cue parameter and metadata parsing module, and reads each component according to the same unified field convention as the sending end. In the cue parameter dequantization and condition assembly module, the quantized low-rank cue parameters are dequantized into a floating-point matrix according to the quantization scale and zero point. The frame-level intent mask sequence is decompressed, and the dequantized low-rank cue parameters, the decompressed frame-level intent mask, and the keyframe index are assembled into the parsing result. Then, it enters the keyframe latent space reconstruction module, where the dequantized low-rank cue parameters are decomposed into a first low-rank matrix and a second low-rank matrix. The product of the two is used as the condition matrix and input into the diffusion generation model. At the same time, the initialized latent variables are read. The initialized latent variables, the condition matrix, and the frame-level intent mask are jointly input into the diffusion generation model. The keyframes corresponding to the keyframe index are denoised iteratively from Gaussian noise along the back diffusion step to the latent variables, and then converted into pixel-domain keyframes by the decoder. In the intermediate time-space interpolation and completion module, a recursive binary interpolation is performed on adjacent pixel-domain keyframes using a learned video interpolation network. This involves first finding the midpoint time frame between the two keyframes, then dividing the time interval into a left and right sub-interval based on the midpoint frame, and repeating the interpolation on both sub-intervals until the entire interval is filled, thus obtaining all intermediate time-space frames. Finally, in the reconstructed video output module, the pixel-domain keyframes and intermediate time-space frames are sorted and concatenated according to the timestamps corresponding to the keyframe indices, outputting a reconstructed video frame sequence with resolution and color space consistent with the geometric normalization convention of the sending end.

[0067] The sender and receiver achieve a closed loop through a compact semantic parameter package: the sender encapsulates geometric normalization parameters, frame-level intent masks, low-rank cue parameters, and keyframe indices in a unified manner, while the receiver decapsulates and restores each component according to the same parsing convention. The sender is responsible for intent mask generation and low-rank inversion optimization, while the receiver is responsible for keyframe latent space reconstruction and intermediate time-space interpolation completion. Both share a diffusion generation model architecture, and through cross-end transmission of low-rank cue parameters and intent masks, end-to-end consistent expression of task intent, region of interest, and pixel coordinates is achieved under constraints of limited bitrate and computing power.

[0068] The following describes an apparatus embodiment of this application, which can be used to execute the semantic video transmission method in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the semantic video transmission method described above.

[0069] This disclosure also provides a semantic video transmission device 400, such as Figure 4 As shown, it includes: The first acquisition module 401 is used to acquire raw video frames and perform preprocessing to obtain a standardized frame sequence. The second acquisition module 402 is used to acquire the task text corresponding to the video processing task and generate a frame-level intent mask that is aligned with the standard frame sequence space based on the task text. The optimization module 403 is used to extract keyframes from the canonical frame sequence and record the keyframe index. It initializes low-rank cue parameters for each keyframe, inputs the low-rank cue parameters into the diffusion generation model to generate reconstruction results, and iteratively optimizes the low-rank cue parameters based on the difference between the reconstruction results and the keyframes under frame-level intent mask weighting to obtain the optimized low-rank cue parameters. The output module 404 is used to encapsulate the optimized low-rank cue parameters, frame-level intent masks and keyframe indices of each keyframe into a compact semantic parameter package and output it to the receiving end. The compact semantic parameter package is used to reconstruct a video frame sequence that is time-aligned with the canonical frame sequence at the receiving end.

[0070] In some optional embodiments, the first acquisition module 401 acquires the original video frames and performs preprocessing to obtain a standardized frame sequence, including: The original video is sampled frame by frame to obtain sampled frames that match the target resolution; The sampled frame is scaled while preserving its aspect ratio to obtain a scaled frame; Edge padding is applied to the scaled frame to obtain a normalized frame; Arrange the standard frames in chronological order to obtain the standard frame sequence.

[0071] In some optional embodiments, the second acquisition module 402 generates a frame-level intent mask spatially aligned with the canonical frame sequence based on the task text, including: For each specification frame, the task text is encoded into text features, and the current specification frame is divided into image blocks and encoded into image features; The text features and image features are fused to obtain a segmentation feature vector. The current canonical frame is then encoded with position preservation to obtain two-dimensional image features. The segmentation feature vector is matched with the two-dimensional image features to generate a continuous mask response map; The continuous mask response map is restored to the current standard frame size, and then the frame-level intent mask is obtained after thresholding. Arrange the frame-level intent masks corresponding to each standard frame in time sequence to obtain a frame-level intent mask sequence that is spatially aligned with the standard frame sequence.

[0072] In some optional embodiments, the optimization module 403 extracts keyframes from the canonical frame sequence and records the keyframe indices. For each keyframe, it initializes low-rank cue parameters, inputs the low-rank cue parameters into a diffusion generation model to generate a reconstruction result, and iteratively optimizes the low-rank cue parameters based on the difference between the reconstruction result and the keyframes under frame-level intent mask weighting, obtaining optimized low-rank cue parameters, including: Keyframes are selected from the standard frame sequence at preset intervals, and their indices are recorded. For each keyframe, low-rank cue parameters are initialized, and the low-rank cue parameters are decomposed into a first low-rank matrix and a second low-rank matrix. The product of the first low-rank matrix and the second low-rank matrix is ​​used as a conditional matrix and input into the diffusion generation model to generate the reconstruction result. The pixel differences between the reconstruction result and the keyframe are calculated, and the difference weights of the corresponding regions of the mask are increased by using frame-level intent masks to obtain weighted differences; The low-rank cue parameters are adjusted in reverse based on the weighted differences. The generation, difference calculation and inverse adjustment are performed iteratively until the convergence condition is met, and the optimized low-rank cue parameters are obtained.

[0073] In some optional embodiments, the output module 404 encapsulates the optimized low-rank cue parameters, frame-level intent masks, and keyframe indices of each keyframe into a compact semantic parameter packet and outputs it to the receiving end, including: The low-rank cue parameters and associated scale information of each keyframe are subjected to octet point quantization and serialization to obtain the cue parameter file; Lossless compression is performed on the frame-level intent mask sequence to obtain compressed mask data; The prompt parameter file, compressed mask data, keyframe index and metadata required for parsing are written into the data structure according to a unified field convention to obtain a compact semantic parameter package; Output the compact semantic parameter packet to the receiving end.

[0074] This disclosure also provides a semantic video transmission device 500, such as... Figure 5 As shown, it includes: The receiving module 501 is used to receive a compact semantic parameter packet from the sending end. The sending end is used to acquire the original video frames and preprocess them to obtain a canonical frame sequence; acquire the task text corresponding to the video processing task, generate a frame-level intent mask that is spatially aligned with the canonical frame sequence based on the task text; extract keyframes from the canonical frame sequence and record the keyframe indices; initialize low-rank cue parameters for each keyframe; input the low-rank cue parameters into the diffusion generation model to generate reconstruction results; iteratively optimize the low-rank cue parameters based on the difference between the reconstruction results and the keyframes under the weighting of the frame-level intent mask to obtain optimized low-rank cue parameters; and encapsulate the optimized low-rank cue parameters, frame-level intent mask, and keyframe indices of each keyframe into a compact semantic parameter packet. The parsing module 502 is used to parse the compact semantic parameter packet to obtain the low-rank cue parameters, frame-level intent mask and key frame index of each key frame; The reconstruction module 503 is used to reconstruct the video based on the low-rank cue parameters, frame-level intent mask and keyframe index of each keyframe, so as to obtain a reconstructed video frame sequence that is time-aligned with the normal frame sequence.

[0075] In some optional embodiments, the parsing module 502 parses the compact semantic parameter packet to obtain the low-rank cue parameters, frame-level intent mask, and keyframe index for each keyframe, including: Read the quantized low-rank hint parameters and their corresponding quantization scale and zeros from the compact semantic parameter package, and dequantize the low-rank hint parameters into a floating-point matrix based on the quantization scale and zeros; Read the compressed frame-level intent mask sequence from the compact semantic parameter packet, decompress the frame-level intent mask sequence, and restore it to a frame-level intent mask that is spatially aligned with the canonical frame sequence. Read the keyframe index and parsing metadata required from the compact semantic parameter package. The parsing metadata includes at least one of frame rate, keyframe interval and rank. The dequantized low-rank cue parameters, the decompressed frame-level intent mask, and the keyframe index are assembled using a unified field to obtain the parsing result.

[0076] In some optional embodiments, the reconstruction module 503 performs video reconstruction based on the low-rank cue parameters, frame-level intent mask, and keyframe index of each keyframe to obtain a reconstructed video frame sequence that is time-aligned with the canonical frame sequence, including: The dequantized low-rank cue parameters are decomposed into a first low-rank matrix and a second low-rank matrix. The product of the first low-rank matrix and the second low-rank matrix is ​​used as the conditional matrix input to the diffusion generation model. Read the initial latent variables, input the initial latent variables, condition matrix and frame-level intent mask into the diffusion generation model, and perform latent space reconstruction on the keyframes corresponding to the keyframe index to obtain pixel domain keyframes. A recursive binary interpolation method is used to perform interpolation on keyframes in two adjacent pixel domains to obtain intermediate time frames. The pixel-domain keyframes and intermediate timeframes are sorted and concatenated according to the timestamps corresponding to the keyframe indices to obtain a reconstructed video frame sequence that is time-aligned with the normalized frame sequence.

[0077] In some optional embodiments, the reconstruction module 503 uses a learned video interpolation network to perform recursive binary interpolation on keyframes in adjacent pixel domains to obtain intermediate timeframes, including: Input two adjacent pixel domain keyframes into a learning video interpolation network to generate the midpoint time frame of the time interval corresponding to the two keyframes. The time interval is divided into a left sub-interval and a right sub-interval, with the midpoint time frame as the boundary; The two adjacent frames on the left are input into the learning video interpolation network to generate the midpoint time frame of the left sub-interval; The two adjacent frames on the right are input into the learning video interpolation network to generate the midpoint time frame of the right sub-interval; Repeat the division and interpolation process for each sub-interval until the entire interval between two keyframes is filled, thus obtaining all intermediate time frames.

[0078] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0079] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0080] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0081] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0082] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0083] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the semantic video transmission method. For example, in some embodiments, the semantic video transmission method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the semantic video transmission method by any other suitable means (e.g., by means of firmware).

[0084] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0085] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0086] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0087] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0088] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0089] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0090] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the disclosed technical solution can be achieved, and this is not limited herein.

[0091] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A semantic video transmission method, wherein, The method, executed by the sending end, includes: The original video frames are acquired and preprocessed to obtain a standardized frame sequence. Obtain the task text corresponding to the video processing task, and generate a frame-level intent mask that is spatially aligned with the standard frame sequence based on the task text; Keyframes are extracted from the canonical frame sequence and keyframe indices are recorded. Low-rank cue parameters are initialized for each keyframe. The low-rank cue parameters are input into the diffusion generation model to generate reconstruction results. Based on the difference between the reconstruction results and the keyframes under the frame-level intent mask weighting, the low-rank cue parameters are iteratively optimized to obtain optimized low-rank cue parameters. The optimized low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index are encapsulated into a compact semantic parameter package and output to the receiving end. The compact semantic parameter package is used at the receiving end to reconstruct a video frame sequence that is time-aligned with the canonical frame sequence.

2. The method according to claim 1, wherein, The process of acquiring and preprocessing the original video frames to obtain a standardized frame sequence includes: The original video is sampled frame by frame to obtain sampled frames that match the target resolution; The sampled frame is scaled while maintaining its aspect ratio to obtain a scaled frame; The scaled frame is edge-padded to obtain a normalized frame; Arrange the standard frames in chronological order to obtain the standard frame sequence.

3. The method according to claim 1, wherein, The step of generating a frame-level intent mask aligned spatially with the canonical frame sequence based on the task text includes: For each specification frame, the task text is encoded into text features, and the current specification frame is divided into image blocks and encoded into image features; The text features and the image features are fused to obtain a segmentation feature vector. The current canonical frame is then encoded with position preservation to obtain two-dimensional image features. The segmentation feature vector is matched with the two-dimensional image features to generate a continuous mask response map; The continuous mask response map is restored to the current standard frame size, and then a frame-level intent mask is obtained after thresholding. Arrange the frame-level intent masks corresponding to each standard frame in chronological order to obtain a frame-level intent mask sequence spatially aligned with the standard frame sequence.

4. The method according to any one of claims 1 to 3, wherein, The process involves extracting keyframes from the canonical frame sequence and recording their indices, initializing low-rank cue parameters for each keyframe, inputting these parameters into a diffusion generation model to generate a reconstruction result, and iteratively optimizing the low-rank cue parameters based on the difference between the reconstruction result and the keyframe under the frame-level intent mask weighting to obtain optimized low-rank cue parameters, including: Keyframes are selected from the standard frame sequence at preset intervals, and their indices are recorded. For each keyframe, a low-rank cue parameter is initialized, and the low-rank cue parameter is decomposed into a first low-rank matrix and a second low-rank matrix. The product of the first low-rank matrix and the second low-rank matrix is ​​used as a conditional matrix and input into the diffusion generation model to generate the reconstruction result. Calculate the pixel difference between the reconstruction result and the key frame, and use the frame-level intent mask to increase the difference weight of the corresponding region of the mask to obtain a weighted difference; The low-rank cue parameters are adjusted in reverse according to the weighted difference. The generation, difference calculation and inverse adjustment are performed iteratively until the convergence condition is met, and the optimized low-rank cue parameters are obtained.

5. The method according to any one of claims 1 to 3, wherein, The step of encapsulating the optimized low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index into a compact semantic parameter packet and outputting it to the receiving end includes: The low-rank cue parameters and associated scale information of each keyframe are subjected to octet point quantization and serialization to obtain the cue parameter file; The frame-level intent mask sequence is losslessly compressed to obtain compressed mask data; The prompt parameter file, the compressed mask data, the keyframe index, and the metadata required for parsing are written into a data structure according to a unified field convention to obtain a compact semantic parameter package. The compact semantic parameter packet is output to the receiving end.

6. A semantic video transmission method, wherein, The method, executed by the receiving end, includes: The system receives a compact semantic parameter packet from a sending end. The sending end is used to acquire and preprocess raw video frames to obtain a canonical frame sequence. It acquires task text corresponding to the video processing task and generates a frame-level intent mask spatially aligned with the canonical frame sequence based on the task text. It extracts keyframes from the canonical frame sequence and records the keyframe indices. For each keyframe, it initializes low-rank cue parameters, inputs the low-rank cue parameters into a diffusion generation model to generate a reconstruction result, and iteratively optimizes the low-rank cue parameters based on the difference between the reconstruction result and the keyframe under the weighting of the frame-level intent mask to obtain optimized low-rank cue parameters. The optimized low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe indices are encapsulated into a compact semantic parameter packet. Parse the compact semantic parameter packet to obtain the low-rank cue parameters of each key frame, the frame-level intent mask, and the key frame index; Based on the low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index, video reconstruction is performed to obtain a reconstructed video frame sequence that is time-aligned with the normalized frame sequence.

7. The method according to claim 6, wherein, The process of parsing the compact semantic parameter packet to obtain the low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index includes: Read the quantized low-rank hint parameters and the corresponding quantization scale and zeros from the compact semantic parameter package, and dequantize the low-rank hint parameters into a floating-point matrix according to the quantization scale and zeros; Read the compressed frame-level intent mask sequence from the compact semantic parameter package, decompress the frame-level intent mask sequence, and restore it to a frame-level intent mask that is spatially aligned with the canonical frame sequence. Read the keyframe index and parsing metadata required from the compact semantic parameter package. The parsing metadata includes at least one of frame rate, keyframe interval, and rank. The dequantized low-rank cue parameters, the decompressed frame-level intent mask, and the keyframe index are assembled using a unified field to obtain the parsing result.

8. The method according to claim 6, wherein, The step of reconstructing the video based on the low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index to obtain a reconstructed video frame sequence that is temporally aligned with the canonical frame sequence includes: The dequantized low-rank cue parameters are decomposed into a first low-rank matrix and a second low-rank matrix. The product of the first low-rank matrix and the second low-rank matrix is ​​used as the conditional matrix input to the diffusion generation model. Read the initial latent variables, and input the initial latent variables, the condition matrix and the frame-level intent mask into the diffusion generation model to perform latent space reconstruction on the key frame corresponding to the key frame index to obtain the pixel domain key frame. A recursive binary interpolation method is used to perform interpolation on keyframes in two adjacent pixel domains to obtain intermediate time frames. The pixel domain keyframes and the intermediate time frames are sorted and concatenated according to the timestamps corresponding to the keyframe indices to obtain a reconstructed video frame sequence that is time-aligned with the normalized frame sequence.

9. The method according to claim 8, wherein, The step of using a learned video interpolation network to perform recursive binary interpolation on keyframes in adjacent pixel domains to obtain intermediate timeframes includes: The keyframes of two adjacent pixel domains are input into a learning video interpolation network to generate the midpoint time frame of the time interval corresponding to the two keyframes. The time interval is divided into a left sub-interval and a right sub-interval, with the midpoint time frame as the boundary. The two adjacent frames on the left are input into the learning-based video interpolation network to generate the midpoint time frame of the left sub-interval; The two adjacent frames on the right are input into the learning-based video interpolation network to generate the midpoint time frame of the right sub-interval; The process of dividing and interpolating each sub-interval is repeated until the entire interval between the two keyframes is filled, thus obtaining all intermediate time frames.

10. A semantic video transmission device, wherein, The device includes: The first acquisition module is used to acquire raw video frames and preprocess them to obtain a standardized frame sequence. The second acquisition module is used to acquire the task text corresponding to the video processing task and generate a frame-level intent mask that is aligned with the space of the standard frame sequence based on the task text. The optimization module is used to extract keyframes from the canonical frame sequence and record keyframe indices, initialize low-rank cue parameters for each keyframe, input the low-rank cue parameters into the diffusion generation model to generate reconstruction results, and iteratively optimize the low-rank cue parameters based on the difference between the reconstruction results and the keyframes under the frame-level intent mask weighting to obtain optimized low-rank cue parameters. The output module is used to encapsulate the optimized low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index into a compact semantic parameter package and output it to the receiving end. The compact semantic parameter package is used to reconstruct a video frame sequence that is time-aligned with the canonical frame sequence at the receiving end.

11. A semantic video transmission device, wherein, The device includes: A receiving module is used to receive a compact semantic parameter packet from a sending end. The sending end is used to acquire raw video frames and preprocess them to obtain a canonical frame sequence; acquire task text corresponding to the video processing task, and generate a frame-level intent mask spatially aligned with the canonical frame sequence based on the task text; extract keyframes from the canonical frame sequence and record the keyframe indices; initialize low-rank cue parameters for each keyframe; input the low-rank cue parameters into a diffusion generation model to generate a reconstruction result; iteratively optimize the low-rank cue parameters based on the difference between the reconstruction result and the keyframe under the weighting of the frame-level intent mask to obtain optimized low-rank cue parameters; and encapsulate the optimized low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe indices into a compact semantic parameter packet. The parsing module is used to parse the compact semantic parameter package to obtain the low-rank cue parameters of each key frame, the frame-level intent mask, and the key frame index. The reconstruction module is used to reconstruct the video based on the low-rank cue parameters of each keyframe, the frame-level intent mask, and the keyframe index, to obtain a reconstructed video frame sequence that is time-aligned with the normalized frame sequence.

12. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-9.

14. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-9.