Cloud edge-end collaborative monitoring video transmission method, system and device based on key semantic guidance video super-resolution and medium
By performing spatial downsampling on terminal devices and extracting key objects on edge servers, combining cloud video reconstruction models, optimizing video transmission bit rate, the problem of large bandwidth consumption of cloud video surveillance system is solved, and efficient video transmission and semantic information preservation in low-bandwidth scenarios are achieved.
Patent Information
- Application Number
- CN202511007574.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-07-22
AI Technical Summary
When transmitting high-resolution video data, cloud video surveillance systems face the problem of high bandwidth consumption and difficulty in reducing the bit rate while maintaining key semantic information, especially in surveillance scenarios with limited bandwidth.
The video super-resolution method based on key semantic guidance is adopted, and spatial downsampling and keyframe extraction are performed through terminal devices, and the edge server performs key object extraction, and the cloud server reconstructs the video reconstruction model based on the key object-guided video, optimizes the transmission bit rate and retains key semantic information.
While reducing the overall system bit rate, the accuracy of object detection is maintained, and bandwidth is significantly saved. It demonstrates the efficiency and scalability in bandwidth-constrained scenarios, and is suitable for low-bandwidth monitoring scenarios.
Smart Images

Figure CN120529104A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of super-resolution, and in particular to a cloud-edge-end collaborative surveillance video transmission method, system, device and medium based on key semantics-guided video super-resolution. Background Art
[0002] With the rapid development of applications such as video on demand (VOD), live streaming, and ultra-low-latency real-time communications, as well as the increasing demand for high-quality, high-fidelity video content, video traffic has become a dominant component of internet data. In this context, distributed surveillance systems play a vital role in various fields, including border patrols, traffic monitoring, environmental monitoring, and smart homes. These systems typically require transmitting video data captured by cameras to cloud servers for processing and analysis, forming cloud video surveillance (CVS) systems.
[0003] However, cloud video surveillance systems, as a key component of distributed surveillance networks, face increasing challenges in transmitting massive amounts of high-resolution video data. The explosive growth of IoT surveillance devices and the continuous improvement of video quality have led to a significant increase in the amount of surveillance video data that needs to be transmitted. While current end-to-cloud collaborative surveillance video transmission methods can reduce bitrates beyond traditional compression techniques, the periodic transmission of high-resolution keyframes in these methods still consumes a significant amount of bandwidth, making it impossible to maintain the key semantic information of the reconstructed video while reducing the video transmission bitrate in bandwidth-constrained surveillance scenarios. Summary of the Invention
[0004] Based on the above technical problems, the present invention provides a cloud-edge-end collaborative surveillance video transmission method, system, device and medium based on key semantics-guided video super-resolution, aiming to overcome the above problems or at least partially solve the above problems.
[0005] A first aspect of the present invention provides a cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution, which is applied to a video transmission system comprising: a terminal device, an edge server, and a cloud server. The method comprises: The terminal device collects a high-resolution video frame sequence, determines a plurality of high-resolution key frames from the high-resolution video frame sequence, divides the high-resolution video sequence between every two high-resolution key frames into a high-resolution video segment, downsamples each high-resolution video segment to obtain a low-resolution video segment, and sends each low-resolution video segment and the corresponding two high-resolution key frames to the edge server; The edge server extracts key objects from each of the multiple high-resolution key frames, removes semantically irrelevant background information from each high-resolution key frame, and retains the semantic information of the key object area to obtain multiple high-resolution key frames after key object extraction. The edge server sends each low-resolution video segment and the corresponding two high-resolution key frames after key object extraction to the cloud server. The cloud server reconstructs each low-resolution video segment and the corresponding high-resolution key frames extracted from the two key objects based on a key-object-guided video reconstruction model to obtain a reconstructed high-resolution video sequence.
[0006] A second aspect of the present invention provides a cloud-edge-end collaborative surveillance video transmission system based on key semantics-guided video super-resolution, the video transmission system comprising: a terminal device, an edge server, and a cloud server; The terminal device is configured to capture a high-resolution video frame sequence, determine a plurality of high-resolution key frames from the high-resolution video frame sequence, divide the high-resolution video sequence between every two high-resolution key frames into a high-resolution video segment, downsample each high-resolution video segment to obtain a low-resolution video segment, and send each low-resolution video segment and the corresponding two high-resolution key frames to the edge server; The edge server is configured to extract key objects from each of the multiple high-resolution key frames, remove semantically irrelevant background information from each high-resolution key frame, and retain semantic information of the key object region to obtain multiple high-resolution key frames after key object extraction. The edge server is configured to send each low-resolution video segment and the corresponding two high-resolution key frames after key object extraction to the cloud server. The cloud server is used to reconstruct a video sequence based on a key object-guided video reconstruction model using each low-resolution video segment and the corresponding high-resolution key frames extracted from the two key objects.
[0007] The third aspect of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the computer program is executed by the processor, it implements the cloud-edge-end collaborative surveillance video transmission method based on key semantics-guided video super-resolution as described in the first aspect of the present invention.
[0008] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the cloud-edge-end collaborative monitoring video transmission method based on key semantic-guided video super-resolution of the first aspect of the present invention is implemented.
[0009] In the cloud-edge-end collaborative surveillance video transmission method based on key semantics-guided video super-resolution proposed in the present invention, on the terminal device side, a high-resolution video frame sequence is collected, spatial downsampling is performed on the high-resolution video frame sequence, and high-resolution key frames are extracted from the collected high-resolution video frame sequence, so that each low-resolution video segment and its corresponding two high-resolution key frames are sent to the edge server; on the edge server side, object detection is performed in the high-resolution key frames, key objects are extracted from the high-resolution key frames, semantically irrelevant background information in the high-resolution key frames is removed, and the semantic information of the key object area in the high-resolution key frames is retained, so as to obtain multiple high-resolution key frames after key object extraction, and each low-resolution video segment and its corresponding two high-resolution key frames after key object extraction are sent to the cloud server. In this way, the present invention avoids consuming a large amount of bandwidth on semantically irrelevant background information, while retaining the semantic information of the key target area, further reducing the bit rate transmitted to the cloud; on the cloud server side, based on the key object-guided video reconstruction model, each low-resolution video segment and the corresponding two high-resolution key frames after key object extraction are used for reconstruction to obtain a reconstructed high-resolution video sequence. In this way, the present invention achieves a balance between bit rate optimization and semantic fidelity (semantic utility) by combining active downsampling on the end side and key object extraction on the edge side to guide the cloud side to perform video reconstruction based on the key object-guided video reconstruction model. Compared with the traditional end-cloud collaborative monitoring video transmission method that transmits complete key frames, the present invention saves a lot of bandwidth and reduces the key frame bit rate. It maintains the accuracy of object detection while significantly reducing the overall system bit rate, highlighting the efficiency and scalability of the present invention in bandwidth-limited monitoring scenarios, and proving that the present invention has strong versatility and practical potential for deployment in low-bandwidth monitoring scenarios. Therefore, in bandwidth-limited monitoring scenarios, it is possible to reduce the video transmission bit rate while ensuring the preservation of key semantic information of the reconstructed video. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0011] Figure 1 This is a flowchart of a method for cloud-edge-device collaborative surveillance video transmission based on key semantics-guided video super-resolution according to an embodiment of the present invention; Figure 2 1 is a schematic structural diagram of a key object guided video reconstruction model according to an embodiment of the present invention; Figure 3 This is a structural block diagram of a cloud-edge-device collaborative surveillance video transmission system based on key semantics-guided video super-resolution provided by one embodiment of the present invention; Figure 4 FIG. 1 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0013] Please refer to Figure 1 , Figure 1 This is a flowchart of a method for cloud-edge-end collaborative surveillance video transmission based on key semantics-guided video super-resolution according to an embodiment of the present invention. Figure 1 As shown, the cloud-edge-end collaborative surveillance video transmission method based on key semantics-guided video super-resolution provided in this embodiment is applied to a video transmission system, which includes: a terminal device, an edge server, and a cloud server. The terminal device collects high-resolution video data, performs spatial downsampling, and selects high-resolution key frames from the high-resolution video data at fixed intervals; the edge server is equipped with a key object extraction algorithm to extract key objects from the high-resolution key frames. The cloud server executes a key object-guided video reconstruction model to reconstruct the high-resolution video.
[0014] In this embodiment, the cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution includes at least the following steps: Step S11: The terminal device collects a high-resolution video frame sequence, determines multiple high-resolution key frames from the high-resolution video frame sequence, divides the high-resolution video sequence between every two high-resolution key frames into a high-resolution video segment, downsamples each high-resolution video segment to obtain a low-resolution video segment, and sends each low-resolution video segment and the corresponding two high-resolution key frames to the edge server.
[0015] In this embodiment, the terminal device may be a surveillance camera device, for example, a remote sensing satellite, an unmanned aerial vehicle, a ground tower, or other platform device. The terminal device acquires a high-resolution video frame sequence, determines multiple high-resolution key frames from the high-resolution video frame sequence, and then, based on the order of the multiple high-resolution key frames in the high-resolution video frame sequence, divides the high-resolution video sequence between each two adjacent high-resolution key frames into a high-resolution video segment. A high-resolution video segment includes two adjacent high-resolution key frames and the high-resolution video sequence between the two adjacent high-resolution key frames.
[0016] Next, each high-resolution video segment obtained by dividing the high-resolution video frame sequence is downsampled to obtain a corresponding low-resolution video segment. The terminal device then sends each low-resolution video segment and its corresponding two high-resolution key frames to the edge server. In an optional example, the terminal device can perform fourfold bicubic spatial downsampling on the high-resolution video segment, resulting in low-resolution video segments with the same frame rate.
[0017] Step S12: The edge server extracts key objects from multiple high-resolution key frames respectively, removes semantically irrelevant background information in each high-resolution key frame, and retains the semantic information of the key object area, to obtain multiple high-resolution key frames after key object extraction, and sends each low-resolution video segment and the corresponding two high-resolution key frames after key object extraction to the cloud server.
[0018] In this embodiment, the edge server can perform key object extraction on each of the multiple received high-resolution key frames. For each high-resolution key frame, this key object extraction removes semantically irrelevant background information from the high-resolution key frame while retaining the semantic information of the key object region within each high-resolution key frame. This results in multiple high-resolution key frames after key object extraction. The edge server then sends each low-resolution video segment and the two high-resolution key frames corresponding to the low-resolution video segment after key object extraction to the cloud server.
[0019] In a specific optional example, key objects may be extracted from a high-resolution key frame to obtain a key object area in the high-resolution key frame, and the RGB values of the remaining areas in the high-resolution key frame except the key object area may be set to 0 to obtain a high-resolution key frame after key object extraction.
[0020] In another optional example, the edge server may encode the low-resolution video segment and the high-resolution keyframes from which the two key objects corresponding to the low-resolution video segment are extracted, and then send the encoded data to the cloud server. For example, the edge server may perform JPEG encoding on the high-resolution keyframes from which the key objects are extracted, and then send the encoded data to the cloud server, while performing H.265 encoding on the low-resolution video segment, and then send the encoded data to the cloud server.
[0021] Step S13: The cloud server reconstructs each low-resolution video segment and the corresponding high-resolution key frames extracted from the two key objects based on the key object-guided video reconstruction model to obtain a reconstructed high-resolution video sequence.
[0022] In this embodiment, a trained key object-guided video reconstruction model is pre-deployed in the cloud server. The cloud server can reconstruct the video based on the key object-guided video reconstruction model using each received low-resolution video segment and the high-resolution key frames extracted from the two key objects corresponding to the low-resolution video segment to obtain a reconstructed high-resolution video sequence.
[0023] The reconstruction process of the key object-guided video reconstruction model is performed on a "video segment" basis. Each time, the model takes as input a low-resolution video segment and two high-resolution keyframes corresponding to the low-resolution segment after key object extraction. The model outputs the reconstructed high-resolution video segment corresponding to the low-resolution segment. This model uses a shifted window-based self-attention mechanism to extract features from the high-resolution keyframes after key object extraction. It then uses optical flow-guided deformable convolution to align the extracted keyframe features with the low-resolution features, thereby achieving high-resolution video reconstruction.
[0024] In this way, the cloud server will input the received multiple low-resolution video segments corresponding to the high-resolution video sequence and the high-resolution key frames extracted from the two key objects corresponding to the multiple low-resolution video segments into the key object-guided video reconstruction model in units of "video segments", and obtain multiple reconstructed high-resolution video segments output in sequence by the key object-guided video reconstruction model, thereby obtaining a reconstructed high-resolution video sequence based on the multiple reconstructed high-resolution video segments.
[0025] In this embodiment, active downsampling is performed on the end side and key objects are extracted on the edge side to guide the cloud side to perform video reconstruction based on the key object-guided video reconstruction model, thereby achieving a balance between bit rate optimization and semantic fidelity. Compared with the traditional end-cloud collaborative monitoring video transmission method that transmits complete key frames, this embodiment saves a lot of bandwidth and reduces the key frame bit rate. It maintains the accuracy of object detection while significantly reducing the overall system bit rate, highlighting the efficiency and scalability of this embodiment in bandwidth-limited monitoring scenarios, and proving that this embodiment has strong versatility and practical potential for deployment in low-bandwidth monitoring scenarios. Therefore, in bandwidth-limited monitoring scenarios, it is possible to reduce the video transmission bit rate while ensuring the preservation of key semantic information of the reconstructed video.
[0026] In combination with the above embodiments, in one embodiment, the present invention further provides a cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution. In this method, in addition to the above steps, step S21 may be included, and the "the edge server extracts key objects from multiple high-resolution key frames" in the above step S12 may include step S22: Step S21: The edge server detects the quality of the communication network with the cloud server.
[0027] In this embodiment, the edge server can detect the quality of the communication network between the edge server and the cloud server. The communication network quality can be characterized by any one or a combination of the following: bandwidth, throughput, transmission rate, etc. In this embodiment, a target threshold is preset, which is the minimum threshold indicating good communication network quality between the edge server and the cloud server.
[0028] Step S22: When the communication network quality is lower than the target threshold, the edge server extracts key objects from multiple high-resolution key frames respectively, and sends each low-resolution video segment and the corresponding two high-resolution key frames after key object extraction to the cloud server.
[0029] In this embodiment, when it is determined that the communication network quality between the edge server and the cloud server is below a target threshold, it is determined that the communication between the edge server and the cloud server is poor. At this time, the edge server performs key object extraction on each of the multiple high-resolution key frames received. For each high-resolution key frame, the key object extraction removes semantically irrelevant background information in each high-resolution key frame, retains the semantic information of the key object area in each high-resolution key frame, thereby obtaining multiple high-resolution key frames after key object extraction, and sends each low-resolution video segment and the two high-resolution key frames after key object extraction corresponding to the low-resolution video segment to the cloud server. In this way, this embodiment avoids consuming excessive bandwidth on semantically irrelevant background information and successfully retains the semantic information of the key target area at a lower bit rate, demonstrating the potential of this embodiment in low-bandwidth scenarios.
[0030] In combination with the above embodiments, the present invention further provides a cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution. In this method, the above step S13 may specifically include steps S31 to S33: Step S31: The cloud server extracts shallow features of each low-resolution video segment and the corresponding high-resolution key frames after the two key objects are extracted through the shallow feature extraction module.
[0031] In this embodiment, the key object-guided video reconstruction model includes at least a shallow feature extraction module for extracting shallow features from the input data in the initial stage. The shallow features are low-level visual information, primarily including basic structural features such as edges, textures, and colors. The cloud server can input each low-resolution video segment and its corresponding two high-resolution key frames after key object extraction into the shallow feature extraction module. The shallow feature extraction module extracts shallow features from each low-resolution video segment and its corresponding two high-resolution key frames after key object extraction using a convolution kernel, thereby obtaining shallow features for each low-resolution video segment and its corresponding two high-resolution key frames after key object extraction.
[0032] In an optional embodiment, a low-resolution video segment is represented as , the high-resolution keyframes after the two key objects are extracted are represented as .in, 、 The low-resolution video frames corresponding to two adjacent high-resolution key frames in a high-resolution video segment; The low-resolution video frames corresponding to the high-resolution video sequence (i.e., multiple high-resolution video frames) between two adjacent high-resolution key frames in a high-resolution video segment; 、 The high-resolution key frames are obtained by extracting the key objects corresponding to two adjacent high-resolution key frames in a high-resolution video segment.
[0033] Step S32: For the j-th low-resolution video frame in each low-resolution video segment, the cloud server calculates the optical flow between the j-th low-resolution video frame and its adjacent low-resolution video frames at multiple scales through the optical flow estimation network, and calculates the optical flow between the j-th low-resolution video frame and the corresponding two high-resolution key frames after key objects are extracted at multiple scales, where j is an integer greater than 0.
[0034] In this embodiment, the cloud server inputs each low-resolution video segment into the optical flow estimation network, and the optical flow estimation network performs optical flow estimation on the low-resolution video segment to obtain optical flows at different scales. Specifically, for the j-th low-resolution video frame ( , j is an integer greater than 0), the cloud server can calculate the optical flow between the j-th low-resolution video frame and its adjacent low-resolution video frames at multiple scales through the optical flow estimation network ( ), and the optical flow between the j-th low-resolution video frame and the corresponding high-resolution key frame after the two key objects are extracted is calculated at multiple scales ( , i, i+k indicates that the terminal device of this embodiment periodically selects high-resolution key frames at a fixed interval, and the fixed interval is k). Among them, the high-resolution key frames after the two key objects corresponding to the j-th low-resolution video frame are extracted , there are two corresponding first low-resolution video frames in the low-resolution video segment (ie, two low-resolution video frames located at both ends of the low-resolution video segment { , }), therefore, the optical flow between the j-th low-resolution video frame and the corresponding high-resolution key frames after the two key objects are extracted is: the optical flow between the j-th low-resolution video frame and the two first low-resolution video frames corresponding to the low-resolution video segment.
[0035] Step S33: The cloud server uses the shallow features of the high-resolution key frames extracted from each low-resolution video segment and the corresponding two key objects, combined with the optical flow output by the optical flow estimation network, to perform attention-based feature extraction, alignment and fusion at multiple scales to obtain the reconstructed high-resolution video sequence.
[0036] In this embodiment, the cloud server can use the shallow features of each low-resolution video segment output by the shallow feature extraction module and the corresponding high-resolution key frames after the two key objects are extracted, and combine the optical flow output by the optical flow estimation network to perform attention-based feature extraction, alignment and fusion at multiple scales, and finally obtain a reconstructed high-resolution video sequence.
[0037] In conjunction with the above embodiments, in one embodiment, the present invention further provides a cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution. In this embodiment, the above step S33, "the cloud server uses the shallow features of the high-resolution keyframe extracted from each low-resolution video segment and the corresponding two key objects to perform attention-based feature extraction at multiple scales," can specifically include the following steps S41 to S42: Step S41: The cloud server performs attention-based feature extraction on the shallow features of each low-resolution video frame in each low-resolution video segment through the multi-scale temporal reciprocal self-attention module to obtain first attention features at multiple scales.
[0038] In this embodiment, the key object guided video reconstruction model also includes: a multi-scale attention-based feature extraction module (FE) for extracting features from low-resolution video frames and high-resolution key frames after key object extraction. The multi-scale attention-based feature extraction module includes: a multi-scale temporal reciprocal self-attention module and a multi-scale self-attention module. It can be understood that each scale of the attention-based feature extraction module consists of two modules: a temporal reciprocal self-attention module (TRSA) and a self-attention module (SA). Among them, the temporal reciprocal self-attention module is applied to low-resolution video frames, and the self-attention module is applied to high-resolution key frames after key object extraction.
[0039] The cloud server can perform attention-based feature extraction on the shallow features of each low-resolution video frame in each low-resolution video segment output by the shallow feature extraction module through a multi-scale temporal reciprocal self-attention module to obtain first attention features at multiple scales.
[0040] Step S42: The cloud server performs attention-based feature extraction on the shallow features of the high-resolution key frame extracted from each key object through the self-attention modules at multiple scales to obtain second attention features at multiple scales.
[0041] In this embodiment, the cloud server can perform attention-based feature extraction on the shallow features of the high-resolution key frame after each key object extraction output by the shallow feature extraction module through self-attention modules at multiple scales to obtain second attention features at multiple scales.
[0042] In this embodiment, feature extraction of low-resolution video frames through the temporal reciprocal self-attention module can achieve joint motion estimation and feature alignment between the features of low-resolution video frames, so that low-resolution video frames can utilize information from other low-resolution video frames without excessively increasing the temporal complexity. Feature extraction of high-resolution key frames after key object extraction through the self-attention module can further extract features of the high-resolution key frames.
[0043] In combination with the above embodiments, in one embodiment, the present invention also provides a cloud-edge-end collaborative surveillance video transmission method based on key semantics-guided video super-resolution. In this method, the above step S33, "the cloud server uses the shallow features of the high-resolution key frames extracted from each low-resolution video segment and the corresponding two key objects, combined with the optical flow output by the optical flow estimation network, to align and fuse them at multiple scales to obtain the reconstructed high-resolution video sequence" can specifically include the following steps S51 to S54: Step S51: The cloud server uses the first attention feature and the second attention feature through the parallel alignment module of the multiple scales, combined with the optical flow output by the optical flow estimation network, to obtain features processed at multiple scales.
[0044] In this embodiment, the key object-guided video reconstruction model also includes a target temporal reciprocal self-attention module and a multi-scale parallel warping module (PW). The multi-scale parallel warping module corresponds to the multi-scale attention-based feature extraction module, i.e., each scale corresponds to one parallel warping module and one attention-based feature extraction module. The parallel warping module is used to align and fuse features using optical flow-guided deformable convolution.
[0045] In this embodiment, the cloud server can align and fuse the features through a parallel alignment module at multiple scales, using the first attention features at multiple scales and the second attention features at multiple scales, and combining the optical flow output by the optical flow estimation network to obtain features processed at multiple scales. The features processed at multiple scales are features of the low-resolution video segment that have been fused with the features of the high-resolution key frames after key object extraction.
[0046] Step S52: The cloud server processes the features processed at multiple scales through the target temporal reciprocal self-attention module to obtain deep features of each low-resolution video segment.
[0047] In this embodiment, the target temporal reciprocal self-attention module is one or more temporal reciprocal self-attention modules that process features processed at multiple scales. The cloud server can process the features processed at multiple scales through the target temporal reciprocal self-attention module to obtain the deep features of each low-resolution video segment. = .in, is the deep feature of the first low-resolution video frame in a low-resolution video segment, is the deep feature of the second low-resolution video frame in a low-resolution video segment, is the deep feature of the third low-resolution video frame in a low-resolution video segment, is the deep feature of the k+1th low-resolution video frame in a low-resolution video segment.
[0048] Step S53: The cloud server adds and upsamples the deep features and shallow features of each low-resolution video segment to obtain intermediate features of each low-resolution video segment.
[0049] In this embodiment, the cloud server adds and upsamples the deep features of each low-resolution video segment and the shallow features of each low-resolution video segment to obtain the intermediate features of each low-resolution video segment. The shallow features extracted by the shallow feature extraction module for each low-resolution video segment can be obtained by first extracting the deep features of each low-resolution video segment. and shallow features of each low-resolution video segment After the addition, a pixel shuffle operation is performed to obtain the intermediate features of each low-resolution video segment.
[0050] Step S54: the cloud server obtains the reconstructed high-resolution video sequence according to the intermediate features of each low-resolution video segment and the upsampling result of each low-resolution video segment.
[0051] In this embodiment, the cloud server will also upsample each low-resolution video segment to obtain an upsampling result of each low-resolution video segment, and then obtain each reconstructed high-resolution video segment based on the intermediate features of each low-resolution video segment and the upsampling result of each low-resolution video segment (for example, the intermediate features of each low-resolution video segment are added to the upsampling result of each low-resolution video segment), thereby obtaining a reconstructed high-resolution video sequence.
[0052] In combination with the above embodiments, in one embodiment, the present invention further provides a cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution. In this method, the above step S51 may specifically include steps S61 to S64: Step S61: The cloud server uses the first attention feature and the second attention feature through the parallel alignment module of the multiple scales, combined with the optical flow output by the optical flow estimation network, to perform a spatial deformation operation to obtain processed low-resolution video frame features and processed high-resolution key frame features.
[0053] In this embodiment, considering that the temporal reciprocal self-attention modules at multiple scales operate within a small spatial window, the small temporal window prevents direct information interaction between locations that are spatially distant, which may be insufficient to handle large motions. Furthermore, to fully utilize the high-frequency information provided by the high-resolution keyframes after key object extraction, the low-resolution video frames must be properly aligned and fused with the high-resolution keyframes after key object extraction. Therefore, the alignment and calibration process in this embodiment applies not only to different low-resolution video frames, but also to low-resolution video frames and high-resolution keyframes after key object extraction.
[0054] This embodiment can use a parallel alignment module of multiple scales, utilize first attention features of multiple scales and second attention features of multiple scales, and combine the optical flow output by the optical flow estimation network to perform a spatial deformation operation to obtain processed low-resolution video frame features and processed high-resolution key frame features.
[0055] In an optional example, for the j-th low-resolution video frame, it is recorded as , the first attention feature extracted at a specific scale is expressed as , the second attention features corresponding to the high-resolution key frames after the two key objects are extracted are recorded as and In order to compare the j-th low-resolution video frame with its previous and next frames ( and ) and the high-resolution key frame after key object extraction, the optical flow estimation network is used to calculate the optical flow between the j-th low-resolution video frame and its adjacent low-resolution video frame, and the optical flow between the j-th low-resolution video frame and the high-resolution key frame after key object extraction. These optical flows are recorded as The calculated optical flow (i.e., the optical flow output by the optical flow estimation network) is then used to warp the features to obtain the processed low-resolution video frame features and the processed high-resolution key frame features, which can be expressed as: ; Where W represents the spatial deformation operation, 、 are the features of the j-1th low-resolution video frame after processing and the features of the j+1th low-resolution video frame after processing; is the optical flow between the j-th low-resolution video frame and its adjacent low-resolution video frames; 、 are the first attention features corresponding to the j-1th low-resolution video frame and the j+1th low-resolution video frame respectively; and The second attention feature corresponding to the high-resolution key frame after extracting the two key objects corresponding to the j-th low-resolution video frame; , are the optical flows between the j-th low-resolution video frame and the corresponding high-resolution key frames after the two key objects are extracted; , The high-resolution key frame features extracted from the two key objects corresponding to the j-th low-resolution video frame after processing.
[0056] Step S62: The cloud server uses the processed low-resolution video frame features and the processed high-resolution key frame features, as well as the first attention features, through the parallel alignment module of the multiple scales, combined with the optical flow output by the optical flow estimation network to obtain the residual of the optical flow and the modulation mask of the deformable convolution.
[0057] In this embodiment, the cloud server can use the obtained processed low-resolution video frame features and processed high-resolution key frame features, as well as the first attention features of multiple scales, through parallel alignment modules at multiple scales, and combine them with the optical flow output by the optical flow estimation network to obtain the residual of the optical flow and the modulation mask of the deformable convolution.
[0058] In an optional example, the j-1th low-resolution video frame feature obtained after processing in the above steps is And the processed j+1th low-resolution video frame features , and the high-resolution key frame features extracted from the two key objects corresponding to the j-th low-resolution video frame after processing 、 , multiple convolutional layers can be applied to compute the residual of the optical flow and the modulation mask of the deformable convolution: ; Among them, c is the splicing operation, C is the multiple convolution layer operations, is the first attention feature corresponding to the j-th low-resolution video frame, are the optical flows between the j-th low-resolution video frame and its adjacent low-resolution video frames, and the optical flows between the j-th low-resolution video frame and the high-resolution key frame after key object extraction; is the residual of optical flow; 、 is the modulation mask of the deformable convolution.
[0059] Step S63: The cloud server uses the parallel alignment module of the multiple scales, the first attention feature, the second attention feature, the residual of the optical flow, and the modulation mask of the deformable convolution, combined with the optical flow output by the optical flow estimation network, to perform a deformable convolution operation to obtain the aligned features.
[0060] In this embodiment, the cloud server can use a parallel alignment module of multiple scales to utilize the first attention features of multiple scales, the second attention features of multiple scales, the residual of the optical flow, and the modulation mask of the deformable convolution, and combine the optical flow output by the optical flow estimation network to perform a deformable convolution operation to obtain the aligned features.
[0061] In an optional example, the features after alignment can be obtained by the following formula: ; Where D is the deformable convolution operation, 、 is the feature after alignment; c is the splicing operation, 、 are the first attention features corresponding to the j-1th low-resolution video frame and the j+1th low-resolution video frame respectively; and The second attention feature corresponding to the high-resolution key frame after extracting the two key objects corresponding to the j-th low-resolution video frame; are the optical flows between the j-th low-resolution video frame and its adjacent low-resolution video frames, and the optical flows between the j-th low-resolution video frame and the high-resolution key frame after key object extraction; is the residual of optical flow; 、 is the modulation mask of the deformable convolution.
[0062] Step S64: The cloud server fuses the aligned features and the first attention features to obtain the features processed at multiple scales.
[0063] In this embodiment, the cloud server can fuse the aligned features with the first attention features at multiple scales to obtain features processed at multiple scales. In an optional example, for the j-th low-resolution video frame, the aligned features obtained in step S63 can be combined with the first attention features at multiple scales. 、 and the first attention features extracted from the j-th low-resolution video frame at multiple scales The features are connected and fused through the multi-layer perceptron (MLP) layer to reduce their dimensionality, and finally the features processed at multiple scales are obtained.
[0064] In one embodiment, if Figure 2 As shown, Figure 2 FIG. 1 is a structural diagram of a key object guided video reconstruction model according to an embodiment of the present invention. Figure 2 In the upper left part, the terminal device collects a high-resolution video frame sequence, and obtains multiple high-resolution video segments based on the high-resolution video frame sequence. For each high-resolution video segment, each low-resolution video segment and its corresponding two high-resolution key frames are obtained and sent to the edge server. Figure 2 In the middle part above, the edge server receives a low-resolution video segment and its corresponding two high-resolution key frames. Taking the two high-resolution key frames as an example, the edge server extracts key objects from the two high-resolution key frames to obtain two high-resolution key frames after key object extraction. Each low-resolution video segment and its corresponding two high-resolution key frames after key object extraction are sent to the cloud server. Figure 2 In the figure, the upper right part shows that the cloud server receives a low-resolution video segment and its corresponding high-resolution key frames after extracting two key objects. Based on the key object-guided video reconstruction model, the low-resolution video segment and the corresponding high-resolution key frames after extracting two key objects are used for reconstruction to obtain a reconstructed high-resolution video segment.
[0065] exist Figure 2The middle part is a schematic diagram of the structure of the key object guided video reconstruction model. The key object guided video reconstruction model at least includes: a shallow feature extraction module (i.e. Figure 2 shallow feature extraction in ), multiple scales of attention-based feature extraction modules (FE) (i.e. Figure 2 Feature extraction in ), parallel alignment module (PW) of multiple scales (i.e. Figure 2 parallel alignment in ) and feature refinement modules (i.e. Figure 2 , such as the target temporal reciprocal self-attention module mentioned above), wherein the attention-based feature extraction module includes: a temporal reciprocal self-attention module and a self-attention module.
[0066] First, for each low-resolution video segment, the optical flow estimation network is used to calculate the optical flow between the low-resolution video frame and its adjacent low-resolution video frames at multiple scales. The optical flow between the low-resolution video frame and the corresponding high-resolution key frames after key object extraction is also calculated at multiple scales. This results in the optical flow output of the optical flow estimation network at multiple scales. Each low-resolution video segment and the corresponding high-resolution key frame after the key objects are extracted are input into the shallow feature extraction module to obtain the shallow features of each low-resolution video segment and the shallow features of the corresponding high-resolution key frame after the key objects are extracted; then the shallow features of the low-resolution video segment and the shallow features of the corresponding high-resolution key frame after the key objects are extracted are input into the first first-scale attention-based feature extraction module, and the shallow features of the low-resolution video segment are extracted by the first-scale temporal reciprocal self-attention module to obtain the first attention feature 1, and the shallow features of the high-resolution key frame after the key objects are extracted are extracted by the first-scale self-attention module to obtain the second attention feature 1, and then the first-scale parallel alignment module is used to align and fuse the features using the first attention feature 1, the second attention feature 1 and the optical flow output by the optical flow estimation network at the first scale to obtain the first-scale fused feature 1.
[0067] The first-scale fusion feature 1 is input into the first second-scale attention-based feature extraction module, and the scale is changed first. Then, the features representing the low-resolution video segment in the first-scale fusion feature 1 are extracted through the second-scale temporal reciprocal self-attention module to obtain the first attention feature 2. The features representing the high-resolution key frame in the first-scale fusion feature 1 are extracted through the second-scale self-attention module to obtain the second attention feature 2. Then, the second-scale parallel alignment module is used to align and fuse the features using the first attention feature 2, the second attention feature 2 and the optical flow output by the optical flow estimation network at the second scale to obtain the second-scale fusion feature 1.
[0068] The second-scale fused features 1 are input into the third-scale attention-based feature extraction module, where they are first scaled. The third-scale temporal reciprocal self-attention module then extracts features representing low-resolution video segments from the second-scale fused features 1 to obtain first-scale attention features 3. The third-scale self-attention module then extracts features representing high-resolution keyframes from the second-scale fused features 1 to obtain second-scale attention features 3. The third-scale parallel alignment module then aligns and fuses the first and second attention features 3 with the optical flow output by the optical flow estimation network at the third scale to obtain third-scale fused features. The scales corresponding to the first, second, and third scales decrease in order.
[0069] The third-scale fusion feature is input into the second second-scale attention-based feature extraction module, and the scale is transformed first. Then, the features representing the low-resolution video segments in the third-scale fusion feature are extracted through the second-scale temporal reciprocal self-attention module to obtain the first attention feature 4. The features representing the high-resolution key frames in the third-scale fusion feature are extracted through the second-scale self-attention module to obtain the second attention feature 4. Then, the first attention feature 4, the second attention feature 4 and the optical flow output by the optical flow estimation network at the second scale are used through the second-scale parallel alignment module to align and fuse the features to obtain the second-scale fusion feature 2.
[0070] The first superimposed feature is obtained by superimposing the second-scale fused feature 2 and the second-scale fused feature 1, and then input into the second first-scale attention-based feature extraction module. The scale is first transformed. Then, the first-scale temporal reciprocal self-attention module is used to extract features representing the low-resolution video segments from the first superimposed feature to obtain the first attention feature 5. The first-scale self-attention module is used to extract features representing the high-resolution keyframes from the first superimposed feature to obtain the second attention feature 5. The first-scale parallel alignment module then uses the first attention feature 5, the second attention feature 5, and the optical flow output by the optical flow estimation network at the first scale to align and fuse the features to obtain the first-scale fused feature 2. The processing process of each parallel alignment module is the same or similar to steps S61 to S64.
[0071] The first-scale fusion feature 1 and the first-scale fusion feature 2 are superimposed to obtain the second superimposed feature, which is then input into the target time reciprocal self-attention module. The second superimposed feature is further refined by the target time reciprocal self-attention module to obtain the deep features of each low-resolution video segment; the deep features of each low-resolution video segment and the shallow features of each low-resolution video segment are then added and upsampled (such as Figure 2), obtain the intermediate features of each low-resolution video segment; finally, superimpose the intermediate features of each low-resolution video segment with the upsampling result obtained by upsampling each low-resolution video segment to obtain the reconstructed high-resolution video segment, that is, the output of the key object guided video reconstruction model.
[0072] exist Figure 2 The left side of the figure below is a schematic diagram of the feature extraction module. The feature extraction module includes a temporal reciprocal self-attention module and a self-attention module. The temporal reciprocal self-attention module is applied to low-resolution video frames to extract features of low-resolution frames, while the self-attention module is applied to high-resolution keyframes after key object extraction to extract features of high-resolution frames.
[0073] exist Figure 2 The right part below is a schematic diagram of the parallel alignment module. The parallel alignment module is used to align and fuse features using the deformable convolution guided by optical flow. Specifically, the parallel alignment module can use the deformable convolution module guided by optical flow to align and fuse the extracted key frame features (such as the second attention feature corresponding to the high-resolution key frame after the two key objects corresponding to the j-th low-resolution video frame are extracted). and ) and the low-resolution feature (the first attention feature corresponding to the j-th low-resolution video frame , the first attention features corresponding to the j-1th low-resolution video frame and the j+1th low-resolution video frame respectively and ) to align and obtain the aligned features (i.e. Figure 2 ).
[0074] In this way, in this embodiment, on the end side, HR videos (high-resolution video frames) are downsampled, and key frames are periodically selected at fixed intervals. On the edge side, object detection is performed to extract key objects from key frames and remove semantically irrelevant background areas. On the cloud side, a key object-guided video reconstruction model is used to reconstruct high-resolution video frames in an object-aware manner. It simultaneously leverages spatial priors (information provided by high-resolution key frames after key object extraction) and spatiotemporal context (information provided by different low-resolution video frames). This effectively protects the semantic integrity of key target areas under limited bitrate conditions, demonstrating the powerful versatility of this embodiment and its practical potential for deployment in low-bandwidth surveillance scenarios.
[0075] In combination with the above embodiments, in one embodiment, the present invention further provides a cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution. In this method, the training process of the key object-guided video reconstruction model may include steps S71 to S74: Step S71: The terminal device collects a sample high-resolution video frame sequence, determines a plurality of sample high-resolution key frames from the sample high-resolution video frame sequence, divides the sample high-resolution video sequence between every two sample high-resolution key frames into a sample high-resolution video segment, downsamples each sample high-resolution video segment to obtain a sample low-resolution video segment, and sends each sample low-resolution video segment and the corresponding two sample high-resolution key frames to the edge server; and sends the sample high-resolution video frame sequence to the cloud server.
[0076] In this embodiment, the terminal device may be a surveillance camera device, for example, a remote sensing satellite, an unmanned aerial vehicle, a ground tower, or other platform device. The terminal device may collect a sample high-resolution video frame sequence, which is a high-resolution video frame sequence used for model training. Multiple sample high-resolution key frames may be determined from the sample high-resolution video frame sequence. The sample high-resolution video sequence between each two adjacent sample high-resolution key frames may be divided into a sample high-resolution video segment according to the order of the multiple sample high-resolution key frames in the sample high-resolution video frame sequence. A sample high-resolution video segment includes two adjacent sample high-resolution key frames and the sample high-resolution video sequence between the two adjacent sample high-resolution key frames.
[0077] Each sample high-resolution video segment is downsampled to obtain a sample low-resolution video segment. Each sample low-resolution video segment and its corresponding two sample high-resolution key frames are sent to the edge server, and the sample high-resolution video frame sequence is sent to the cloud server. In an optional example, the sample high-resolution video frame sequence can be sent to the cloud server in the form of multiple sample high-resolution video segments.
[0078] Step S72: The edge server extracts key objects from multiple sample high-resolution key frames respectively, removes semantically irrelevant background information in each sample high-resolution key frame, and retains the semantic information of the sample key object area, to obtain multiple sample high-resolution key frames after key objects are extracted, and sends each sample low-resolution video segment and the corresponding two sample high-resolution key frames after key objects are extracted as a training sample to the cloud server.
[0079] In this embodiment, the edge server can perform key object extraction on each of the received sample high-resolution key frames. For each sample high-resolution key frame, key object extraction removes semantically irrelevant background information from each sample high-resolution key frame, while retaining the semantic information of the key object region within each sample high-resolution key frame. This results in multiple sample high-resolution key frames after key object extraction. The edge server then sends each sample low-resolution video segment and the two sample high-resolution key frames corresponding to that sample low-resolution video segment after key object extraction as a training sample to the cloud server.
[0080] Step S73: The cloud server obtains a reconstructed sample high-resolution video sequence based on the multiple training samples through the model to be trained.
[0081] In this embodiment, the cloud server can process the multiple training samples received through the model to be trained to obtain multiple reconstructed sample high-resolution video segments, and thus obtain a reconstructed sample high-resolution video sequence corresponding to the sample high-resolution video sequence based on the multiple reconstructed sample high-resolution video segments.
[0082] Step S74: The cloud server updates the model parameters of the to-be-trained model according to the sample high-resolution video sequence and the reconstructed sample high-resolution video sequence to obtain the key object-guided video reconstruction model.
[0083] In this embodiment, the cloud server can update the model parameters of the to-be-trained model based on the sample high-resolution video sequence and the reconstructed sample high-resolution video sequence corresponding to the sample high-resolution video sequence until the to-be-trained model converges, thereby obtaining a key object-guided video reconstruction model. In an optional example, a loss calculation can be performed based on a reconstructed sample high-resolution video segment obtained by the to-be-trained model for each training sample, as well as the sample high-resolution video segment corresponding to the training sample, to obtain a reconstruction loss value. The model parameters of the to-be-trained model are then updated based on the reconstruction loss value until the to-be-trained model converges, thereby obtaining a key object-guided video reconstruction model.
[0084] It should be noted that for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0085] Based on the same inventive concept, an embodiment of the present invention provides a cloud-edge-device collaborative surveillance video transmission system based on key semantics-guided video super-resolution. Figure 3 , Figure 3 This is a structural block diagram of a cloud-edge-end collaborative surveillance video transmission system based on key semantics-guided video super-resolution provided by an embodiment of the present invention. Figure 3 As shown, the video transmission system includes: a terminal device, an edge server and a cloud server; The terminal device is configured to capture a high-resolution video frame sequence, determine a plurality of high-resolution key frames from the high-resolution video frame sequence, divide the high-resolution video sequence between every two high-resolution key frames into a high-resolution video segment, downsample each high-resolution video segment to obtain a low-resolution video segment, and send each low-resolution video segment and the corresponding two high-resolution key frames to the edge server; The edge server is configured to extract key objects from each of the multiple high-resolution key frames, remove semantically irrelevant background information from each high-resolution key frame, and retain semantic information of the key object region to obtain multiple high-resolution key frames after key object extraction. The edge server is configured to send each low-resolution video segment and the corresponding two high-resolution key frames after key object extraction to the cloud server. The cloud server is used to reconstruct a video sequence based on a key object-guided video reconstruction model using each low-resolution video segment and the corresponding high-resolution key frames extracted from the two key objects.
[0086] Optionally, the edge server is further configured to detect the quality of the communication network with the cloud server; The edge server is specifically used to extract key objects from multiple high-resolution key frames when the quality of the communication network is lower than the target threshold, so as to send each low-resolution video segment and the corresponding two high-resolution key frames after key object extraction to the cloud server.
[0087] Optionally, the key object guided video reconstruction model includes at least a shallow feature extraction module; the cloud server is specifically configured to: Extracting shallow features of each low-resolution video segment and the corresponding high-resolution key frame after the two key objects are extracted by the shallow feature extraction module; For the j-th low-resolution video frame in each low-resolution video segment, the cloud server calculates the optical flow between the j-th low-resolution video frame and its adjacent low-resolution video frames at multiple scales through the optical flow estimation network, and calculates the optical flow between the j-th low-resolution video frame and the corresponding two high-resolution key frames after key object extraction at multiple scales, where j is an integer greater than 0; By utilizing shallow features of high-resolution key frames extracted from each low-resolution video segment and the corresponding two key objects, combined with the optical flow output by the optical flow estimation network, attention-based feature extraction, alignment and fusion are performed at multiple scales to obtain the reconstructed high-resolution video sequence.
[0088] Optionally, the key object guided video reconstruction model further includes: a multi-scale temporal reciprocal self-attention module and a multi-scale self-attention module; the cloud server is specifically configured to: For shallow features of each low-resolution video frame in each low-resolution video segment, performing attention-based feature extraction through the multiple-scale temporal reciprocal self-attention modules to obtain first attention features at multiple scales; For the shallow features of the high-resolution key frame after each key object is extracted, attention-based feature extraction is performed through the self-attention modules of multiple scales to obtain second attention features of multiple scales.
[0089] Optionally, the key object guided video reconstruction model further includes: a target temporal reciprocal self-attention module and a multi-scale parallel alignment module; the cloud server is specifically configured to: By using the parallel alignment module of the multiple scales, the first attention feature and the second attention feature are combined with the optical flow output by the optical flow estimation network to obtain features processed at multiple scales; Processing the features processed at the multiple scales through the target temporal reciprocal self-attention module to obtain deep features of each low-resolution video segment; The deep features and shallow features of each low-resolution video segment are added and upsampled to obtain the intermediate features of each low-resolution video segment; The reconstructed high-resolution video sequence is obtained according to the intermediate features of each low-resolution video segment and the upsampling result of each low-resolution video segment.
[0090] Optionally, the cloud server is specifically used to: Performing a spatial deformation operation using the first attention feature and the second attention feature in combination with the optical flow output by the optical flow estimation network through the parallel alignment module at multiple scales to obtain processed low-resolution video frame features and processed high-resolution key frame features; The parallel alignment module at multiple scales utilizes the processed low-resolution video frame features, the processed high-resolution key frame features, and the first attention features, combined with the optical flow output by the optical flow estimation network, to obtain an optical flow residual and a modulation mask of a deformable convolution; Performing a deformable convolution operation on the first attention feature, the second attention feature, the residual of the optical flow, and the modulation mask of the deformable convolution in combination with the optical flow output by the optical flow estimation network to obtain aligned features through the parallel alignment module of the multiple scales; The aligned features and the first attention features are fused to obtain the features processed at multiple scales.
[0091] Optionally, the terminal device is further configured to collect a sample high-resolution video frame sequence, determine a plurality of sample high-resolution key frames from the sample high-resolution video frame sequence, divide the sample high-resolution video sequence between every two sample high-resolution key frames into a sample high-resolution video segment, downsample each sample high-resolution video segment to obtain a sample low-resolution video segment, and send each sample low-resolution video segment and the corresponding two sample high-resolution key frames to the edge server; and send the sample high-resolution video frame sequence to the cloud server; The edge server is further configured to extract key objects from the multiple sample high-resolution key frames, remove semantically irrelevant background information from each sample high-resolution key frame, and retain semantic information of the sample key object region to obtain multiple sample high-resolution key frames after key object extraction. Each sample low-resolution video segment and the corresponding two sample high-resolution key frames after key object extraction are sent to the cloud server as a training sample. The cloud server is also used to obtain a reconstructed sample high-resolution video sequence based on multiple training samples through the model to be trained; based on the sample high-resolution video sequence and the reconstructed sample high-resolution video sequence, the model parameters of the model to be trained are updated to obtain the key object guided video reconstruction model.
[0092] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the cloud-edge-end collaborative monitoring video transmission method based on key semantic-guided video super-resolution as described in any of the above embodiments of the present invention.
[0093] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, such as Figure 4 As shown, Figure 4 This is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed, the processor implements the steps of the cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution as described in any of the above embodiments of the present invention.
[0094] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0095] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0096] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, embodiments of the present invention may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0097] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0098] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0100] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0101] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0102] The above is a detailed introduction to the cloud-edge-end collaborative surveillance video transmission method, system, equipment and medium based on key semantic-guided video super-resolution provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution, characterized by: Applied to a video transmission system, the video transmission system includes: a terminal device, an edge server and a cloud server, the method includes: The terminal device collects a high-resolution video frame sequence, determines a plurality of high-resolution key frames from the high-resolution video frame sequence, divides the high-resolution video sequence between every two high-resolution key frames into a high-resolution video segment, downsamples each high-resolution video segment to obtain a low-resolution video segment, and sends each low-resolution video segment and the corresponding two high-resolution key frames to the edge server; The edge server extracts key objects from each of the multiple high-resolution key frames, removes semantically irrelevant background information from each high-resolution key frame, and retains the semantic information of the key object area to obtain multiple high-resolution key frames after key object extraction. The edge server sends each low-resolution video segment and the corresponding two high-resolution key frames after key object extraction to the cloud server. The cloud server reconstructs each low-resolution video segment and the corresponding high-resolution key frames extracted from the two key objects based on a key-object-guided video reconstruction model to obtain a reconstructed high-resolution video sequence.
2. The cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution according to claim 1 is characterized in that: The method further comprises: The edge server detects the quality of the communication network between the edge server and the cloud server; The edge server extracts key objects from the multiple high-resolution key frames respectively, including: When the communication network quality is lower than the target threshold, the edge server extracts key objects from multiple high-resolution key frames respectively, and sends each low-resolution video segment and the corresponding two high-resolution key frames after key object extraction to the cloud server.
3. The cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution according to claim 1 is characterized in that: The key object guided video reconstruction model includes at least a shallow feature extraction module; the cloud server reconstructs each low-resolution video segment and the corresponding two high-resolution key frames extracted from the key objects based on the key object guided video reconstruction model to obtain a reconstructed high-resolution video sequence, including: The cloud server extracts shallow features of each low-resolution video segment and the corresponding high-resolution key frame after the two key objects are extracted through the shallow feature extraction module; For the j-th low-resolution video frame in each low-resolution video segment, the cloud server calculates the optical flow between the j-th low-resolution video frame and its adjacent low-resolution video frames at multiple scales through the optical flow estimation network, and calculates the optical flow between the j-th low-resolution video frame and the corresponding two high-resolution key frames after key object extraction at multiple scales, where j is an integer greater than 0; The cloud server uses shallow features of high-resolution key frames extracted from each low-resolution video segment and the corresponding two key objects, combined with the optical flow output by the optical flow estimation network, to perform attention-based feature extraction, alignment and fusion at multiple scales to obtain the reconstructed high-resolution video sequence.
4. The cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution according to claim 3 is characterized in that: The key object guided video reconstruction model further includes: a multi-scale temporal reciprocal self-attention module and a multi-scale self-attention module; the cloud server utilizes shallow features of each low-resolution video segment and the high-resolution key frame extracted from the corresponding two key objects to perform attention-based feature extraction at multiple scales, including: The cloud server performs attention-based feature extraction on shallow features of each low-resolution video frame in each low-resolution video segment through the multiple-scale temporal reciprocal self-attention module to obtain first attention features at multiple scales; The cloud server performs attention-based feature extraction on the shallow features of the high-resolution key frame after extracting each key object through the self-attention modules of the multiple scales to obtain second attention features of the multiple scales.
5. The method for cloud-edge-device collaborative surveillance video transmission based on key semantics-guided video super-resolution according to claim 4 is characterized in that: The key object guided video reconstruction model further includes: a target temporal reciprocal self-attention module and a multi-scale parallel alignment module; the cloud server utilizes shallow features of each low-resolution video segment and the high-resolution key frames extracted from the corresponding two key objects, combined with the optical flow output by the optical flow estimation network, to perform alignment and fusion at multiple scales to obtain the reconstructed high-resolution video sequence, including: The cloud server uses the first attention feature and the second attention feature through the parallel alignment module of the multiple scales, combined with the optical flow output by the optical flow estimation network, to obtain features processed at multiple scales; The cloud server processes the features processed at multiple scales through the target temporal reciprocal self-attention module to obtain deep features of each low-resolution video segment; The cloud server adds and upsamples the deep features and shallow features of each low-resolution video segment to obtain intermediate features of each low-resolution video segment; The cloud server obtains the reconstructed high-resolution video sequence according to the intermediate features of each low-resolution video segment and the upsampling result of each low-resolution video segment.
6. The cloud-edge-device collaborative surveillance video transmission method based on key semantics-guided video super-resolution according to claim 5 is characterized in that: The cloud server uses the first attention feature and the second attention feature through the parallel alignment module of the multiple scales, combined with the optical flow output by the optical flow estimation network, to obtain features processed at multiple scales, including: The cloud server uses the first attention feature and the second attention feature in combination with the optical flow output by the optical flow estimation network through the parallel alignment module of the multiple scales to perform a spatial deformation operation to obtain a processed low-resolution video frame feature and a processed high-resolution key frame feature; The cloud server uses the processed low-resolution video frame features and the processed high-resolution key frame features, as well as the first attention features, through the parallel alignment module at multiple scales, and combines the optical flow output by the optical flow estimation network to obtain an optical flow residual and a modulation mask of a deformable convolution; The cloud server uses the parallel alignment module at multiple scales to perform a deformable convolution operation on the first attention feature, the second attention feature, the residual of the optical flow, and the modulation mask of the deformable convolution, combined with the optical flow output by the optical flow estimation network, to obtain aligned features; The cloud server fuses the aligned features and the first attention features to obtain the features processed at multiple scales.
7. The method for cloud-edge-device collaborative surveillance video transmission based on key semantics-guided video super-resolution according to any one of claims 1 to 6, characterized in that: The training process of the key object guided video reconstruction model includes: The terminal device collects a sample high-resolution video frame sequence, determines a plurality of sample high-resolution key frames from the sample high-resolution video frame sequence, divides the sample high-resolution video sequence between every two sample high-resolution key frames into a sample high-resolution video segment, downsamples each sample high-resolution video segment to obtain a sample low-resolution video segment, and sends each sample low-resolution video segment and the corresponding two sample high-resolution key frames to the edge server; and sends the sample high-resolution video frame sequence to the cloud server; The edge server extracts key objects from the multiple sample high-resolution key frames, removes semantically irrelevant background information from each sample high-resolution key frame, and retains the semantic information of the sample key object area to obtain multiple sample high-resolution key frames after key object extraction. Each sample low-resolution video segment and the corresponding two sample high-resolution key frames after key object extraction are sent to the cloud server as a training sample; The cloud server obtains a reconstructed sample high-resolution video sequence based on the plurality of training samples through the to-be-trained model; The cloud server updates the model parameters of the model to be trained according to the sample high-resolution video sequence and the reconstructed sample high-resolution video sequence to obtain the key object guided video reconstruction model.
8. A cloud-edge-device collaborative surveillance video transmission system based on key semantics-guided video super-resolution, characterized by: The video transmission system includes: a terminal device, an edge server and a cloud server; The terminal device is configured to capture a high-resolution video frame sequence, determine a plurality of high-resolution key frames from the high-resolution video frame sequence, divide the high-resolution video sequence between every two high-resolution key frames into a high-resolution video segment, downsample each high-resolution video segment to obtain a low-resolution video segment, and send each low-resolution video segment and the corresponding two high-resolution key frames to the edge server; The edge server is configured to extract key objects from each of the multiple high-resolution key frames, remove semantically irrelevant background information from each high-resolution key frame, and retain semantic information of the key object region to obtain multiple high-resolution key frames after key object extraction. The edge server is configured to send each low-resolution video segment and the corresponding two high-resolution key frames after key object extraction to the cloud server. The cloud server is used to reconstruct a video sequence based on a key object-guided video reconstruction model using each low-resolution video segment and the corresponding high-resolution key frames extracted from the two key objects.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the cloud-edge-end collaborative surveillance video transmission method based on key semantic-guided video super-resolution as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cloud-edge-end collaborative surveillance video transmission method based on key semantic-guided video super-resolution as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Video compression method and device based on key frame guidance super-resolution
CN112019861A
Cloud-side collaborative video monitoring method, device and system
CN116264610A
End-cloud combined super-resolution video reconstruction method and system based on key frame
CN116523758A
Extremely low-bit-rate security scene monitoring video encoding and decoding method and system
CN116634178A
Multi-attention-based video super-resolution reconstruction network construction method and application thereof
CN116993585A
Cited By
Video stream data processing method and device, equipment and storage medium
CN120812316A
A method, apparatus, device and storage medium for processing video stream data.
CN120812316B