A security video monitoring method and system based on cloud-edge collaborative computing

CN122476180BActive Publication Date: 2026-09-15CHENGDU QUDIAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610978435.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-15
Estimated Expiration
2046-07-02

AI Technical Summary

Technical Problem

[0003]本申请提供一种基于云边协同计算的安防视频监控方法及系统,解决了现有技术存在资源受限场景下因缺乏内容感知能力而导致高价值安防视频数据处理质量显著降低的技术问题

Benefits of technology

[0016] This application addresses the technical problems of existing technologies where computation scheduling and transmission encoding are disconnected and lack content awareness by constructing a video semantic importance assessment mechanism at the edge node side and using the quantified semantic weights as core decision factors to generate a joint scheduling strategy that includes computation task offloading granularity and data encoding parameters. The core principle lies in using semantic weights to establish a unified priority scale across levels, enabling the system to automatically identify and lock high-value security event areas such as intrusion targets and abnormal behaviors when resources are limited. This allows for precise allocation of limited edge computing power and uplink bandwidth to these areas: for high semantic weight areas, fine-grained offloading and high-bitrate ROI encoding are used to ensure the accuracy and visual quality of security target recognition; for low semantic weight areas, coarse-grained offloading and low-bitrate background compression are used to release resources. This content-value-based differentiated resource allocation mechanism significantly improves the accuracy of identifying key security events and the real-time response capability in security scenarios compared to the traditional blind adjustment method that relies solely on load or bandwidth, achieving global optimization of security video data processing performance under a cloud-edge collaborative architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122476180B_ABST
    Figure CN122476180B_ABST
Patent Text Reader

Abstract

The application provides a security video monitoring method and system based on cloud edge collaborative computing, relates to the technical field of image communication, and solves the technical problem that in the prior art, in a resource-limited scene, high-value security video data processing quality is significantly reduced due to the lack of content perception capability. The method specifically comprises the following steps: acquiring a video data stream; performing semantic importance evaluation on the video data stream to obtain a semantic weight; generating a joint scheduling strategy based on the semantic weight and a current node resource state; the joint scheduling strategy comprises a calculation task offloading granularity and a data coding parameter; the current node resource state comprises a local available calculation resource, a cloud available calculation resource and a network transmission bandwidth; and performing calculation task offloading and data coding on the video data stream according to the joint scheduling strategy. The application is used for security video monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image communication technology, and in particular to a security video surveillance method and system based on cloud-edge collaborative computing. Background Technology

[0002] In cloud-edge collaborative security video surveillance applications, existing data processing architectures typically manage computational task scheduling and video encoding and transmission as two independent control loops. This means the computing side only offloads tasks based on node load, and the transmission side only adjusts the encoding bitrate based on network bandwidth. However, this fragmented control mechanism ignores the semantic value differences within the video content itself. Consequently, in resource-constrained scenarios such as network congestion or limited edge computing power, the system cannot distinguish between critical security events like intrusion targets and abnormal behavior and background noise. It often applies a uniform degradation strategy to the entire screen, resulting in decreased accuracy in identifying high-value security targets or delayed response to critical security events. Therefore, existing technologies suffer from a significant reduction in the processing quality of high-value security video data in resource-constrained scenarios due to a lack of content awareness. Summary of the Invention

[0003] This application provides a security video surveillance method and system based on cloud-edge collaborative computing, which solves the technical problem that the lack of content awareness capabilities in resource-constrained scenarios leads to a significant reduction in the processing quality of high-value security video data in existing technologies.

[0004] To achieve the above objectives, this application adopts the following technical solution: Firstly, a method is provided, comprising: acquiring a video data stream; evaluating the semantic importance of the video data stream to obtain semantic weights; generating a joint scheduling strategy based on the semantic weights and the current node resource status; the joint scheduling strategy includes computation task offloading granularity and data encoding parameters; the current node resource status includes locally available computing resources, cloud-available computing resources, and network transmission bandwidth; and performing computation task offloading and data encoding on the video data stream according to the joint scheduling strategy.

[0005] The above-mentioned solution introduces a video semantic importance assessment mechanism at the edge and uses the assessed semantic weights as the core decision factors. This breaks the traditional situation where computation scheduling and transmission coding are independent of each other. The system can dynamically generate a joint strategy that includes task offloading granularity and coding parameters based on the actual value of the video content. This allows the system to prioritize the processing quality and transmission efficiency of high semantic value areas under limited resource conditions. It solves the technical problem that existing technologies suffer from a significant reduction in the processing quality of high-value security video data due to a lack of content awareness in resource-constrained scenarios.

[0006] In conjunction with the first aspect mentioned above, in one possible implementation, semantic importance assessment of the video data stream is performed to obtain semantic weights, including: extracting visual features of video frames in the video data stream; identifying target entities in the video frames and the image regions where the target entities are located based on the visual features; and calculating the semantic weights of each image region according to the entity type of the target entity and the regional attributes of the image region.

[0007] In conjunction with the first aspect mentioned above, in one possible implementation, a joint scheduling strategy is generated based on semantic weights and the current node resource status, including: determining the target processing priority of each image region based on semantic weights; allocating the corresponding computing task offloading granularity according to the target processing priority, local available computing resources, cloud available computing resources, and network transmission bandwidth; and allocating the corresponding data encoding parameters according to the target processing priority, local available computing resources, cloud available computing resources, and network transmission bandwidth.

[0008] In conjunction with the first aspect above, in one possible implementation, the semantic weights of each image region are calculated based on the entity type of the target entity and the region attributes of the image region. This includes: obtaining preset entity type base weights and region location base weights; detecting the motion state of the target entity and generating motion state compensation coefficients based on the motion state, where the motion state includes at least one of motion speed and motion direction; and calculating the semantic weights of each image region according to the following formula:

[0009] in, For semantic weights, As a semantic foundation component, =α× +β× , As the basic weight for entity type, As the basic weight for regional location, Let α and β be the motion state compensation coefficients. and The corresponding preset weighting coefficient, λ is the preset motion modulation intensity parameter.

[0010] In conjunction with the first aspect mentioned above, in one possible implementation, a corresponding computational task offloading granularity is allocated based on the target processing priority, locally available computing resources, cloud-available computing resources, and network transmission bandwidth. This includes: when the target processing priority is higher than a preset priority threshold and the network transmission bandwidth meets a preset bandwidth condition, a feature-level offloading granularity or a model-level offloading granularity is allocated as the computational task offloading granularity; when the target processing priority is lower than or equal to the preset priority threshold, or the network transmission bandwidth does not meet the preset bandwidth condition, a video frame-level offloading granularity is allocated as the computational task offloading granularity; wherein, the feature-level offloading granularity indicates that the extracted feature data is offloaded to the cloud, the model-level offloading granularity indicates that the intermediate layer output data of the neural network model is offloaded to the cloud, and the video frame-level offloading granularity indicates that the original video frame or compressed video frame is offloaded to the cloud.

[0011] In conjunction with the first aspect above, in one possible implementation, corresponding data encoding parameters are allocated based on the target processing priority, locally available computing resources, cloud-available computing resources, and network transmission bandwidth. This includes: identifying image regions with a target processing priority higher than a preset priority threshold as regions of interest, and allocating a first bitrate and a first resolution to the regions of interest; identifying image regions with a target processing priority lower than or equal to the preset priority threshold as background regions, and allocating a second bitrate and a second resolution to the background regions; wherein the first bitrate is greater than the second bitrate, and the first resolution is greater than or equal to the second resolution.

[0012] In conjunction with the first aspect mentioned above, in one possible implementation, the video data stream is subjected to computational task offloading and data encoding according to a joint scheduling strategy, including: dividing the video data stream into local processing data and cloud offloading data according to the granularity of computational task offloading; performing inference computation on the local processing data at the edge node to obtain local computation results; encoding the cloud offloading data according to data encoding parameters to obtain an encoded data stream, and sending the encoded data stream to the cloud through a transmission channel.

[0013] In conjunction with the first aspect mentioned above, in one possible implementation, the encoded data stream is sent to the cloud via a transmission channel, including: real-time monitoring of the network connectivity status of the transmission channel; when an interruption is detected in the network connectivity status, the encoded data stream is cached in the local storage of the edge node, and a degradation processing strategy is triggered, which includes reducing the sampling rate of subsequent video frames or reducing the encoding resolution; when the network connectivity status is detected to be restored to a connection, the cached encoded data stream is incrementally retransmitted to the cloud based on the timestamp or data version number.

[0014] In conjunction with the first aspect above, in one possible implementation, after acquiring the video data stream, the method further includes: performing edge-side preprocessing on the video data stream, wherein the edge-side preprocessing includes at least one of video denoising, dehazing, super-resolution reconstruction, video frame sampling, and keyframe extraction.

[0015] Secondly, a security video surveillance system based on cloud-edge collaborative computing is provided, comprising: a data acquisition module, a semantic evaluation module, a joint scheduling module, and a collaborative execution module; the data acquisition module is used to acquire video data streams; the semantic evaluation module is used to evaluate the semantic importance of the video data streams and obtain semantic weights; the joint scheduling module is used to generate a joint scheduling strategy based on the semantic weights and the current node resource status; the joint scheduling strategy includes the granularity of computation task offloading and data encoding parameters; the current node resource status includes locally available computing resources, cloud-available computing resources, and network transmission bandwidth; the collaborative execution module is used to perform computation task offloading and data encoding on the video data streams according to the joint scheduling strategy.

[0016] This application addresses the technical problems of existing technologies where computation scheduling and transmission encoding are disconnected and lack content awareness by constructing a video semantic importance assessment mechanism at the edge node side and using the quantified semantic weights as core decision factors to generate a joint scheduling strategy that includes computation task offloading granularity and data encoding parameters. The core principle lies in using semantic weights to establish a unified priority scale across levels, enabling the system to automatically identify and lock high-value security event areas such as intrusion targets and abnormal behaviors when resources are limited. This allows for precise allocation of limited edge computing power and uplink bandwidth to these areas: for high semantic weight areas, fine-grained offloading and high-bitrate ROI encoding are used to ensure the accuracy and visual quality of security target recognition; for low semantic weight areas, coarse-grained offloading and low-bitrate background compression are used to release resources. This content-value-based differentiated resource allocation mechanism significantly improves the accuracy of identifying key security events and the real-time response capability in security scenarios compared to the traditional blind adjustment method that relies solely on load or bandwidth, achieving global optimization of security video data processing performance under a cloud-edge collaborative architecture.

[0017] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0018] Figure 1 A flowchart illustrating a security video surveillance method based on cloud-edge collaborative computing, provided for an embodiment of this application; Figure 2 A flowchart illustrating another security video surveillance method based on cloud-edge collaborative computing provided in Embodiment 2 of this application; Figure 3 This is a flowchart illustrating another security video surveillance method based on cloud-edge collaborative computing provided in Embodiment 3 of this application; Figure 4 This is a flowchart illustrating another security video surveillance method based on cloud-edge collaborative computing provided in Embodiment 4 of this application; Figure 5 This is a system architecture diagram of a security video surveillance system based on cloud-edge collaborative computing, provided for an embodiment of this application. Detailed Implementation

[0019] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0020] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0021] Example 1: like Figure 1 As shown, this embodiment provides a cloud-edge collaborative video data processing method applied to edge nodes. This method achieves efficient processing and transmission of video data through content awareness and resource collaboration at the edge, specifically including the following steps.

[0022] S101, Obtain video data stream.

[0023] Among them, video data stream refers to a continuous sequence of images generated by the front-end acquisition device and transmitted to the edge node. It can be an unprocessed raw bitstream or a standard video stream that has undergone preliminary compression.

[0024] In this embodiment, after acquiring the video data stream, edge-side preprocessing can be performed on the video data stream. Edge-side preprocessing includes at least one of video denoising, dehazing, super-resolution reconstruction, video frame sampling, and keyframe extraction. Specifically, after receiving the video data stream, one or more preprocessing operators can be selectively activated according to the current environmental conditions or preset configuration. For example, when the image visibility is detected to be below a threshold, a dehazing algorithm module is automatically loaded to enhance the video frames; or when computing power is limited, redundant static background frames are filtered out by a keyframe extraction algorithm, retaining only frames containing dynamic changes or specific events for subsequent processing. This preprocessing mechanism, as an optional pre-enhancement method, can be flexibly adjusted according to the actual scenario requirements, rather than being a fixed and unchanging necessary step.

[0025] It should be noted that the purpose of edge-side preprocessing is to improve the quality of input data or reduce the total amount of data, thereby providing a clearer and more compact data foundation for subsequent semantic evaluation. However, its specific implementation is not limited to the types listed above, and any image processing technology that can improve the usability of video data can be applied here.

[0026] Based on the above steps, by introducing a flexible edge-side preprocessing mechanism in the data acquisition stage, the environmental adaptability and information density of the video data stream are effectively improved, avoiding subsequent semantic misjudgment or resource waste caused by low quality of raw data, and ensuring the reliability of the input end of the entire processing link.

[0027] S102. Perform semantic importance assessment on the video data stream to obtain semantic weights.

[0028] Semantic weight is a quantitative numerical indicator used to characterize the semantic importance of each image region in a video data stream. It reflects the value density of information contained in different regions of the video image for specific business objectives. High-weight regions usually correspond to people, vehicles, objects, or abnormal events that need to be focused on, while low-weight regions are mostly background or irrelevant interference.

[0029] In this embodiment, video frames from a video data stream are input into a lightweight semantic evaluation model deployed at edge nodes. Visual features are extracted, and target entities and their locations are identified. Then, by integrating multi-dimensional information such as entity type and region attributes, structured semantic data containing location coordinates and corresponding weight values ​​is calculated and generated. This process achieves an abstract transformation from low-level pixel information to high-level semantic value. The output semantic weights serve as the core basis for subsequent scheduling decisions, directly guiding the differentiated allocation of resources.

[0030] As an example, the semantic evaluation module outputs a set of metadata in JSON format, which records that there is a pedestrian target in the center area of ​​the screen, with a corresponding semantic weight of 0.85, while the semantic weight of the surrounding background area is 0.12. This metadata is passed to the downstream scheduling module along with the video frame.

[0031] Based on the above steps, by constructing a content-aware semantic evaluation process, the abstract value of video content is transformed into calculable quantitative indicators, breaking the blindness of "data equality" in traditional video processing. This provides objective and real-time decision input for subsequent precise resource allocation, enabling the system to understand the content of the video.

[0032] S103. Based on semantic weights and the current node resource status, generate a joint scheduling strategy. The joint scheduling strategy includes computation task offloading granularity and data encoding parameters. Then, according to the joint scheduling strategy, perform computation task offloading and data encoding on the video data stream.

[0033] The joint scheduling strategy refers to the coordinated control instructions that simultaneously constrain the allocation of computing resources and the allocation of transmission bandwidth under a unified priority framework. Its core lies in associating semantic weights with available local computing resources, available cloud computing resources, and network transmission bandwidth to form an optimal resource allocation scheme across levels.

[0034] In this embodiment, the semantic weights obtained in S102 and the real-time monitored node resource status are read to determine the target processing priority for each image region. Subsequently, based on this priority and resource constraints, it is dynamically determined which data should be processed locally and which should be offloaded to the cloud (i.e., the granularity of computational task offloading), and what bitrate and resolution should be used for encoding different regions (i.e., data encoding parameters). Finally, the collaborative execution module strictly follows the generated strategy instructions to divide the video data stream into locally processed data and cloud-offloaded data, performing inference computation on the local portion and encoding the cloud portion according to specified parameters before sending it. The entire process is completed in a closed loop at the edge node side; the cloud only acts as a data receiver and environmental parameter provider, without participating in strategy generation and execution.

[0035] As an example, when a region is detected to have extremely high semantic weight and the current uplink bandwidth is sufficient, the joint scheduling strategy instructs that the feature data of that region be directly offloaded to the cloud for in-depth analysis, while allocating a high bitrate for lossless encoding; when bandwidth is limited, the strategy automatically switches to offloading only keyframes and reducing the bitrate of the background region in order to maintain the transmission of core data.

[0036] Based on the above steps, by establishing a joint mapping mechanism between semantic weights and resource states, the collaborative generation and execution of computation task offloading and data encoding parameters are realized, solving the problem of fragmented scheduling of computation and transmission. This ensures that, under limited edge computing power and network bandwidth, high semantic value areas can always obtain matching processing quality and transmission efficiency, significantly improving the overall data processing efficiency under the cloud-edge collaborative architecture.

[0037] Based on the above technical solution, this embodiment constructs a complete edge-side semantic-driven video processing flow through three core steps: acquiring video streams and optionally performing edge preprocessing, performing semantic importance assessment to obtain quantized weights, and generating and executing a joint scheduling strategy based on weights and resource states. This flow not only ensures the quality of input data through preprocessing but also achieves an automated closed loop from content awareness to resource action through the close integration of semantic assessment and joint scheduling. This lays a solid logical foundation and architectural support for the detailed development of assessment algorithms, scheduling rules, and anomaly handling mechanisms in subsequent embodiments.

[0038] Example 2: In some embodiments, such as Figure 2 As shown, semantic importance is evaluated on the video data stream to obtain semantic weights, including S201, S202, S203, and S204, which are explained in detail below: S201. Extract the visual features of video frames from the video data stream.

[0039] Visual features refer to structured information vectors that can be abstracted from raw pixel data to characterize the attributes of image content. They not only include low-level statistical information such as color and brightness, but also cover mid-level semantic cues such as texture, shape, and edge distribution. They are the cornerstone of subsequent target recognition and semantic understanding.

[0040] In this embodiment, video frames are input into a lightweight convolutional neural network or Transformer encoder deployed at edge nodes. Feature maps are extracted layer by layer through multi-layer convolution and pooling operations. These feature maps compress redundant pixel data while preserving spatial structure information, forming a high-dimensional and compact feature representation. To adapt to the limited computing resources on the edge side, depthwise separable convolution or channel pruning techniques can be used to lightweight the backbone network, significantly reducing computational complexity while ensuring feature representation capabilities.

[0041] It should be noted that visual feature extraction is not limited to a single general feature extractor. Multi-branch parallel extraction strategies can also be adopted according to specific scenario requirements. For example, global scene features and local detail features can be extracted and fused at the same time to enhance adaptability to complex environments. The specific network architecture for feature extraction should not be regarded as a limitation of the present invention.

[0042] As an example, MobileNetV3 is used as the backbone network. The input video frames have a resolution of 320×320. After processing through 12 bottleneck layers, the output feature tensor has a size of 20×20×96. This tensor is the input data for the subsequent recognition module.

[0043] Based on the above steps, by efficiently extracting visual features from video frames at the edge, the initial condensation of key semantic information from massive pixel data is achieved, providing a high-quality data foundation for subsequent accurate identification of target entities and their attributes, and effectively balancing the contradiction between feature representation capability and edge computing overhead.

[0044] S202. Based on visual features, identify the target entity in the video frame and the image region where the target entity is located.

[0045] In this context, the target entity refers to an independent object in the video frame that has clear business interest value, such as pedestrians, vehicles, specific equipment, or abnormal objects, while the image region refers to the spatial range occupied by the entity on the two-dimensional image plane, usually represented in the form of a bounding box or segmentation mask. Together, they constitute the spatial anchor point for semantic evaluation.

[0046] In this embodiment, the visual features obtained in S201 are input into the target detection head or instance segmentation head. The category label of the target entity and its corresponding spatial coordinate information are output synchronously through regression and classification branches, thereby establishing an accurate mapping relationship between the entity type and the image region. For overlapping or occluded targets, non-maximum suppression or association prediction mechanisms can be introduced for post-processing optimization to ensure that each effective target can be accurately captured and given a unique region identifier.

[0047] Based on the above steps, by transforming abstract visual features into specific target entities and their spatial locations, semantic evaluation is realized from "feature space" to "physical space". This allows subsequent weight calculations to be applied precisely to the actual objects in the image, avoiding the waste of resources and semantic ambiguity caused by uniform evaluation of the entire image.

[0048] S203. Obtain the preset entity type basic weight and region location basic weight, detect the motion state of the target entity, and generate motion state compensation coefficient based on the motion state.

[0049] Among them, the entity type basic weight reflects the prior importance of different types of targets in a specific business scenario, the region location basic weight represents the difference in attention to different spatial locations in the image, and the motion state compensation coefficient is an adjustment factor generated in real time based on the target's dynamic behavior. The three together constitute the multi-dimensional decision basis for semantic weight.

[0050] In this embodiment, an entity type weight lookup table and a region location weight heatmap are pre-configured. The entity type weight lookup table uses the category label of the target entity as an index and the corresponding basic weight value as the table entry. The region location weight heatmap uses the image spatial coordinates as an index and the attention weight value of the corresponding location as the pixel value. When a target entity is identified, the basic weight of the entity type is retrieved from the lookup table using the category label of the target entity as the key. The basic weights of the region location are obtained by sampling the center coordinates of the image region where the target entity is located in the heatmap. Simultaneously, the displacement vector of the target between consecutive frames is tracked using optical flow or Kalman filtering. Specifically, for optical flow, the region where the target is located is used as a template in two adjacent frames. The displacement vector of each pixel in that region between the two frames is calculated, and the average value is taken as the inter-frame displacement vector of the target. For Kalman filtering, a linear motion model is established using the position and velocity of the target in the previous frame as state variables. The position of the target in the current frame is estimated through two recursive steps of prediction and observation update. The difference between the estimated position in the current frame and the position in the previous frame is the inter-frame displacement vector. Then, based on the displacement vectors of multiple consecutive frames, the magnitude and rate of change of the target's motion velocity are calculated. The magnitude and rate of change of the motion velocity and direction are normalized to the range of 0 to 1 and then weighted and fused to obtain the motion state compensation coefficient. The larger the coefficient, the higher the target's dynamic activity.

[0051] As an example, in a traffic monitoring scenario, the basic weight of the entity type "ambulance" is set to 0.9, and that of "ordinary car" is set to 0.4; the basic weight of the regional position in the center of the screen is 0.8, and that of the edge area is 0.3; if an ambulance is detected to be driving towards the center of the screen at a speed of 60 km / h, the generated motion state compensation coefficient is 0.85, while the coefficient for a stationary vehicle of the same type is only 0.1.

[0052] Based on the above steps, by integrating static prior knowledge with dynamic real-time perception, a multi-dimensional and adaptive semantic evaluation input system was constructed. This system can both inherit the stable guidance of business rules and keenly capture sudden dynamic events, significantly improving the sensitivity and accuracy of semantic weights in responding to changes in real scenarios.

[0053] S204. Calculate the semantic weight of each image region based on the entity type of the target entity and the region attributes of the image region.

[0054] The calculation of semantic weights is not a simple linear superposition, but rather a coupling operation between basic semantic components and dynamic motion factors through a specific nonlinear modulation mechanism, in order to achieve a significant enhancement of high-value dynamic events.

[0055] Optionally, in this embodiment of the application, the semantic weights of each image region are... Satisfy the following formula:

[0056] in, =α× +β× exp(·) is an exponential function with the natural constant e as its base; The semantic weights of the final output. As a semantic foundation component, As the basic weight for entity type, As the basic weight for regional location, Let α and β be the motion state compensation coefficients. and The corresponding preset weighting coefficients, The preset motion modulation intensity parameter; the core of this formula is to use an exponential function to modulate the basic components. The modulation intensity is determined by the product of the motion state compensation coefficient and the semantic basic components. This means that only when the target itself has high basic value and is in an active motion state will a significant weight amplification effect be triggered. For low-value targets or static high-value targets, the weight growth remains gradual.

[0057] It should be noted that the key reason for using exponential modulation instead of nonlinear weighting is to simulate the nonlinear focusing characteristics of human visual attention. In security or emergency scenarios, the information increment brought by a high-risk target moving at high speed is several times or even tens of times greater than that of a stationary target under the same conditions. Linear models cannot reflect this "qualitative change" in priority, while exponential functions can precisely characterize this semantic salience that amplifies sharply with the accumulation of risk. In addition, the parameter λ can be dynamically adjusted according to the actual business sensitivity. The larger the value of λ, the stronger the amplification effect of motion on the weight, and vice versa, thus giving the system flexible scenario adaptability.

[0058] Based on the above steps, by introducing an exponential modulation mechanism and a multi-parameter coupling formula, the semantic weights are deeply integrated with the static attributes and dynamic behaviors of the target and achieve nonlinear response. This enables the system to automatically highlight the event targets that are truly valuable for emergency response in complex scenarios, overcoming the sluggishness and averaging defects of the traditional linear weighting method in the perception of emergencies. This provides a more discriminative and practical decision-making basis for subsequent joint scheduling.

[0059] Example 3: In some embodiments, such as Figure 3 As shown, a joint scheduling strategy is generated based on semantic weights and the current node resource status, including S301, S302, and S303, which are explained in detail below: S301. Based on semantic weights, determine the target processing priority for each image region.

[0060] Among them, target processing priority is a key intermediate variable that maps continuous semantic weight values ​​to discrete or hierarchical resource scheduling instructions. Its role is to transform abstract content value into a basis for system-executable queue sorting and resource preemption.

[0061] In this embodiment of the application, after reading the semantic weights of each image region calculated in Embodiment 2, a dynamic threshold division or normalized sorting algorithm is used to divide each region into multiple levels such as high priority, medium priority and low priority, or directly generate a continuous priority score from 0 to 100. This priority not only reflects the semantic importance of the target entity itself, but also implies the urgency of the current resource environment. For example, when computing power is tight, even if the semantic weights are the same, regions in motion or in key positions may be given a higher actual processing priority.

[0062] It should be noted that priority determination does not simply rely on the static mapping of semantic weights. A feedback adjustment mechanism can also be introduced to dynamically adjust the priority threshold of the current cycle based on the completion rate or delay of the task execution in the previous cycle, so as to avoid system jitter caused by long-term starvation of low-priority tasks or excessive resource consumption by high-priority tasks.

[0063] As an example, semantic weights greater than 0.8 are set as high priority, 0.3 to 0.8 as medium priority, and less than 0.3 as low priority. When a pedestrian target is detected with a semantic weight of 0.85, its area is marked as the first position in the high priority queue, while the background tree area with a weight of 0.15 is assigned to the low priority buffer queue.

[0064] Based on the above steps, by constructing a standardized conversion mechanism from semantic weights to processing priorities, a unified benchmark is provided for joint decision-making in the subsequent computation and transmission dimensions, ensuring the consistency and fairness of cross-level resource allocation.

[0065] S302. Allocate the corresponding computing task offloading granularity based on the target processing priority, local available computing resources, cloud available computing resources, and network transmission bandwidth.

[0066] Among them, the granularity of computing task offloading refers to the data abstraction level when video data is divided into tasks between edge nodes and the cloud, which directly determines the amount of data transmitted and the upper limit of the accuracy of cloud analysis.

[0067] In this embodiment, when the target processing priority is higher than a preset priority threshold and the network transmission bandwidth meets the preset bandwidth condition, feature-level offloading granularity or model-level offloading granularity is assigned as the computation task offloading granularity; when the target processing priority is lower than or equal to the preset priority threshold, or the network transmission bandwidth does not meet the preset bandwidth condition, video frame-level offloading granularity is assigned as the computation task offloading granularity. Feature-level offloading granularity indicates that the extracted feature data is offloaded to the cloud, model-level offloading granularity indicates that the intermediate layer output data of the neural network model is offloaded to the cloud, and video frame-level offloading granularity indicates that the original video frame or compressed video frame is offloaded to the cloud. Specifically, for high-priority targets, if bandwidth is sufficient, feature-level or model-level offloading is preferred because such fine-grained data has a small volume but high information density, which can save bandwidth and retain deep semantics for high-precision analysis in the cloud; while for low-priority targets or when bandwidth is limited, it falls back to video frame-level offloading. Although the data volume is large, no complex preprocessing is required on the edge side, which can reduce edge computing power consumption and ensure basic visibility.

[0068] It should be noted that the switching between the three offloading granularities is not a hard jump. In actual deployment, a hybrid granularity mode can also be supported. For example, feature-level offloading can be used for high-priority areas within the same frame, while frame-level compression transmission can be used for low-priority background areas, thereby achieving the ultimate optimization of resource utilization at the single-frame scale.

[0069] As an example, when a high-priority "fire smoke" target is identified and the current uplink bandwidth is 20Mbps (above the threshold of 10Mbps), the system chooses to upload the output tensor (model-level data) of the third convolutional layer of the ResNet50 network to the cloud for confirmation; while when the bandwidth drops to 5Mbps, it automatically switches to uploading the H.265 compressed bitstream of the keyframe.

[0070] Based on the above steps, by establishing a multi-level offloading granularity adaptive switching mechanism under the dual constraints of priority and bandwidth, a dynamic balance between computational accuracy and transmission overhead is achieved, ensuring that high-value targets can always obtain the optimal analysis path under limited resources.

[0071] S303. Allocate corresponding data encoding parameters based on the target processing priority, available local computing resources, available cloud computing resources, and network transmission bandwidth.

[0072] Among them, data encoding parameters refer to the quality control indicators such as bit rate, resolution, and frame rate used by the video encoder when processing different image regions. Its allocation strategy directly determines the visual fidelity and bandwidth utilization efficiency of the video stream.

[0073] In this embodiment, image regions with a target processing priority higher than a preset priority threshold are identified as regions of interest (ROIs), and a first bitrate and a first resolution are assigned to them. Image regions with a target processing priority lower than or equal to the preset priority threshold are identified as background regions, and a second bitrate and a second resolution are assigned to them. Specifically, the first bitrate is greater than the second bitrate, and the first resolution is greater than or equal to the second resolution. This step shares the same priority determination result as S302, but operates on the transport layer rather than the computation layer. For high-priority ROIs, the encoder is configured to use a higher quantization parameter (QP) lower bound and a finer motion estimation search range to ensure clear details. For low-priority background regions, the quantization step size is relaxed and the spatial sampling rate is reduced to significantly compress redundant information.

[0074] It should be noted that this differentiated coding configuration is not executed in isolation, but rather forms a synergistic enhancement effect with the offloading granularity. When a high-priority region adopts feature-level offloading, its corresponding ROI coding parameters can be appropriately reduced (because the feature data already carries core information). When frame-level offloading is adopted, the ROI coding quality needs to be further improved to make up for the lack of semantic information. This linkage adjustment mechanism enables the joint scheduling strategy to truly achieve the system efficiency of "1+1>2".

[0075] As an example, for the high-priority vehicle license plate area, a first bitrate of 4Mbps and a resolution of 1080P are allocated; for the low-priority road background, a second bitrate of 500Kbps and a resolution of 480P are allocated, and the two are multiplexed in the same video stream.

[0076] Based on the above steps, by accurately mapping semantic priorities to differentiated coding resource configurations, the visual quality and machine readability of key targets are significantly improved without increasing the overall bandwidth burden, verifying the effectiveness and necessity of the joint scheduling strategy in the transmission dimension.

[0077] Based on the above technical solution, this embodiment breaks through the barrier of independent optimization of computing and transmission in traditional solutions by constructing a unified priority system and synchronously driving the adaptive allocation of computing task offloading granularity and data encoding parameters. This mechanism enables the system to complete cross-level collaborative decision-making within milliseconds based on real-time changes in content value and resource status. This ensures both the processing accuracy and transmission quality of high-value targets and avoids the ineffective occupation of valuable resources by low-value areas, fully demonstrating the technical advantages of semantic-driven joint scheduling under the cloud-edge collaborative architecture.

[0078] Example 4: In some embodiments, such as Figure 4 As shown, according to the joint scheduling strategy, computational task offloading and data encoding are performed on the video data stream, including S401, S402, and S403, which are explained in detail below: S401. Based on the granularity of the computation task offloading, the video data stream is divided into local processing data and cloud offloading data, and inference computation is performed on the local processing data at the edge node to obtain the local computation result.

[0079] Data partitioning refers to assigning different data copies or references of the same video frame to different processing pipelines in memory or video memory based on the offloading granularity identifier determined in the joint scheduling strategy. Locally processed data usually refers to the original frames or intermediate features that need to be analyzed in a closed loop at the edge, while cloud-offloaded data refers to data objects that need to be uploaded to the cloud for in-depth mining or persistent storage.

[0080] In this embodiment, when the joint scheduling strategy indicates feature-level offloading, the video frame is input into the local feature extraction network, the generated feature tensor is marked as cloud offloading data and sent to the encoding queue, and the feature tensor or its corresponding original ROI region is also copied and sent to the local classifier as local processing data; when the strategy indicates frame-level offloading, the original video frame as a whole is marked as cloud offloading data, and only low-resolution thumbnails or keyframes are extracted as local processing data for lightweight event trigger detection; after the local inference calculation is completed, the output local calculation result not only includes the current identification label and confidence, but can also be used as temporal context to feed back to the semantic evaluation module of the next frame, forming a self-reinforcing closed loop of edge perception.

[0081] It should be noted that data partitioning is not a physical, forced separation. In actual system implementation, zero-copy technology or shared memory mechanism can be used to allow local processing and cloud offloading to share the same underlying data buffer, distinguishing processing paths only through metadata tags, thereby avoiding additional latency and bandwidth consumption caused by data copying.

[0082] As an example, for high-priority pedestrian targets, the strategy specifies model layer unloading. The system directly maps the output feature map (size 14×14×512) of the ResNet50 network Stage 3 to the sending buffer as cloud unloading data, while keeping the output feature map of Stage 4 in the local GPU memory to continue to perform subsequent classification head operations, and obtain the local calculation result of "pedestrian-male-backpack". The whole process does not require the CPU to participate in data transfer.

[0083] Based on the above steps, through refined data partitioning and local inference closed loop, the low latency requirement for real-time response on the edge side is guaranteed, while high-information-density input data is provided for cloud analysis, achieving optimal allocation of computing resources in the spatiotemporal dimension.

[0084] S402. Encode the data unloaded from the cloud according to the data encoding parameters to obtain an encoded data stream, and send the encoded data stream to the cloud through the transmission channel.

[0085] The implementation of data encoding parameters refers to dynamically injecting the control variables such as bit rate, resolution, and frame rate allocated for each image region in the joint scheduling strategy into the configuration interface of the video encoder, so that the encoder can perceive and execute differentiated compression strategies when processing each frame.

[0086] In this embodiment, when calling the hardware encoder or software encoding library, the corresponding quantization parameters and motion estimation range are set according to the coordinates of the region of interest and the background region defined in the strategy. For high-priority regions of interest, a smaller quantization step size is used to preserve texture details, while for low-priority background regions, the quantization step size is increased and some B-frame encoding is skipped to reduce bitrate usage. The encoded bitstream is encapsulated into a transmission packet with a timestamp and sequence number, and sent to the cloud server via a transmission channel through RTSP, HTTP-FLV, or a private TCP / UDP protocol stack. During the transmission process, a network congestion control algorithm can also be combined to dynamically fine-tune the transmission rate based on the real-time round-trip latency and packet loss rate to avoid network jitter caused by sudden traffic surges.

[0087] It should be noted that the activation of encoding parameters has frame-level or GOP-level granularity. When the joint scheduling strategy changes drastically between consecutive frames, the encoder can synchronize the new parameter configuration by inserting IDR frames or SEI messages to ensure that the cloud decoding end can correctly parse the differentiated encoded video stream and prevent decoding screen tearing or semantic misalignment caused by parameter switching.

[0088] As an example, the encoder receives a strategy instruction and assigns QP=18 and 1080P resolution to the license plate ROI area in the center of the image, and QP=38 and 480P resolution to the surrounding background area. The size of the encoded single frame bitstream is 120KB, which saves 65% of the bandwidth compared to the 350KB of uniform encoding of the whole image at 1080P (QP=28). Moreover, the license plate characters are still clearly distinguishable after being decoded in the cloud.

[0089] Based on the above steps, by accurately mapping semantically driven encoding parameters to the encoder execution unit, the visual fidelity and machine readability of key targets are significantly improved without increasing the total bandwidth overhead, verifying the effective implementation of the joint scheduling strategy at the transport layer.

[0090] S403: Monitor the network connectivity status of the transmission channel in real time, and perform buffer downgrade or incremental retransmission operations based on the connectivity status.

[0091] Among them, monitoring and handling network connectivity status is a key defense mechanism to ensure data integrity and service continuity of the cloud-edge collaborative system in extreme environments such as weak network or network outage. Its core lies in establishing an adaptive fault-tolerant link of "detection-caching-degradation-retransmission".

[0092] In this embodiment, a heartbeat detection or data packet ACK timeout mechanism is used to determine in real time whether the transmission channel is interrupted. When an interruption is detected in the network connectivity status, the encoded data stream to be sent is immediately written into the local circular buffer of the edge node, and a degradation processing strategy is triggered simultaneously. The degradation processing strategy includes reducing the sampling rate of subsequent video frames or reducing the encoding resolution, such as reducing the frame rate from 25fps to 5fps or the resolution from 1080P to 720P, in order to reduce the write pressure on local storage and extend the data recording time during the network outage. When the network connectivity status is detected to be restored, the data in the local buffer is read, sorted and deduplicated based on the timestamp or data version number, and the cached encoded data stream is incrementally retransmitted to the cloud. During the retransmission process, the sending rate can be dynamically adjusted according to the current restored bandwidth status to avoid the retransmission traffic impacting the normal business flow.

[0093] It should be noted that the execution of the degradation processing strategy should be delayed and smooth, that is, degradation should only be triggered after the network interruption lasts for more than a preset threshold (such as 2 seconds) to avoid frequent switching of image quality due to instantaneous network jitter. At the same time, the incremental retransmission logic should prioritize the transmission of high semantic weight data. Low priority background data can be retransmitted when bandwidth is sufficient or discarded according to the FIFO strategy when the local cache is full, thereby maximizing the retention rate of high-value data with limited storage space.

[0094] As an example, in a tunnel monitoring scenario, when a vehicle enters the tunnel, causing a complete loss of 4G signal, the system detects three consecutive heartbeat timeouts within 200ms. It then activates the local NVMe hard drive cache and reduces the encoding frame rate from 30fps to 10fps and the resolution to D1 format. When the vehicle leaves the tunnel and the signal is restored, the system finds that 120 frames of data are missing based on the data packet sequence number. It then prioritizes retransmitting 30 key event segments with a semantic weight higher than 0.7 at twice the speed. The remaining background frames are slowly retransmitted during idle periods in the background, ensuring that the video of the traffic accident that occurred in the tunnel is uploaded completely without blocking the subsequent real-time monitoring stream.

[0095] Based on the above steps, by constructing a network outage self-governance mechanism with the capabilities of perception, caching, degradation and intelligent retransmission, the shortcomings of traditional video transmission solutions in the event of permanent data loss or service interruption when the network is unstable are effectively overcome, and the robustness and data reliability of the cloud-edge collaborative system in complex real-world environments are significantly improved.

[0096] Example 5: This embodiment provides a cloud-edge collaborative video data processing system deployed on an edge node. This system is used to execute the cloud-edge collaborative video data processing method described in any one of embodiments 1 to 4 above. Figure 5As shown, the system's logical architecture includes a data acquisition module 501, a semantic evaluation module 502, a joint scheduling module 503, and a collaborative execution module 504. These modules interact and transmit instructions via an internal bus or shared memory mechanism, collectively forming an intelligent processing entity operating in a closed loop at the edge. It should be noted that although this embodiment describes the system as functional modules, in actual physical implementation, these modules can be integrated into a single software program running on the edge computing device's processor, or they can be split into multiple independent microservice containers managed through orchestration tools such as Kubernetes. Furthermore, some modules (such as data acquisition and encoding) can even be carried out by dedicated hardware acceleration cards.

[0097] The data acquisition module 501 is used to acquire video data streams. Specifically, this module serves as the system's input interface, supporting the access and parsing of various mainstream video transmission protocols (such as RTSP, ONVIF, GB / T28181, etc.). It can pull or receive video streams pushed from front-end acquisition devices in real time. Simultaneously, this module integrates an optional preprocessing unit, capable of performing operations such as denoising, dehazing, super-resolution reconstruction, or keyframe extraction on the original video stream according to configuration instructions. It then encapsulates the processed high-quality video frames and their metadata into standardized internal data objects and passes them to downstream modules. The data acquisition module 501 not only solves the compatibility issues of heterogeneous device access but also improves the accuracy of subsequent semantic analysis through pre-processing quality enhancement methods, ensuring the robustness of the system input.

[0098] The semantic evaluation module 502 is used to evaluate the semantic importance of the video data stream and obtain semantic weights. These semantic weights characterize the semantic importance of each image region in the video data stream. Specifically, this module incorporates a lightweight deep learning inference engine (such as TensorRT, NCNN, or OpenVINO) and loads a pre-trained object detection and semantic segmentation model. Upon receiving a video frame, the module automatically calls the inference engine to extract visual features and identify target entities. Combining a preset entity type weight table, a region location weight map, and real-time calculated motion state compensation coefficients, it calculates the semantic weight value for each region of interest according to the exponential modulation formula disclosed in Example 2. The output of this module is a set of structured semantic metadata containing region coordinates, category labels, and quantized weight values. This data directly reflects the value density of the image content, providing a precise decision-making basis for subsequent differentiated resource allocation and realizing the automated conversion from pixel-level information to semantic-level value.

[0099] The joint scheduling module 503 generates a joint scheduling strategy based on semantic weights and the current node resource status. This strategy includes the granularity of task offloading and data encoding parameters. Specifically, this module is the core of the system's decision-making process. Internally, it maintains a real-time resource monitoring probe that periodically collects status information such as local CPU / GPU utilization, memory usage, available computing power feedback from the cloud, and uplink network bandwidth. Based on this, the module maps the semantic weights output by the semantic evaluation module 502 to target processing priorities. Then, according to the matching relationship between priorities and resource status, it dynamically looks up tables or calculates to generate joint scheduling strategy instructions that include the selection of offloading granularity (feature level / model level / frame level) and encoding parameter configuration (ROI bitrate / background bitrate / resolution). By establishing a unified priority scale to simultaneously drive resource configuration in both computing and transmission dimensions, this module breaks the traditional architecture's separation of computing power scheduling and video encoding, ensuring that high-value targets always receive the optimal processing path and transmission quality under limited resources.

[0100] The collaborative execution module 504 is used to offload computational tasks and encode data from the video data stream according to the joint scheduling strategy. Specifically, this module is the final executor of the strategy, and it includes a local inference unit, a video encoding unit, and a network transmission unit. The local inference unit performs classification or detection tasks on the data that needs to be processed locally according to the offloading granularity specified by the strategy. The video encoding unit receives the differentiated encoding parameters in the strategy and performs ROI adaptive encoding on the data offloaded from the cloud. The network transmission unit is responsible for sending the encoded bitstream to the cloud and monitoring the link status in real time. When the network is interrupted, it automatically triggers local caching and degradation strategies, and performs incremental retransmission based on timestamps after the network is restored. By transforming the abstract scheduling strategy into specific hardware operation instructions and supplementing it with a robust network outage self-governance mechanism, this module ensures the service continuity and data integrity of the system under complex network environments and dynamic load conditions, enabling the cloud-edge collaborative video data processing performance to be truly realized.

[0101] Based on the above technical solution, the system provided in this embodiment fully implements a semantic-driven cloud-edge collaborative video processing method through a modular architecture. The modules are tightly coupled and work collaboratively, achieving a closed-loop process from data access, content understanding, policy generation to task execution at the edge nodes. This system not only possesses real-time perception capabilities of the value of video content but also dynamically optimizes the joint configuration of computing and transmission resources accordingly. This effectively solves the technical problem of degraded processing quality of high-value targets in resource-constrained scenarios, while maintaining the independence and integrity of edge-side entities, facilitating flexible deployment and expansion in practical security monitoring, industrial quality inspection, and other scenarios.

[0102] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0103] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.

Claims

1. A security video surveillance method based on cloud-edge collaborative computing, applied to edge nodes, characterized in that, include: Acquire video data stream; The semantic importance of the video data stream is evaluated to obtain semantic weights. This step includes: Extract the visual features of video frames from the video data stream; Based on the visual features, identify the target entity in the video frame and the image region where the target entity is located; Based on the entity type of the target entity and the region attributes of the image region, the semantic weight of each image region is calculated, including: Obtain the preset entity type base weight and region location base weight; The motion state of the target entity is detected, and a motion state compensation coefficient is generated based on the motion state, wherein the motion state includes at least one of motion speed and motion direction; The semantic weights of each image region are calculated according to the following formula: ; in, The semantic weight, As a semantic foundation component, , The basic weight for the entity type, The basic weights for the location of the region are... The motion state compensation coefficient is... , They are respectively and The corresponding preset weighting coefficients, The preset motion modulation intensity parameters; Based on the semantic weights and the current node resource status, a joint scheduling strategy is generated; the joint scheduling strategy includes the granularity of computing task offloading and data encoding parameters; the current node resource status includes locally available computing resources, cloud-available computing resources, and network transmission bandwidth; According to the joint scheduling strategy, computational task offloading and data encoding are performed on the video data stream.

2. The method according to claim 1, characterized in that, The step of generating a joint scheduling strategy based on the semantic weights and the current node resource status includes: Based on the semantic weights, the target processing priority of each image region is determined; Based on the target processing priority, the available local computing resources, the available cloud computing resources, and the network transmission bandwidth, allocate the corresponding computing task offloading granularity; The corresponding data encoding parameters are allocated based on the target processing priority, the available local computing resources, the available cloud computing resources, and the network transmission bandwidth.

3. The method according to claim 2, characterized in that, The step of allocating the corresponding computing task offloading granularity based on the target processing priority, the locally available computing resources, the cloud-available computing resources, and the network transmission bandwidth includes: When the target processing priority is higher than the preset priority threshold and the network transmission bandwidth meets the preset bandwidth condition, the feature-level offloading granularity or the model-level offloading granularity is assigned as the computation task offloading granularity. When the target processing priority is lower than or equal to the preset priority threshold, or the network transmission bandwidth does not meet the preset bandwidth condition, the video frame-level offloading granularity is assigned as the computing task offloading granularity. Specifically, the feature-level unloading granularity indicator unloads the extracted feature data to the cloud, the model-level unloading granularity indicator unloads the intermediate layer output data of the neural network model to the cloud, and the video frame-level unloading granularity indicator unloads the original video frame or compressed video frame to the cloud.

4. The method according to claim 2, characterized in that, The step of allocating corresponding data encoding parameters based on the target processing priority, the locally available computing resources, the cloud-available computing resources, and the network transmission bandwidth includes: The image regions whose target processing priority is higher than a preset priority threshold are identified as regions of interest, and a first bitrate and a first resolution are assigned to the regions of interest; Image regions whose target processing priority is lower than or equal to the preset priority threshold are identified as background regions, and a second bitrate and a second resolution are assigned to the background regions. Wherein, the first bit rate is greater than the second bit rate, and the first resolution is greater than or equal to the second resolution.

5. The method according to claim 1, characterized in that, The step of performing computational task offloading and data encoding on the video data stream according to the joint scheduling strategy includes: Based on the computational task offloading granularity, the video data stream is divided into local processing data and cloud offloading data; The locally processed data is used to perform inference calculations at the edge node to obtain local calculation results; The cloud-unloaded data is encoded according to the data encoding parameters to obtain an encoded data stream, and the encoded data stream is sent to the cloud through a transmission channel.

6. The method according to claim 5, characterized in that, Sending the encoded data stream to the cloud via a transmission channel includes: Real-time monitoring of the network connectivity status of the transmission channel; When the network connectivity is detected to be interrupted, the encoded data stream is cached in the local storage of the edge node, and a degradation processing strategy is triggered. The degradation processing strategy includes reducing the sampling rate of subsequent video frames or reducing the encoding resolution. When the network connectivity is detected to be restored, the cached encoded data stream is incrementally retransmitted to the cloud based on the timestamp or data version number.

7. The method according to claim 1, characterized in that, After acquiring the video data stream, the method further includes: The video data stream is subjected to edge-side preprocessing, which includes at least one of video denoising, dehazing, super-resolution reconstruction, video frame sampling, and keyframe extraction.

8. A security video surveillance system based on cloud-edge collaborative computing, deployed at edge nodes, characterized in that, The system includes: a data acquisition module, a semantic evaluation module, a joint scheduling module, and a collaborative execution module; The data acquisition module is used to acquire video data streams; The semantic evaluation module is used to evaluate the semantic importance of the video data stream and obtain semantic weights. This step includes: extracting visual features of video frames in the video data stream; identifying target entities in the video frames and the image regions where the target entities are located based on the visual features; calculating the semantic weights of each image region according to the entity type of the target entity and the region attributes of the image region, which includes: obtaining preset entity type base weights and region location base weights; detecting the motion state of the target entity and generating motion state compensation coefficients based on the motion state, wherein the motion state includes at least one of motion speed and motion direction; the calculation of the semantic weights of each image region satisfies the following formula: ; in, The semantic weight, As a semantic foundation component, , The basic weight for the entity type, The basic weights for the location of the region are... The motion state compensation coefficient is... , They are respectively and The corresponding preset weighting coefficients, The preset motion modulation intensity parameters; The joint scheduling module is used to generate a joint scheduling strategy based on the semantic weights and the current node resource status; the joint scheduling strategy includes the granularity of computing task offloading and data encoding parameters; the current node resource status includes locally available computing resources, cloud-available computing resources, and network transmission bandwidth; The collaborative execution module is used to perform computational task offloading and data encoding on the video data stream according to the joint scheduling strategy.

Citation Information

Patent Citations

  • Complexity and semantic perception video analysis method and system based on edge cloud collaboration

    CN119402681A

  • Video stream real-time compression method based on edge calculation

    CN122027804A