A cloud-edge fusion intelligent security video monitoring method and system
Patent Information
- Application Number
- CN202611151037.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]但是,现有智能安防视频监控技术仍存在一定不足
本发明通过构建融合时序行为分析、边缘状态感知以及事件关联推理的边缘智能分析模型,实现了对监控对象行为、场景状态以及边缘计算能力的联合分析,相较于传统仅依赖视频图像内容进行安防事件识别的方法,能够充分利用连续关键帧之间的时序关联信息,提高复杂环境下异常行为识别的准确性。具体而言,本发明通过对关键帧视觉特征与目标编号、目标类别、目标位置以及检测置信度等目标信息进行融合编码,形成能够表征目标连续行为变化过程的时序行为特征,有效降低单帧误检、目标遮挡以及短时姿态变化对事件识别结果的影响。进一步地,通过引入边缘监测上下文向量,将边缘计算设备的资源状态、网络状态以及健康状态作为场景状态评估因素,使事件分析过程能够根据当前边缘运行环境动态调整,提高边缘智能分析模型对不同计算条件的适应能力。此外,通过事件关联推理模块融合目标行为特征、场景等级特征、风险特征以及边缘设备状态信息,生成候选安防事件及对应边缘决策信息,使系统能够根据事件风险程度和边缘计算能力动态确定后续处理策略,避免高风险事件因计算资源不足导致响应延迟,同时降低低风险事件对边缘资源的占用。通过上述方式,在保证安防事件识别准确性的同时,提高智能安防系统的实时响应能力、资源调度效率以及云边协同处理能力。
Smart Images

Figure CN122824876A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of intelligent security and cloud-edge fusion, and particularly to an intelligent security video surveillance method and system that integrates cloud and edge. Background Technology
[0002] With the rapid development of smart cities, smart parks, and digital management systems, video surveillance systems have become crucial infrastructure for public safety, industrial production, traffic management, and the protection of key areas. Traditional video surveillance systems primarily rely on video acquisition terminals to continuously acquire image data from the monitored area and then centrally upload the collected video data to cloud servers for storage and analysis. However, with the increasing number of monitored areas and the large-scale application of high-definition video and intelligent analysis algorithms, the scale of video data is growing exponentially. This leads to problems such as high data transmission pressure, high computing resource consumption, and significant event response delays in the centralized cloud processing model. Especially in security scenarios with unstable network environments or high real-time requirements, relying solely on cloud analysis is insufficient for timely detection and rapid response to abnormal behavior. Therefore, introducing edge computing capabilities into video surveillance systems, enabling real-time data analysis through edge devices, and combining this with the global computing and knowledge management capabilities of the cloud to achieve cloud-edge integrated intelligent security video surveillance has become an important development direction in the current intelligent security field.
[0003] In existing technologies, research on intelligent video surveillance analysis mainly focuses on deep learning-based methods for video target detection, behavior recognition, and event analysis. For example, some existing technologies deploy convolutional neural network models in monitoring servers to perform target detection and behavior classification on acquired video images, enabling automatic identification of people, vehicles, and abnormal behaviors. Some studies further combine temporal analysis models to model the motion trajectories of targets in continuous video frames, thereby improving behavior recognition capabilities in complex environments. Furthermore, with the development of edge computing technology, some intelligent security systems deploy lightweight artificial intelligence models to cameras or edge computing nodes to perform target detection and abnormal event screening locally, reducing the pressure on video data uploads and improving response speed. Simultaneously, some cloud-edge collaborative solutions utilize large-scale cloud computing resources to further analyze abnormal events uploaded from the edge, improving event judgment capabilities through cloud model optimization and historical data learning.
[0004] However, existing intelligent security video surveillance technologies still have certain shortcomings. First, most existing video analysis-based methods only focus on information from the video images themselves, primarily judging events based on the target's appearance, location, or motion state. They lack comprehensive consideration of the resource status of edge computing devices, network status, and device health status, resulting in the analysis model's inability to dynamically adjust processing strategies according to the actual operating environment. Second, existing edge intelligent analysis methods typically only complete target detection or simple event classification, lacking joint modeling of continuous video temporal information, scene environment information, and historical monitoring knowledge. This makes them susceptible to target occlusion, changes in viewing angle, and interference from complex scenes, leading to false positives or false negatives of abnormal events. Finally, in existing cloud-edge collaborative security systems, the edge side and the cloud typically use a fixed data upload strategy, failing to dynamically allocate computing tasks based on the event risk level, edge computing capabilities, and network communication status. This results in untimely responses to high-risk events or the consumption of excessive computing and communication resources for low-risk events.
[0005] To address this issue, this invention proposes a cloud-edge fusion intelligent security video surveillance method. By constructing a data collaboration mechanism between video acquisition terminals, edge computing devices, and the cloud, it enables edge perception, edge analysis, cloud verification, and collaborative handling of security incidents. Summary of the Invention
[0006] This invention proposes a cloud-edge fusion intelligent security video surveillance method and system. Step S1 involves collecting video perception data from the monitored area and operational data from edge devices, and then standardizing both the video data and device operational data to construct a time-series video data stream of the monitored object containing behavioral information of the monitored object and operational status information of the edge devices, as well as an edge monitoring context vector. This achieves a unified representation of video perception information and edge operational status information, improving the correlation analysis capability and state description completeness of multi-source security data. It solves the technical problems in traditional security systems where video data and edge device status data are independent, data formats are inconsistent, and it is difficult to provide a unified data basis for subsequent cloud-edge collaborative analysis. Step S2 utilizes an edge intelligent analysis model to jointly analyze the time-series video data stream of the monitored object and the edge monitoring context vector, sequentially performing target behavior parsing, scene state assessment, and event correlation reasoning. Combining continuous behavioral changes of the monitored object and the real-time operational status of the edge devices, it generates a security event candidate set and edge decision information. This system enables rapid edge-side screening and real-time response to security incidents, addressing the issues of high response latency due to reliance on centralized cloud analysis in traditional video surveillance, insufficient event judgment dimensions due to video content recognition alone, and the inability of edge computing resources to participate in event analysis. Step S3 determines the cloud-edge collaborative processing strategy based on edge decision information and utilizes a cloud-based intelligent model that integrates historical monitoring knowledge to review and assess the uploaded security incident candidate set. It then generates event confirmation results and collaborative handling instructions by combining current event characteristics with historical monitoring patterns, achieving dynamic task allocation between rapid edge detection and in-depth cloud analysis. Step S4 generates intelligent security system linkage control strategies based on the collaborative handling instructions for candidate security incidents and drives video acquisition terminals, edge computing devices, and related security equipment to perform device linkage control, dynamic scheduling of monitoring resources, and closed-loop event management. This achieves full-process collaborative control from event discovery and risk confirmation to response execution, resolving the problem of separation between alarm information and device control in traditional security systems.
[0007] To achieve the above objectives, the present invention provides a cloud-edge fusion intelligent security video surveillance method, comprising the following steps: S1: Collect video sensing data and edge device operation data of the monitoring area, and perform data normalization processing on the video sensing data and edge device operation data respectively to obtain the time-series video data stream of the monitored object and the edge monitoring context vector. S2: Using the edge intelligence analysis model, target behavior analysis, scene state assessment and event correlation reasoning are performed on the time-series video data stream of the monitored object and the edge monitoring context vector to generate a security event candidate set and edge decision information; S3: Based on the edge decision information, determine the cloud-edge collaborative processing strategy. According to the cloud-edge collaborative processing strategy, use the cloud-based intelligent model that integrates historical monitoring knowledge to perform event review and risk assessment on the security event candidate set, and generate event confirmation results and collaborative handling instructions for the candidate security events in the security event candidate set. S4: Based on the collaborative handling instructions of the candidate security events, generate a linkage control strategy for the intelligent security system, and perform device linkage control, dynamic scheduling of monitoring resources, and closed-loop management of events on the intelligent security system based on the linkage control strategy.
[0008] As a further improvement of the present invention: Furthermore, the video perception data consists of a sequence of monitoring video frames collected from the monitoring area at continuously collected timestamps, and the edge device operation data is the operation status information of the edge computing device, wherein the operation status information includes the device identifier, operation status monitoring cycle, device temperature, device online status, available bandwidth, network latency, CPU utilization, GPU utilization, and memory usage of the edge computing device.
[0009] Furthermore, step S1, which involves data normalization processing of the video sensing data to obtain a time-series video data stream of the monitored object, also includes: S11: Acquire video perception data. The edge computing device performs unified encoding format, unified image resolution, and unified timestamp format processing on the monitoring video frames in the video perception data to obtain video perception data with standardized format. S12: Perform key frame extraction processing on the formatted video perception data at fixed time intervals, and sequentially perform image quality enhancement, motion region filtering and target detection processing on the extracted key frames to generate target information of the extracted key frames, wherein the target information includes target number, target category, target location and detection confidence. S13: The keyframes that have undergone image quality enhancement processing and the target information corresponding to the keyframes are spliced together to form keyframe information. S14: Sort the keyframe information according to the order in which the timestamps appear after the timestamp format is unified, and use it as the time-series video data stream of the monitored object.
[0010] Furthermore, step S1, which involves performing data normalization processing on the edge device operating data to obtain an edge monitoring context vector, also includes: S15: Based on the CPU utilization, GPU utilization, and memory usage in the edge device's operating data, calculate the resource status indicators of the edge computing device during the operating status monitoring cycle; S16: Based on the available bandwidth and network latency in the edge device operation data, normalize the available bandwidth and network latency respectively, and calculate the network status index of the edge computing device in the operation status monitoring period based on the normalized available bandwidth and network latency. S17: Normalize the device temperature in the edge device operation data, and calculate the health status index of the edge computing device during the operation status monitoring cycle based on the normalized device temperature and the online status of the device. S18: The resource status indicators, network status indicators, and health status indicators of the edge computing device during the operation status monitoring cycle are concatenated to form the edge monitoring context vector of the edge computing device during the operation status monitoring cycle.
[0011] Furthermore, the edge intelligent analysis model in step S2 is deployed on edge computing devices using a lightweight modular deployment method. The edge intelligent analysis model includes a temporal behavior encoding module, a scene state fusion module, and an event association reasoning module.
[0012] Furthermore, step S2 utilizes an edge intelligence analysis model to perform target behavior parsing, scene state assessment, and event correlation reasoning on the time-series video data stream of the monitored object and the edge monitoring context vector, and also includes: S21: The edge computing device extracts the time-series video data stream of the monitored object to be analyzed and the edge monitoring context vector in the current operating state monitoring cycle, and inputs the extracted time-series video data stream of the monitored object and the edge monitoring context vector into the edge intelligent analysis model. S22: The temporal behavior encoding module in the edge intelligent analysis model extracts visual features from keyframes and encodes the target information of the keyframes. It then correlates and fuses the visual features of the keyframes with the encoded target information to obtain temporal behavior features characterizing the continuous behavioral changes of the monitored object. An event category recognition model is used to identify the event category corresponding to the temporal behavior features. The formula for generating the temporal behavior features is: ; ; ; in, Represents temporal behavioral characteristics, This represents the behavioral characteristics of the nth keyframe in the time-series video data stream of the monitored object. N represents the total number of keyframes. Represents the visual features of the nth keyframe. This represents the encoded target information of the nth keyframe. Represents the fusion coefficient. Represents the temporal fusion coefficients. This represents the fusion result of the visual features of the nth keyframe and the encoded target information; S23: The scene state fusion module calculates the state influence weight of the edge computing device on the current monitoring object's time-series video data stream based on the edge detection context vector, and uses the state influence weight to adaptively weight and fuse the time-series behavior features to obtain a fused behavior representation. The fused behavior representation is then weighted and concatenated with the edge detection context vector as a scene state feature. S24: The event association reasoning module extracts scene state features and temporal behavior features, calculates event association values and event scores. If the event score is higher than a preset score threshold, the temporal video data stream of the monitored object and the edge monitoring context vector are used as candidate security events. Event features of the candidate security events are extracted, and edge decision information of the candidate security events is calculated based on the event features. The formula for calculating the event score is: ; in, Indicates the associated value of the event. Indicates the event rating. Represents scene state characteristics. Both represent correlation coefficients. This represents the trainable convolutional weight matrix. Represents an exponential function with the natural constant as its base; S25: Construct a security event candidate set by combining all candidate security events and their event characteristics.
[0013] Further, in step S3, determining the cloud-edge collaborative processing strategy based on the edge decision information includes: The cloud-edge collaborative processing strategy is a method for processing candidate security events. Based on the edge decision information of candidate security events in the candidate security event set, if the edge decision information is lower than a preset decision threshold, the candidate security event and its event characteristics are uploaded to the cloud for processing; otherwise, the edge computing device directly processes the candidate security event.
[0014] Furthermore, in step S3, based on the cloud-edge collaborative processing strategy, a cloud-based intelligent model integrating historical monitoring knowledge is used to perform event review and risk assessment on the security event candidate set, generating event confirmation results and collaborative handling instructions for the security events in the candidate set. This also includes: The cloud-based intelligent model, deployed in the cloud, receives candidate security events and their event characteristics uploaded to the cloud. It then uses the cloud-based intelligent model, which integrates historical monitoring knowledge, to perform event review and risk assessment on the candidate security events. The cloud-based intelligent model extracts historical event features of the monitoring areas associated with candidate security events from historical monitoring knowledge, calculates the event correlation degree between the historical event features and the event features of the candidate security events, and performs a weighted fusion of the event correlation degree and event score of the candidate security events to obtain the event credibility of the candidate security events. If the event credibility is higher than a preset credibility threshold, it means that the candidate security event has passed the event review and is marked as a valid security event; otherwise, the candidate security event is marked as an event to be observed for manual review. The formula for calculating the correlation degree of the events is: ; in, Indicates the degree of relevance of events. Indicates the characteristics of the event, Indicates the characteristics of historical events, Represents the L2 norm; Based on the event characteristics, calculate the security risk value of the effective security event; The event type, monitoring area, collection timestamp range, event credibility, and security risk value of the effective security event are concatenated to form the event confirmation result of the effective security event and are stored. Based on the security risk value, a collaborative handling instruction is generated and sent to the video acquisition terminal and edge computing device.
[0015] The present invention also provides a cloud-edge integrated intelligent security system, characterized in that the intelligent security system includes a video acquisition terminal, an edge computing device, and a cloud, to realize the technical steps of the cloud-edge integrated intelligent security video surveillance method described above.
[0016] Compared with existing technologies, this invention proposes a cloud-edge fusion intelligent security video surveillance method and system, which has the following beneficial effects: This invention constructs an edge intelligent analysis model that integrates temporal behavior analysis, edge state perception, and event correlation reasoning. This model enables joint analysis of monitored object behavior, scene state, and edge computing capabilities. Compared to traditional methods that rely solely on video image content for security event recognition, this invention fully utilizes the temporal correlation information between consecutive keyframes, improving the accuracy of abnormal behavior recognition in complex environments. Specifically, this invention fuses keyframe visual features with target information such as target number, target category, target location, and detection confidence level to form temporal behavior features that characterize the continuous behavioral changes of the target. This effectively reduces the impact of single-frame false detections, target occlusion, and short-term pose changes on event recognition results. Furthermore, by introducing an edge monitoring context vector, the resource status, network status, and health status of the edge computing device are used as scene state evaluation factors. This allows the event analysis process to dynamically adjust according to the current edge operating environment, improving the edge intelligent analysis model's adaptability to different computing conditions. Furthermore, by integrating target behavior characteristics, scene level characteristics, risk characteristics, and edge device status information through the event correlation reasoning module, candidate security events and corresponding edge decision information are generated. This enables the system to dynamically determine subsequent processing strategies based on the event risk level and edge computing capabilities, avoiding response delays caused by insufficient computing resources for high-risk events, while reducing the occupation of edge resources by low-risk events. Through these methods, the real-time response capability, resource scheduling efficiency, and cloud-edge collaborative processing capability of the intelligent security system are improved while ensuring the accuracy of security event identification. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a cloud-edge fusion intelligent security video surveillance method according to an embodiment of the present invention.
[0018] Figure 2 This is a monitoring diagram of a monitoring area provided in an embodiment of the present invention.
[0019] Figure 3 This is a structural diagram of an edge intelligent analysis model provided in an embodiment of the present invention. Detailed Implementation
[0020] The realization of the objectives, functional characteristics, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0021] This invention provides a cloud-edge integrated intelligent security video surveillance method and system. The executing entity of this cloud-edge integrated intelligent security video surveillance method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this invention: a server, a terminal, etc. In other words, the cloud-edge integrated intelligent security video surveillance method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0022] Reference Figure 1 , Figure 2 as well as Figure 3 Embodiment 1 of the present invention is as follows: A cloud-edge fusion intelligent security video surveillance method, the method comprising: S1: Collect video sensing data and edge device operation data of the monitoring area, and perform data normalization processing on the video sensing data and edge device operation data respectively to obtain the time-series video data stream of the monitored object and the edge monitoring context vector. S2: Using the edge intelligence analysis model, target behavior analysis, scene state assessment and event correlation reasoning are performed on the time-series video data stream of the monitored object and the edge monitoring context vector to generate a security event candidate set and edge decision information; S3: Based on the edge decision information, determine the cloud-edge collaborative processing strategy. According to the cloud-edge collaborative processing strategy, use the cloud-based intelligent model that integrates historical monitoring knowledge to perform event review and risk assessment on the security event candidate set, and generate event confirmation results and collaborative handling instructions for the candidate security events in the security event candidate set. S4: Based on the collaborative handling instructions of the candidate security events, generate a linkage control strategy for the intelligent security system, and perform device linkage control, dynamic scheduling of monitoring resources, and closed-loop management of events on the intelligent security system based on the linkage control strategy.
[0023] In step S1, the video perception data consists of a sequence of monitoring video frames with continuous timestamps collected in the monitoring area. The edge device operation data is the operation status information of the edge computing device, which includes the device identifier, operation status monitoring cycle, device temperature, device online status, available bandwidth, network latency, CPU utilization, GPU utilization, and memory usage of the edge computing device.
[0024] Step S1 involves performing data normalization processing on the video sensing data to obtain a time-series video data stream of the monitored object, and also includes: S11: Acquire video perception data. The edge computing device performs unified encoding format, unified image resolution, and unified timestamp format processing on the monitoring video frames in the video perception data to obtain video perception data with standardized format. S12: Perform key frame extraction processing on the formatted video perception data at fixed time intervals, and sequentially perform image quality enhancement, motion region filtering and target detection processing on the extracted key frames to generate target information of the extracted key frames, wherein the target information includes target number, target category, target location and detection confidence. S13: The keyframes that have undergone image quality enhancement processing and the target information corresponding to the keyframes are spliced together to form keyframe information. S14: Sort the keyframe information according to the order in which the timestamps appear after the timestamp format is unified, and use it as the time-series video data stream of the monitored object.
[0025] Step S1 involves performing data normalization processing on the edge device operating data to obtain the edge monitoring context vector, and also includes: S15: Based on the CPU utilization, GPU utilization, and memory usage in the edge device's operating data, calculate the resource status indicators of the edge computing device during the operating status monitoring cycle; S16: Based on the available bandwidth and network latency in the edge device operation data, normalize the available bandwidth and network latency respectively, and calculate the network status index of the edge computing device in the operation status monitoring period based on the normalized available bandwidth and network latency. S17: Normalize the device temperature in the edge device operation data, and calculate the health status index of the edge computing device during the operation status monitoring cycle based on the normalized device temperature and the online status of the device. S18: The resource status indicators, network status indicators, and health status indicators of the edge computing device during the operation status monitoring cycle are concatenated to form the edge monitoring context vector of the edge computing device during the operation status monitoring cycle.
[0026] As an embodiment of the present invention, refer to as follows Figure 2 The diagram shows a monitoring area. Video acquisition terminals and edge computing devices are deployed within the monitoring area. The types of video acquisition terminals include fixed cameras, dome cameras, and infrared cameras. The video acquisition terminals continuously output video streams and the acquisition timestamps of each monitoring video frame in the video stream at a fixed sampling frequency. The fixed sampling frequency is set to 25fps by default. The video acquisition terminals synchronously transmit the output video stream to the edge computing devices. The edge computing device uses the video streams output by each video acquisition terminal as video sensing data and synchronously collects operational status information during the operational status monitoring cycle. The default length of the operational status monitoring cycle is set to 5 minutes. Specifically, the operational status information of the edge computing device is constructed by collecting the device temperature, device online status, available bandwidth, network latency, CPU utilization, GPU utilization, and memory usage during the operational status monitoring cycle. The collected device temperature, device online status, available bandwidth, network latency, CPU utilization, GPU utilization, and memory usage are all average operating values of various indicators of the edge computing device within the operational status monitoring cycle. For the device online status, 1 indicates that the device is online, and 0 indicates that the device is offline. The device being online indicates that the edge computing device is in a working state.
[0027] Specifically, in step S11, the monitoring video frames are uniformly converted to H.265 encoding format, and the image resolution of the monitoring video frames is uniformly adjusted to 1920×1080 pixels. When the image resolution is higher than 1920×1080, bilinear interpolation is used for scaling; when the image resolution is lower than 1920×1080, edge-aware adaptive interpolation is used for resampling. The acquisition timestamp is uniformly recorded in UTC time format with a time accuracy of milliseconds. As an embodiment of the present invention, the interpolation formula of the edge-aware adaptive interpolation method is as follows: ; ; ; ; in, Indicates the pixel coordinates after interpolation Pixel value at that location, Represents the pixel coordinates in the surveillance video frame Pixel value at that location, This represents the interpolation kernel function. Represents pixel coordinates Edge adaptive weighting factor, Indicates the window radius (default setting is 3). This represents the nuclear decay coefficient (default setting is 0.15). Represents the sine function. It represents 180 degrees. This represents an exponential function with the natural constant as its base. This represents the edge enhancement factor (default setting is 0.5). Represents the pixel coordinates in the surveillance video frame pixel gradient at that location Indicates the selection of a set The maximum value in, Represents the Sinc function; It should be noted that traditional interpolation methods use a uniform interpolation strategy in image edge regions and regions with complex textures, which can easily lead to smoothing of target edge details, loss of texture information, and local ringing artifacts. Therefore, the edge-aware adaptive interpolation method proposed in this invention introduces edge-adaptive weights and a distance exponential decay mechanism on the basis of traditional interpolation kernels, so that the contribution of neighboring pixels is not only related to spatial distance, but also to local pixel gradient features. Specifically, for target edge regions, pixels with larger gradients receive higher interpolation weights, which can effectively maintain the continuity of target contours and edge clarity; for flat background regions, the influence of high-frequency information is automatically reduced, suppressing ringing artifacts in the interpolation process. At the same time, the exponential decay term further reduces the interference of distant pixels, improving local interpolation stability. Therefore, compared with traditional interpolation methods (such as the Lanczos interpolation method), the method proposed in this invention can better maintain target details and texture features in security monitoring scenarios during image resampling, improve the accuracy of subsequent target detection, target tracking, and behavior recognition, and is more suitable for cloud-edge fusion intelligent security video monitoring scenarios.
[0028] Furthermore, the formatted video perception data is processed by extracting keyframes by reserving one frame every three frames as a keyframe. For the extracted keyframes, the CLAHE (Contrast Limiting Adaptive Histogram Equalization) algorithm is used to enhance the local contrast in dark areas, where the limiting factor can be set to 2.0 and the window size can be set to 8×8 pixels; then, a bilateral filtering algorithm is used for noise suppression to obtain keyframes after image enhancement, where the spatial domain standard deviation in the bilateral filtering algorithm can be set to 5 and the grayscale domain standard deviation can be set to 30. For the keyframes that have undergone image enhancement processing, a static background model of the monitoring area is established using a background modeling method. The pixel difference between the keyframe and the background model is calculated. When the pixel grayscale change exceeds 25 and the area of continuous change is greater than 200 pixels, the corresponding area in the keyframe with grayscale change is identified as a candidate motion region. Morphological opening and closing operations are used to eliminate isolated noise and holes. The YOLOv8-N lightweight detection network is used to perform target detection on the candidate motion regions after morphological opening and closing operations. The target category, target position, and detection confidence are output. When the detection confidence is not lower than 0.60, the corresponding detection result is retained; otherwise, it is regarded as background interference and is removed. The target tracking algorithm is used to correlate the target detection results of adjacent keyframes and assign a unique target number to the same target; The formula for calculating the resource status indicators of edge computing devices during the operational status monitoring cycle in step S1 is as follows: ; in, Indicators representing resource status. These represent, in order, the CPU utilization, GPU utilization, and memory usage in the edge device's operational data. These represent the weighting coefficients for CPU utilization, GPU utilization, and memory usage, respectively, with the default settings. The values are 0.4, 0.4, and 0.2 respectively. The normalization methods for network latency and device temperature are as follows: during the operation of the edge computing device, the maximum and minimum values are selected, and the minimum-maximum normalization method is used to normalize the network latency and device temperature; the available bandwidth is converted into the proportion of the remaining bandwidth, and the proportion of the remaining bandwidth is used as the available bandwidth after normalization. The calculation formulas for the network status indicators and health status indicators of the edge computing device during the operation status monitoring cycle are as follows: ; ; in, Indicates network status metrics, Indicators representing health status This indicates the available bandwidth after normalization. This represents the network latency after normalization. This indicates the equipment temperature after normalization. Indicates the device's online status. All represent control weights, default settings. The values are 0.6 and 0.5 respectively.
[0029] It should be noted that this invention constructs an edge monitoring context vector by performing multi-dimensional quantitative analysis of the resource status, network status, and health status of edge computing devices during operation. Compared to traditional methods that rely solely on single device load parameters for status judgment, this invention can more comprehensively and accurately reflect the real-time operational capabilities of edge computing devices during intelligent security task execution. Specifically, this invention calculates resource status indicators by integrating CPU utilization, GPU utilization, and memory usage, enabling resource consumption to be represented in a unified quantitative form, providing a basis for subsequent task scheduling and model inference strategy optimization. By normalizing available bandwidth and network latency, and combining the influence of bandwidth resources and communication latency to calculate network status indicators, this invention reduces the impact of differences in data dimensions under different network environments on status assessment results, improving the stability of network status evaluation. By integrating device temperature and online status to calculate health status indicators, this invention can promptly reflect the risk of abnormal operation of edge devices, avoiding monitoring task interruptions due to overheating, offline issues, etc. Furthermore, by concatenating resource status indicators, network status indicators, and health status indicators to form a unified edge monitoring context vector, the subsequent edge intelligent analysis model can dynamically adjust event analysis strategies and resource allocation strategies based on the device operating environment. This improves the computational adaptability and system reliability in the security event identification process, reduces cloud data transmission pressure, and achieves cloud-edge collaborative optimization of the intelligent security system.
[0030] The edge intelligence analysis model in step S2 is deployed on edge computing devices using a lightweight and modular deployment approach. The edge intelligence analysis model includes a temporal behavior encoding module, a scene state fusion module, and an event association reasoning module.
[0031] Specifically, refer to, for example Figure 3 The edge intelligence analysis model structure diagram shown is illustrated. The temporal behavior encoding module includes a frame feature encoding unit, a target information encoding unit, a temporal association unit, and a behavior feature output unit. The frame feature encoding unit is used to extract visual features from consecutive keyframes. The target information encoding unit is used to uniformly encode the target number, target category, target location, and detection confidence. The temporal association unit is used to associate and fuse the visual features with the encoded target information in chronological order to establish the behavioral evolution relationship of the monitored object between consecutive keyframes. The behavior feature output unit is used to output temporal behavior features that can characterize the continuous behavioral change process of the monitored object and use an event category recognition model to identify the event category corresponding to the temporal behavior features. The scene state fusion module includes a context parsing unit, a state weight generation unit, a feature fusion unit, and a scene state output unit. The context parsing unit is used to parse the resource state indicators, network state indicators, and health state indicators in the edge monitoring context vector. The state weight generation unit is used to calculate the state influence weight of the edge computing device on the current monitoring object's time-series video data stream based on each state indicator. The feature fusion unit is used to adaptively weight and fuse the time-series behavioral features using the state influence weight. The scene state output unit is used to output scene state features that comprehensively reflect the overall operating status of the monitoring area and the behavioral status of the monitoring object. The event association reasoning module includes a behavior association unit, an event scoring unit, a candidate event generation unit, and an edge decision generation unit. The behavior association unit is used to establish behavior association relationships between monitored objects and between monitored objects and the scene by combining scene state characteristics. The event scoring unit is used to calculate the event score and risk level of each candidate event based on the behavior association relationship. The candidate event generation unit is used to filter candidate events that meet preset conditions based on the event score to form a security event candidate set. The edge decision generation unit is used to combine the event score and the edge monitoring context vector to generate edge decision information to determine the edge processing mode of the current security event and the subsequent cloud-edge collaborative processing method.
[0032] In step S2, the temporal behavior encoding module in the edge intelligent analysis model extracts visual features from keyframes and encodes the target information of the keyframes. It then correlates and fuses the visual features of the keyframes and the encoded target information to obtain temporal behavior features that characterize the continuous behavioral changes of the monitored object. An event category recognition model is then used to identify the event category corresponding to these temporal behavior features. The formula for generating the temporal behavior features is as follows: ; ; ; in, Represents temporal behavioral characteristics, This represents the behavioral characteristics of the nth keyframe in the time-series video data stream of the monitored object. N represents the total number of keyframes. Represents the visual features of the nth keyframe. This represents the encoded target information of the nth keyframe. This represents the fusion coefficient (default setting is 0.6). This represents the timing fusion coefficient (default setting is 0.1). This represents the fusion result of the visual features of the nth keyframe and the encoded target information; Specifically, the event category recognition model is a trainable convolutional neural network structure; It should be noted that this invention utilizes continuous keyframes to construct a unified behavior representation, enabling the complete expression of the motion trajectory, behavior changes, and temporal correlation of the monitored object, effectively reducing the impact of single-frame detection errors and short-term target occlusion on the behavior analysis results; Specifically, the formula for calculating the state influence weight is as follows: ; in, This indicates that the state affects the weight. All represent the weighting coefficients of each state index in the edge detection context vector, with the default settings. The values are 0.45, 0.3, and 0.25 respectively. The temporal behavior features described by the state influence weight L The formula for adaptive weighted summation is: The construction method of the scene state features is as follows: ,in This represents the edge detection context vector. This indicates the splicing coefficient (default setting is 0.7). Specifically, due to the characteristics of temporal behavior If the feature length is higher than the edge detection context vector, then during the construction of scene state features, the edge detection context vector is padded with zeros so that the temporal behavior features have the same length as the edge detection context vector. It should be noted that since the edge monitoring context vector corresponds to the entire time-series video data stream, the overall fusion approach can avoid the same context information from being repeatedly used in calculations in consecutive keyframes, reducing computational complexity. At the same time, it enables the scene state to synchronously reflect changes in the behavior of the monitored object and the operating status of the edge nodes, improving the consistency of the scene state description.
[0033] Specifically, the preset scoring threshold is set to 0.75 by default, and the event features of the candidate security events include target behavior features, scene level features, risk features, and the edge detection context vector of the edge computing device associated with the candidate security event; the default setting is... They are 0.4 and 0.6 respectively; In one embodiment of the present invention, the rate of change of the target position in consecutive keyframes is calculated as the movement speed of the target person in the candidate security event, and the movement speed of the target person is used as the target behavior feature; by querying the regional attributes of the monitoring area, the scene level feature of the monitoring area is quantified, wherein the value range of the scene level feature is 1-10, and the higher the scene level feature, the higher the monitoring importance of the monitoring area; by mapping the event category to a risk value, the product between the risk value and the event score is calculated as the risk feature, for example, the risk value of event category - loitering is 0.4, the risk value of event category - illegal entry is 0.8, and the risk value of event category - violent behavior is 0.95; Furthermore, the calculation formula for edge decision information of candidate security events based on event features is as follows: ; in, This represents edge decision information for candidate security events. This indicates the weight of the state of the edge computing devices associated with the candidate security event. The target behavioral characteristics, scene level characteristics, and risk characteristics of the candidate security events are represented in that order, respectively. All represent feature correlation coefficients, default settings. The values were 0.4, 0.35, and 0.25, respectively. This indicates the control parameter (default setting is 0.01). This represents the security risk value of candidate security events; the higher the edge decision information, the lower the security risk value / the higher the state influence weight, where a higher state influence weight indicates stronger computing power and more abundant computing resources of the edge computing device; Specifically, the edge computing device directly processes candidate security events as follows: Video segments within a preset time window before and after the candidate security event are associated and saved. For example, video segments from 30 seconds before the candidate security event to 60 seconds after the candidate security event are extracted as event evidence data. Based on the video acquisition terminal associated with the candidate security event, terminal control parameters are generated, and control commands are sent to the video acquisition terminal to enable it to perform enhanced monitoring of the target area, including adjusting the camera angle, increasing the acquisition frequency, expanding the target area resolution, and continuously tracking the target location. Simultaneously, a reporting information is sent to the management personnel. This reporting information includes the terminal number of the video acquisition terminal associated with the candidate security event and the acquisition timestamp range of the candidate security event. The preset decision threshold is set to 0.6 by default.
[0034] Specifically, step S3, which determines the cloud-edge collaborative processing strategy based on the edge decision information, includes: The cloud-edge collaborative processing strategy is a method for processing candidate security events. Based on the edge decision information of candidate security events in the candidate security event set, if the edge decision information is lower than a preset decision threshold, the candidate security event and its event characteristics are uploaded to the cloud for processing; otherwise, the edge computing device directly processes the candidate security event.
[0035] Step S3, based on the cloud-edge collaborative processing strategy, utilizes a cloud-based intelligent model that integrates historical monitoring knowledge to perform event review and risk assessment on the security event candidate set, generating event confirmation results and collaborative handling instructions for the security events in the candidate set. It also includes: The cloud-based intelligent model, deployed in the cloud, receives candidate security events and their event characteristics uploaded to the cloud. It then uses the cloud-based intelligent model, which integrates historical monitoring knowledge, to perform event review and risk assessment on the candidate security events. The cloud-based intelligent model extracts historical event features of the monitoring areas associated with candidate security events from historical monitoring knowledge, calculates the event correlation degree between the historical event features and the event features of the candidate security events, and performs a weighted fusion of the event correlation degree and event score of the candidate security events to obtain the event credibility of the candidate security events. If the event credibility is higher than a preset credibility threshold, it means that the candidate security event has passed the event review and is marked as a valid security event; otherwise, the candidate security event is marked as an event to be observed for manual review. Based on the event characteristics, calculate the security risk value of the effective security event; The event type, monitoring area, collection timestamp range, event credibility, and security risk value of the effective security event are concatenated to form the event confirmation result of the effective security event and are stored. Based on the security risk value, a collaborative handling instruction is generated and sent to the video acquisition terminal and edge computing device.
[0036] Specifically, the historical event characteristics of the monitored area mentioned in step S4 are the average value of event characteristics uploaded to the cloud by the monitored area in the past. The preset trust threshold is set to 0.7 by default, and the weighted fusion coefficients of event correlation and event score are set to 0.6 and 0.4 by default, respectively. The event type, monitoring area, collection timestamp range, event credibility, and security risk value of the effective security event are concatenated to form the event confirmation result of the effective security event and are stored. Based on the security risk value, a collaborative handling instruction is generated and sent to the video acquisition terminal and edge computing device.
[0037] As an embodiment of the present invention, candidate security events are classified according to the security risk values output by the cloud-based intelligent model, and corresponding collaborative handling instructions are generated based on different risk levels. When the security risk value is greater than or equal to 0.85, the current valid security event is determined to be a high-risk event, and a Level 1 collaborative handling instruction is generated: the associated video acquisition terminal is controlled to increase the video acquisition frequency of the target area, improve the image resolution, and adjust the pan-tilt unit to the target position for continuous tracking. At the same time, the edge computing device is controlled to increase the analysis priority of the corresponding video stream, increase the target behavior analysis frequency, and upload the event video clips, target trajectory information, and edge analysis results to the cloud in real time. When the security risk value is greater than or equal to 0.6 but less than 0.85, the current valid security event is determined to be a medium-risk event, and a Level 2 collaborative handling instruction is generated: the video acquisition terminal is controlled to maintain normal acquisition status and perform local enhanced acquisition of the target area. At the same time, the edge computing device is controlled to increase the computing resource allocation ratio of the event area, continuously analyze the target behavior, and upload event summary information at preset time intervals. When the security risk value is less than 0.6, the current event is determined to be a low-risk event, and a Level 3 collaborative handling instruction is generated. The video acquisition terminal is controlled to maintain the current acquisition parameters, and the edge computing device adopts the conventional analysis mode, only saving the event feature information and uploading it to the cloud for historical data updates in subsequent periods.
[0038] By employing a tiered collaborative response strategy based on security risk values, video acquisition terminals, edge computing devices, and cloud-based intelligent models can dynamically adjust the allocation of monitoring resources according to the severity of events. This enables rapid response to high-risk events and low-resource-consumption processing of low-risk events, thereby improving the real-time performance and resource utilization efficiency of the cloud-edge integrated intelligent security system.
[0039] Example 2 As another embodiment of the present invention, the intelligent security system includes a video acquisition terminal, an edge computing device, and a cloud, to realize the technical steps of the cloud-edge fusion intelligent security video surveillance method as described in Embodiment 1; As the data sensing entry point of the intelligent security system, the video acquisition terminal is used to continuously acquire video information within the monitored area, extract key frames according to a preset sampling strategy, and perform preliminary detection on the monitored objects in the key frames to obtain the target information corresponding to the monitored objects. The target information includes the target number, target category, target location, and detection confidence level. Furthermore, the video acquisition terminal adjusts the acquisition parameters according to the collaborative processing instructions issued by the cloud or edge side, including video acquisition frequency, image resolution, acquisition angle, and target tracking mode, to achieve dynamic enhanced monitoring of key areas and high-risk targets.
[0040] As an intermediate processing node connecting a video acquisition terminal and a cloud, the edge computing device is configured to receive the time-series video data stream of a monitored object uploaded by the video acquisition terminal, construct an edge monitoring context vector in combination with its own operating status information, use an edge intelligent analysis model deployed on the edge side to perform real-time analysis on the behavior of the monitored object, scene status assessment and event association reasoning, and generate a candidate set of security events and edge decision information; meanwhile, the edge computing device executes local rapid response according to the event risk level and its own computing resource status, including continuous target tracking, video enhanced analysis and device linkage control, and uploads candidate security events that require further confirmation by the cloud and corresponding event features to the cloud.
[0041] As a global analysis and collaborative decision-making center of the intelligent security system, the cloud is configured to receive candidate security events, event features and historical monitoring data uploaded by the edge computing device, integrate historical monitoring knowledge through a cloud intelligent model deployed on the cloud, perform event review, risk assessment and global association analysis on the candidate security events, and generate security event confirmation results and collaborative disposal instructions; meanwhile, the cloud issues a collaborative control strategy corresponding to the collaborative disposal instructions to the video acquisition terminal and the edge computing device, so as to realize multi-device linkage response, dynamic scheduling of monitoring resources and closed-loop management of security events.
[0042] Specifically, the present invention realizes cloud-edge integration processing of the intelligent security system through data collaborative interaction among the video acquisition terminal, the edge computing device and the cloud intelligent model. Specifically, the video acquisition terminal is responsible for acquiring video perception data of a monitoring area, extracting a time-series video data stream of a monitored object and then sending the data stream to the edge computing device; the edge computing device constructs an edge monitoring context vector in combination with its own operating status data, and uses an edge intelligent analysis model deployed on the edge side to perform real-time behavior analysis, scene status assessment and event association reasoning on the time-series video data stream of the monitored object, so as to generate candidate security events and corresponding edge decision information, and realize local processing of security tasks with low delay and high real-time requirement; for candidate security events meeting the cloud review conditions, the edge computing device uploads the candidate security events and their event features to the cloud, the cloud intelligent model further integrates historical monitoring knowledge, performs event review, risk assessment and global association analysis on the candidate security events, generates event confirmation results and collaborative disposal instructions, and issues the collaborative disposal instructions to the video acquisition terminal and the edge computing device, so as to realize acquisition parameter adjustment, edge resource scheduling and multi-device linkage control. Through the above data closed loop of terminal perception, rapid edge analysis, in-depth cloud decision-making and cloud-edge collaborative feedback, the intelligent security system can dynamically allocate computing tasks according to the event risk level and the operating status of devices, which reduces the video data transmission pressure and cloud computing load, and improves the real-time performance and accuracy of security event detection and the overall operating efficiency of the system.
[0043] Example 3 As an embodiment of the present invention, historical surveillance video data, target detection data, and corresponding edge device operating status data in different security scenarios are collected. The collected data are then processed by time alignment, target association, and feature organization to construct a training dataset containing time-series video data of monitored objects and edge monitoring context vectors. The training dataset is manually labeled, and the manually labeled content includes event type, labeling results of whether it is a candidate security event, and cloud-edge collaborative processing strategy of candidate security events. The training loss function is constructed with the goal of maximizing the manually labeled content and the labeling results of event type, whether it is a candidate security event, and cloud-edge collaborative processing strategy of candidate security events output by the model. The model parameters are iteratively updated using an optimization algorithm based on gradient backpropagation. By adjusting the network weights, the model output results gradually approach the real labeled results. After training, independent test sets were used to input collected data that was not used in training for testing and verification. According to the test and verification results, the event type recognition accuracy reached 95.2%, which was used to evaluate the model's ability to classify different security event types such as illegal intrusion, abnormal stay, crowd gathering, and dangerous behavior; the candidate security event detection accuracy reached 94.6%, which was used to evaluate the model's ability to filter valid security events from continuous monitoring data; the candidate security event false negative rate was 2.1%, which was used to evaluate the model's omission of real security events; the candidate security event false positive rate was 3.3%, which was used to evaluate the model's ability to reduce the misjudgment of normal behavior as abnormal events; the cloud-edge collaborative processing strategy decision accuracy reached 92.8%, which was used to evaluate the accuracy of the model in selecting edge processing, cloud-edge collaborative processing, or cloud processing modes based on the event risk level and the resource status of edge devices; the average latency of a single inference by the model was 38ms, which meets the processing requirements of edge computing devices for real-time security analysis tasks. The above metrics demonstrate that the trained edge intelligence analysis model can accurately analyze the behavior of monitored objects, security events, and cloud-edge collaboration strategies in different security scenarios, thereby improving the real-time performance, accuracy, and resource scheduling efficiency of intelligent security systems.
[0044] It should be noted that the terms "comprising," "including," or any other variations thereof used herein are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0045] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0046] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A cloud-edge fusion intelligent security video surveillance method, characterized in that, The method includes: S1: Collect video sensing data and edge device operation data of the monitoring area, and perform data normalization processing on the video sensing data and edge device operation data respectively to obtain the time-series video data stream of the monitored object and the edge monitoring context vector. S2: Using the edge intelligence analysis model, target behavior analysis, scene state assessment and event correlation reasoning are performed on the time-series video data stream of the monitored object and the edge monitoring context vector to generate a security event candidate set and edge decision information; S3: Based on the edge decision information, determine the cloud-edge collaborative processing strategy. According to the cloud-edge collaborative processing strategy, use the cloud-based intelligent model that integrates historical monitoring knowledge to perform event review and risk assessment on the security event candidate set, and generate event confirmation results and collaborative handling instructions for the candidate security events in the security event candidate set. S4: Based on the collaborative handling instructions of the candidate security events, generate a linkage control strategy for the intelligent security system, and perform device linkage control, dynamic scheduling of monitoring resources, and closed-loop management of events on the intelligent security system based on the linkage control strategy.
2. The cloud-edge fusion intelligent security video surveillance method as described in claim 1, characterized in that, In step S1, the video perception data consists of a sequence of monitoring video frames collected from the monitoring area at continuous timestamps. The edge device operation data is the operation status information of the edge computing device, which includes the device identifier, operation status monitoring cycle, device temperature, device online status, available bandwidth, network latency, CPU utilization, GPU utilization, and memory usage of the edge computing device.
3. The cloud-edge fusion intelligent security video surveillance method as described in claim 2, characterized in that, Step S1 involves performing data normalization processing on the video sensing data to obtain a time-series video data stream of the monitored object, and also includes: S11: Acquire video perception data. The edge computing device performs unified encoding format, unified image resolution, and unified timestamp format processing on the monitoring video frames in the video perception data to obtain video perception data with standardized format. S12: Perform key frame extraction processing on the formatted video perception data at fixed time intervals, and sequentially perform image quality enhancement, motion region filtering and target detection processing on the extracted key frames to generate target information of the extracted key frames, wherein the target information includes target number, target category, target location and detection confidence. S13: The keyframes that have undergone image quality enhancement processing and the target information corresponding to the keyframes are spliced together to form keyframe information. S14: Sort the keyframe information according to the order in which the timestamps appear after the timestamp format is unified, and use it as the time-series video data stream of the monitored object.
4. The cloud-edge fusion intelligent security video surveillance method as described in claim 2, characterized in that, Step S1 involves performing data normalization processing on the edge device operating data to obtain the edge monitoring context vector, and also includes: S15: Based on the CPU utilization, GPU utilization, and memory usage in the edge device's operating data, calculate the resource status indicators of the edge computing device during the operating status monitoring cycle; S16: Based on the available bandwidth and network latency in the edge device operation data, normalize the available bandwidth and network latency respectively, and calculate the network status index of the edge computing device in the operation status monitoring period based on the normalized available bandwidth and network latency. S17: Normalize the device temperature in the edge device operation data, and calculate the health status index of the edge computing device during the operation status monitoring cycle based on the normalized device temperature and the online status of the device. S18: The resource status indicators, network status indicators, and health status indicators of the edge computing device during the operation status monitoring cycle are concatenated to form the edge monitoring context vector of the edge computing device during the operation status monitoring cycle.
5. The cloud-edge fusion intelligent security video surveillance method as described in claim 1, characterized in that, The edge intelligent analysis model in step S2 is deployed on edge computing devices using a lightweight modular deployment method. The edge intelligent analysis model includes a temporal behavior encoding module, a scene state fusion module, and an event association reasoning module.
6. The cloud-edge fusion intelligent security video surveillance method as described in claim 5, characterized in that, Step S2 utilizes an edge intelligence analysis model to perform target behavior parsing, scene state assessment, and event correlation reasoning on the time-series video data stream of the monitored object and the edge monitoring context vector. It also includes: S21: The edge computing device extracts the time-series video data stream of the monitored object to be analyzed and the edge monitoring context vector in the current operating state monitoring cycle, and inputs the extracted time-series video data stream of the monitored object and the edge monitoring context vector into the edge intelligent analysis model. S22: The temporal behavior encoding module in the edge intelligent analysis model extracts visual features from keyframes and encodes the target information of the keyframes. It then correlates and fuses the visual features of the keyframes with the encoded target information to obtain temporal behavior features characterizing the continuous behavioral changes of the monitored object. An event category recognition model is used to identify the event category corresponding to the temporal behavior features. The formula for generating the temporal behavior features is: ; ; ; in, Represents temporal behavioral characteristics, This represents the behavioral characteristics of the nth keyframe in the time-series video data stream of the monitored object. N represents the total number of keyframes. Represents the visual features of the nth keyframe. This represents the encoded target information of the nth keyframe. Represents the fusion coefficient. Represents the temporal fusion coefficients. This represents the fusion result of the visual features of the nth keyframe and the encoded target information; S23: The scene state fusion module calculates the state influence weight of the edge computing device on the current monitoring object's time-series video data stream based on the edge detection context vector, and uses the state influence weight to adaptively weight and fuse the time-series behavior features to obtain a fused behavior representation. The fused behavior representation is then weighted and concatenated with the edge detection context vector as a scene state feature. S24: The event association reasoning module extracts scene state features and temporal behavior features, calculates event association values and event scores. If the event score is higher than a preset score threshold, the temporal video data stream of the monitored object and the edge monitoring context vector are used as candidate security events. Event features of the candidate security events are extracted, and edge decision information of the candidate security events is calculated based on the event features. The formula for calculating the event score is: ; in, Indicates the associated value of the event. Indicates the event rating. Represents scene state characteristics. Both represent correlation coefficients. This represents the trainable convolutional weight matrix. Represents an exponential function with the natural constant as its base; S25: Construct a security event candidate set by combining all candidate security events and their event characteristics.
7. The cloud-edge fusion intelligent security video surveillance method as described in claim 1, characterized in that, The S3 step, which determines the cloud-edge collaborative processing strategy based on the edge decision information, includes: The cloud-edge collaborative processing strategy is a method for processing candidate security events. Based on the edge decision information of candidate security events in the candidate security event set, if the edge decision information is lower than a preset decision threshold, the candidate security event and its event characteristics are uploaded to the cloud for processing; otherwise, the edge computing device directly processes the candidate security event.
8. The cloud-edge fusion intelligent security video surveillance method as described in claim 7, characterized in that, Step S3, based on the cloud-edge collaborative processing strategy, utilizes a cloud-based intelligent model that integrates historical monitoring knowledge to perform event review and risk assessment on the security event candidate set, generating event confirmation results and collaborative handling instructions for the security events in the candidate set. It also includes: The cloud-based intelligent model, deployed in the cloud, receives candidate security events and their event characteristics uploaded to the cloud. It then uses the cloud-based intelligent model, which integrates historical monitoring knowledge, to perform event review and risk assessment on the candidate security events. The cloud-based intelligent model extracts historical event features of the monitoring areas associated with candidate security events from historical monitoring knowledge, calculates the event correlation degree between the historical event features and the event features of the candidate security events, and performs a weighted fusion of the event correlation degree and event score of the candidate security events to obtain the event credibility of the candidate security events. If the event credibility is higher than a preset credibility threshold, it means that the candidate security event has passed the event review and is marked as a valid security event; otherwise, the candidate security event is marked as an event to be observed for manual review. Based on the event characteristics, calculate the security risk value of the effective security event; The event type, monitoring area, collection timestamp range, event credibility, and security risk value of the effective security event are concatenated to form the event confirmation result of the effective security event and are stored. Based on the security risk value, a collaborative handling instruction is generated and sent to the video acquisition terminal and edge computing device.
9. The cloud-edge fusion intelligent security system as described in claim 1, characterized in that, The intelligent security system includes a video acquisition terminal, an edge computing device, and a cloud, to implement the technical steps of the cloud-edge fusion intelligent security video surveillance method as described in any one of claims 1-8.