Air-ground collaborative automatic alarm method and system for park security and medium
By employing digital twin technology and edge computing in the park's security system, the system can identify security incident targets in real time and conduct multi-node collaborative deployment, solving the problems of cloud latency and low computing power utilization in existing technologies, and achieving efficient security incident handling and tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU BAITONG COMM TECH CO LTD
- Filing Date
- 2026-06-24
- Publication Date
- 2026-07-24
AI Technical Summary
Existing air-ground collaborative security technologies suffer from high cloud processing latency, low edge computing power utilization, and slow cross-node collaborative response, leading to untimely handling of security incidents and a high rate of false alarms and missed reports.
The system employs a digital twin-based edge perception matrix network, which identifies security event targets in real time through edge nodes, predicts spatiotemporal trajectories using a geographic model, triggers multi-node collaborative deployment by a dynamic master control node, and utilizes drones for adaptive alarms, thereby achieving cloud-edge hybrid control and multi-source data fusion.
It reduces cloud bandwidth pressure, shortens cross-node collaborative response time, realizes proactive predictive security, and significantly improves identification accuracy and tracking reliability.
Smart Images

Figure CN122454681A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent security technology, specifically to an air-ground collaborative automatic alarm method, system, and medium for park security. Background Technology
[0002] With the comprehensive advancement of smart park construction, park security systems are rapidly evolving from traditional manual monitoring to intelligent, unmanned, and integrated systems. The deep integration of edge computing, digital twins, and drone technology provides crucial technical support for addressing long-standing pain points such as incomplete security coverage and delayed emergency response in large-area parks. The air-ground collaborative security model, combining the wide-area continuous coverage of ground monitoring with the mobile and close-range advantages of drones, has become the core development direction of the next generation of park security systems, attracting widespread attention and application promotion within the industry.
[0003] However, existing air-ground collaborative security technologies have not yet broken through the fundamental limitations of traditional centralized architectures. They cannot achieve dynamic autonomous collaboration between edge nodes and fine-grained scheduling of computing power. As a result, when facing complex security scenarios in large-area parks, the systems generally suffer from a series of derivative problems such as high emergency response delays, low collaborative deployment efficiency, severe fragmentation of evidence chains, and low utilization of computing resources. These issues make it difficult to meet the core requirements of smart parks for the real-time performance, reliability, and intelligence of security systems. Summary of the Invention
[0004] This application provides an air-ground collaborative automatic alarm method, system, and medium for park security, aiming to solve the technical problems of high cloud processing latency, low edge computing power utilization, and slow cross-node collaborative response in existing centralized security systems, which lead to untimely handling of security incidents and high false alarm rates.
[0005] In view of the above problems, this application provides an air-ground coordinated automatic alarm method, system and medium for park security.
[0006] Firstly, this application provides an air-ground collaborative automatic alarm method for park security. The method includes: when a first edge risk identification node receives the original video stream of a first physical monitoring area transmitted back by a first edge perception matrix, and drives a differentiated identification task model to identify a security event target, it performs multimodal feature extraction by segmenting event-related video segments from the original video stream to obtain target visual features and spatiotemporal motion features; the first edge risk identification node performs spatiotemporal trajectory prediction based on a park geographic model and the spatiotemporal motion features to locate N security-related areas; using the first edge risk identification node as a dynamic master control node, it sends a linkage control command package and the target visual features to N associated risk identification nodes in the N security-related areas to trigger collaborative deployment; the first edge risk identification node, based on the N node tracking data transmitted back by the N associated risk identification nodes, coordinates an alarm drone to adaptively alarm the security event target under dynamic trajectory replanning.
[0007] Secondly, this application provides an air-ground collaborative automatic alarm system for park security. The system includes: a multimodal feature extraction module, used to extract multimodal features from event-related video segments of the original video stream when a first edge risk identification node receives the original video stream of a first physical monitoring area transmitted back by a first edge perception matrix, driving a differentiated identification task model to identify a security event target, thereby obtaining target visual features and spatiotemporal motion features; an associated area positioning module, used by the first edge risk identification node to perform spatiotemporal trajectory prediction based on a park geographic model and the spatiotemporal motion features, locating N security-related areas; a collaborative deployment triggering module, used by the first edge risk identification node as a dynamic master control node to send a linkage control command package and the target visual features to N associated risk identification nodes in the N security-related areas, triggering collaborative deployment; and an adaptive alarm module, used by the first edge risk identification node to trigger an alarm drone to adaptively alarm the security event target under dynamic trajectory replanning, based on the N node tracking data transmitted back by the N associated risk identification nodes.
[0008] Thirdly, this application provides a computer-readable storage medium storing a computer program for executing the air-ground collaborative automatic alarm method for campus security provided in this application.
[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages: By adopting a technical solution based on digital twin-based edge perception matrix networking, regional customized differentiated identification task model, dynamic master control node cross-node collaborative deployment, cloud-edge hybrid control of UAV air-ground collaborative alarm, and multi-source data fusion trajectory prediction, this solution solves the technical problems of high cloud processing latency, low edge computing power utilization, slow cross-node collaborative response, untimely handling of security incidents, and high false alarm and missed alarm rates in existing centralized security systems. It achieves the technical effects of reducing cloud bandwidth pressure, shortening cross-node collaborative response time, realizing proactive predictive security, and significantly improving identification accuracy and tracking reliability.
[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0011] Figure 1 A flowchart illustrating an air-ground coordinated automatic alarm method for park security is provided for embodiments of this application.
[0012] Figure 2 This application provides a schematic diagram of the structure of an air-ground collaborative automatic alarm system for park security.
[0013] Figure labeling: Multimodal feature extraction module 11, associated region localization module 12, collaborative deployment triggering module 13, adaptive alarm module 14. Detailed Implementation
[0014] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0015] The overall concept of the technical solution provided in this application is as follows: This application provides an air-ground collaborative automatic alarm method, system, and medium for park security. It utilizes a park geographic model to deploy edge nodes in a network, employs a differentiated identification model to identify security targets in real time, extracts spatiotemporal motion features to predict trajectories and delineate security-related areas, coordinates multiple nodes for collaborative deployment via a master control node, and uses drones for dynamic tracking to achieve adaptive alarms. After an event, it integrates multi-source video to form an evidence chain for archiving, realizing integrated operation of intelligent early warning, tracking and handling, and source tracing and archiving for park security.
[0016] After introducing the basic principles of this application, various non-limiting embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0017] Example 1, as Figure 1 As shown in the embodiment of this application, an air-ground coordinated automatic alarm method for park security is provided, the method including: Step S100: When the first edge risk identification node receives the original video stream of the first physical monitoring area transmitted back by the first edge perception matrix, and drives the differential identification task model to identify the security event target, it performs multimodal feature extraction by segmenting the event-related video segments from the original video stream to obtain the target's visual features and spatiotemporal motion features.
[0018] Specifically, the first edge risk identification node refers to an industrial-grade edge computing device deployed near a specific physical monitoring area within the park. The first edge perception matrix is a distributed acquisition network composed of multiple cameras within the same physical monitoring area, responsible for collecting raw monitoring data of the area around the clock. The first physical monitoring area refers to an independent geospatial unit controlled by a single edge node, defined based on a digital twin model. The differentiated identification task model is a multi-model parallel inference system customized according to the functional attributes of the area, with differentiated computing power priorities allocated to models of different risk types. Security event targets refer to targets or events detected by the model that conform to preset risk rules. For example: an open flame in a warehouse corner, a person climbing over a fence, or an unauthorized vehicle entering without authorization. Target visual features are feature vectors used to describe the static appearance of the target, including color, texture, and shape. Spatiotemporal motion features are feature vectors used to describe the dynamic changes of the target, including position, speed, and direction.
[0019] Specifically, when the first edge risk identification node continuously receives the original video stream of the first physical monitoring area transmitted back by the first edge perception matrix, it drives the differentiated identification task model to perform real-time parallel inference. When the model detects a security event target that meets the preset risk rules and reaches the preset confidence threshold, a time analysis window is constructed based on the identification trigger time. At the same time, a region of interest (ROI) is constructed based on the spatial bounding box of the security event target. Based on the time analysis window and ROI, the event-related video segments are accurately segmented from the original video stream. The model type of the target is identified as the risk identification context, and multimodal feature extraction is performed to obtain the target's visual features and spatiotemporal motion features. Specifically, when the differential recognition task model detects a security event target and reaches a preset confidence threshold, a time analysis window is constructed based on the recognition trigger time, extending backward by a preset duration. A region of interest (ROI) is constructed using the spatial bounding box of the security event target. Based on the time analysis window and ROI, event-related video segments containing only the complete process before and after the event and the surrounding area of the target are accurately segmented from the original video stream. Irrelevant background areas and time segments are removed to reduce invalid data. The segmented event-related video segments undergo frame sampling and normalization preprocessing to extract continuous temporal frame sequences and unify their size and pixel distribution. A pre-trained ResN algorithm is then used. The et-50 visual backbone network extracts static visual features frame by frame, generates single-frame feature vectors through global average pooling, and fuses multi-frame visual information through a temporal attention module to obtain target visual features. Simultaneously, LiteFlowNet3 is used to calculate dense optical flow maps between adjacent frames, which are input into a 3D convolutional network to extract inter-frame motion features. Combined with the temporal coordinate changes of the target's spatial bounding box, velocity, acceleration, and motion direction parameters are calculated to generate spatiotemporal motion features. Based on the model type of the identified target as the risk recognition context, the two types of features are weighted and enhanced, and fused into a unified multimodal feature vector through a feature splicing layer, which is then output to the subsequent spatiotemporal trajectory prediction and collaborative deployment modules.
[0020] Preferably, the pre-trained ResNet-50 visual backbone network adopts a bottleneck residual block core architecture with a total of 50 layers, consisting of an input layer, four convolutional stages, and a global average pooling layer. The input layer receives a 224×224 RGB image, which enters the convolutional stage after 7×7 convolution and 3×3 max pooling. The four stages contain 3, 4, 6, and 3 bottleneck residual blocks, respectively. Each residual block consists of a 1×1 dimensionality reduction convolution, a 3×3 spatial convolution, a 1×1 dimensionality increase convolution, and cross-layer residual connections, effectively solving the gradient vanishing problem in deep networks and outputting multi-scale feature maps at scales of 1 / 4, 1 / 8, 1 / 16, and 1 / 32. The training is divided into two stages: The first stage is to perform full pre-training on the ImageNet-1K general dataset, using the cross-entropy loss function, AdamW optimizer, initial learning rate of 1e-3, cosine annealing decay, and training for 100 epochs to learn general visual features; The second stage is to fine-tune on a domain dataset containing 800,000 campus security images, freezing the parameters of the first three convolutional stages and only fine-tuning the fourth stage to enhance the robustness of feature extraction for security targets such as people, vehicles, and open flames.
[0021] This step offloads video stream processing and inference to local execution on edge nodes, avoiding network latency and bandwidth consumption caused by uploading raw video to the cloud, and keeping event recognition response time within milliseconds. Precise video segmentation based on time windows and ROI eliminates a large amount of irrelevant background data, reducing the amount of data required for subsequent feature extraction and transmission. Simultaneously, combined with multimodal feature extraction from the risk recognition context, it not only accurately describes the target's static appearance but also captures its dynamic change patterns, providing a high-precision feature foundation for subsequent spatiotemporal trajectory prediction, multi-node collaborative recognition, and UAV approach locking, thus improving the overall system's recognition accuracy and tracking reliability.
[0022] Step S200: The first edge risk identification node performs spatiotemporal trajectory prediction based on the park's geographical model and the spatiotemporal motion characteristics to locate N security-related areas.
[0023] Specifically, the park's geographic model is a three-dimensional digital geospatial model of the park built based on digital twin technology. Security-related areas refer to adjacent physically monitored areas that a target may enter or pass through in the future, and which are independently controlled by other edge risk identification nodes.
[0024] Specifically, the first edge risk identification node loads a 3D geographic model of the park, extracts spatial constraint information of the target's current location and surrounding area, including building boundaries, prohibited areas, road directions, entrance and exit locations, and obstacle distribution, and constructs a motion spatial constraint matrix that conforms to physical laws; the extracted spatiotemporal motion feature vector containing the target's continuous multi-frame spatial coordinates, instantaneous velocity, motion direction, and acceleration parameters is input into the trajectory prediction model, and the motion state of the target in the next 30 to 60 seconds is iteratively deduced in combination with the spatial constraint matrix to generate multiple candidate motion trajectories with occurrence probabilities; high-probability trajectories with occurrence probabilities exceeding a preset threshold are selected, and the corresponding adjacent physical monitoring areas independently controlled by other edge risk identification nodes are matched according to the geographical range covered by these trajectories. After deduplication, N security-related areas are finally located.
[0025] This step, by combining precise trajectory prediction with geospatial constraints, achieves a shift from passive response to proactive prediction. It can lock in areas where targets may spread in advance, allowing sufficient time for subsequent multi-node collaborative deployment, and effectively eliminates monitoring blind spots and tracking breakpoints that exist in traditional security systems.
[0026] Step S300: Using the first edge risk identification node as the dynamic master control node, send the linkage control command package and the target visual features to the N associated risk identification nodes of the N security associated areas to trigger collaborative deployment.
[0027] Specifically, the dynamic master control node is automatically assumed by the edge risk identification node that first detects a security event. It is fully responsible for the coordinated command, data aggregation, and drone control of this event. After the event is handled, it automatically releases its master control authority and reverts to a normal edge node. The linkage control instruction package refers to a standardized inter-edge communication data packet, which contains core information such as a unique event ID, target type, risk level, deployment duration, timestamp, collaborative tracking requirements, and data feedback rules.
[0028] Specifically, after completing the location of N security-related areas, the first edge risk identification node is upgraded to the sole dynamic master control node for this security event, generates a globally unique event identifier ID, constructs a linkage control command package, and simultaneously sends the extracted target visual features as attachments to the N associated risk identification nodes corresponding to the N security-related areas via a low-latency P2P communication link between edge nodes. Upon receiving the command, each associated risk identification node marks the event as the highest priority task, temporarily suspends low-priority secondary risk identification and inference tasks, allocates more than 80% of the available computing resources to the target matching and tracking module, calls the cosine similarity matching algorithm, and performs frame-by-frame rapid comparison of the received target visual features with the real-time video stream of its own monitoring area. Once a target with a matching degree exceeding a preset threshold is found, continuous tracking is immediately initiated, automatically adjusting the pan-tilt angle and optical focal length of the corresponding camera to ensure that the target is always in the center of the image, and transmitting the target's spatial coordinates, motion status, and confidence score back to the dynamic master control node in real time at preset time intervals, forming a collaborative deployment network with multi-node cross-coverage.
[0029] This step, through a dynamic master control node mechanism and direct P2P communication between edge nodes, eliminates the central node bottleneck and single point of failure risk of traditional centralized collaborative architecture, shortens the response time of cross-node collaborative deployment, and achieves seamless collaboration of multiple nodes and full-domain tracking without blind spots.
[0030] Step S400: The first edge risk identification node, based on the N node tracking data returned by the N associated risk identification nodes, coordinates with the alarm drone to adaptively alarm the security event target under dynamic trajectory replanning.
[0031] Specifically, node tracking data consists of real-time target status data packets transmitted back by each associated risk identification node. These packets include the target's precise spatiotemporal coordinates, identification confidence level, corresponding camera number, tracking status identifier, and motion parameter update timestamps. Adaptive alarm refers to an intelligent alarm mechanism where the drone automatically adjusts the alarm mode, intensity, and frequency based on target type, relative distance, speed, and risk level.
[0032] Specifically, the first edge risk identification node, acting as a dynamic master control node, continuously receives node tracking data from N associated risk identification nodes, performs spatiotemporal alignment and outlier filtering on all data; using the fused target's real-time position and speed as input, combined with obstacle distribution information from the park's geographical model, it dynamically replans the drone's trajectory, adjusting the drone's flight speed, altitude, and heading in real time; simultaneously, based on the target type and the real-time relative distance between the drone and the target, it triggers corresponding adaptive alarm strategies; during the alarm process, the first edge risk identification node continuously receives updated tracking data from associated nodes, iteratively optimizing the drone's trajectory and alarm parameters until the security incident is resolved.
[0033] This step solves the problem of easily losing targets in traditional fixed-track tracking by fusing multi-source tracking data and performing millisecond-level dynamic replanning of the trajectory. It achieves continuous and stable tracking of high-speed moving targets. At the same time, the adaptive alarm mechanism based on target characteristics improves the effectiveness and accuracy of alarms and avoids invalid alarms from interfering with the normal order of the park.
[0034] Furthermore, the method also includes: the first edge risk identification node using the initial spatiotemporal coordinates of the security event target as a reference waypoint to initiate a drone scheduling request to the cloud management platform; the cloud management platform constructs an initial reference flight path based on the initial spatiotemporal coordinates, schedules the alarm drone to the initial spatiotemporal coordinates, and delegates the flight control authority of the alarm drone to the first edge risk identification node.
[0035] Specifically, the initial spatiotemporal coordinates refer to the three-dimensional spatial coordinates and precise timestamp corresponding to the first detection of a security event target by the differentiated identification task model and the achievement of a pre-set confidence threshold. These coordinates are calculated by matching the camera calibration parameters of the edge perception matrix with the park's geographical model, and include latitude, longitude, altitude, and the event trigger time. The reference waypoint refers to the core reference point for UAV flight mission planning, i.e., the initial target position that the UAV needs to reach first, serving as the benchmark for subsequent dynamic adjustments to the flight path. The cloud-based management platform is the central system responsible for the unified management of all security resources in the park, undertaking the core functions of UAV cluster status monitoring, task scheduling, flight path planning, and flight permission control. The initial reference flight path is a preliminary flight path generated by the cloud based on the UAV's current position and the reference waypoint, including the takeoff point, waypoints, target point, flight altitude, cruising speed, and no-fly zone avoidance rules. The alarm UAV is a small multi-rotor UAV equipped with an audible and visual alarm device, a high-definition airborne camera, a GPS positioning module, and an autonomous flight control system, specifically designed for proximity alarms and target tracking in security events.
[0036] Specifically, after the first edge risk identification node completes the initial identification of the security event target and obtains its initial spatiotemporal coordinates, it uses these coordinates as the reference waypoint for the UAV mission and sends a scheduling request to the cloud management platform, including the target type, risk level, reference waypoint, and the number of requested UAVs. After receiving the request, the cloud management platform queries the current location, remaining battery power, mission status, and equipment health of all online alarm UAVs in real time, selects the UAVs that are closest to the reference waypoint and meet the mission requirements, and generates an initial reference flight path that avoids all obstacles by combining the no-fly zone, building height, and obstacle distribution information in the park's geographical model. It then sends a takeoff command to the selected UAV and distributes the initial reference flight path to the UAV flight control system. At the same time, the cloud temporarily releases global flight control of the UAV and delegates real-time flight control authority to the first edge risk identification node that initiated the request, establishing an end-to-end low-latency communication link between the edge node and the UAV, allowing the edge node to subsequently send dynamic replanning commands for the flight path based on the target's motion status.
[0037] This step utilizes a hybrid control mode that combines global resource scheduling in the cloud with real-time control permissions delegated to the edge. This achieves optimal global allocation of drone cluster resources and solves the network latency problem associated with remotely controlling drones from the cloud.
[0038] Furthermore, the method also includes: dividing the three-dimensional geographic space of the park into multiple physical monitoring areas based on a digital twin model, and networking multiple edge risk identification nodes deployed in the multiple physical monitoring areas according to the topology of the park's monitoring resources and the physical spatial coverage relationship of the multiple physical monitoring areas to obtain multiple edge perception matrices; during the process of receiving and locally storing the original video stream returned by the corresponding edge perception matrix, each edge risk identification node performs real-time risk identification through the differentiated identification task model configured by its own node, outputs a status summary heartbeat, and periodically uploads the status summary heartbeat to the cloud management platform through a 5G communication link.
[0039] Specifically, the physical monitoring area is an independent geospatial unit divided based on a digital twin model. The size of each area is determined by the computing power capacity of a single edge node and the effective coverage of the camera, and is uniformly managed by an edge risk identification node. The monitoring resource topology is a collection of the physical deployment locations, network connections, and data transmission paths of all monitoring cameras, edge risk identification nodes, communication base stations, and other equipment within the park. The edge perception matrix refers to a distributed acquisition network composed of all monitoring cameras deployed within the same physical monitoring area and uniformly managed by the corresponding edge risk identification node, responsible for all-weather video data acquisition in that area. Local circular storage refers to the first-in, first-out (FIFO) storage mechanism adopted by the edge nodes. When the storage capacity reaches a preset threshold, the oldest video data is automatically overwritten, achieving continuous local storage of the original video stream. The status summary heartbeat refers to the lightweight status information generated by the edge nodes, including core parameters such as node operating status, computing power utilization, storage utilization, risk identification result statistics, and device health.
[0040] Specifically, during the system deployment phase, a digital twin model of the park is loaded. Combining the maximum computing power capacity of a single edge node and the effective monitoring coverage radius of a single camera, the three-dimensional geographic space of the park is evenly divided into multiple non-overlapping and clearly defined physical monitoring areas, ensuring that the data volume of all cameras in each area does not exceed the processing capacity of the corresponding edge node. Based on the topology of the park's monitoring resources and the spatial adjacency and coverage overlap of each physical monitoring area, all monitoring cameras deployed in each physical monitoring area are bound one-to-one with the corresponding edge risk identification nodes to form a network, creating multiple independently operating edge perception matrices. After the system is officially launched, each edge risk identification node continuously receives the original video streams transmitted back from the corresponding edge perception matrix. The video streams are stored in real time using a local circular storage mechanism. At the same time, the pre-configured differentiated identification task models of each node are driven to perform real-time risk identification inference frame by frame on the video stream. A status summary heartbeat containing node operating status, computing power utilization, risk identification statistics, and device health is generated every 5 seconds. The status summary heartbeat is periodically uploaded to the cloud management platform through a 5G communication link, realizing global monitoring of the entire park's security system operating status from the cloud.
[0041] This step achieves optimal configuration and load balancing of campus security resources through refined area division and distributed networking architecture based on digital twins. At the same time, it reduces the computing and bandwidth pressure on the cloud through local processing at edge nodes and a lightweight heartbeat synchronization mechanism, ensuring the high reliability and scalability of the system.
[0042] Furthermore, the method also includes: performing regional risk profile analysis based on the park functional attribute information of the first physical monitoring area to determine the dominant security risk type; retrieving historical security event logs of the first physical monitoring area and performing multi-dimensional risk aggregation to obtain multiple historical occurrence frequencies of multiple secondary security risk types; retrieving the dominant risk pre-trained model and multiple secondary risk pre-trained models from the basic model library according to the dominant security risk type and multiple secondary security risk types, and performing local fine-tuning training to obtain the dominant risk fine-tuning model and multiple secondary risk fine-tuning models; and further, for the dominant... After setting the highest priority computing power for the primary risk fine-tuning model and allocating differentiated computing power scheduling priorities to the secondary risk fine-tuning models based on the frequency of occurrence of multiple historical events, the differentiated identification task model is constructed by connecting the primary risk fine-tuning model and the multiple secondary risk fine-tuning models in parallel. Based on a preset time window, the original video stream transmitted back from the first edge perception matrix is slid-segmented to obtain multiple time-series segments, which are input into the differentiated identification task model. Parallel risk identification reasoning is performed through the primary risk fine-tuning model and the multiple secondary risk fine-tuning models to output the security event target.
[0043] Specifically, the functional attribute information of the park is the basic information describing the purpose and security level of the physical monitoring area, including area type, personnel density, material storage type, access permissions, and security management requirements. The dominant security risk type refers to the security risk type that poses the greatest threat to the area, has the highest probability of occurrence, or is the most serious. Secondary security risk types refer to other security risk types that may occur in the area besides the dominant risk. The basic model library refers to a resource library uniformly maintained in the cloud that contains pre-trained models for various security risks; all models have been fully trained on a general security dataset.
[0044] Specifically, the first edge risk identification node acquires the functional attribute information of the park in the first physical monitoring area, and performs regional risk profiling analysis in conjunction with security industry standards and park security management regulations to assess the probability and severity of various security risks, and determine the dominant security risk type that poses the greatest security threat to the area. Next, it retrieves the historical security event logs of the area over the past 12 months, performs multi-dimensional risk aggregation statistics by event type, occurrence time, and severity level, filters out extremely low-risk types with occurrence frequencies below a preset threshold, and obtains multiple secondary security risk types and their corresponding historical occurrence frequencies. A model retrieval request is sent to the cloud management platform to retrieve the corresponding pre-trained model for the dominant risk and multiple pre-trained models for the secondary risks from the cloud-based basic model library. Lightweight fine-tuning training is performed locally on the edge node using a transfer learning strategy, freezing most parameters of the model's backbone network and only fine-tuning the final detection head and classification. First, a dominant risk fine-tuning model and multiple secondary risk fine-tuning models adapted to the specific scenario of the area are obtained. Then, the dominant risk fine-tuning model is given the highest computing power scheduling priority and a fixed proportion of core computing power quota is allocated. The risk weight of each secondary security risk type is calculated based on its historical occurrence frequency and hazard level. Differentiated dynamic computing power quotas are allocated to the corresponding secondary risk fine-tuning models. By connecting all fine-tuning models in parallel and completing interface adaptation, the final construction of the differentiated identification task model is completed. Based on the preset time window length and sliding step size, the original video stream returned by the first edge perception matrix is continuously slid segmented to generate multiple overlapping time-series video segments. These segments are then input into the differentiated identification task model in sequence. Parallel risk identification inference is performed through the dominant risk fine-tuning model and multiple secondary risk fine-tuning models, and the security event target identification result containing target category, confidence score, and spatial bounding box coordinates is output.
[0045] The differential recognition task model adopts a four-layer modular loosely coupled architecture consisting of an input preprocessing layer, a multi-model parallel inference layer, a computing power dynamic scheduling layer, and a result fusion output layer. All modules support hot-swapping and independent upgrades. The input preprocessing layer is responsible for uniformly receiving the raw video stream returned by the edge perception matrix, performing H.265 / H.264 hardware decoding, uniform scaling to 1080P resolution, normalizing pixel values to the [0, 1] interval, and performing sliding time window segmentation according to a preset 1-second window length and 0.5-second step size to generate continuous temporal video segments for input to the subsequent inference layer. The multi-model parallel inference layer consists of one dominant risk fine-tuning model and 3-5 secondary risk fine-tuning models. The dominant risk model adopts a precision-priority medium-sized network architecture, such as YOLOv8m or RT-DETR-L, while the secondary risk models adopt a speed-priority lightweight network architecture, such as YOLOv8n or MobileNetV3-SSD. Each model has an independent inference engine instance and isolated memory partitions to avoid resource contention between different models. The dynamic computing power scheduling layer, as the core control unit, runs a priority-based weighted round-robin scheduling algorithm. It monitors the inference latency, GPU utilization, and detection confidence of each model in real time. When the inference latency of the dominant risk model exceeds 50ms or the detection confidence is below 0.7, it automatically allocates 20%-30% of the computing power quota from lower-priority secondary models. When a secondary risk model has no effective output for 10 consecutive time windows, its computing power quota is automatically reduced to the minimum threshold. The result fusion output layer receives the parallel inference results from all models, performs category-level non-maximum suppression to remove duplicate detection boxes, and performs confidence-weighted fusion according to model priority. Specifically, the dominant model has a weight of 0.7, and the secondary models have a weight of 0.3, ultimately outputting the recognition result.
[0046] The differentiated identification task model employs a full-process construction and training method involving cloud-based pre-training, regional customization, edge-based lightweight fine-tuning, and dynamic iteration. Functional attribute labels, physical boundary information, distribution of sensitive facilities, and personnel and vehicle traffic data of the target physical monitoring area are extracted from the park's digital twin system. Combined with security industry standards and park security management regulations, regional risk profiling analysis is conducted to determine the dominant security risk type with the greatest impact on the area's security. Historical security event logs from the past 12 months are retrieved for multi-dimensional data cleaning and aggregation statistics. The frequency, average duration, and severity level of each secondary security risk type are calculated, filtering out extremely low-risk types with an occurrence frequency below 1%, thus determining the number and types of secondary risk models to be deployed. Pre-trained models for the corresponding categories are retrieved from the cloud-based basic model library. These pre-trained models adopt an RT-DETR-L end-to-end target detection architecture, consisting of a backbone network, a Transformer encoder-decoder, and a detection head. The backbone network uses ResNet-50 to extract multi-scale visual features and fuses semantic information from different levels through a feature pyramid network. The Transformer encoder performs global context modeling on the feature sequence, and the decoder focuses on potential target regions through a deformable attention mechanism. The detection head directly outputs the target category and bounding box coordinates without the need for non-maximum suppression post-processing. Training is divided into two stages: the first stage is full pre-training on a general security dataset containing 1.2 million images to learn the basic features of common targets such as people, vehicles, open flames, and smoke; the second stage is adaptive pre-training on a domain dataset containing 500,000 images of specific scenes in the park to optimize robustness to changes in park environment, lighting, and occlusion. Training employed the AdamW optimizer with an initial learning rate of 1e-4, combining mixed-precision training and gradient accumulation strategies. A total of 100 training epochs were conducted. The model achieved a mAP@0.5 of 96.2% on the park security test set, demonstrating strong basic feature extraction capabilities. Lightweight fine-tuning was performed locally at edge nodes using transfer learning, freezing the parameters of the first 80% of the backbone network and fine-tuning only the final detection and classification heads. Simultaneously, data augmentation operations such as random cropping, flipping, brightness adjustment, and noise addition were performed on local historical security data. The computational priority of each model was calculated using the formula "Risk Priority = Frequency of Occurrence × Hazard Level Coefficient." The dominant risk model was allocated the highest priority and 50%-60% of the fixed computational power, while secondary risk models were allocated dynamic computational power proportional to their risk priority. All fine-tuned models were connected in parallel and interface adaptation was completed to generate the final differentiated recognition task model. The model automatically underwent incremental fine-tuning monthly, incorporating newly added security event data to continuously optimize recognition accuracy.
[0047] This step, through regionally customized model construction and differentiated computing power scheduling mechanisms, ensures both high identification accuracy and low response latency for primary security risks and comprehensive coverage of secondary security risks, while reducing false alarm and false negative rates, even under the condition of limited computing power at edge nodes.
[0048] Furthermore, the method also includes: constructing a time analysis window based on the identification trigger time of the security event target, and constructing a region of interest (ROI) based on the spatial bounding box of the security event target; segmenting the event-related video segments from the original video stream based on the time analysis window and the ROI; performing multimodal feature extraction on the event-related video segments using the output model type corresponding to the security event target as the risk identification context to obtain the target's visual features and spatiotemporal motion features; and performing spatiotemporal trajectory prediction based on the park geographic model and spatiotemporal motion features, using the output model type as the trajectory prediction constraint, to locate the N security-related areas.
[0049] Specifically, the trigger moment refers to the precise timestamp when the differentiated identification task model detects a security event target and the confidence level reaches a preset threshold; it serves as the benchmark for event time dimension analysis. The time analysis window, centered on the trigger moment, extends backward and forward by a preset duration to capture video content encompassing the complete process before and after the event. The spatial bounding box refers to the rectangular bounding box output by the target detection model that completely encloses the security event target, including the upper left and lower right pixel coordinates of the target in the video frame. The output model type indicates the specific model category for detecting the security event target, corresponding to different security risk types, such as open flame / smoke models, personnel intrusion models, and vehicle illegal parking models.
[0050] Specifically, when the differential recognition task model detects a security event target and triggers the recognition event, the first edge risk recognition node constructs a time analysis window that looks back 10 seconds and extends forward 5 seconds, based on the recognition trigger time. Simultaneously, it expands outward by 20% from the target spatial bounding box output by the differential recognition task model to form a Region of Interest (ROI). Based on the time analysis window and the ROI, it accurately segments event-related video clips containing only the events before and after the event and the surrounding area of the target from the locally stored original video stream, removing all irrelevant time clips and background areas. Using the output model type of the detected target as the risk recognition context, it performs multimodal feature extraction to process the event-related video clips, dynamically adjusting the extraction weights of visual and motion features according to different model types to obtain targeted enhanced target visual features and spatiotemporal motion features. Using the output model type as a trajectory prediction constraint, combined with spatial constraint information in the park's geographical model and the extracted spatiotemporal motion features, it predicts the target's future trajectory based on the trajectory prediction model, selecting adjacent physical monitoring areas with high probability trajectory coverage. After deduplication, it locates N security-related areas. In this process, the output model type is used as a key constraint for trajectory prediction: the trajectory prediction model dynamically adjusts its prediction preference for different motion modes, such as people walking, vehicles driving, and fire spread, based on the output model type, so that the generated candidate trajectories are more consistent with the actual behavioral characteristics of the current security event targets, thereby improving the targeting of security-related area positioning.
[0051] The trajectory prediction model employs a target type-aware, geographically constrained, bidirectional GRU attention architecture, consisting of a spatiotemporal feature encoding layer, a geographic constraint embedding layer, a bidirectional GRU temporal inference layer, a target type attention layer, and a probability output layer. The spatiotemporal feature encoding layer encodes the target's coordinates, velocity, acceleration, and other spatiotemporal motion features into a 64-dimensional vector. The geographic constraint embedding layer converts spatial information such as park roads, restricted areas, and buildings into a 32-dimensional constraint vector, which is then concatenated with the motion features and input to the temporal inference layer. The bidirectional GRU layer captures the continuity of target motion through forward and backward modeling. The target type attention layer dynamically weights the motion state at different historical moments based on the differentiated identification task model type, strengthening the motion pattern that matches the target characteristics. The probability output layer generates five candidate trajectories and their probabilities of occurrence. The target type attention layer takes the output model type as input, encodes it into a type-aware vector, and interacts with the spatiotemporal features for weighting. This explicitly uses the output model type as a constraint condition for trajectory prediction, achieving differentiated enhancement of trajectory prediction for targets with different security risks, and providing a more accurate spatiotemporal basis for subsequently locating N security-related areas. During training, the dataset was first pre-trained on a public dataset containing 1.2 million general personnel and vehicle trajectories, using a negative log-likelihood loss function. Then, it was fine-tuned in batches according to target type on a historical security trajectory dataset of the park, using the AdamW optimizer with an initial learning rate of 1e-4 and 50 training rounds.
[0052] This step, through precise data cropping in both time and space dimensions and enhanced feature and trajectory prediction based on risk context, reduces the amount of invalid data processing while improving the targeting of feature extraction and the accuracy of trajectory prediction, laying a solid data foundation for subsequent multi-node collaborative deployment and integrated air-ground tracking.
[0053] Furthermore, the first edge risk identification node, based on the N node tracking data returned by the N associated risk identification nodes, coordinates with the alarm drone to adaptively alarm the security event target under dynamic trajectory replanning. The method includes: the first edge risk identification node receiving the N node tracking data returned by the N associated risk identification nodes and updating the target prediction coordinates of the security event target; performing dynamic replanning of the initial reference trajectory based on the target prediction coordinates to obtain an optimized tracking trajectory, and sending it to the alarm drone; the alarm drone autonomously flying based on the optimized tracking trajectory and approaching the security event target, and then continuously tracking the security event target under close-range locking based on the target's visual characteristics, and continuously executing security alarms.
[0054] Specifically, node tracking data refers to the real-time status information output by associated nodes after receiving linkage instructions and target visual features, performing target re-identification in their own monitoring video streams, continuously locking onto the target, and outputting this information. Target prediction coordinates are obtained by the first edge risk identification node by fusing its own tracking data with that of N associated nodes from multiple sources.
[0055] Specifically, the first edge risk identification node, acting as the dynamic master node for this event, continuously receives node tracking data from N associated risk identification nodes. It performs spatiotemporal alignment and 3σ criterion outlier filtering on all data, eliminating invalid data with a confidence level below 0.6. A Kalman filter algorithm is used to fuse multi-source valid data, updating the target's motion state in real time and predicting its next target coordinates. Based on the target's predicted coordinates, and combined with obstacle distribution information such as buildings, trees, and power lines in the park's geographical model, as well as performance constraints such as the UAV's maximum flight speed, turning radius, and endurance, the initial reference trajectory is dynamically replanned 10 times per second. This generates an optimized tracking trajectory containing a continuous waypoint sequence, flight speed, and altitude, and is then transmitted via 5G low-frequency radio. The delayed communication link sends real-time data to the alarm drone. After receiving the optimized tracking trajectory, the alarm drone immediately adjusts its flight attitude and flies autonomously along the trajectory. When it approaches the target at a preset warning distance of 15-20 meters, it activates the onboard high-definition camera to collect real-time video streams, extracts visual features frame by frame, and performs cosine similarity matching with the received target visual features. When the matching degree exceeds the 0.85 threshold, it completes the close-range locking of the security event target. The PID control algorithm automatically adjusts the onboard gimbal parameters to maintain continuous tracking of the target, and triggers corresponding adaptive security alarms according to the target type and real-time relative distance. For personnel intrusion targets, it uses graded audio and visual alarms; for illegally parked vehicles, it uses voice alarms; and for open flame targets, it uses strong light warning alarms, until the security event is handled.
[0056] This step realizes a closed loop of air-ground collaborative alarm from static coordinate guidance to dynamic visual locking, which significantly improves the ability to continuously, accurately, and automatically track and deal with fast-moving or deliberately escaping targets within the park.
[0057] Furthermore, after the tracking and alarm of the security event target is completed: the first edge risk identification node retrieves the associated monitoring video clips from the corresponding associated risk identification node based on the target's predicted coordinates; the first edge risk identification node retrieves the airborne tracking video clips and original alarm information from the alarmed drone; the first edge risk identification node spatially splices the event-related video clips, associated monitoring video clips, and airborne tracking video clips into key process data, and associates them with the original alarm information to generate a structured alarm evidence chain, which is then compressed and uploaded to the cloud management platform for alarm archiving.
[0058] Specifically, associated surveillance video clips refer to video clips captured by ground cameras deployed in the associated security area while a target passes through the area, including the target's movement trajectory and behavior records from different ground perspectives. Airborne tracking video clips refer to close-range video clips captured by alarm drones using their onboard high-definition cameras after the drones have approached and locked onto the target.
[0059] Specifically, when a security incident target is successfully handled or leaves the park's monitoring range, the first edge risk identification node determines that the tracking and alarm task has ended. Based on the target's complete spatiotemporal trajectory, it sends a data retrieval command to all associated risk identification nodes to retrieve associated monitoring video clips during the target's stay time in each associated area. Simultaneously, it sends a data transmission command to the alarm drone to retrieve the airborne tracking video clips and original alarm information collected by the drone throughout the tracking and alarm process. Subsequently, the first edge risk identification node performs unified preprocessing on all collected video data, aligning the event-related video clips, associated monitoring video clips, and airborne tracking video clips based on the park's geographical model and precise timestamps. It then performs multi-view spatial stitching according to the time sequence and spatial location of the event to generate key process data covering the entire event process. The event ID, target type, risk level, handling time, participation records of each node, and original alarm information are associated and bound with the key process data to generate a structured alarm evidence chain containing complete metadata. The evidence chain is compressed using H.265 hard compression technology and uploaded to the cloud management platform via a 5G communication link to complete the alarm archiving of this security incident.
[0060] This step automatically completes the spatiotemporal stitching of multi-source video data and the generation of structured evidence chains at the edge, replacing the tedious process of manually retrieving and organizing evidence in the traditional way. This ensures the integrity, objectivity, and immutability of the evidence, and improves the standardization of security incident handling and the efficiency of post-incident traceability.
[0061] Furthermore, after the first edge risk identification node archives the alarm, it cancels the collaborative deployment command for the N associated risk identification nodes and requests the cloud management platform to reclaim the flight control authority over the alarmed drone.
[0062] Specifically, once the first edge risk identification node completes the compressed upload of the structured alarm evidence chain and receives the archiving confirmation receipt from the cloud management platform, it terminates its role as the dynamic master control node. Specifically, it broadcasts a collaborative deployment deactivation command to all N associated risk identification nodes. This command includes an event end timestamp and resource reset requirements. Upon receiving the deactivation command, each associated risk identification node immediately ceases its dedicated tracking task for the target of this event, releases the temporarily occupied 80% of computing resources, restarts the suspended secondary risk identification inference task, resets the pan-tilt angle and optical focal length of the corresponding camera to the preset normal monitoring position, and sends a notification to the first edge risk identification node. The first edge risk identification node sends a resource reset completion confirmation; subsequently, the first edge risk identification node sends a flight control permission revocation request to the cloud management platform, which includes the drone number, mission completion time, and flight data summary of this mission; after receiving the request and verifying its accuracy, the cloud management platform immediately revokes the flight control permission of the alarmed drone, sends a return-to-home command to the drone, and the drone automatically returns to the preset takeoff point and performs landing and self-check procedures; the first edge risk identification node sends the final handling report of this incident to the cloud and also restores itself to the normal edge risk identification state, and the entire system returns to the normal full-domain security monitoring mode.
[0063] This step, through a standardized task completion process and an automated resource reset mechanism, enables rapid recovery of system resources and self-healing, ensuring that all participating nodes and devices can be restored to normal monitoring status after the incident is resolved. This effectively avoids system performance degradation caused by long-term resource occupation and guarantees the continuous and stable operation of the park's security system.
[0064] In summary, the air-ground coordinated automatic alarm method for park security provided in this application has the following technical effects: 1. By employing digital twins to divide the park's monitoring areas and completing edge node networking, a customized regional identification model is deployed to identify security targets in real time. Combined with spatiotemporal features, movement trajectories are predicted to lock onto associated areas. Multi-node collaborative deployment and drone dispatch are used for intelligent tracking and alarm activation. Afterwards, evidence archiving and device permission reset are completed. This overall solution abandons the traditional centralized cloud processing model, rationally allocates edge computing power, reduces data transmission latency, and comprehensively improves the park's full-process control capabilities for security incidents, from early warning and response to tracing.
[0065] 2. By designating the edge node that first detects a risk as a dynamic master control node, and issuing linkage command packets and unified target visual features to surrounding related nodes, a comprehensive collaborative deployment mechanism is quickly triggered. Relying on direct communication at the edge to achieve rapid command transmission and on unified visual features to achieve rapid target matching across nodes, the collaborative response time is significantly shortened, establishing a comprehensive, blind-spot-free collaborative security monitoring network.
[0066] 3. By guiding the drone to autonomously approach the target along an optimized tracking trajectory, and relying on pre-stored target visual features to complete close-range lock-on tracking, adaptive hierarchical security alarms are executed simultaneously based on the on-site scenario. This enables the drone to operate autonomously without close-range manual control, and allows for flexible switching of alarm modes to adapt to different security events. It ensures stable and unwavering target tracking while significantly improving the actual effectiveness of on-site warnings and deterrents.
[0067] Example 2, based on the same inventive concept as the air-ground collaborative automatic alarm method for park security in the aforementioned examples, such as... Figure 2 As shown in the embodiment of this application, an air-ground collaborative automatic alarm system for park security is provided. The system includes: a multimodal feature extraction module 11, used to extract multimodal features from the event-related video segments of the original video stream when the first edge risk identification node receives the original video stream of the first physical monitoring area returned by the first edge perception matrix and drives the differentiated identification task model to identify the security event target, thereby obtaining the target's visual features and spatiotemporal motion features; an associated area positioning module 12, used by the first edge risk identification node to perform spatiotemporal trajectory prediction based on the park geographical model and the spatiotemporal motion features, and locate N security-related areas; a collaborative deployment triggering module 13, used by the first edge risk identification node as the dynamic master control node to send a linkage control command package and the target's visual features to the N associated risk identification nodes of the N security-related areas to trigger collaborative deployment; and an adaptive alarm module 14, used by the first edge risk identification node to trigger an alarm drone to adaptively alarm the security event target under dynamic replanning of the flight path based on the N node tracking data returned by the N associated risk identification nodes.
[0068] Furthermore, the system is also used to perform the following steps: the first edge risk identification node uses the initial spatiotemporal coordinates of the security event target as the reference waypoint and initiates a drone scheduling request to the cloud management platform; the cloud management platform constructs an initial reference trajectory based on the initial spatiotemporal coordinates, schedules the alarm drone to the initial spatiotemporal coordinates, and delegates the flight control authority of the alarm drone to the first edge risk identification node.
[0069] Furthermore, the system is also used to perform the following steps: dividing the three-dimensional geographic space of the park into multiple physical monitoring areas based on the digital twin model, and according to the topology of the park's monitoring resources and the physical spatial coverage relationship of the multiple physical monitoring areas, networking and connecting multiple edge risk identification nodes deployed in the multiple physical monitoring areas to obtain multiple edge perception matrices; during the process of receiving and locally storing the original video stream returned by the corresponding edge perception matrix, each edge risk identification node performs real-time risk identification through the differentiated identification task model configured by its own node, outputs a status summary heartbeat, and periodically uploads the status summary heartbeat to the cloud management platform through the 5G communication link.
[0070] Furthermore, the system is also used to perform the following steps: based on the park functional attribute information of the first physical monitoring area, perform regional risk profile analysis to determine the dominant security risk type; retrieve historical security event logs of the first physical monitoring area, perform multi-dimensional risk aggregation, and obtain multiple historical occurrence frequencies of multiple secondary security risk types; based on the dominant security risk type and multiple secondary security risk types, retrieve the corresponding pre-trained models of the dominant risk and multiple secondary risk from the basic model library, and perform local fine-tuning training to obtain the fine-tuned model of the dominant risk and multiple secondary risk; and perform local fine-tuning training on the dominant risk and multiple secondary risk. The dominant risk fine-tuning model is given the highest priority computing power, and the secondary risk fine-tuning models are assigned differentiated computing power scheduling priorities based on the multiple historical occurrence frequencies. The differentiated identification task model is constructed by connecting the dominant risk fine-tuning model and the multiple secondary risk fine-tuning models in parallel. The original video stream returned by the first edge perception matrix is divided by sliding segmentation based on a preset time window to obtain multiple time-series segments, which are input into the differentiated identification task model. Parallel risk identification reasoning is performed by the dominant risk fine-tuning model and the multiple secondary risk fine-tuning models to output the security event target.
[0071] Furthermore, the system is also used to perform the following steps: constructing a time analysis window based on the identification trigger time of the security event target, and constructing a region of interest (ROI) based on the spatial bounding box of the security event target; segmenting the event-related video segments from the original video stream based on the time analysis window and the ROI; performing multimodal feature extraction on the event-related video segments using the output model type corresponding to the security event target as the risk identification context to obtain the target's visual features and spatiotemporal motion features; and performing spatiotemporal trajectory prediction based on the park's geographical model and spatiotemporal motion features, using the output model type as the trajectory prediction constraint, to locate the N security-related areas.
[0072] Furthermore, the system is also used to perform the following steps: the first edge risk identification node receives N node tracking data returned by the N associated risk identification nodes, and updates the target prediction coordinates of the security event target; based on the target prediction coordinates, the initial reference trajectory is dynamically replanned to obtain an optimized tracking trajectory, which is then sent to the alarm drone; the alarm drone autonomously flies based on the optimized tracking trajectory and approaches the security event target, and then continuously tracks the security event target under close-range locking based on the target's visual characteristics, and continuously executes security alarms.
[0073] Furthermore, the system is also used to perform the following steps: after the tracking and alarm of the security event target ends: the first edge risk identification node retrieves the associated monitoring video clip from the corresponding associated risk identification node based on the target's predicted coordinates; the first edge risk identification node retrieves the airborne tracking video clip and the original alarm information from the alarmed drone; the first edge risk identification node spatially splices the event-related video clip, the associated monitoring video clip, and the airborne tracking video clip into key process data, and associates them with the original alarm information to generate a structured alarm evidence chain, which is then compressed and uploaded to the cloud management platform for alarm archiving.
[0074] Furthermore, the system is also used to perform the following steps: after the first edge risk identification node archives the alarm, it releases the collaborative deployment command for the N associated risk identification nodes and requests the cloud management platform to revoke the flight control authority over the alarmed drone.
[0075] Example 3: Based on the same inventive concept as the air-ground coordinated automatic alarm method for park security in the foregoing examples, this application also provides a computer-readable storage medium storing a computer program, which, when executed, implements the steps of the air-ground coordinated automatic alarm method for park security described in any one of the above examples.
[0076] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. An air-ground coordinated automatic alarm method for park security, characterized in that, The method includes: When the first edge risk identification node receives the original video stream of the first physical monitoring area transmitted back by the first edge perception matrix, and drives the differential identification task model to identify the security event target, it performs multimodal feature extraction by segmenting the event-related video segments from the original video stream to obtain the target's visual features and spatiotemporal motion features. The first edge risk identification node performs spatiotemporal trajectory prediction based on the park's geographical model and the spatiotemporal motion characteristics to locate N security-related areas; Using the first edge risk identification node as the dynamic master control node, a linkage control instruction packet and the target visual features are sent to the N associated risk identification nodes of the N security-related areas to trigger collaborative deployment; The first edge risk identification node, based on the N node tracking data returned by the N associated risk identification nodes, coordinates with the alarm drone to adaptively alarm the security event target under dynamic trajectory replanning.
2. The air-ground coordinated automatic alarm method for park security as described in claim 1, characterized in that, The method further includes: The first edge risk identification node uses the initial spatiotemporal coordinates of the security event target as the reference waypoint to initiate a drone scheduling request to the cloud management platform; The cloud-based management platform constructs an initial reference trajectory based on the initial spatiotemporal coordinates, dispatches the alarm drone to the initial spatiotemporal coordinates, and delegates the flight control authority of the alarm drone to the first edge risk identification node.
3. The air-ground coordinated automatic alarm method for park security as described in claim 2, characterized in that, The method further includes: Based on the digital twin model, the three-dimensional geographic space of the park is divided into multiple physical monitoring areas. According to the topology of the park's monitoring resources and the physical spatial coverage relationship of the multiple physical monitoring areas, multiple edge risk identification nodes deployed in the multiple physical monitoring areas are networked and connected to obtain multiple edge perception matrices. During the process of receiving and locally storing the original video stream returned by the corresponding edge perception matrix, each edge risk identification node performs real-time risk identification through the differentiated identification task model configured by its own node, outputs a status summary heartbeat, and periodically uploads the status summary heartbeat to the cloud management platform through the 5G communication link.
4. The air-ground coordinated automatic alarm method for park security as described in claim 3, characterized in that, The method further includes: Based on the park's functional attribute information of the first physical monitoring area, a regional risk profile analysis is conducted to determine the dominant security risk type; Retrieve historical security event logs from the first physical monitoring area, perform multi-dimensional risk aggregation, and obtain multiple historical occurrence frequencies of multiple secondary security risk types; Based on the dominant security risk type and multiple secondary security risk types, the dominant risk pre-training model and multiple secondary risk pre-training models are retrieved from the basic model library and then fine-tuned locally to obtain the dominant risk fine-tuning model and multiple secondary risk fine-tuning models. After setting the highest priority computing power for the dominant risk fine-tuning model and allocating differentiated computing power scheduling priorities to the multiple secondary risk fine-tuning models based on the multiple historical occurrence frequencies, the construction of the differentiated identification task model is completed by connecting the dominant risk fine-tuning model and the multiple secondary risk fine-tuning models in parallel. The original video stream transmitted back by the first edge perception matrix is segmented based on a preset time window to obtain multiple time-series segments. These segments are then input into the differentiated identification task model. Parallel risk identification reasoning is performed through the dominant risk fine-tuning model and multiple secondary risk fine-tuning models to output the security event target.
5. The air-ground coordinated automatic alarm method for park security as described in claim 4, characterized in that, The method further includes: A time analysis window is constructed based on the identification trigger time of the security event target, and a region of interest (ROI) is constructed based on the spatial bounding box of the security event target. Based on the time analysis window and the region of interest (ROI), the event-related video segments are segmented from the original video stream; Using the output model type corresponding to the security event target as the risk identification context, multimodal feature extraction is performed on the event-related video segments to obtain the target's visual features and spatiotemporal motion features; Using the output model type as the trajectory prediction constraint, spatiotemporal trajectory prediction is performed based on the park geographical model and spatiotemporal motion characteristics to locate the N security-related areas.
6. The air-ground coordinated automatic alarm method for park security as described in claim 2, characterized in that, The first edge risk identification node, based on the N node tracking data returned by the N associated risk identification nodes, coordinates with the alarm drone to adaptively alarm the security event target under dynamic trajectory replanning. The method includes: The first edge risk identification node receives N node tracking data returned by the N associated risk identification nodes and updates the target prediction coordinates of the security event target; Based on the predicted target coordinates, the initial reference trajectory is dynamically replanned to obtain an optimized tracking trajectory, which is then sent to the alarm drone. The alarm drone autonomously flies based on the optimized tracking trajectory and approaches the security event target. It then continuously tracks the security event target based on the target's visual characteristics and continuously executes security alarms.
7. The air-ground coordinated automatic alarm method for park security as described in claim 6, characterized in that, After the tracking and alarm for the aforementioned security incident target has ended: The first edge risk identification node retrieves the associated monitoring video clip from the corresponding associated risk identification node based on the target predicted coordinates; The first edge risk identification node retrieves onboard tracking video clips and original alarm information from the alarmed drone; The first edge risk identification node spatially splices the event-related video clips, associated monitoring video clips, and airborne tracking video clips into key process data, and associates them with the original alarm information to generate a structured alarm evidence chain, which is then compressed and uploaded to the cloud management platform for alarm archiving.
8. The air-ground coordinated automatic alarm method for park security as described in claim 2, characterized in that, After the first edge risk identification node archives the alarm, it releases the collaborative deployment command for the N associated risk identification nodes and requests the cloud management platform to reclaim the flight control authority over the alarmed drone.
9. An air-ground coordinated automatic alarm system for park security, characterized in that: The system is used to execute the air-ground coordinated automatic alarm method for park security as described in any one of claims 1 to 8, the system comprising: The multimodal feature extraction module is used to extract multimodal features from the original video stream of the first physical monitoring area transmitted back by the first edge perception matrix when the first edge risk identification node receives the original video stream of the first physical monitoring area and drives the differential identification task model to identify the security event target. This module then extracts the target visual features and spatiotemporal motion features by segmenting the event-related video segments from the original video stream. The associated area positioning module is used by the first edge risk identification node to perform spatiotemporal trajectory prediction based on the park geographical model and the spatiotemporal motion features, and to locate N security associated areas. The collaborative deployment triggering module is used to send a linkage control instruction package and the target visual features to the N associated risk identification nodes of the N security-related areas, with the first edge risk identification node as the dynamic master control node, to trigger collaborative deployment. The adaptive alarm module is used by the first edge risk identification node to coordinate with the alarm drone to adaptively alarm the security event target based on the N node tracking data returned by the N associated risk identification nodes and under the dynamic replanning of the flight path.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which is used to execute the air-ground collaborative automatic alarm method for park security as described in any one of claims 1 to 8.