Low-altitude security control method and device, medium and equipment

Through the "sky-air-ground" collaborative perception and intelligent decision-making system, the entire chain of collaborative linkage for low-altitude targets has been realized, solving the problems of large detection blind spots, weak information fusion, and insufficient collaborative linkage in low-altitude security, and improving the detection, tracking and handling capabilities of the low-altitude security system.

CN121921670APending Publication Date: 2026-04-24XIAN CHENHANG EXCELLENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN CHENHANG EXCELLENCE TECH CO LTD
Filing Date
2026-01-07
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing low-altitude security measures suffer from large detection blind spots, weak information fusion capabilities, and insufficient coordination, making it difficult to effectively detect, track, and deal with low-altitude, slow-moving, and small targets.

Method used

By constructing an integrated system for collaborative perception and intelligent decision-making across space, air, and ground, and utilizing space-based, airborne, and ground-based sensing equipment for wide-area remote sensing image scanning, target detection, tracking, positioning, and identification, combined with multi-source information fusion and risk assessment models, a full-chain collaborative linkage for low-altitude targets can be achieved.

Benefits of technology

It enhances the wide-area detection, precise tracking, and intelligent identification capabilities for low-altitude targets, improves the coverage, real-time response, and collaborative handling of low-altitude security systems, and effectively overcomes the shortcomings of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921670A_ABST
    Figure CN121921670A_ABST
Patent Text Reader

Abstract

The invention discloses a low-altitude security control method and device, a medium and equipment, and belongs to the technical field of low-altitude security and control, and the method comprises the steps: carrying out the periodic scanning of a control region through space-based sensing equipment, and obtaining a wide-area remote sensing image; analyzing the remote sensing image based on a pre-trained target detection model, and detecting a potential target; tracking and positioning the target through the air-based sensing equipment and the tracking and positioning model to obtain coordinates; locking a target and acquiring an image according to the coordinates through foundation sensing equipment, and processing the image through a target recognition model to obtain a target category probability distribution vector; fusing the remote sensing image, the target coordinate and the category probability distribution vector to generate a low-altitude security situation map; and constructing a risk assessment model, assessing a target threat level in real time based on the situation map, and dynamically generating an optimal disposal scheme. According to the invention, the coverage range, the response speed and the co-processing capability of the low-altitude security and protection system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of low-altitude safety and control technology, specifically relating to a low-altitude safety and control method, device, medium and equipment. Background Technology

[0002] With the increasing prevalence of low-altitude aircraft, including drones, they bring convenience to logistics, agriculture, surveying and mapping, but also pose serious challenges to low-altitude security in key areas (such as airports, nuclear power plants, military bases, and major event venues). These "low, slow, and small" targets are characterized by being difficult to detect, track, and handle.

[0003] Existing low-altitude security measures suffer from the following main shortcomings: 1. Limited detection methods: They typically rely on single ground-based radars or optoelectronic equipment, resulting in blind spots and insufficient detection capabilities against ultra-low-altitude, slow-moving, small targets. 2. Weak information fusion capabilities: Detection information from different sources (such as radar tracks, video images, and ADS-B signals) is isolated from each other, forming information silos and failing to create a unified and comprehensive air situation awareness.

[0004] 3. Lack of coordination and collaboration: Space-based (satellites), air-based (early warning drones), and ground-based (monitoring points) platforms operate independently, failing to form an effective collaborative detection and joint response capability.

[0005] In view of the above problems, there is an urgent need for a low-altitude security control solution that can integrate multi-dimensional resources of air, space, and ground to achieve coordinated linkage across the entire chain from perception and decision-making to handling. Summary of the Invention

[0006] In view of the shortcomings of the existing technology, this application aims to provide a low-altitude security control method, device, medium and equipment with wide coverage and efficient collaborative handling capabilities. This application aims to improve the ability to detect, accurately track, intelligently identify and handle dynamic threats to "low, slow and small" targets in a wide area.

[0007] To achieve the above objectives, this application provides the following technical solution: A low-altitude security control method includes: periodically scanning the controlled area using space-based sensing equipment to acquire wide-area remote sensing images; analyzing the wide-area remote sensing images based on a pre-trained target detection model to detect potential targets in the wide-area remote sensing images; tracking and locating the detected potential targets using space-based sensing equipment based on a pre-trained tracking and positioning model to obtain the coordinates of the potential targets; locking onto the potential targets and acquiring images using ground-based sensing equipment based on the coordinates, and analyzing and processing the images using a pre-trained target recognition model to obtain a target category probability distribution vector; fusing the wide-area remote sensing images, the coordinates of the potential targets, and the target category probability distribution vector to obtain a low-altitude security situation map; constructing a risk assessment model, and assessing the threat level of potential targets in real time based on the low-altitude security situation map, and dynamically generating the optimal response plan based on the threat level.

[0008] Optionally, the target detection model includes: a pulse-triggered convolutional network for detecting the sparsity of moving targets and abnormal regions in wide-area remote sensing images; a dynamic computation sub-network for removing the influence of confusing factors on target recognition through causal reasoning; a causal intervention reasoning module for performing multi-path feature fusion and lightweight classification regression on the region of interest output by the pulse-triggered convolutional network to complete target localization and preliminary recognition; and a spatiotemporal memory module for recording and associating historical target trajectories and features.

[0009] Optionally, the pulse-excited convolutional network includes: a multi-scale downsampling layer, a pulse-excited layer, and a pulse fusion and coding layer. The multi-scale downsampling layer is used to perform multi-scale downsampling on wide-area remote sensing images to extract multi-scale primary features. The pulse-excited layer is used to perform sparsification filtering on the multi-scale primary features to generate a pulse sparse pulse feature map. The pulse fusion and coding layer is used to fuse the sparse pulse feature map to obtain a pulse coding map that marks the region of interest.

[0010] Optionally, the dynamic computation sub-network includes: a dynamic routing encoder, an adaptive feature refinement module, and a lightweight classification and regression head. The dynamic routing encoder is used to select multiple convolutional paths based on the pulse code map for weighted fusion to extract multi-scale and multi-receptive field feature representations. The adaptive feature refinement module is used to perform soft selection and enhancement on the multi-scale and multi-receptive field feature representations to obtain refined feature maps. The lightweight classification and regression head is used to classify and review the refined feature maps and output the target class probability and bounding box coordinates.

[0011] Optionally, the tracking and positioning model includes: a multi-source feature input layer, a temporal feature fusion layer, a trajectory memory and query layer, a trajectory prediction and association layer, an adaptive filtering positioning layer, and a fusion output layer connected in sequence. The multi-source feature input layer receives and processes heterogeneous feature data from space-based, air-based, and ground-based sensors to achieve unified mapping across modal feature spaces. The temporal feature fusion layer captures the dynamic continuity and multi-scale pattern features of target motion based on heterogeneous feature data. The trajectory memory and query layer achieves target identity association and trajectory segment matching across time periods. The trajectory prediction and association layer predicts the future motion state of the target and distinguishes between multiple targets based on the current state and historical information provided by the multi-source feature input layer, the temporal feature fusion layer, and the trajectory memory and query layer. The adaptive filtering localization layer is used to perform adaptive state estimation and coordinate smoothing on the distinguished multi-targets; the fusion output layer is used to integrate the output information of the multi-source feature input layer, the temporal feature fusion layer, the trajectory memory and query layer, the trajectory prediction and association layer, and the adaptive filtering localization layer, and generate the final localization result.

[0012] Optionally, the step of fusing the wide-area remote sensing image, the coordinates of potential targets, and the probability distribution vector of target categories to obtain a low-altitude security situation map includes: aligning and standardizing the wide-area remote sensing image, the coordinates of potential targets, and the probability distribution vector of target categories; constructing a situation layer based on the standardized wide-area remote sensing image, the coordinates of potential targets, and the probability distribution vector of target categories; and rendering the constructed situation layer to obtain a low-altitude security situation map.

[0013] Optionally, the risk assessment model includes: a multi-dimensional feature fusion module, a dynamic threat calculation engine, and an interpretable output module. The multi-dimensional feature fusion module is used to uniformly encode and fuse multi-source heterogeneous features from the low-altitude security situation map; the dynamic threat calculation engine is used to comprehensively analyze the target's intent, capabilities, opportunities, and collaborative behaviors to calculate an interpretable threat level score; and the interpretable output module is used to implement threat score level mapping, decision support, and output.

[0014] This application also provides a low-altitude security control device, the device comprising: a first acquisition module, used to periodically scan the control area using space-based sensing equipment to acquire wide-area remote sensing images; a target detection module, used to analyze the wide-area remote sensing images based on a pre-trained target detection model to detect potential targets in the wide-area remote sensing images; a target localization module, used to track and locate the detected potential targets using space-based sensing equipment based on a pre-trained tracking and localization model to obtain the coordinates of the potential targets; a second acquisition module, used to lock onto the potential targets based on the coordinates using ground-based sensing equipment and acquire images, and to analyze and process the images using a pre-trained target recognition model to obtain a target category probability distribution vector; a fusion module, used to fuse the wide-area remote sensing images, the coordinates of the potential targets, and the target category probability distribution vector to obtain a low-altitude security situation map; and an evaluation module, used to construct a risk assessment model, and to evaluate the threat level of potential targets in real time based on the low-altitude security situation map, and to dynamically generate the optimal handling plan based on the threat level.

[0015] This application also provides a storage medium including instructions that, when executed on a computer, cause the computer to perform the method as described in the preceding claim.

[0016] This application also provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any of the preceding claims.

[0017] Compared with the prior art, the beneficial effects of this application are as follows: This application constructs an integrated system for collaborative perception and intelligent decision-making across the "sky-air-ground" domain, enabling wide-area detection, precise tracking, intelligent identification, and dynamic risk assessment of low-altitude targets. It effectively overcomes the shortcomings of traditional methods, such as large blind spots, weak information fusion, and poor collaborative capabilities, and enhances the overall coverage, real-time response, and collaborative handling of low-altitude security systems. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a low-altitude security control method according to an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a low-altitude security control device provided in another embodiment of this application; Figure 3 This is a schematic diagram of the structure of a storage medium provided in another embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0020] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0021] In this invention, unless otherwise explicitly specified and limited, the terms "connection," "fixed," etc., should be interpreted broadly. For example, "fixed" can mean a fixed connection, a detachable connection, or an integral part; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0022] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the meaning of "and / or" throughout the text includes three parallel solutions; for example, "A and / or B" includes solution A, solution B, or a solution where both A and B are satisfied simultaneously. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0023] Figure 1 This application provides an embodiment of a low-altitude security control method that integrates air, space, and ground operations. Figure 1 As shown, the method includes the following steps: S100: Periodically scans the controlled area using space-based sensing equipment to acquire wide-area remote sensing images; S200: Analyze the wide-area remote sensing image based on a pre-trained target detection model to detect potential targets in the wide-area remote sensing image; S300: The airborne sensing device tracks and locates detected potential targets based on a pre-trained tracking and localization model to obtain the coordinates of the potential targets; S400: It uses ground-based sensing devices to lock onto potential targets based on coordinates and acquire images, and uses a pre-trained target recognition model to analyze and process the images to obtain the target category probability distribution vector; S500: The wide-area remote sensing image, the coordinates of potential targets, and the probability distribution vector of target categories are fused to obtain a low-altitude security situation map; S600: Construct a risk assessment model and assess the threat level of potential targets in real time based on the low-altitude security situation map, and dynamically generate the optimal response plan according to the threat level.

[0024] This embodiment achieves a closed-loop process of space-based wide-area scanning, target detection, airborne tracking and positioning, ground-based fine identification, multi-source information fusion, and real-time risk assessment and handling for low-altitude targets. This enables full-chain, multi-level collaborative perception and intelligent decision-making for low-altitude targets, from macroscopic discovery to microscopic confirmation, from trajectory tracking to intent understanding. It can effectively improve the low-altitude security system's ability to continuously track targets, accurately identify them, and dynamically respond and handle them in complex environments.

[0025] In another exemplary embodiment, in step S200, the target detection model includes: a pulse-excited convolutional network, a dynamic computation sub-network, a causal intervention inference module, and a spatiotemporal memory module; wherein, the pulse-excited convolutional network is used to detect the sparsity of moving targets and abnormal regions in wide-area remote sensing images, and the network specifically includes a multi-scale downsampling layer, a pulse-excited layer, and a pulse fusion and encoding layer, wherein the multi-scale downsampling layer uses three cascaded dilated separable convolutions to perform multi-scale downsampling on the wide-area remote sensing image to extract multi-scale primary features. Features (e.g., including a large-hole-rate coarse sampling layer, a medium-hole-rate feature extraction layer, and a small-hole-rate refinement layer, wherein the large-hole-rate coarse sampling layer has a hole rate of 6 or 8 to capture large-area motion-blurred regions or significant anomalous patches in wide-area remote sensing images; the medium-hole-rate feature extraction layer has a hole rate of 3 or 4 to fuse the output of the large-hole-rate coarse sampling layer and initially distinguish high-frequency background noise in the target domain; the small-hole-rate refinement layer has a hole rate of 1 or 2 to refine features and calibrate spatial information from the output of the medium-hole-rate feature extraction layer).

[0026] The impulse firing layer first performs global average pooling on the multi-scale primary features (dimensions [C, H, W], where C is the number of channels, H is the height, and W is the width) to compress all spatial information within each channel into a single scalar value, thereby generating a channel description vector with C channels. Further, a dual-path attention mechanism is implemented after the global average pooling layer. This mechanism generates an adaptive impulse firing threshold for each channel in the channel description vector. One path uses a lightweight sub-network (e.g., a fully connected layer) that takes the channel description vector as input and, through learned weight parameters, independently and dynamically calculates and outputs a base firing threshold for each channel, allowing the model to adaptively select based on the semantic importance of the current feature. The other path captures dynamic information by combining the channel description vector of the current frame with the impulse residual information of the previous frame to generate a temporal adjustment. This adjustment is used to correct the thresholds of each channel, making the model more sensitive to changes in channels (e.g., by lowering their thresholds to promote firing). The outputs of the two paths (base threshold and temporal adjustment) are fused through weighted aggregation to generate a final adaptive activation threshold for each of the C channels. Finally, the original multi-scale primary features are compared channel-by-channel and element-by-element with the generated C-dimensional threshold vector. For each spatial location (h, w) and each channel, the location is retained only if its activation value is higher than the threshold corresponding to channel C; otherwise, its value is suppressed to zero. This process ultimately outputs a sparse impulse feature map, which retains only the impulse response regions above the threshold to reduce the computational burden on subsequent operations.

[0027] The pulse fusion and coding layer is used to perform cross-scale pulse fusion on sparse pulse feature maps through a lightweight Transformer-style cross-attention mechanism to obtain pulse responses at different resolutions. Finally, a pulse coding map is output, which can mark multiple regions of interest, including multiple features such as motion suspicion, abnormal texture, and small target candidates.

[0028] The dynamic computation subnetwork is used to perform multi-path feature fusion and lightweight classification and regression on the region of interest output by the pulse-fired convolutional network to complete target localization and preliminary recognition. This network includes a dynamic routing encoder, an adaptive feature refinement module, and a lightweight classification and regression head. The dynamic routing encoder takes the pulse-coded map with marked regions of interest output by the pulse fusion and coding layer as input and uses 1×1 convolution to generate routing weights. ,in, Indicates the first The number of convolutional paths is set to, for example, 4. Each path is a depthwise separable convolution with different sizes and dilation rates: Path1: 3×3DWConv,dilation=1, Path2: 3×3DWConv,dilation=2, Path3: 5×5DWConv,dilation=1, Path4: 5×5DWConv,dilation=2. Ultimately, this encoder can extract multi-scale, multi-receptive-field feature representations. :

[0029] in, This represents the input pulse code diagram; This indicates the preset number of convolution paths; This represents a depthwise separable convolution operator.

[0030] The adaptive feature refinement module introduces a learnable gating mechanism and calculates the feature representation output by the dynamic routing encoder using the following formula. Channel importance vector :

[0031] in, Indicates the global pooling operator; Represents the learnable weight matrix; express Activation function.

[0032] Furthermore, regarding the channel importance vector Perform soft selection to obtain refined feature maps. :

[0033] The lightweight classification and regression head first processes the refined feature map of the input. Global average pooling is performed to compress the data into channel feature vectors, which serve as the common input for two subsequent parallel branching steps. These two branches are a classification branch and a regression branch. The classification branch employs a two-layer fully connected network (the first layer maps the channel feature vectors to 128 dimensions and introduces non-linearity using the GELU activation function; the second layer further maps the 128-dimensional features to the number of target classes and outputs a k-dimensional vector representing the predicted probability of each class). The regression branch inputs the channel feature vectors into another independent two-layer fully connected network (also first mapped to 128 dimensions and activated by GELU) and outputs the four-dimensional coordinate parameters of the target bounding box (e.g., including the x-coordinate of the center point, the y-coordinate of the center point, width, and height).

[0034] The causal intervention reasoning module is used to isolate the influence of confounding factors on target identification through causal reasoning. Its working principle includes: First, this module internally constructs a simple structured causal model. The nodes in the model include factors such as target type, appearance features, movement trajectory, and background environment, with pre-defined causal relationships between these factors. During operation, the model actively engages in counterfactual thinking through "hypothetical interventions." For example, the model might set (intervene) the appearance features of a target (such as wing texture) to the form of a "typical drone," while keeping other variables in the causal model unchanged (such as the current movement trajectory and background), and then reassess the probability distribution of the target category. Alternatively, it might intervene by setting the target's movement pattern to a "hovering state" and then observe the change in the probability of it being classified as a "bird." By systematically comparing the differences in the model's predictions before and after such interventions, this module can analyze the strength of the causal effect of different features on the final classification, thus shifting the decision-making basis from superficial statistical correlation to deep causal mechanisms. This mechanism allows the model to filter out confounding factors and focus on causally decisive features. Therefore, even when facing adversarial samples, rare scenarios, or the presence of strong correlation artifacts, it can make more stable and interpretable judgments.

[0035] The spatiotemporal memory module records and associates historical target trajectories and features. This module employs an updatable circular storage structure to continuously encode and store dynamic information of each target detected over several past scan cycles, including its geographical location, deep appearance features, and complete motion trajectory fragments. When processing the current frame, the model actively queries this memory unit through attention mechanisms or similarity retrieval to achieve the following functions: First, by utilizing the spatiotemporal continuity of historical trajectories, it predicts the region where each target is most likely to appear in the current frame, thereby guiding front-end perception modules such as the pulse excitation layer to perform focused calculations and improve the recall rate for small or occluded targets. Second, by associating and matching the features of currently detected targets with historical trajectories in memory, it achieves cross-cycle target identity association and trajectory continuation, and preliminarily infers their movement patterns and potential intentions (such as loitering, traversing, and gathering), providing crucial inputs with long-term context for subsequent risk assessment and behavior analysis models.

[0036] In another exemplary embodiment, in step S300, the tracking and localization model includes a multi-source feature input layer, a temporal feature fusion layer, a trajectory memory and query layer, a trajectory prediction and association layer, an adaptive filtering localization layer, and a fusion output layer connected in sequence. The following, this embodiment will describe each of the above layers in detail.

[0037] The multi-source feature input layer receives and processes heterogeneous feature data from different sensors (such as optical, infrared, and radar) from space-based, air-based, and ground-based sources, achieving unified mapping across modal feature spaces. This layer contains a lightweight 1×1 convolutional layer to initially map features from different modalities to a unified channel dimension. Following the 1×1 convolutional layer is a cross-modal multi-head attention layer, which calculates the interdependencies between features from different modalities. For example, it learns how to use precise distance information provided by radar to enhance and calibrate the target position in the visual image, thereby outputting a fused multi-source feature tensor that is aligned in both feature space and time step.

[0038] The temporal feature fusion layer is used to capture the dynamic continuity and multi-scale pattern features of target motion based on the input heterogeneous feature data. This layer includes a bidirectional gated recurrent unit (BRN), which takes the aligned features output from the cross-modal multi-head attention layer as input and encodes the target's past and future contextual information in the forward and backward directions, respectively, to model the target's motion trend and inertia. A temporal convolutional layer is cascaded after the BRN, which uses one-dimensional dilated convolution to perform strided sampling in the temporal dimension, aiming to extract motion pattern features (such as constant speed, acceleration, and turning) at different time scales. The two layers work together to output an enhanced temporal feature rich in both long-term and short-term motion information.

[0039] The trajectory memory and query layer is used to achieve target identity association and trajectory segment matching across time periods. This layer includes a dynamically updatable memory matrix cache layer to store the feature encoding and motion state of each tracked target in historical processing cycles. When processing the current frame, a similarity-based attention layer is activated, which uses the features of the currently detected target as the "query" to perform efficient retrieval and matching in the memory matrix to find the historical trajectory segments most likely belonging to the same target.

[0040] The trajectory prediction and association layer is used to predict the future motion state of targets and perform multi-target differentiation based on the current state and historical information provided by the aforementioned layers. This layer consists of a lightweight attention layer, a trajectory prediction module, and a trajectory matching layer. The lightweight self-attention layer analyzes the interaction relationships (such as aggregation and following) between all targets in the scene to understand the dynamics of group movement. Subsequently, the trajectory prediction module (e.g., using a small fully connected network) calculates the possible positions of each target in the next few frames, forming a predicted trajectory. The trajectory matching layer then completes the cross-frame multi-target data association by calculating the correlation between the predicted trajectory and the actual detection result in the next frame (e.g., using the Hungarian algorithm combined with motion and appearance similarity costs), ensuring the continuity of the trajectory.

[0041] The adaptive filtering localization layer is used for adaptive state estimation and coordinate smoothing of the differentiated multi-target data. This layer employs a parameter-learnable Kalman filter variant layer, where the weights of the state transition matrix and observation matrix are not fixed but fine-tuned through backpropagation to better adapt to the specific motion models of targets such as UAVs. This layer receives predicted trajectory points and actual observation points to perform optimal state estimation. Simultaneously, the adaptive filtering localization layer also includes an occlusion-aware weighted layer running in parallel with the Kalman filter variant layer. This layer evaluates the target's confidence level in real time (e.g., based on feature quality or visibility score) and automatically adjusts the confidence weights of the filtering process on the predicted values ​​when the target encounters temporary occlusion leading to missing observations or quality degradation, thereby smoothing the output coordinates and suppressing jitter and drift.

[0042] The fusion output layer is responsible for integrating the output information of all the above layers and generating the final localization result. First, it uses a multi-scale feature fusion convolutional layer to convolve and fuse the coordinate sequence output from the adaptive filtering layer, the visual features of the current frame, and the trajectory confidence. Finally, a coordinate regression fully connected layer maps the fused features to the final three-dimensional spatial coordinates of the target in the current frame, thus completing the closed loop from multi-source perception to precise localization.

[0043] In another exemplary embodiment, in step S400, the target recognition model is represented by the following formula:

[0044]

[0045]

[0046]

[0047]

[0048] in, This represents the target class probability distribution vector output by the model; This represents a visible light feature map from a ground-based optical sensor, including spatial information such as color and texture; This represents a thermal imaging feature map from an infrared sensor; This represents the scene context vector, which includes encoded environment information such as time, weather, and geographical region. Represents a mixed feature encoding function; and They are respectively and Abbreviation; This represents a cross-modal attention layer, used to calculate the interaction weights between visible light and infrared features, achieving feature alignment and complementary enhancement. This indicates element-wise addition; This indicates a channel-level concatenation operation, used to concatenate attention-weighted features with features that are directly added together; This represents a 3×3 depthwise separable convolution; This represents a causal perception inference function; Indicates from fusion feature map; Indicates a budget, based on right Make hypothetical modifications to generate counterfactual features; Represents a multilayer perceptron; Represents the learnable scalar coefficients; This represents the adaptive focusing attention function; Indicates from Causal enhancement feature map; Indicates global average pooling; This represents a 1×1 convolutional layer; This represents the trajectory attention prior, used to indicate the spatial region where the target may appear; This represents the Sigmoid activation function.

[0049] In another exemplary embodiment, step S500 involves fusing the wide-area remote sensing image, the coordinates of potential targets, and the probability distribution vector of target categories to obtain a low-altitude security situation map, including the following steps: S501: Align and standardize the coordinates and target category probability distribution vectors of wide-area remote sensing images and potential targets; This step first requires spatiotemporal consistency processing of the wide-area remote sensing imagery, the coordinates of potential targets, and the probability distribution vectors of target categories. Specifically, this includes: firstly, unifying the wide-area remote sensing imagery to a standard geographic coordinate system (such as WGS-84) through image registration and geographic correction methods; secondly, aligning the continuous coordinate sequence of potential targets with the acquisition time of the remote sensing image using a timestamp synchronization mechanism to eliminate temporal deviations caused by sensor response delays or communication latency; and thirdly, normalizing and smoothing the target category probability distribution vectors to ensure numerical stability and conformity to probability distribution characteristics. Finally, all data is converted into a unified spatiotemporal grid representation, laying the foundation for subsequent multi-source fusion.

[0050] S502: Construct a situational imagery layer based on standardized wide-area remote sensing images, coordinates of potential targets, and probability distribution vectors of target categories; In this step, after data alignment and standardization, a multi-dimensional situational awareness layer is constructed using a layered fusion architecture. The bottom layer is the geographic background layer, composed of standardized wide-area remote sensing imagery overlaid with vector geographic information (such as no-fly zones and outlines of key facilities). The middle layer is the target dynamic layer, which draws real-time movement trajectories based on target coordinate sequences and labels their speed, altitude, heading, and other status attributes. The top layer is the semantic enhancement layer, which performs type labeling and intent inference for each target based on target category probability distribution vectors and generates semantic tags such as "suspected reconnaissance" and "loitering surveillance" by combining historical behavior patterns. Furthermore, the layers complement each other through spatial association and attribute mapping, forming a structured situational awareness representation.

[0051] S503: Render the constructed situational awareness layer to obtain a low-altitude security situational awareness map.

[0052] In this step, this embodiment renders and dynamically synthesizes the constructed situational awareness layer using a real-time visualization engine. Specifically, it employs a GPU-accelerated rendering pipeline to overlay, color, and annotate the geographic background layer, target dynamic layer, and semantic enhancement layer, generating a highly interactive two-dimensional or three-dimensional situational awareness map. The rendering process supports interactive functions such as viewpoint switching, scale zooming, trajectory playback, and target filtering, and can visually enhance targets based on threat levels (e.g., color, flashing cues). The final output low-altitude security situational awareness map is published in image, vector, or streaming data format, supporting real-time retrieval and command and decision-making applications across multiple terminals.

[0053] In another exemplary embodiment, in step S600, the risk assessment model includes: a multi-dimensional feature fusion module, a dynamic threat calculation engine, and an interpretable output module. The multi-dimensional feature fusion module is used to uniformly encode and fuse multi-source heterogeneous features from the low-altitude security situation map; the dynamic threat calculation engine is used to comprehensively analyze the target's intent, capabilities, opportunities, and collaborative behaviors to calculate an interpretable threat level score; and the interpretable output module is used to implement threat score level mapping, decision support, and output.

[0054] Specifically, the multi-dimensional feature fusion module consists of three parallel processing units: a spatiotemporal encoding unit, a behavioral feature extraction unit, and an environmental feature encoding unit. The spatiotemporal encoding unit employs a simplified attention mechanism combined with a lightweight LSTM to extract spatiotemporal features from the target trajectory sequence and capture motion patterns and spatial distribution regularities. The behavioral feature extraction unit, based on a pre-trained lightweight classifier, semantically encodes target type, motion parameters, and historical behavioral patterns to generate a 128-dimensional behavioral feature vector. The environmental feature encoding unit maps discrete and continuous features such as regional sensitivity, time factors, weather conditions, and defense status into a compact environmental vector through regularization processing. Finally, the outputs of all these units are concatenated through channels and subjected to 1×1 convolution for feature alignment and dimensionality reduction, ultimately outputting a comprehensive feature tensor that provides normalized multi-dimensional input for subsequent threat assessment.

[0055] The dynamic threat calculation engine employs a four-dimensional parallel computing and adaptive fusion architecture, specifically comprising an intent risk assessment subnetwork, a capability risk assessment subnetwork, an opportunity window assessment subnetwork, and a cooperative threat detection subnetwork. The intent risk assessment subnetwork analyzes the deviation, dwell behavior, and approach rate of target movement patterns using trajectory anomaly detection algorithms (such as density-based Local Outlier Factor (LOF)). The capability risk assessment subnetwork calculates the physical threat potential of a target based on its type, maneuver parameters, and payload characteristics using a weighted formula. The opportunity window assessment subnetwork identifies threat availability in the current spatiotemporal context by combining environmental cover conditions and defense system status in real time. The cooperative threat detection subnetwork analyzes the spatiotemporal correlation and interaction patterns between multiple targets using lightweight graph convolution, identifying cooperative behaviors such as formation and relay. Finally, the outputs of each subnetwork are dynamically adjusted by scene-adaptive weights and converged into an initial threat value using a weighted fusion formula. A cooperative enhancement factor is then introduced to weight and boost the group threat, ultimately outputting a continuous threat score within the range of 0 to 1.

[0056] The interpretable output module first maps continuous threat values ​​to five discrete threat levels (observation, warning, alert, threat, and emergency) using a piecewise threshold function, and binds preset handling suggestions and a visual coding scheme (color, flashing frequency) to each level. Subsequently, based on the current threat level, target type, environmental situation, and available handling resources, the module dynamically generates the optimal handling plan through a lightweight decision tree and rule engine. The plan covers handling methods (such as voice alerts, electronic interference, navigation deception, physical interception, etc.), response unit scheduling instructions, and collaborative handling procedures. At the same time, the module generates structured interpretable output, including the main components of the threat (such as "abnormal intent 85%, medium capability 60%)", key evidence summaries (such as "abnormal trajectory approach + stay in sensitive areas"), assessment confidence level, and explanation of the basis for the handling plan. Finally, it is pushed synchronously to the command terminal, handling unit, and log system through a standardized data interface, supporting human-machine collaborative decision-making and closed-loop management of handling.

[0057] In another exemplary embodiment, this application also provides a low-altitude security control device, such as... Figure 2 As shown, the device includes: a first acquisition module 100, used to periodically scan the controlled area using space-based sensing equipment to acquire wide-area remote sensing images; a target detection module 200, used to analyze the wide-area remote sensing images based on a pre-trained target detection model to detect potential targets in the wide-area remote sensing images; a target localization module 300, used to track and locate the detected potential targets using space-based sensing equipment based on a pre-trained tracking and localization model to obtain the coordinates of the potential targets; a second acquisition module 400, used to lock onto potential targets based on coordinates using ground-based sensing equipment and acquire images, and analyze and process the images using a pre-trained target recognition model to obtain a target category probability distribution vector; a fusion module 500, used to fuse the wide-area remote sensing images, the coordinates of potential targets, and the target category probability distribution vector to obtain a low-altitude security situation map; and an evaluation module 600, used to construct a risk assessment model, and based on the low-altitude security situation map, to evaluate the threat level of potential targets in real time, and dynamically generate the optimal response plan based on the threat level.

[0058] Based on the above embodiments, refer to Figure 3 The computer-readable storage medium of exemplary embodiments of this application will be described below. Please refer to [link / reference]. Figure 3The computer-readable storage medium shown is an optical disc 40, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it implements the steps described in the above-described method implementation. For example, it periodically scans the controlled area using space-based sensing equipment to acquire wide-area remote sensing images; analyzes the wide-area remote sensing images based on a pre-trained target detection model to detect potential targets in the wide-area remote sensing images; uses an air-based sensing device to track and locate the detected potential targets based on a pre-trained tracking and positioning model to obtain the coordinates of the potential targets; uses a ground-based sensing device to lock onto the potential targets based on the coordinates and acquire images, and analyzes and processes the images using a pre-trained target recognition model to obtain a target category probability distribution vector; it fuses the wide-area remote sensing images, the coordinates of the potential targets, and the target category probability distribution vector to obtain a low-altitude security situation map; it constructs a risk assessment model and assesses the threat level of potential targets in real time based on the low-altitude security situation map, and dynamically generates the optimal response plan based on the threat level. The specific implementation methods of each step will not be repeated here.

[0059] It should be noted that the computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be described in detail here.

[0060] Based on the above embodiments, this application also provides an electronic device, which is described below with reference to... Figure 4 An electronic device for file downloading according to an exemplary embodiment of this application will be described.

[0061] Figure 4 A block diagram is shown of an exemplary electronic device 50 suitable for implementing embodiments of the present application. The electronic device 50 may be a computer system or a cloud server. Figure 4 The electronic device 50 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0062] like Figure 4 As shown, the electronic device 50 includes, but is not limited to: one or more processors or processing units 501, system memory 502, and bus 503 connecting different system components (including system memory 502 and processing unit 501).

[0063] Electronic device 50 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 50, including volatile and non-volatile media, removable and non-removable media.

[0064] System memory 502 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 5021 and / or cache memory 5022. Electronic device 50 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, ROM 5023 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 4 (Not shown in the image, usually referred to as "hard drive"). Although not shown in... Figure 4 The diagram illustrates that a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to bus 503 via one or more data media interfaces. System memory 502 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this application.

[0065] A program / utility 5025 having a set (at least one) of program modules 5024 may be stored, for example, in system memory 502, and such program modules 5024 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment. Program modules 5024 typically perform the functions and / or methods described in the embodiments of this application.

[0066] Electronic device 50 can also communicate with one or more external devices 504 (such as a keyboard, pointing device, display, etc.). This communication can be performed through input / output (I / O) interface 505. Furthermore, electronic device 50 can also communicate with one or more networks (such as local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via network adapter 506. Figure 4 As shown, network adapter 506 communicates with other modules of electronic device 50 (such as processing unit 501) via bus 503. It should be understood that, although... Figure 4 As not shown, it can be used in conjunction with electronic device 50 with other hardware and / or software modules.

[0067] The processing unit 501 executes various functional applications and data processing by running programs stored in the system memory 502. For example, it periodically scans the controlled area using space-based sensing devices to acquire wide-area remote sensing images; it analyzes the wide-area remote sensing images based on a pre-trained target detection model to detect potential targets in the wide-area remote sensing images; it tracks and locates the detected potential targets using air-based sensing devices based on a pre-trained tracking and positioning model to obtain the coordinates of the potential targets; it locks onto potential targets based on coordinates using ground-based sensing devices and acquires images, and analyzes and processes the images using a pre-trained target recognition model to obtain target category probability distribution vectors; it fuses the wide-area remote sensing images, the coordinates of potential targets, and the target category probability distribution vectors to obtain a low-altitude security situation map; it constructs a risk assessment model and assesses the threat level of potential targets in real time based on the low-altitude security situation map, and dynamically generates the optimal response plan based on the threat level.

[0068] The specific implementation methods of each step will not be repeated here. It should be noted that although several units / modules or sub-units / sub-modules of the file concurrent download device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0069] In the description of this application, it should be noted that the terms "first", "second", and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0070] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0071] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0072] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0073] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0074] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions that enable a computer device (which may be a personal computer, a cloud server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0075] The above embodiments are only for illustrating the technical concept and features of this application, and are intended to enable those skilled in the art to understand the content of this application and implement it accordingly. They should not be construed as limiting the scope of protection of this application. All equivalent changes or modifications made in accordance with the spirit and essence of this application should be included within the scope of protection of this application.

[0076] Finally, it should be noted that the above descriptions are merely optional examples of this application and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A low-altitude security control method, characterized in that, The method includes: The controlled area is periodically scanned using space-based sensing equipment to acquire wide-area remote sensing images; The wide-area remote sensing image is analyzed based on a pre-trained target detection model to detect potential targets in the wide-area remote sensing image. The airborne sensing device tracks and locates detected potential targets based on a pre-trained tracking and localization model to obtain the coordinates of the potential targets; The ground-based sensing device locks onto potential targets based on coordinates and acquires images. The images are then analyzed and processed using a pre-trained target recognition model to obtain the target category probability distribution vector. The wide-area remote sensing image, the coordinates of potential targets, and the probability distribution vector of target categories are fused to obtain a low-altitude security situation map; A risk assessment model is constructed, and the threat level of potential targets is assessed in real time based on the low-altitude security situation map. The optimal response plan is then dynamically generated based on the threat level.

2. The method according to claim 1, characterized in that, The target detection model includes: Pulse-excited convolutional networks are used to detect sparsity of moving targets and anomalous regions in wide-area remote sensing images; A dynamic computational subnetwork is used to isolate the influence of confounding factors on target recognition through causal reasoning. The causal intervention reasoning module is used to perform multi-path feature fusion and lightweight classification regression on the region of interest output by the pulse-excited convolutional network in order to complete target localization and preliminary identification. The cross-temporal memory module is used to record and associate historical target trajectories and features.

3. The method according to claim 2, characterized in that, The pulse-excited convolutional network includes: The layers consist of a multi-scale downsampling layer, a pulse excitation layer, and a pulse fusion and coding layer. The multi-scale downsampling layer is used to perform multi-scale downsampling on wide-area remote sensing images and extract multi-scale primary features. The pulse excitation layer is used to perform sparsification screening of multi-scale primary features to generate a pulse sparse pulse feature map. The pulse fusion and coding layer is used to fuse sparse pulse feature maps to obtain pulse-coded maps that label regions of interest.

4. The method according to claim 2, characterized in that, The dynamic computing subnetwork includes: The system comprises a dynamic routing encoder, an adaptive feature refinement module, and a lightweight classification and regression head. The dynamic routing encoder is used to select multiple convolutional paths based on the pulse coding map for weighted fusion, and extract feature representations with multiple scales and receptive fields. The adaptive feature refinement module is used to perform soft selection and enhancement on feature representations with multiple scales and multiple receptive fields to obtain refined feature maps. The lightweight classification and regression head is used to classify and review the refined feature maps, and output the target class probability and bounding box coordinates.

5. The method according to claim 1, characterized in that, The tracking and localization model comprises: a multi-source feature input layer, a temporal feature fusion layer, a trajectory memory and query layer, a trajectory prediction and association layer, an adaptive filtering localization layer, and a fusion output layer, connected sequentially. The multi-source feature input layer is used to receive and process heterogeneous feature data from space-based, air-based, and ground-based sensors, and to achieve unified mapping across modal feature spaces; The temporal feature fusion layer is used to capture the dynamic continuity and multi-scale pattern features of target motion based on heterogeneous feature data; The trajectory memory and query layer is used to achieve target identity association and trajectory segment matching across time periods; The trajectory prediction and association layer is used to predict the future motion state of the target and distinguish multiple targets based on the current state and historical information provided by the multi-source feature input layer, the temporal feature fusion layer and the trajectory memory and query layer. The adaptive filtering localization layer is used to perform adaptive state estimation and coordinate smoothing on the distinguished multi-targets; The fusion output layer integrates the output information from the multi-source feature input layer, the temporal feature fusion layer, the trajectory memory and query layer, the trajectory prediction and association layer, and the adaptive filtering localization layer to generate the final localization result.

6. The method according to claim 1, characterized in that, The process of fusing the wide-area remote sensing imagery, the coordinates of potential targets, and the probability distribution vector of target categories to obtain a low-altitude security situation map includes: Alignment and standardization are performed on the coordinates and target category probability distribution vectors of wide-area remote sensing images and potential targets; A situational awareness layer is constructed based on standardized wide-area remote sensing imagery, the coordinates of potential targets, and the probability distribution vectors of target categories. The constructed situational awareness layer is rendered to obtain a low-altitude security situational awareness map.

7. The method according to claim 1, characterized in that, The risk assessment model includes: The module includes a multi-dimensional feature fusion module, a dynamic threat calculation engine, and an interpretable output module. The multi-dimensional feature fusion module is used to uniformly encode and fuse multi-source heterogeneous features from the low-altitude security situation map; The dynamic threat calculation engine is used to synthesize the target's intent, capabilities, opportunities, and collaborative behaviors to calculate an interpretable threat level score. The interpretable output module is used to implement threat score level mapping, decision support, and output.

8. A low-altitude security control device, characterized in that, The device includes: The first acquisition module is used to periodically scan the controlled area through space-based sensing equipment to acquire wide-area remote sensing images. The target detection module is used to analyze the wide-area remote sensing image based on a pre-trained target detection model in order to detect potential targets in the wide-area remote sensing image. The target localization module is used to track and locate detected potential targets using an airborne sensing device based on a pre-trained tracking and localization model, so as to obtain the coordinates of the potential targets. The second acquisition module is used to lock onto potential targets based on coordinates using ground-based sensing devices and acquire images, and to analyze and process the images using a pre-trained target recognition model to obtain the target category probability distribution vector. The fusion module is used to fuse the wide-area remote sensing image, the coordinates of potential targets, and the probability distribution vector of target categories to obtain a low-altitude security situation map. The assessment module is used to build a risk assessment model, assess the threat level of potential targets in real time based on the low-altitude security situation map, and dynamically generate the optimal response plan according to the threat level.

9. A storage medium, characterized in that, It includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, The electronic device includes: Memory, processor, and computer programs stored in memory and executable on the processor, wherein, When the processor executes the program, it implements the method as described in any one of claims 1 to 7.