Autonomous defense decision and response method and system for tower-held unmanned aerial vehicle based on artificial intelligence

By using a policy generation network based on a metacognitive reinforcement learning framework and cognitive distillation technology, a closed loop of perception-decision-execution-learning-evolution for a drone defense system was constructed, which solved the problem of poor adaptability and achieved efficient interception of intelligent threats.

CN122047488APending Publication Date: 2026-05-15GUANGZHOU HUACHUANG XINNENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU HUACHUANG XINNENG TECHNOLOGY CO LTD
Filing Date
2026-02-03
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing drone defense systems have poor adaptability when facing unknown escape strategies or complex interference, their decision-making and execution are disconnected, they lack a continuous security evolution mechanism, and they are unable to cope with intelligent threats.

Method used

A policy generation network based on a metacognitive reinforcement learning framework is adopted to construct a unified situational representation through multi-source perception data, generate cooperative interception instructions, and make dynamic adjustments during execution. Combined with cognitive distillation technology and online evolutionary verification, the system can achieve continuous security evolution.

Benefits of technology

It enhances the intelligence and robustness of the drone defense system in dynamic combat environments, and significantly improves the success rate of coordinated interception of highly mobile and intelligent intrusion targets and the overall combat effectiveness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122047488A_ABST
    Figure CN122047488A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence-based autonomous defense decision and response method and system for a tower-held unmanned aerial vehicle. According to the method, unified situation representation including the physical state and tactical intention of an invading unmanned aerial vehicle is constructed, and a strategy generation network based on meta-cognitive reinforcement learning training is utilized to generate a collaborative interception instruction. According to the network integration opponent strategy inference and multi-step situation prediction module, an inner layer generates a current optimal strategy through adversarial simulation deduction, and an outer layer improves the adaptive capacity by learning an adversarial mode evolution rule. After the instruction is converted into control parameters, the multiple defense unmanned aerial vehicles are driven to execute a collaborative interception task. During execution, the system continuously collects confrontation interaction trajectory data, extracts new features and evolution laws of a confrontation mode through a cognitive distillation technology, and updates the network cognitive module. Meanwhile, an online evolution verification mechanism is established, and safe and stable closed-loop cognitive evolution is achieved. The system can significantly improve the autonomous adaptation and cooperative countermeasure ability of the unmanned aerial vehicle defense system to unknown threats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drone defense technology, specifically to an artificial intelligence-based autonomous defense decision-making and response method and system for tower-based drones, which is particularly suitable for collaborative interception and intelligent countermeasures against low-altitude, slow-speed, and small drone targets. Background Technology

[0002] With the rapid development and widespread adoption of drone technology, the risk of drones being used for illegal reconnaissance, smuggling, harassment, and even terrorist attacks is increasing, posing a serious challenge to low-altitude security in key areas and critical infrastructure. Traditional drone defense methods, such as radio jamming, navigation deception, and physical interception, mostly rely on preset rules or human intervention. When facing drones with intelligent maneuverability, they often suffer from limitations such as delayed response, limited strategies, and difficulty in dealing with coordinated intrusions.

[0003] Currently, research on drone countermeasures based on artificial intelligence technology is booming. Schemes using deep reinforcement learning to train single or multiple drones to perform pursuit missions have been reported. However, most existing methods are trained in relatively closed, rule-fixed simulation environments. The learned strategies are usually highly bound to the training environment, and their performance often declines significantly when facing new, untrained escape strategies emerging in real-world adversarial situations or when subjected to complex environmental interference. This stems from the fact that their models typically lack the ability to effectively infer the opponent's strategic intentions and cannot continuously learn and evolve during adversarial processes. Furthermore, most schemes focus on the decision-making algorithm itself and fail to deeply integrate with a complete engineering chain, including multi-source perception fusion, real-time collaborative control, adversarial experience extraction, and system security evolution, making it difficult to meet the practical defense requirements of high reliability and high autonomy.

[0004] Therefore, how to build an intelligent defense system with real-time situational understanding, adversarial strategy inference, forward-looking decision-making, collaborative execution, and continuous security evolution in adversarial situations has become a key technical challenge for improving the effectiveness of low-altitude active defense and responding to future intelligent threats. Summary of the Invention

[0005] The purpose of this invention is to provide an artificial intelligence-based autonomous defense decision-making and response method and system for tower-based unmanned aerial vehicles (UAVs). This invention aims to solve the problems of poor adaptability, disconnect between decision-making and execution, and lack of continuous security evolution mechanism in existing UAV defense systems when facing unknown escape strategies or complex interference. It aims to achieve a complete intelligent defense closed loop from perception to cognition and from decision-making to evolution.

[0006] In a first aspect, embodiments of this application provide an artificial intelligence-based autonomous defense decision-making and response method for tower-stationed unmanned aerial vehicles (UAVs), the method comprising: By acquiring multi-source sensing data through low-altitude situational awareness units, a unified situational representation is constructed that includes inferences about the physical state and tactical intentions of intruding drones. The unified situational representation is input into a policy generation network trained based on a metacognitive reinforcement learning framework. The policy generation network generates cooperative interception instructions by integrating an adversary policy inference module and a multi-step situational prediction module. The metacognitive reinforcement learning framework includes two learning mechanisms: the inner layer generates the current optimal policy through adversarial simulation and the outer layer learns the evolution law of adversarial mode through cognitive modeling. The coordinated interception command is converted into motion control parameters to drive multiple defense drones to perform coordinated interception tasks, and dynamic adjustments are made based on the real-time situation during the execution process; During the interception mission, data on adversarial interaction trajectories are continuously collected; Based on the adversarial interaction trajectory data, new features and evolutionary patterns of the adversarial patterns are extracted using cognitive distillation technology, which are then used to update the cognitive modeling module of the policy generation network. Establish an online evolution verification mechanism to perform performance verification before deploying the updated policy generation network, ensuring that the defense system achieves secure and stable closed-loop cognitive evolution.

[0007] Secondly, embodiments of this application provide an artificial intelligence-based autonomous defense decision-making and response system for tower-stationed unmanned aerial vehicles (UAVs), applied to the artificial intelligence-based autonomous defense decision-making and response method for tower-stationed UAVs as described in the first aspect, the system comprising: The low-altitude situational awareness unit is used to acquire multi-source sensing data and construct a unified situational representation that includes inferences about the physical state and tactical intentions of intruding drones. The strategy generation network, trained based on a metacognitive reinforcement learning framework, is used to receive the unified situation representation and generate cooperative interception instructions through its integrated adversary strategy inference module and multi-step situation prediction module. The metacognitive reinforcement learning framework includes two learning mechanisms: the inner layer generates the current optimal strategy through adversarial simulation and the outer layer learns the evolution law of adversarial mode through cognitive modeling. The command execution and dynamic control unit is used to convert the cooperative interception command into motion control parameters, drive multiple defense drones to perform cooperative interception tasks, and make dynamic adjustments based on the real-time situation during the execution process; The adversarial interaction data acquisition module is used to continuously collect adversarial interaction trajectory data during the execution of interception tasks; The cognitive evolutionary learning module is used to extract new features and evolutionary patterns of adversarial patterns based on the adversarial interaction trajectory data through cognitive distillation technology, and to update the cognitive modeling module of the policy generation network. The online evolution verification unit is used to perform performance verification before deploying the updated policy generation network, ensuring that the defense system achieves secure and stable closed-loop cognitive evolution.

[0008] Thirdly, embodiments of this application provide an electronic device, including: processor; Memory used to store processor-executable instructions; The processor is configured to implement the AI-based autonomous defense decision-making and response method for tower-stationed unmanned aerial vehicles as described in the first aspect when executing the instructions.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program that instructs a device to execute the AI-based autonomous defense decision-making and response method for tower-stationed unmanned aerial vehicles as described in the first aspect.

[0010] The beneficial effects of this invention are as follows: By introducing a metacognitive reinforcement learning framework and an adversary strategy inference module, the system can not only generate the current optimal interception strategy, but also understand and learn the evolutionary laws of adversarial modes, possessing the ability to predict and adapt to unknown threats, and achieving a leap from state-action mapping to cognitive-evolutionary understanding. This invention organically integrates multi-source perception fusion, intent understanding, strategy generation, collaborative control, experience learning, and security verification, constructing a complete technical chain of perception-decision-execution-learning-evolution, solving the problem of the disconnect between the decision model and the engineering system in traditional solutions. Through cognitive distillation technology and online evolutionary verification mechanisms, the system can continuously refine new knowledge and securely update its own model in real-world adversarial situations, forming an intelligent life form that can continuously enhance itself along with the evolution of threats, fundamentally improving the long-term effectiveness and robustness of the defense system. Through dynamic role allocation and distributed collaborative control mechanisms, multiple defense drones can form an organic whole, flexibly changing formations and tactics according to the real-time situation, significantly improving the success rate of collaborative interception of highly mobile and intelligent intrusion targets and the overall combat effectiveness of the system. Through the above-mentioned technical solution, this invention significantly improves the intelligence level, adaptability and overall combat effectiveness of the UAV autonomous defense system in dynamic and uncertain confrontation environments, and provides an effective technical path for building a next-generation intelligent low-altitude security defense system. Attached Figure Description

[0011] Figure 1 This is a schematic diagram of an AI-based autonomous defense decision-making and response method for tower-stationed unmanned aerial vehicles (UAVs) provided in an embodiment of this application.

[0012] Figure 2 This is an architecture diagram of an AI-based autonomous defense decision-making and response system for tower-stationed unmanned aerial vehicles (UAVs) provided in one embodiment of this application.

[0013] Figure 3 A schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0015] It should be noted that in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.

[0016] Based on the embodiments described in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] Example 1

[0018] Figure 1 This is a schematic flowchart illustrating an artificial intelligence-based autonomous defense decision-making and response method for tower-stationed unmanned aerial vehicles (UAVs) according to an embodiment of this application. Figure 1 As shown, an artificial intelligence-based autonomous defense decision-making and response method for tower-stationed unmanned aerial vehicles (UAVs) includes: S1. Acquire multi-source sensing data through a low-altitude situational awareness unit to construct a unified situational representation that includes inferences about the physical state and tactical intentions of the intruding UAV. Multi-source information fusion and tactical intention perception. This step serves as the data input and preprocessing stage of the entire method. Its core function is to align, correlate, and fuse raw detection data from multiple heterogeneous sensors in a spatiotemporal manner using specific algorithms (such as the spatiotemporal graph attention mechanism) to construct a unified, structured representation of the environmental state. This representation transcends purely physical states (such as position and velocity), incorporating inferences about the intruding target's behavioral patterns and tactical intentions through computational analysis, providing input rich in high-level semantic information for subsequent intelligent decision-making.

[0019] Specifically, in this embodiment, the step of constructing a unified situational representation that includes inferences about the physical state and tactical intentions of the intruding drone specifically includes: Step 1: Analyze the maneuver sequence of the intruding drone using a spatiotemporal graph attention mechanism to infer its tactical intent and generate a probability distribution of intent categories. A spatiotemporal graph is a data structure used to represent the dynamic behavior of a target in space and time. The drone's state (e.g., position, velocity) at each time point is a node in the graph, and the lines (edges) connecting the nodes represent the transition relationships between states at different time points, forming a time series. The maneuver sequence of the intruding drone refers to the position, velocity, and acceleration changes of the target over a period of time (e.g., the past 3 seconds) continuously acquired by sensors. Intent categories are predefined high-level labels describing the drone's purpose, such as: circling reconnaissance, close-range penetration, flanking maneuver, and rapid withdrawal. The function of this step is to perform high-level semantic interpretation of historical behavior data. It takes the target's raw, low-level motion data as input, uses the spatiotemporal graph attention mechanism to automatically identify key patterns and temporal dependencies in the maneuver sequence, and then outputs a probability distribution. For example, the system might output: circling reconnaissance: 0.7, close-range penetration: 0.2, rapid withdrawal: 0.1. This means the system has a 70% confidence level in determining that the target is conducting reconnaissance rather than a direct attack. This step provides crucial information about the adversary's intentions for subsequent decision-making.

[0020] Step 2 involves establishing a semantic encoder for the multi-source sensing data, converting radio detection signals, radar point cloud data, and electro-optical tracking information into unified tactical feature vectors. The function of this step is to address the problem of data heterogeneity and extract high-level features. Data from different sensors has completely different formats, dimensions, and physical meanings, making direct mathematical operations impossible. The semantic encoder acts as a translator, converting each type of raw data into a single language—a unified-dimensional feature vector. These feature vectors no longer represent raw pixels or signal strength, but rather high-level, tactically relevant abstract attributes such as target mobility, target contour features, and target electromagnetic activity. This allows subsequent steps to fuse and compare information from these different sources.

[0021] Step 3: Calculate the threat entropy value of the current adversarial situation based on the intention probability distribution and tactical feature vector to quantify environmental uncertainty. This invention proposes a threat entropy (Geometric Threat Entropy) measure based on information geometry. Compared with traditional Shannon entropy, it can more sensitively capture subtle changes in the intention probability distribution on the probability manifold, thus providing earlier warnings of highly uncertain deceptive maneuvers. The function of this step is to quantify the deceptiveness or risk ambiguity of the current situation. It not only assesses the magnitude of the threat but also our grasp of the threat assessment. If the intention probability distribution is highly concentrated (e.g., [0.95, 0.05, 0.0]), it indicates that the system is very confident in the opponent's intentions; in this case, the entropy value is low, and the uncertainty is small. If the intention probability distribution is very even (e.g., [0.33, 0.33, 0.34]), it indicates that the system is completely unable to determine the opponent's intentions; in this case, the entropy value is high, and the uncertainty is large. A high entropy value indicates to the decision network that the current situation is complex and difficult to discern, requiring a more cautious and inclusive strategy. Threat entropy is the application of Shannon entropy in information theory to military situation assessment.

[0022] Intent uncertainty quantification based on information geometry can be used to deepen the calculation of threat entropy, upgrading it from traditional information entropy to measuring the confidence of intent inference on a probabilistic manifold. The formula is as follows: , in, It is the probability distribution vector of intent. Indicates the first The probability of a certain tactical intention. It represents the number of intent categories. It is the Fisher information metric tensor defined on the probabilistic simplex. It is the logarithmic gradient of the probability component. This is geometric threat entropy. A higher value indicates greater uncertainty in the system's intent inference, and this uncertainty is measured based on the geometric properties of the probability distribution. It is more sensitive than traditional Shannon entropy in capturing subtle but crucial changes in the probability distribution. It upgrades simple entropy calculation to a confidence metric based on information geometry. When an intruding drone performs deceptive maneuvers (such as rapidly oscillating intent probabilities between several options), traditional Shannon entropy may not change significantly, but geometric entropy can keenly capture such oscillations on the manifold, thus providing earlier warnings of high situational uncertainty and driving the metacognitive scheduling module to adopt a more cautious reasoning mode.

[0023] Step 4 integrates physical state information, intent probability distribution, threat entropy value, and environmental constraint information into a machine-readable multidimensional cognitive tensor, serving as the unified situational awareness representation. Physical state information refers to the fundamental physical quantities of all targets, such as the position and velocity of all drones (intruders and defenders). The intent probability distribution P originates from Step 1. The threat entropy value H originates from Step 3. Environmental constraint information includes things like electronic fence boundaries, no-fly zone coordinates, building and obstacle locations, wind direction and speed, etc. The multidimensional cognitive tensor can be understood as a structured multidimensional array or table, organizing all the above information according to a predetermined format. The function of this step is information integration and standardized output. It is the final step in the entire situational awareness construction process, packaging and integrating the different types and dimensions of information (basic data, high-level inferences, uncertainty measures, environmental knowledge) produced in the previous steps into a unified, fixed-format data structure—the multidimensional cognitive tensor. This tensor is a complete cognitive snapshot of the current battlefield environment for the entire system and serves as the unique and standardized input interface for the subsequent policy generation network. This design ensures that the information received by the decision-making module is rich, structured, and consistent.

[0024] S2. The unified situational representation is input into a policy generation network trained based on a metacognitive reinforcement learning framework. This network generates cooperative interception instructions by integrating an adversary policy inference module and a multi-step situational prediction module. The metacognitive reinforcement learning framework includes two learning mechanisms: the inner layer generates the current optimal policy through adversarial simulation, and the outer layer learns the evolutionary patterns of adversarial modes through cognitive modeling. Predictive cooperative policy generation is the core of this method's intelligent decision-making. Its function is to receive the unified situational representation output from step 1 and process it using the policy generation network trained based on the metacognitive reinforcement learning framework. This network has a dual-layer optimization mechanism: the inner mechanism performs adversarial simulation to generate the immediate optimal tactical action sequence based on the current situation; the outer mechanism focuses on learning the dynamic evolutionary patterns of the adversarial environment itself. By integrating the adversary policy inference and multi-step situational prediction modules, the network can comprehensively evaluate the current and future situations, ultimately outputting a coordinated and consistent cooperative interception instruction set for multiple agents, realizing the transformation from situational awareness to cooperative action plans. The evolutionary patterns of adversarial modes learned by the outer layer will serve as a dynamic environmental model or strategy evolution prior, and will be provided to the adversarial simulation and deduction process of the inner layer in real time. This will enable the inner layer strategy generation to be optimized in a simulation environment that is closer to the real adversarial trend, thereby improving the adaptability and foresight of the strategy.

[0025] Specifically, in this embodiment, the policy generation network includes the following modules: The metacognitive scheduling module dynamically configures the network's attention allocation strategy and inference computation depth based on the characteristics of the adversarial situation. Metacognition, in cognitive science, refers to the ability to monitor, regulate, and evaluate the cognitive process itself. In this technical solution, the metacognitive scheduling module is a control unit responsible for managing and optimizing the allocation of computational resources within the policy generation network. The characteristics of the adversarial situation refer to the various types of information contained in the unified situational representation constructed earlier, particularly threat entropy (a measure of uncertainty), intent probability distribution (target intent clarity), and tactical feature vectors (target capability characteristics). The attention allocation strategy refers to the mechanism by which the neural network decides which part of the information to focus on when processing input data. For example, when the target intent is ambiguous (high entropy), it may be necessary to simultaneously focus on the features corresponding to multiple possible intents; when the target intent is clear, the focus is on the typical behavioral patterns of that intent. The inference computation depth refers to the number of simulation steps or recursive layers used by the network when performing policy inference. For example, for simple and direct targets, shallow inference can be used for a fast response; for complex and cunning targets, multi-step deep inference is required to predict their subsequent changes. This module is the computational resource scheduling center of the policy generation network. It dynamically adjusts its internal computational strategy based on the complexity of the current situation (such as the level of uncertainty and the clarity of intent): In low-uncertainty scenarios (low threat entropy, clear intent): it reduces the inference depth, adopts a fast response mode, and concentrates computational resources on generating the current optimal strategy. In high-uncertainty scenarios (high threat entropy, ambiguous intent): it increases the inference depth, explores multiple possible future evolution paths in parallel, and adopts a more conservative and inclusive strategy.

[0026] The adversary strategy inference module is used to identify the potential escape strategy categories and their switching probabilities of intruding drones in real time. Escape strategy categories are predefined patterns describing typical evasive behaviors of intruding drones. Examples include: straight-line acceleration escape (flying in a straight line at maximum speed in the current direction); serpentine maneuvering evasion (periodically turning left and right to increase interception difficulty); altitude change escape (rapidly climbing or diving to change altitude); and terrain utilization concealment (maneuvering towards buildings or terrain-covered areas). The switching probability refers to the probability matrix of a target switching from one escape strategy to another. This module is the adversary behavior analyzer of the strategy generation network. Based on historical motion data and the current situation, it infers in real time what escape strategy the intruding drone is currently employing and predicts what strategy it might change to next. This allows the defense system to anticipate the adversary's behavior, rather than just react. Given a set of escape strategy categories, the module outputs two key pieces of information: the current strategy probability distribution and the strategy transition probability matrix, where each element represents the probability of switching from one strategy state to another.

[0027] Specifically, the adversary strategy inference module includes: a variational autoencoder unit for encoding and clustering the motion patterns of the intruding drone; a strategy memory unit for storing prototypes of escape strategies encountered in the past and their evaluation of their effectiveness against the drone; an attention retrieval unit for retrieving the historical strategies most relevant to the current situation from the strategy memory; and a strategy prediction unit for combining the current encoding results with the historical retrieval results to output the probability distribution of the strategy category and the transition probability matrix.

[0028] In a specific embodiment, the adversary strategy inference module first encodes and performs pattern matching on the real-time motion data of the intruding drone. This module uses a variational autoencoder to compress multi-dimensional motion features (such as position, velocity, and acceleration sequences) within a continuous time window into a set of low-dimensional latent encoded vectors representing its inherent maneuvering patterns. Simultaneously, the module calls a pre-built strategy memory—a database storing various escape strategy prototypes from historical adversarial scenarios. Each prototype contains its motion pattern encoded vector, statistical features, and historical adversarial effects. Using an attention retrieval mechanism, the system calculates the comprehensive similarity between the current encoding and all prototypes in the memory. This similarity is determined by the encoding vector similarity, motion feature matching degree, and the context of the current defensive posture. After retrieving the several historical strategy prototypes with the highest similarity, they are weighted and fused using attention weights. Finally, the strategy prediction unit outputs a quantified probability distribution, explicitly indicating the most likely escape strategy type the intruding drone will currently adopt (e.g., a 70% probability of serpentine maneuver evasion) and its probability matrix of potential switching to other strategies, providing crucial prior input for subsequent situation prediction and decision generation.

[0029] This module possesses continuous evolution capabilities, the core of which lies in its online update mechanism for the policy memory. When the system's predictions deviate significantly from the actual behavior of the intruding drone, or when the matching degree between the current motion pattern and all historical prototypes falls below a threshold, a new knowledge generation and integration process is triggered. At this point, the system generates new codes for the currently uninterpreted motion sequences using a variational autoencoder, combines this with the final result of the adversarial exercise (such as successful escape or interception), creates a new policy prototype, and stores it in the memory. Simultaneously, the system dynamically adjusts the weights or confidence levels of existing prototypes based on the actual results of each adversarial exercise: prototypes that successfully aided prediction are strengthened, while prototypes that repeatedly failed to predict are weakened. Through this closed loop of perception-inference-verification-update, the policy inference module can continuously absorb new adversarial experience, expanding its policy recognition boundaries, thereby maintaining high inference accuracy and adaptability even when facing constantly evolving and unprecedented escape tactics.

[0030] The policy prediction unit can employ continuous-time policy transition prediction based on neural differential equations, enabling it to predict policy evolution over continuous time, rather than transitions at discrete time steps. The formula is as follows: , in, It is the latent state vector of the policy, which changes continuously over time. It is a moment A unified situational representation. It is a situation encoder network that encodes high-dimensional situations. Mapped to a conditional vector. It is defined by the drift function and diffusion function parameterized by the neural network, which together define a neural stochastic differential equation. It is a standard Wiener process (Brownian motion) used to model the inherent randomness in policy evolution. ξ is a learnable latent policy style variable used to distinguish the tactical preferences of different aggressor individuals. This upgrades policy inference from a discrete probability matrix to a continuous-time stochastic process. The model not only outputs the policy transition probability at the next time step but also generates the complete policy state distribution at any future time point. This is crucial for predicting hybrid or asymptotic policy switching with continuously changing characteristics, and provides more accurate continuous-time conditional inputs for multi-step situation prediction.

[0031] The multi-step situation prediction module is used to predict the situation evolution over multiple decision cycles based on the current state and the identified escape strategy; this module is the future inference engine of the policy generation network. It takes the current state and the adversary's policy inference as input, and through physical simulation or learning models, predicts in parallel the possible situations at multiple future time points under different defense strategies. This enables the system to assess the long-term consequences of different action plans. Conditional variational autoencoders or deep probabilistic prediction models can be used. For example, given the current state and the assumption of a straight-line escape strategy, the module might predict: 1 second later: the target will move forward 20 meters, and the defending drone A will be 15 meters away from the target; 3 seconds later: the target will move forward 60 meters, the defending drone A will be 10 meters away from the target, and drone B will be 25 meters away from the target; 5 seconds later: the target will approach the boundary, and the defending drone A may have a 70% probability of successfully intercepting it.

[0032] Specifically, the multi-step situation prediction module includes: a sequence generator based on the Transformer architecture, used to generate probabilistic escape trajectories of intrusion targets in the future, based on the current state of the defense formation and the identified escape strategies; a multi-scale prediction unit, used to generate trajectory predictions at different time granularities in parallel, including precise short-term maneuver details, medium-term probabilistic trajectories representing tactical branches, and long-term situations indicating strategy-level evolution trends; a trajectory confidence evaluation unit, used to calculate a confidence score for each generated trajectory that integrates physical feasibility, historical frequency, and situation matching degree; and a trajectory clustering and filtering unit, used to cluster a large number of generated probabilistic trajectories into a finite number of typical scenarios.

[0033] The multi-step situation prediction module first uses a sequence generator based on the Transformer architecture. It takes the current unified situation representation and the escape strategy probabilities output by the adversary strategy inference module as inputs. Through its unique self-attention mechanism, it captures the spatiotemporal interaction between the defensive formation and the invading target, and in parallel deduces dozens or even hundreds of probabilistic escape trajectories that the invading target may take over multiple decision cycles in the future. These trajectories are not single, fixed paths, but are presented as probability distributions, each with a probability weight. Simultaneously, the multi-scale prediction unit within the module is activated, performing differentiated predictions for different time scales: for short-term scales of 0.5 to 2 seconds, predictions focus on precise maneuver details, such as specific acceleration changes and turning radii; for medium-term scales of 2 to 10 seconds, predictions characterize possible tactical branching points, such as the probability of the target choosing to flank to the left or attack to the right at a specific location; and for long-term scales exceeding 10 seconds, predictions emphasize strategy-level situational evolution trends, such as whether the target as a whole is attempting to break through the encirclement or retreat to a specific area. All generated trajectories are sent to the trajectory confidence evaluation unit. This unit uses a trained evaluation network to comprehensively calculate the physical feasibility (whether it meets the dynamic constraints), historical frequency of occurrence (the degree of commonness in the policy memory), and matching degree with the current real-time situation of each trajectory, and finally assigns a confidence score between 0 and 1 to each trajectory.

[0034] After generating a large number of possible trajectories with probability weights and confidence scores, the trajectory clustering and filtering unit begins its work to extract actionable prediction scenarios. This unit first employs a spatiotemporal feature-based clustering algorithm (such as DBSCAN) to group all trajectories according to the similarity of their spatial direction and time points, classifying hundreds of scattered trajectories into a limited number of typical clusters, each representing a major future trend. Next, the system selects the most representative trajectory from each cluster (usually the one with the highest combined probability weight and confidence score) and calculates the total probability of that cluster (the sum of probabilities of all trajectories within the cluster) as the significance of that typical scenario. For example, after clustering and filtering, three typical scenarios might be output: Scenario A (probability 40%, confidence 0.85) indicates a high probability that the target will accelerate and break through in a southeast direction; Scenario B (probability 35%, confidence 0.78) indicates that the target may use a serpentine maneuver to maneuver westward; and Scenario C (probability 25%, confidence 0.65) indicates a low probability that the target will attempt to climb and escape. These aggregated and filtered typical scenarios, along with their significance and confidence levels, are ultimately output to the strategy optimization generation module as the core input for multi-step forward-looking decision-making and collaborative strategy optimization, effectively balancing the richness of predictions with the feasibility of decisions.

[0035] Specifically, the following formula can be used: a trajectory predictor based on physical information neural operators is used as a sequence generator for the multi-step situation prediction module to ensure that the generated trajectory strictly follows the physical conservation laws.

[0036] , in, It is a neural operator with parameters as follows: It learns the mapping from input to output. It is the initial state tensor of all intelligent agents (intrusion and defense). It is the sequence of coordinated control instructions of the defender over the time interval [0, T]. C is the escape strategy condition code obtained from the adversary strategy inference module. It is a predicted sequence of future trajectories. The constraint is a partial differential equation of physical information (in this case, the continuity equation of the law of conservation of mass), where L is the Lagrangian density of the scene and v is the velocity field. This constraint is hard- or soft-embedded into the neural operator through a penalty term. During training, physical conservation laws are introduced as strong constraints into the deep learning model. This ensures that the generated predicted trajectories are not only reasonable in terms of data distribution but also feasible in terms of physical dynamics (e.g., energy is roughly conserved, and there are no abrupt changes that violate dynamics). This greatly improves the confidence of the generated trajectories, especially for long-term situations requiring long-term predictions, avoiding the physically absurd outputs that might occur from purely data-driven models.

[0037] The strategy optimization and generation module is used to generate cooperative interception commands by integrating prediction results and to verify the logical consistency and feasibility of the generated strategies. The integrated prediction results refer to predictions of multiple future scenarios from the multi-step situation prediction module. The cooperative interception command is a cooperative action plan for multiple defense drones, including parameters such as the target position, speed, and time for each drone. Logical consistency verification checks for contradictions within the generated strategy (e.g., two drones being assigned the same spatiotemporal location). Feasibility verification checks whether the strategy meets physical constraints (e.g., maximum speed and acceleration limits). This module is the decision optimization and verification center of the strategy generation network. It integrates all prediction information, generates the optimal cooperative strategy through optimization algorithms, and ensures that the strategy is logically and physically feasible.

[0038] S3. The coordinated interception commands are converted into motion control parameters to drive multiple defense drones to perform coordinated interception tasks, and dynamic adjustments are made based on real-time situational awareness during execution. Distributed coordinated control and dynamic adaptation. This step is responsible for converting high-level strategy commands into specific executable actions. Its functions include solving abstract coordinated commands into low-level motion control parameters for each defense drone platform, and driving multiple platforms to form a formation to perform interception tasks. The key function is dynamic adjustment based on real-time situational awareness; that is, the execution unit does not statically execute preset commands, but has online feedback and adjustment capabilities. It can adaptively adjust control parameters, formation configuration, and even task roles according to real-time environmental changes and platform status during task execution to ensure the robustness and completion of the task in a dynamic confrontation environment.

[0039] Specifically, in this embodiment, the step of converting the cooperative interception command into motion control parameters to drive multiple defense drones to perform cooperative interception tasks includes: a dynamic role allocation step: assigning dynamic roles, including interception, blocking, and decoy, to each defense drone based on the real-time situation, and predicting the role cooperative evolution path; a control command mapping step: mapping high-level cooperative strategies to low-level control parameters of each drone through a motion primitive library; a distributed cooperative control step: establishing a local communication and consensus mechanism among defense drones to achieve formation maintenance and task coordination; and an anomaly handling step: monitoring execution deviations and system anomalies in real time, triggering trajectory replanning or mode switching to ensure task continuity.

[0040] During the execution phase of the coordinated interception command, the system first performs dynamic role allocation. Based on a real-time updated unified situational awareness, this step assigns dynamic tactical roles to each defensive UAV in the formation. For example, UAV A is designated as the primary interceptor responsible for frontal approach, UAVs B and C as flanking blockers responsible for compressing the target's escape space, while UAV D may temporarily serve as a decoy or reserve to respond to unforeseen changes. This allocation is not static but based on a real-time predictive model. This model not only assigns current roles but also predicts how the roles of each UAV might evolve collaboratively over several decision cycles as the target maneuvers and the situation evolves (e.g., when the target suddenly turns, the original flanking blocker B needs to switch to the primary interceptor). Subsequently, the control command mapping step is triggered. The system uses a predefined kinematic primitive library to decompose high-level, semantically defined collaborative strategies (such as A and B forming a pincer movement) into a low-level, temporally ordered sequence of motion control parameters executable by each UAV, including specific roll angles, pitch angles, throttle inputs, and corresponding execution times, ensuring that high-level tactical intentions are accurately translated into low-level flight maneuvers.

[0041] The dynamic role allocation step includes: extracting spatial, kinematic, and cooperative relationship features related to roles from the unified situational awareness using a graph neural network; predicting the role evolution sequence for multiple decision cycles based on a sequence prediction model; solving for the optimal role allocation scheme under multiple constraints; and designing a smooth switching controller to ensure the continuity and stability of motion during role transitions. The dynamic role allocation step begins with deep feature extraction from the unified situational awareness using a graph neural network. This network abstracts the defensive formation and intrusion targets as nodes in a graph, and their spatial relationships and kinematic interactions as edges, thereby extracting multidimensional features highly correlated with role competence, including the azimuth angle of each UAV relative to the target, the rate of change of distance, remaining energy, and cooperative constraints such as the unobstructed line of sight between UAVs. Based on these features, a sequence prediction model based on a long short-term memory network and an attention mechanism is activated. This model not only evaluates the optimal role configuration at the current moment (e.g., determining which UAV is most suitable as the primary interceptor), but also predicts how the role configuration of the entire formation will dynamically evolve over the next 5 to 10 decision cycles as the target maneuvers and energy is consumed, outputting a probability sequence of role evolution. Subsequently, a multi-objective optimization solver runs in real time, simultaneously considering multiple constraints such as maximum interception success rate, minimum overall energy consumption, communication load balancing, and avoiding role conflicts. By solving this optimization problem, it assigns a Pareto-optimal role allocation scheme for the current moment. Finally, to ensure flight safety and formation stability of the UAVs when executing role switching commands, the system invokes a specially designed smooth switching controller. This controller, based on model predictive control principles, calculates an optimal trajectory for each UAV to smoothly transition from its current state to the state required by its new role. This trajectory strictly satisfies dynamic constraints and avoids conflicts with teammates, thereby achieving motion continuity and overall formation stability during the tactical role reassignment process.

[0042] Based on commutative game theory, the optimal role allocation scheme is solved under multiple constraints. It is modeled as a cooperative game with transition utility, as shown in the following formula:

[0043] , , Where D represents the collection of defensive drones, He was one of them. It is a role allocation scheme. Indicates drone The assigned role, Π, is the set of all possible options. It is a drone In the allocation scheme The local utility function under the condition of expected interception benefit Costs related to energy, communication, etc. constitute. It is a cooperative gain term. It is a drone and The synergy coefficient matrix between them (calculated from graph neural network features). It is an indicator function that measures the synergistic effect of pairing two machine characters. It is the core solution set of the cooperative game, ensuring that the allocation scheme is stable for the alliance (i.e., no subset of drones can gain higher benefits by deviating from the current scheme). λ and β are the balance coefficients that adjust the weights of interception benefits and cooperative gains. It indicates that the condition is met, or that it is subject to or constrained by. Representing the collective of all participants It is a characteristic function defined on all subsets. It is the expected value function.

[0044] This model elevates role allocation from a static optimization problem to a dynamic cooperative game problem. It not only maximizes global efficiency but also ensures the stability and fairness of the allocation through the concept of a core solution, preventing any drone or drone group from deviating from its cooperative strategy due to dissatisfaction with the allocation. This is crucial for maintaining long-term cooperation among autonomous individuals in a distributed system.

[0045] After the command is issued, the distributed collaborative control steps begin to operate to maintain the robustness of the formation coordination. Each UAV does not rely solely on centralized commands from the ground station, but rather establishes a local communication network with each other. By exchanging their own status, intentions, and local observation information, and running lightweight distributed consensus algorithms (such as consensus protocols based on neighboring node states), they collaboratively maintain the overall interception formation (such as a diamond encirclement) and dynamically coordinate the mission rhythm, maintaining basic cooperation even during brief interruptions of the central communication link. Simultaneously, an anomaly handling step continues to operate as a safety guarantee. This step monitors execution deviations (such as exceeding position error limits or response delays) in real time by comparing the actual motion state of each UAV with the expected command trajectory; it also monitors the system's health status (such as battery level, abnormal sensor data, or loss of contact with teammates). Once a predefined abnormal condition is detected, the system immediately triggers the corresponding autonomous recovery mechanism. For example, for minor trajectory deviations, local trajectory replanning is performed; for single-UAV failures or strong interference, dynamic mode switching is initiated, directing nearby UAVs to take over their mission roles, or the entire formation switches to a degraded but safe collaborative mode, thereby ensuring that the continuity of the interception mission is not unexpectedly interrupted.

[0046] S4. During the interception mission, continuously collect adversarial interaction trajectory data. Adversarial interaction data collection and recording. The function of this step is to systematically capture and store complete interaction information throughout the entire adversarial process. This information constitutes the original experience dataset upon which subsequent learning and evolutionary stages rely.

[0047] Specifically, in this embodiment, during the execution of the cooperative interception mission, the system continuously and synchronously records multi-dimensional adversarial interaction trajectory data through a distributed data acquisition framework. This data is stored in a structured manner with timestamp alignment. Its content not only includes high-frequency state snapshots of all UAVs (including intrusion targets and defense formations) in each decision cycle (such as 0.1-second intervals), such as precise three-dimensional position, velocity vector, attitude angle, and raw sensor readings, but also fully records the cooperative interception commands output by the strategy generation network at the corresponding moment, the actual execution control parameters of each UAV, and the fusion perception results from the low-altitude situational awareness unit. In addition, the system will specially mark key event points, such as the starting moment of the defender's maneuver interception, the moment of identification when the intrusion target changes its escape strategy, the alarm moment when both sides approach the preset danger distance, and the final outcome of the mission (successful interception, target escape, or mission termination). This forms a multimodal, full-element interaction trajectory archive containing spatiotemporal dynamics, decision logic, execution feedback, and adversarial results, providing a high-quality, traceable original data source for subsequent cognitive distillation and model evolution.

[0048] S5. Based on the adversarial interaction trajectory data, new features and evolutionary patterns of adversarial modes are extracted using cognitive distillation technology, which is then used to update the cognitive modeling module of the policy generation network. Adversarial knowledge extraction and incremental model update. This step is crucial for the evolutionary capability of this method. Its function is to automatically extract potential, generalizable new adversarial mode features and their evolutionary patterns from the large amount of high-dimensional raw interaction data collected in step 4 using machine learning techniques such as cognitive distillation. Then, this extracted structured knowledge (rather than raw data) is incrementally learned and updated into the cognitive modeling module of the policy generation network, thereby improving the network's understanding and adaptation to similar or new variant adversarial strategies, and achieving iterative enhancement of model performance.

[0049] Specifically, in this embodiment, the extraction of new features and evolutionary patterns of adversarial modes using cognitive distillation technology to update the cognitive modeling module of the strategy generation network includes: automatically identifying new adversarial mode features from adversarial interaction trajectory data; converting the identified adversarial modes into abstract knowledge representations that can be understood by the cognitive model; and fusing the newly extracted knowledge representations with the existing knowledge base of the system to resolve knowledge conflicts and redundancy issues. The fusion steps include: evaluating the compatibility of new knowledge with existing knowledge in multiple dimensions, including logical consistency, empirical support, and practical utility value; initiating a multi-strategy conflict resolution process based on evidence theory, case reasoning, and priority rules when knowledge conflicts are detected; constructing a traceable knowledge evolution graph to record the complete path and verification effect of knowledge updates; testing the robustness and generalization ability of the fused knowledge system in a diverse adversarial simulation environment; and updating the fused knowledge to the cognitive modeling module of the strategy generation network using incremental learning.

[0050] A specific implementation of cognitive distillation technology begins with the system automatically mining patterns from a large amount of adversarial interaction trajectory data. Using unsupervised clustering and anomaly detection algorithms, it identifies statistically significant new adversarial pattern features not yet covered by existing policy memories, such as a hybrid escape strategy combining sudden deceleration and irregular zigzag turns. Next, the system uses a semantic abstraction network to transform these identified new patterns, represented as low-level motion sequences, into high-level, structured knowledge representations that the cognitive model can directly understand and process. These representations typically employ a rule-based form similar to IF-THEN or a probabilistic graphical model, for example, encoding a high probability that the intruder will adopt strategy X when the defender is in a tight encirclement, characterized by attributes Y and Z. The core knowledge fusion process then begins: the system first conducts a quantitative assessment of the compatibility of new and old knowledge from three dimensions: logical consistency (whether the new rule contradicts the existing rule base), empirical support (the frequency and stability of the new pattern in historical data), and practical utility value (the historical win rate of using this pattern). Once a conflict is detected, such as the new rule strategy X being effective while the old rule strategy X is easily countered, the system initiates a multi-strategy conflict resolution process based on evidence theory, multi-source confidence, reference to the solutions of similar historical cases, and following preset expert priority rules to decide which knowledge to adopt, revise, or discard. During this process, a traceable knowledge evolution graph is dynamically updated, fully recording the triggering data, decision basis, fusion results, and subsequent test performance in the verification environment. To ensure reliability, the fused new knowledge system will be subjected to rigorous stress testing in a high-fidelity simulation environment containing multiple unknown threat variants to evaluate its robustness and generalization ability. Only knowledge that passes all tests will be safely integrated into the cognitive modeling module of the policy generation network through incremental learning by fine-tuning network parameters or updating the rule base. This will enable continuous learning and adaptation to unknown adversarial modes without compromising existing core capabilities.

[0051] The formula for knowledge fusion conflict detection based on topological data analysis is as follows. This formula is used for conflict detection in the knowledge fusion process, identifying fundamental contradictions between old and new knowledge from the perspective of data topology.

[0052] , in: These represent the newly extracted knowledge representation and the existing knowledge base of the system, respectively. It is a dataset of related adversarial interaction trajectories used for verification. It is a mapping function from knowledge-data pairs to point cloud data. Specifically, it applies a knowledge base to the data. On top of that, a set of high-dimensional feature vectors (such as prediction error, state distribution, etc.) are generated to form a point cloud. It calculates the dimensionality of the point cloud. The first The order Betti number is a topological invariant representing the order of Betti numbers in a point cloud. The number of voids (e.g.) It is the number of connected components. (It is the number of ring structures). The weights of Betti numbers in different dimensions are typically assigned to higher-order holes, as they represent more complex structures. The conflict index measures the topological differences in data interpretation resulting from new and old knowledge. A higher index indicates fundamentally different interpretations of the same data by new and old knowledge (e.g., one views the data as multiple discrete clusters, while the other views it as a continuous manifold), foreshadowing a fundamental conflict. This formula provides a deeper conflict detection method based on topological essence. Traditional methods may only detect direct contradictions in rule formulations, while this formula can uncover deep contradictions in the understanding of data generation mechanisms between new and old knowledge. Even if two pieces of knowledge do not have direct logical conflicts, forced fusion can lead to system performance collapse if their implicit data topological models are inconsistent. This greatly enhances the security and reliability of the knowledge fusion process.

[0053] S6. Establish an online evolution verification mechanism to perform performance verification before deploying the updated policy generation network, ensuring that the defense system achieves a safe and stable closed-loop cognitive evolution. This step provides security and quality control for the entire learning and evolution process. Its function is to comprehensively test and evaluate the performance, stability, and security of the updated network in a controlled verification environment, based on a preset multi-dimensional performance indicator system, before deploying the policy generation network updated in step 5 to the actual operating environment. Only updated versions that pass verification are allowed to be deployed, thus ensuring that each self-update of the system leads to performance improvement or at least maintenance, without introducing uncontrollable degradation or risks, ultimately achieving a safe, stable, and closed-loop system cognitive evolution process.

[0054] Specifically, in this embodiment, the steps of establishing an online evolution verification mechanism include: establishing a multi-dimensional evaluation index system encompassing interception success rate, response speed, resource efficiency, and strategy novelty; designing a progressive evolution strategy, decomposing system updates into multiple controllable stages and verifying them step by step; conducting comprehensive testing and comparative verification of the new version strategy generation network in a digital twin simulation environment, with test scenarios specifically including novel adversarial modes that trigger this knowledge update, to verify the system's absorption effect of newly learned knowledge and its performance improvement in corresponding scenarios; and establishing an intelligent version management and automatic rollback mechanism, automatically restoring to a stable version when performance degradation exceeds a safety threshold. A specific implementation of the online evolution verification mechanism begins with establishing a refined multi-dimensional evaluation index system. This system not only includes the core interception success rate and average response speed but also incorporates resource efficiency indicators (such as average energy consumption per interception and communication bandwidth utilization) and strategy novelty scores (used to quantify the difference and creativity of new strategies relative to the historical strategy library), thereby providing a comprehensive profile of the strategy generation network's performance. Based on this framework, the system employs a progressive evolutionary strategy, breaking down a complete model update into multiple controllable stages. For example, the first stage only updates the parameters of the opponent policy inference module and verifies it in limited scenarios. The second stage gradually introduces updates to the multi-step prediction module and expands the testing scope, ensuring that each change can be isolated, verified, and evaluated. Subsequently, in a highly realistic digital twin simulation environment, the system conducts large-scale, multi-round comprehensive testing of the new version of the policy generation network. This environment accurately simulates real-world physical dynamics, sensor noise, and various complex adversarial scenarios. The testing process includes not only A / B comparisons with the current online version but also horizontal comparisons with a series of historical baseline versions on tasks of different difficulty levels to accurately quantify its performance gains or losses. To ensure the safety boundaries of the evolution process, the system simultaneously runs an intelligent version management and automatic rollback mechanism. This mechanism continuously monitors the performance of new versions across all evaluation dimensions. Once it detects that the decline of any key indicator (such as the interception success rate) exceeds the preset safety threshold, or that a fatal flaw occurs in a specific high-risk scenario, it will automatically trigger the rollback process, seamlessly restoring the system to the previous fully validated stable version and generating a detailed diagnostic report for subsequent iterative improvements. This achieves closed-loop operation of model evolution under the premise of controllable risk and guaranteed performance.

[0055] Example 2

[0056] like Figure 2As shown, this application provides an architecture diagram of an AI-based autonomous defense decision-making and response system for tower-stationed unmanned aerial vehicles (UAVs), which is applied to the AI-based autonomous defense decision-making and response system for tower-stationed UAVs as described in Embodiment 1. The system includes: a low-altitude situational awareness unit 210, a strategy generation network 220, an instruction execution and dynamic control unit 230, an adversarial interaction data acquisition module 240, a cognitive evolutionary learning module 250, and an online evolutionary verification unit 260.

[0057] The low-altitude situational awareness unit 210 is used to acquire multi-source perception data and construct a unified situational representation that includes inferences about the physical state and tactical intentions of the intruding UAV.

[0058] The policy generation network 220, trained based on a metacognitive reinforcement learning framework, is used to receive the unified situation representation and generate cooperative interception instructions through its integrated adversary policy inference module and multi-step situation prediction module. The metacognitive reinforcement learning framework includes two learning mechanisms: the inner layer generates the current optimal policy through adversarial simulation and the outer layer learns the evolution law of adversarial mode through cognitive modeling.

[0059] The instruction execution and dynamic control unit 230 is used to convert the cooperative interception instruction into motion control parameters, drive multiple defense drones to perform cooperative interception tasks, and make dynamic adjustments based on the real-time situation during the execution process.

[0060] The adversarial interaction data acquisition module 240 is used to continuously collect adversarial interaction trajectory data during the execution of interception tasks.

[0061] The cognitive evolutionary learning module 250 is used to extract new features and evolutionary patterns of adversarial patterns based on the adversarial interaction trajectory data through cognitive distillation technology, and to update the cognitive modeling module of the policy generation network.

[0062] The online evolution verification unit 260 is used to perform performance verification before deploying the updated policy generation network, ensuring that the defense system achieves secure and stable closed-loop cognitive evolution.

[0063] Figure 3 This is an electronic device provided in one embodiment of this application. For example... Figure 3 As shown, the electronic device includes at least the following components: processor 301 and memory 300, communication interface 303, and bus 302.

[0064] In this embodiment of the application, memory 300 is used to store executable instructions of processor 301, which, when configured to execute instructions, implements the method as described in the first aspect.

[0065] In embodiments of this application, a computer-readable storage medium includes instructions that instruct a device to perform the method as described in the first aspect. For example, the instructions instruct the device to perform... Figure 1 The method is shown in the process steps.

[0066] In one embodiment of this application, the program operating in the electronic device may be a program that controls the central processing unit (CPU) to achieve the functions of the above-described embodiments of the present invention (a program that enables the computer to function). Information processed by these systems is then temporarily stored in random access memory (RAM) during processing, and subsequently stored in various ROMs, including read-only memory (FlashROM) and hard disk drives (HDDs), and read, corrected, and written by the CPU as needed.

[0067] It should be noted that a portion of the electronic device described in the above embodiments can also be implemented using a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the program recorded on the recording medium can be read into the computer and executed.

[0068] It should be noted that the computer mentioned here refers to a computer built into an electronic device, employing hardware including an operating system and peripheral devices. Furthermore, computer-readable recording media refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage systems including hard drives built into the computer.

[0069] Furthermore, computer-readable recording media can include: media that dynamically stores programs for short periods of time, such as communication lines used when transmitting programs via networks including the Internet or telephone lines; and media that store programs for fixed periods of time, such as volatile memory inside a computer serving as a server or client in this case. In addition, the aforementioned program can be a program used to implement the above-mentioned functions, or it can be a program that can implement the above-mentioned functions by combining them with programs already recorded in the computer.

[0070] Furthermore, the electronic device in the above embodiments can also be implemented as an assembly (system group) composed of multiple systems. Each system constituting the system group can possess some or all of the functions or functional blocks of the electronic device in the above embodiments. As a system group, it is sufficient to have all the functions or functional blocks of the electronic device.

[0071] Those skilled in the art should recognize that the above embodiments are only used to illustrate this application and are not intended to limit this application. Any appropriate changes and variations made to the above embodiments within the essential spirit and scope of this application fall within the scope of protection claimed in this application.

Claims

1. An artificial intelligence-based autonomous defense decision-making and response method for tower-stationed unmanned aerial vehicles (UAVs), characterized in that, include: By acquiring multi-source sensing data through low-altitude situational awareness units, a unified situational representation is constructed that includes inferences about the physical state and tactical intentions of intruding drones. The unified situational representation is input into a policy generation network trained based on a metacognitive reinforcement learning framework. The policy generation network generates cooperative interception instructions by integrating an adversary policy inference module and a multi-step situational prediction module. The metacognitive reinforcement learning framework includes two learning mechanisms: the inner layer generates the current optimal policy through adversarial simulation and the outer layer learns the evolution law of adversarial mode through cognitive modeling. The coordinated interception command is converted into motion control parameters to drive multiple defense drones to perform coordinated interception tasks, and dynamic adjustments are made based on the real-time situation during the execution process; During the interception mission, data on adversarial interaction trajectories are continuously collected; Based on the adversarial interaction trajectory data, new features and evolutionary patterns of the adversarial patterns are extracted using cognitive distillation technology, which are then used to update the cognitive modeling module of the policy generation network. Establish an online evolution verification mechanism to perform performance verification before deploying the updated policy generation network, ensuring that the defense system achieves secure and stable closed-loop cognitive evolution.

2. The method according to claim 1, characterized in that, The steps for constructing a unified situational representation that includes inferences about the physical state and tactical intentions of intruding drones specifically include: By analyzing the maneuver sequences of intruding drones using a spatiotemporal graph attention mechanism, their tactical intentions can be inferred and probability distributions of intention categories can be generated. Establish a semantic encoder for multi-source sensing data to convert radio detection signals, radar point cloud data, and electro-optical tracking information into a unified tactical feature vector; Based on the intention probability distribution and tactical feature vector, the threat entropy value of the current confrontation situation is calculated to quantify environmental uncertainty; Physical state information, intent probability distribution, threat entropy value, and environmental constraint information are fused into a machine-readable multidimensional cognitive tensor, which serves as the unified situation representation.

3. The method according to claim 1, characterized in that, The policy generation network includes the following modules: The metacognitive scheduling module is used to dynamically configure the attention allocation strategy and inference computation depth within the network based on the characteristics of the adversarial situation. The adversary strategy inference module is used to identify the potential escape strategy categories of intruding drones and their switching probabilities in real time. The multi-step situation prediction module is used to predict the situation evolution over multiple decision cycles based on the current state and the identified escape strategy. The strategy optimization and generation module is used to generate collaborative interception instructions by comprehensively predicting the results, and to verify the logical consistency and feasibility of the generated strategy.

4. The method according to claim 3, characterized in that, The opponent strategy inference module includes: The variational autoencoder unit is used to encode and cluster the motion patterns of the intruding drone; The strategy memory unit is used to store prototypes of escape strategies encountered in the past and their evaluation of their effectiveness against them. An attention retrieval unit is used to retrieve the historical strategy most relevant to the current situation from the strategy memory. The strategy prediction unit combines the current encoding results with historical retrieval results to output the probability distribution and transition probability matrix of the strategy category.

5. The method according to claim 3, characterized in that, The multi-step situation prediction module includes: A sequence generator based on the Transformer architecture is used to generate a probabilistic escape trajectory of an intrusion target in the future, based on the current state of the defense formation and the identified escape strategy. Multi-scale prediction units are used to generate trajectory predictions at different time granularities in parallel, including precise short-term maneuver details, medium-term probabilistic trajectories characterizing tactical branches, and long-term situations identifying strategy-level evolution trends. The trajectory confidence assessment unit is used to calculate a confidence score for each generated trajectory, taking into account the overall physical feasibility, historical frequency, and situational matching degree. The trajectory clustering and filtering unit is used to cluster a large number of probabilistic trajectories into a finite number of typical scenarios.

6. The method according to claim 1, characterized in that, The step of converting the cooperative interception command into motion control parameters to drive multiple defense drones to perform cooperative interception tasks specifically includes: Dynamic role allocation steps: Assign dynamic roles, including interception, blocking, and decoy, to each defense drone based on the real-time situation, and predict the collaborative evolution path of the roles; Control command mapping steps: Map the high-level cooperative strategy to the low-level control parameters of each UAV through the motion primitive library; Distributed collaborative control steps: Establish local communication and consensus mechanisms among defense drones to achieve formation maintenance and task coordination; Anomaly handling steps: Monitor execution deviations and system anomalies in real time, and trigger trajectory replanning or mode switching to ensure task continuity.

7. The method according to claim 6, characterized in that, The dynamic role allocation step includes: Extract role-related spatial, motion, and cooperative relationship features from a unified situational representation using graph neural networks; Predicting role evolution sequences across multiple decision-making cycles based on sequence prediction models; Solve for the optimal role allocation scheme under multiple constraints; The design of the smooth switching controller ensures the continuity and stability of movement during character transitions.

8. The method according to claim 1, characterized in that, The method of extracting new features and evolutionary patterns of adversarial modes through cognitive distillation, and using them to update the cognitive modeling module of the policy generation network, includes: Automatically identify new adversarial pattern features from adversarial interaction trajectory data; The identified adversarial patterns are converted into abstract knowledge representations that can be understood by the cognitive model; The newly extracted knowledge representation is integrated with the existing knowledge base of the system to resolve knowledge conflicts and redundancy issues. The integration steps include: The system assesses the compatibility of new knowledge with existing knowledge across multiple dimensions, including logical consistency, empirical support, and practical utility. When knowledge conflicts are detected, a multi-strategy conflict resolution process is initiated based on evidence theory, case reasoning, and priority rules. A traceable knowledge evolution graph is constructed to record the complete path of knowledge updates and verification results. The system tests the robustness and generalization ability of the fused knowledge system in diverse adversarial simulation environments. The fused knowledge is updated to the cognitive modeling module of the policy generation network using an incremental learning approach.

9. The method according to claim 1, characterized in that, The steps for establishing the online evolutionary verification mechanism specifically include: Establish a multi-dimensional evaluation index system that includes interception success rate, response speed, resource efficiency, and strategy novelty; Design a progressive evolution strategy to break down system updates into multiple controllable stages and verify them step by step; The new version of the policy generation network was comprehensively tested and compared in a digital twin simulation environment; Establish an intelligent version management and automatic rollback mechanism to automatically revert to a stable version when performance degradation exceeds a safety threshold.

10. An AI-based autonomous defense decision-making and response system for tower-stationed unmanned aerial vehicles (UAVs), applied to the method described in any one of claims 1 to 9, characterized in that, The system includes: The low-altitude situational awareness unit is used to acquire multi-source sensing data and construct a unified situational representation that includes inferences about the physical state and tactical intentions of intruding drones. The strategy generation network, trained based on a metacognitive reinforcement learning framework, is used to receive the unified situation representation and generate cooperative interception instructions through its integrated adversary strategy inference module and multi-step situation prediction module. The metacognitive reinforcement learning framework includes two learning mechanisms: the inner layer generates the current optimal strategy through adversarial simulation and the outer layer learns the evolution law of adversarial mode through cognitive modeling. The command execution and dynamic control unit is used to convert the cooperative interception command into motion control parameters, drive multiple defense drones to perform cooperative interception tasks, and make dynamic adjustments based on the real-time situation during the execution process; The adversarial interaction data acquisition module is used to continuously collect adversarial interaction trajectory data during the execution of interception tasks; The cognitive evolutionary learning module is used to extract new features and evolutionary patterns of adversarial patterns based on the adversarial interaction trajectory data through cognitive distillation technology, and to update the cognitive modeling module of the policy generation network. The online evolution verification unit is used to perform performance verification before deploying the updated policy generation network, ensuring that the defense system achieves secure and stable closed-loop cognitive evolution.