An anti-unmanned aerial vehicle simulation system and method fusing a warhead damage model
Patent Information
- Application Number
- CN202610767854.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-08-21
AI Technical Summary
[0004]为了弥补以上不足,本发明提供了一种融合战斗部毁伤模型的反无人机仿真系统及方法,旨在改善现有反无人机仿真中高保真毁伤物理仿真与实时飞行控制仿真在时间尺度上不可调和,导致无法将精确毁伤评估接入闭环强化学习训练的问题
[0052]1、本发明中,通过构建离线高保真物理仿真与在线轻量化代理推断的跨时间尺度解耦架构,以预先训练的轻量化多毁伤元代理模型替代传统的显式动力学实时计算,使得带有高保真毁伤特性判定的反无人机仿真能够流畅运行,实现了将战斗部终点毁伤效应的高精度评估无缝嵌入无人机飞行控制闭环控制算法训练。
Smart Images

Figure CN122616318A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of real-time damage simulation and intelligent decision-making technology, and in particular to an anti-drone simulation system and method that integrates a warhead damage model. Background Technology
[0002] Micro-sized drone swarms pose an increasingly serious threat in low-altitude penetration. Terminal autonomous interception using interceptor drones equipped with directional explosive warheads has become an important development direction in the counter-drone field. To train interceptors to possess end-to-end autonomous flight and detonation control capabilities, extensive closed-loop training in high-fidelity simulation environments is essential. However, existing counter-drone simulation platforms generally employ simplified kinematic rendezvous criteria or overall damage probability assessment methods based on empirical formulas, making it difficult to achieve a balance between simulation accuracy and computational efficiency.
[0003] While realistic explosion fragment penetration and fluid-structure interaction explicit dynamic simulations (such as in LS-DYNA) can provide high-fidelity damage assessment results, a single calculation can take several days or even weeks. This makes it difficult to integrate with the physical flight engine of a UAV, which requires a real-time operating frequency of over 100Hz, for closed-loop reinforcement learning training. Consequently, high-fidelity damage physics simulation and real-time flight control simulation are irreconcilable on a time scale. Summary of the Invention
[0004] To overcome the above shortcomings, this invention provides an anti-drone simulation system and method that integrates a warhead damage model. It aims to improve the problem that the high-fidelity damage physics simulation and real-time flight control simulation are incompatible in terms of time scale in existing anti-drone simulations, which makes it impossible to integrate accurate damage assessment into closed-loop reinforcement learning training.
[0005] In a first aspect, the present invention provides the following technical solution: an anti-drone simulation system integrating a warhead damage model, comprising:
[0006] The multimodal feature acquisition and dimensionality reduction module is used to acquire the multimodal rendezvous state features of the interceptor and target aircraft in the real-time dynamic simulation engine.
[0007] The fuze determination and rendezvous transient extraction module is used to extract the six-degree-of-freedom pose parameters of the detonation transient when the interceptor and the target aircraft meet the preset virtual proximity fuze triggering conditions, and to vector-superimpose the static fragmentation velocity of the warhead with the terminal velocity of the interceptor to reconstruct the dynamic spatial kill zone.
[0008] The multi-destruction element surrogate model inference module is used to input at least the six-degree-of-freedom pose parameters of the detonation transient as transient intersection features into the pre-trained lightweight multi-destruction element surrogate model, and output the micro-failure probability sequence of each key component inside the target aircraft.
[0009] The dynamic damage tree logic mapping module is used to input the micro-failure probability sequence into the embedded dynamic damage tree logic mapping model, and output the macro-system-level functional loss state of the target aircraft through multi-level spatial and logic gate aggregation operations.
[0010] The closed-loop execution and strategy feedback module is used to transform the macro-system-level functional loss state into an adaptive endpoint feedback reward for the control algorithm, and drive the target aircraft to execute the physical degradation response or crash response corresponding to the macro-system-level functional loss state in the real-time dynamic simulation engine, forming an end-to-end closed-loop training.
[0011] Preferably, in the multimodal feature acquisition and dimensionality reduction module, the step of acquiring the multimodal rendezvous state features of the interceptor and target aircraft specifically includes:
[0012] At long range, the three-dimensional coordinates of the target aircraft with Gaussian white noise are acquired as radar guidance status features to guide the interceptor aircraft to align its course.
[0013] During the close encounter phase, the visual sensor node is activated, and a lightweight target detection algorithm is called to extract the two-dimensional bounding box parameters of the target aircraft from the simulated camera image. The offset vector of the high-value components inside the target aircraft relative to the center of the two-dimensional bounding box is semantically segmented as a visual dimensionality reduction feature.
[0014] The radar guidance state features and the visual dimensionality reduction features are fused into a one-dimensional state vector, which serves as the multimodal intersection state features.
[0015] Preferably, in the fuse determination and rendezvous transient extraction module, the step of extracting the six-degree-of-freedom pose parameters of the detonation transient specifically includes:
[0016] The relative state between the interceptor aircraft and the target aircraft is monitored in the real-time dynamic simulation engine.
[0017] When the relative state satisfies the preset detonation distance and azimuth geometric threshold of the directional warhead carried by the interceptor aircraft, the detonation signal is triggered;
[0018] Based on the detonation signal, extract the six-degree-of-freedom pose parameters of the detonation transient.
[0019] Preferably, in the fuze determination and rendezvous transient extraction module, the step of reconstructing the dynamic spatial kill zone specifically includes:
[0020] Obtain the terminal velocity of the interceptor vehicle from the six-degree-of-freedom pose parameters of the detonation transient;
[0021] The static fragmentation velocity of the warhead is vector-superimposed with the terminal velocity of the interceptor vehicle to obtain the dynamic fragmentation velocity field.
[0022] The dynamic spatial kill zone is reconstructed based on the dynamic fragment dispersion velocity field.
[0023] Preferably, in the multi-damage surrogate model inference module, the step of outputting the microscopic failure probability sequence of each key component inside the target aircraft specifically includes:
[0024] The six-degree-of-freedom pose parameters of the detonation transient are used as the transient intersection features and input into the pre-trained lightweight multi-damage surrogate model.
[0025] The lightweight multi-damage surrogate model performs forward inference based on the transient intersection features and outputs the microscopic failure probability sequence of each key component inside the target aircraft.
[0026] Preferably, the step of outputting the microscopic failure probability sequence of each key component inside the target aircraft by forward inference based on the transient rendezvous features using the lightweight multi-damage surrogate model specifically includes:
[0027] The lightweight multi-damage proxy model receives the transient intersection features as input;
[0028] The lightweight multi-damage element proxy model, based on the transient rendezvous characteristics, calculates the kinetic energy failure probability of each key component inside the target aircraft under the action of fragment armor-piercing damage elements, and the overload failure probability under the action of shock wave damage elements.
[0029] The kinetic failure probability and the overload failure probability are combined for probability calculation, and the microscopic failure probability sequence of each key component inside the target aircraft is output.
[0030] Preferably, in the dynamic damage tree logic mapping module, the step of outputting the macroscopic system-level functional loss state of the target aircraft specifically includes:
[0031] Receive the microscopic failure probability sequence;
[0032] The failure probabilities of each key component in the micro-failure probability sequence are input into the dynamic damage tree logic mapping model. According to the preset spatial adjacency effect cascade judgment mechanism and series-parallel logic gate structure in the dynamic damage tree logic mapping model, probability aggregation calculation is performed layer by layer from the bottom component node to the top system node.
[0033] Based on the result of the probability aggregation operation, the macroscopic system-level functional loss state of the target aircraft is output, which includes mission damage state and maneuver damage state.
[0034] Preferably, in the closed-loop execution and policy feedback module, the step of transforming the macroscopic system-level functional loss state into an adaptive endpoint feedback reward for the control algorithm specifically includes:
[0035] Receive the macroscopic system-level functional loss state output by the dynamic damage tree logic mapping module;
[0036] When the macro-system-level functional loss state is a task damage state, a first-order basic reward value is generated.
[0037] When the macro-system-level functional loss state is a motor damage state, a second-order higher-order reward value is generated, which is greater than the first-order basic reward value.
[0038] When the dynamic damage tree logic mapping model determines that the preset high-value weak component is covered by the reconstructed directional fragmentation field with a probability higher than the set threshold, a precision strike tactical reward value is generated by superimposing it on the first-order basic reward value or the second-order higher-order reward value.
[0039] The first-order basic reward value, the second-order higher-order reward value, or the reward value obtained by superimposing the precision strike tactical reward value are used as the adaptive endpoint feedback reward of the control algorithm.
[0040] Preferably, in the closed-loop execution and strategy feedback module, the step of driving the target aircraft to execute the physical degradation response or crash response corresponding to the macroscopic system-level functional loss state in the real-time dynamic simulation engine specifically includes:
[0041] Receive the macroscopic system-level functional loss state output by the dynamic damage tree logic mapping module;
[0042] When the macro-system level function loss state is a mission damage state, the robot operating system command communication node sends a wandering command or return command to the flight control node of the target aircraft, driving the target aircraft to perform a physical degradation response in the real-time dynamic simulation engine;
[0043] When the macro-system level loss state is a maneuver damage state, a forced power-off and propeller lock command is issued to the flight control node of the target aircraft through the robot operating system command communication node, driving the target aircraft to execute a crash response in the real-time dynamic simulation engine;
[0044] The physical degradation response or crash response, together with the adaptive endpoint feedback reward, constitutes the end-to-end closed-loop training.
[0045] Secondly, the present invention provides the following technical solution: a method for simulating anti-drone damage by integrating a warhead destruction model, the method comprising:
[0046] In the real-time dynamic simulation engine, the multimodal rendezvous state characteristics of the interceptor aircraft and the target aircraft are obtained;
[0047] When the interceptor and the target aircraft meet the preset virtual proximity fuse triggering conditions, the six-degree-of-freedom pose parameters of the detonation transient are extracted, and the static fragmentation velocity of the warhead is vector-superimposed with the terminal velocity of the interceptor to reconstruct the dynamic spatial kill zone.
[0048] At least the six-degree-of-freedom pose parameters of the detonation transient are used as transient intersection features and input into a pre-trained lightweight multi-damage surrogate model to output the microscopic failure probability sequence of each key component inside the target aircraft.
[0049] The micro-failure probability sequence is input into the embedded dynamic damage tree logic mapping model, and the macro-system-level functional loss state of the target aircraft is output through multi-level spatial and logic gate aggregation operations.
[0050] The macroscopic system-level functional loss state is transformed into an adaptive endpoint feedback reward for the control algorithm, and the target aircraft is driven to execute a physical degradation response or crash response corresponding to the macroscopic system-level functional loss state in the real-time dynamic simulation engine, forming an end-to-end closed-loop training.
[0051] The present invention has the following beneficial effects:
[0052] 1. In this invention, by constructing a cross-timescale decoupled architecture of offline high-fidelity physical simulation and online lightweight agent inference, and replacing the traditional explicit dynamic real-time calculation with a pre-trained lightweight multi-damage agent model, the anti-UAV simulation with high-fidelity damage characteristic determination can run smoothly, and realizes the seamless embedding of high-precision evaluation of the warhead terminal damage effect into the training of UAV flight control closed-loop control algorithm.
[0053] 2. In this invention, through the synergistic effect of the multimodal visual perception feature dimensionality reduction mechanism and the directional reward feedback mechanism based on dynamic damage tree, the trained end-to-end interception strategy no longer stops at simple close-range collision, but can autonomously emerge to fly around the target's weak point location and find the best detonation attitude angle to achieve advanced tactical behavior of directional and precise damage to high-value components.
[0054] 3. In this invention, by converting the macroscopic system-level functional loss state output by the damage tree logic mapping model into the adaptive endpoint feedback reward of the control algorithm and the physical degradation or crash response command of the target aircraft in the simulation engine, a complete end-to-end closed-loop training system of perception-rendezvous-damage-feedback-disabling is constructed, providing a directly driven simulation training environment for the iterative optimization of anti-UAV interception strategies. Attached Figure Description
[0055] Figure 1 This is an architecture diagram of an anti-drone simulation system that integrates a warhead damage model, as proposed in this invention.
[0056] Figure 2 This is a flowchart of the multimodal perception and feature dimensionality reduction of an anti-UAV simulation system that integrates a warhead damage model, as proposed in this invention.
[0057] Figure 3 This is a schematic diagram of a dynamic damage tree structure with a control algorithm reward mapping for an anti-UAV simulation system that integrates a warhead damage model, as proposed in this invention.
[0058] Figure 4 This is a schematic diagram of the network structure of a lightweight multi-damage agent model for an anti-UAV simulation system that integrates a warhead damage model, as proposed in this invention. Detailed Implementation
[0059] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] Example 1:
[0061] In a first embodiment of the present invention, the present invention provides an anti-drone simulation system that integrates a warhead damage model, such as... Figure 1 As shown, it includes:
[0062] The multimodal feature acquisition and dimensionality reduction module is used to acquire the multimodal rendezvous state features of the interceptor and target aircraft in the real-time dynamic simulation engine.
[0063] Furthermore, in the multimodal feature acquisition and dimensionality reduction module, the steps for acquiring the multimodal rendezvous state features of the interceptor and target aircraft specifically include:
[0064] At long range, the three-dimensional coordinates of the target aircraft with Gaussian white noise are acquired as radar guidance status characteristics to guide the interceptor aircraft to align its course.
[0065] During the close-range rendezvous phase, the visual sensor node is activated, and a lightweight target detection algorithm is called to extract the two-dimensional bounding box parameters of the target aircraft from the simulated camera image. The offset vector of the high-value components inside the target aircraft relative to the center of the two-dimensional bounding box is semantically segmented and used as a visual dimensionality reduction feature.
[0066] The radar guidance state features and visual dimensionality reduction features are fused into a one-dimensional state vector, which serves as the state feature for multimodal intersection.
[0067] Specifically, when the relative distance between the interceptor and the target aircraft exceeds a preset distance threshold, the system obtains the true coordinates of the target aircraft in three-dimensional space through an airborne radar simulation node. Gaussian white noise is injected into these coordinates to simulate the calculation error of the miniature radar, generating noisy three-dimensional coordinates as radar guidance state characteristics. The interceptor's flight control node performs a heading alignment maneuver based on these radar guidance state characteristics, causing the interceptor to approach the target aircraft. As a preferred embodiment, the preset distance threshold is set to 50 meters.
[0068] When the relative distance between the interceptor and the target aircraft enters a preset close-range encounter threshold range, the system activates the onboard visual sensor simulation node. The visual sensor simulation node acquires image frames in real-time from the simulated camera's rendered image, processes the image frames using a lightweight target detection algorithm, and outputs the target aircraft's two-dimensional bounding box parameters. Specifically, the two-dimensional bounding box parameters are the coordinates of the bounding box center in the pixel coordinate system, as well as the width and height of the bounding box. As a preferred embodiment, the preset close-range encounter threshold range is less than 15 meters.
[0069] Based on the two-dimensional bounding box obtained by the lightweight target detection algorithm, the semantic segmentation branch is further invoked to perform pixel-level classification of the region inside the bounding box, identifying high-value components inside the target aircraft, and calculating the bias vector of the geometric center of each high-value component relative to the center of the two-dimensional bounding box. (Bias vector) Specifically, it is expressed as follows:
[0070] ;
[0071] Where n is the number of high-value components identified. Let be the offset between the geometric center of the i-th high-value component and the center of the two-dimensional bounding box in the horizontal direction of the pixel coordinate system. Let be the offset between the geometric center of the i-th high-value component and the center of the two-dimensional bounding box in the vertical direction of the pixel coordinate system. High-value components inside the target aircraft refer to those that play a crucial role in enabling the target aircraft to complete its predetermined mission or maintain its flight attitude, including the flight controller, battery, motors, navigation module, and communication antenna. The specific types of these high-value components can be pre-configured according to the configuration of the target aircraft.
[0072] The aforementioned two-dimensional bounding box parameters and bias vector together constitute the visual dimensionality reduction features.
[0073] The system performs temporal alignment and data association between radar guidance state features acquired at long range and visual dimensionality reduction features acquired at close range. It then employs a Kalman filter algorithm for multi-sensor data fusion to generate a one-dimensional state vector. This one-dimensional state vector is composed of at least the target aircraft's three-dimensional spatial position estimate, target aircraft's three-dimensional velocity estimate, two-dimensional bounding box parameters, and the bias vectors of each high-value component, concatenated sequentially. This one-dimensional state vector serves as the input to the online control network for the multimodal intersection state features. The specific implementation process of the aforementioned multimodal feature acquisition and dimensionality reduction module is as follows: Figure 2 As shown.
[0074] The fuze determination and rendezvous transient extraction module is used to extract the six-degree-of-freedom pose parameters of the detonation transient when the interceptor and target aircraft meet the preset virtual proximity fuze trigger conditions, and to vector superimpose the static fragmentation velocity of the warhead with the terminal velocity of the interceptor aircraft to reconstruct the dynamic spatial kill zone.
[0075] Furthermore, in the fuse determination and rendezvous transient extraction module, the specific steps for extracting the six-degree-of-freedom pose parameters of the detonation transient include:
[0076] In the real-time dynamic simulation engine, the relative state between the interceptor aircraft and the target aircraft is monitored;
[0077] When the relative state meets the preset detonation distance and azimuth geometric threshold of the directional warhead carried by the interceptor aircraft, the detonation signal is triggered.
[0078] Based on the detonation signal, extract the six-degree-of-freedom pose parameters of the detonation transient.
[0079] Furthermore, in the fuse determination and rendezvous transient extraction module, the steps for reconstructing the dynamic spatial kill zone specifically include:
[0080] Obtain the terminal velocity of the interceptor vehicle from the six-DOF pose parameters during the detonation transient.
[0081] The static fragmentation velocity field of the warhead is vector-superimposed with the terminal velocity of the interceptor vehicle to obtain the dynamic fragmentation velocity field.
[0082] Based on the dynamic fragment dispersion velocity field, the dynamic spatial kill zone is reconstructed.
[0083] Specifically, the fuze determination and rendezvous transient extraction module is used to extract the six-degree-of-freedom pose parameters of the detonation transient when the interceptor and target aircraft meet the preset virtual proximity fuze triggering conditions. It then vector-superimposes the static fragmentation velocity of the warhead with the terminal velocity of the interceptor aircraft to reconstruct the dynamic spatial kill zone. The implementation process consists of two stages: trigger determination and parameter extraction, and vector superposition and kill zone reconstruction.
[0084] In the real-time dynamic simulation engine, the relative state between the interceptor and the target aircraft is monitored at a preset simulation step size. The relative state includes the relative distance, relative azimuth angle, and relative velocity vector between the interceptor and the target aircraft. When the relative state meets the preset detonation distance and azimuth angle geometric thresholds of the directional warhead carried by the interceptor, the system triggers a detonation signal. The detonation distance and azimuth angle geometric thresholds are predetermined by the fragmentation cone angle and kill radius of the directional warhead. The triggering condition can be expressed as:
[0085] ;
[0086] in, To determine the current relative distance between the interceptor and the target aircraft, To preset the detonation distance threshold, The current azimuth angle of the target aircraft relative to the aiming axis of the interceptor aircraft's directional warhead. A preset detonation azimuth threshold is used. As a preferred embodiment, The value ranges from 3 meters to 8 meters. The value ranges from ±15° to ±30°.
[0087] Based on the detonation signal, the system extracts the six-DOF pose parameters during the detonation transient. These parameters include the interceptor's three-dimensional position coordinates, three-dimensional attitude angles, three-dimensional linear velocity vector, and three-dimensional angular velocity vector at the moment of detonation. These six-DOF pose parameters are directly read and frozen from the aircraft status interface of the real-time dynamic simulation engine.
[0088] Obtain the terminal velocity of the interceptor in the six-DOF pose parameters during the detonation transient. The terminal velocity is the three-dimensional linear velocity vector of the interceptor at the instant of detonation, expressed as... The dynamic fragmentation velocity field is obtained by vector superposition of the static fragmentation velocity field of the warhead and the terminal velocity of the interceptor. The static fragmentation velocity field of the warhead is the initial fragmentation velocity vector of each fragment when the warhead detonates in a static state, which is determined in advance by the warhead charge parameters, fragment pre-forming parameters, and shell structure. For the j-th fragment in the fragmentation field, the vector superposition is calculated using the following formula:
[0089] ;
[0090] in, Let be the dynamic dispersion velocity vector of the j-th fragment. Let be the static fragment field scattering velocity vector of the j-th fragment. This is the terminal velocity vector of the interceptor aircraft. The above vector superposition is performed on all fragments in the fragmentation field to obtain the dynamic fragment dispersion velocity field.
[0091] Based on the dynamic fragment dispersion velocity field, the dynamic spatial kill zone is reconstructed. The reconstruction process is as follows: taking the three-dimensional position coordinates of the detonating transient interceptor as the origin and the aiming axis direction of the interceptor's directional warhead as the central axis, according to the fragment dispersion angle range and initial velocity distribution of the static fragment field of the directional warhead, combined with the dynamic fragment dispersion velocity field after the above vector superposition, the dynamic dispersion trajectory envelope of each fragment is determined in three-dimensional space. The spatial region enclosed by this envelope is the dynamic spatial kill zone.
[0092] Through the above two stages, the system completes the entire process from fuse triggering determination and detonation transient parameter fixing to dynamic kill zone reconstruction, providing transient intersection feature input for the subsequent multi-damage element proxy model inference module.
[0093] The multi-destruction surrogate model inference module is used to input at least the six-degree-of-freedom pose parameters of the detonation transient as transient intersection features into a pre-trained lightweight multi-destruction surrogate model, and output the micro-failure probability sequence of each key component inside the target aircraft.
[0094] Furthermore, in the multi-damage surrogate model inference module, the steps for outputting the microscopic failure probability sequence of each key component inside the target aircraft specifically include:
[0095] The six-degree-of-freedom pose parameters of the detonation transient are used as transient intersection features and input into a pre-trained lightweight multi-damage surrogate model.
[0096] The lightweight multi-damage surrogate model performs forward inference based on transient intersection features, and outputs the microscopic failure probability sequence of each key component inside the target aircraft.
[0097] Furthermore, the steps of using a lightweight multi-damage surrogate model to perform forward inference based on transient intersection features and output the microscopic failure probability sequence of each key component inside the target aircraft specifically include:
[0098] The lightweight multi-damage surrogate model receives transient intersection features as input;
[0099] The lightweight multi-damage element surrogate model is based on transient rendezvous characteristics to calculate the kinetic energy failure probability of each key component inside the target aircraft under the action of fragment armor-piercing damage elements, and the overload failure probability under the action of shock wave damage elements.
[0100] The kinetic failure probability and the overload failure probability are jointly calculated to output the microscopic failure probability sequence of each key component inside the target aircraft.
[0101] Specifically, the multi-destruction element surrogate model inference module is used to input at least the six-DOF pose parameters of the detonation transient as transient intersection features into a pre-trained lightweight multi-destruction element surrogate model, and output the microscopic failure probability sequence of each key component inside the target aircraft. Its implementation process is divided into two stages: surrogate model input and forward inference.
[0102] The six-DOF pose parameters of the detonation transient extracted by the fuse determination and rendezvous transient extraction module are used as transient rendezvous features. Specifically, these transient rendezvous features include the relative distance, relative azimuth, relative velocity vector magnitude, and detonation offset angle between the interceptor and target aircraft at the moment of detonation. These features are directly derived from the six-DOF pose parameters or obtained through geometric transformation. The transient rendezvous features are organized into a fixed-dimensional input feature vector and input into a pre-trained lightweight multi-damage surrogate model. The lightweight multi-damage surrogate model is a pre-trained offline multilayer perceptron neural network. The number of nodes in its input layer equals the dimension of the transient rendezvous features, the number of nodes in its output layer equals the number of key components inside the target aircraft, and the hidden layers use a predetermined number of fully connected layers. Its network structure is as follows: Figure 4 As shown. Specifically, the lightweight multi-damage surrogate model is a multilayer perceptron neural network, including: an input layer with the number of nodes equal to the dimension of the transient intersection features, used to receive transient intersection features such as relative distance, relative azimuth angle, relative velocity vector magnitude, and detonation offset angle; a first hidden layer, a fully connected layer using the ReLU activation function, used to perform nonlinear mapping on the input features; a second hidden layer, also a fully connected layer using the ReLU activation function, used to reduce and fuse the features; and an output layer with the number of nodes equal to the number of key components inside the target aircraft, using the Sigmoid activation function, used to output a sequence of microscopic failure probabilities for each key component with values ranging from 0 to 1.
[0103] The offline training process of the lightweight multi-damage surrogate model is as follows: A parameterized script is used to generate batch explicit dynamic collision simulation datasets containing different rendezvous velocities, detonation offset angles, and miss distances. In the explicit dynamic simulation, the complex heterogeneous components of the target aircraft are equivalent to standard aluminum alloy target plates for fragment penetration calculation. The equivalent target plate method is used to extract the remaining kinetic energy and ultimate penetration velocity of the fragments. The penetration kinetic energy corresponding to the remaining kinetic energy of the fragments being greater than the ultimate penetration velocity is used as the criterion to calculate the fragment penetration failure probability of each key component inside the target aircraft. Simultaneously, the explosive load function is applied to the target... Overpressure is applied to the aircraft body, and the peak overpressure and arrival impulse of each key component structural node are recorded. The overload failure probability of the shock wave to the target structural node is evaluated based on the overpressure-impulse criterion. That is, when the combination point of the peak overpressure and arrival impulse is above the preset overpressure-impulse damage threshold curve, the component is determined to have overload failure. Using multidimensional intersection parameters as input features and the joint probability results of fragment penetration failure probability and shock wave overload failure probability as supervision labels, a multilayer perceptron network is trained to obtain a lightweight multi-damage surrogate model that can respond to the micro-failure probability sequence in milliseconds.
[0104] The lightweight multi-damage element surrogate model receives transient intersection features as input and performs forward inference based on these features to calculate the kinetic energy failure probability of each critical component inside the target aircraft under the action of fragmentation armor-piercing damage elements, and the overload failure probability under the action of shock wave damage elements. For the k-th critical component, the kinetic energy failure probability is expressed as... The overload failure probability is expressed as The kinetic failure probability and overload failure probability are jointly calculated to output the microscopic failure probability sequence of each critical component inside the target aircraft. For the k-th critical component, its microscopic failure probability... Calculate using the following formula:
[0105] ;
[0106] in, Let be the kinetic energy failure probability of the k-th critical component under the sole action of a fragmentation armor-piercing damage element. Let be the overload failure probability of the k-th critical component under the action of a single shock wave damage element. Let be the combined failure probability of the k-th critical component under the combined action of the two damaging elements.
[0107] The lightweight multi-damage surrogate model outputs a sequence of microscopic failure probabilities for each critical component, expressed as follows: , where K is the total number of critical components assessed inside the target aircraft. This micro-failure probability sequence serves as the input to the subsequent dynamic damage tree logic mapping module.
[0108] The dynamic damage tree logic mapping module is used to input the micro-failure probability sequence into the embedded dynamic damage tree logic mapping model, and output the macro-system-level functional loss state of the target aircraft through multi-level spatial and logic gate aggregation operations.
[0109] Furthermore, in the dynamic damage tree logic mapping module, the steps for outputting the macroscopic system-level functional loss state of the target aircraft specifically include:
[0110] Receive micro-failure probability sequence;
[0111] The failure probabilities of each key component in the micro-failure probability sequence are input into the dynamic damage tree logic mapping model. According to the preset spatial adjacency effect cascade judgment mechanism and series-parallel logic gate structure in the dynamic damage tree logic mapping model, probability aggregation calculation is performed layer by layer from the bottom component node to the top system node.
[0112] Based on the results of the probability aggregation operation, the macroscopic system-level functional loss state of the target aircraft is output. The macroscopic system-level functional loss state includes mission damage state and maneuver damage state.
[0113] Specifically, the dynamic damage tree logic mapping module is used to input the micro-failure probability sequence into the embedded dynamic damage tree logic mapping model. Through multi-level spatial and logic gate aggregation operations, it outputs the macro-level system-level functional loss state of the target aircraft. Its implementation process is divided into two stages: input reception and spatial cascade processing, and layer-by-layer probability aggregation and macro-level state output.
[0114] The system receives the microscopic failure probability sequence output by the multi-damage surrogate model inference module. This sequence includes the failure probabilities of each critical component within the target aircraft, denoted as... Where K is the total number of critical components. The failure probability of each critical component in the micro-failure probability sequence is input into the dynamic damage tree logic mapping model. The dynamic damage tree logic mapping model pre-defines the component-level damage tree structure of the target aircraft. This damage tree structure uses each critical component as the bottom-level component node, each functional subsystem as the intermediate-level subsystem node, and the mission damage state and maneuver damage state as the top-level system node. The nodes at each level are connected through series and parallel logic gates according to the functional logic relationship of the target aircraft. Before aggregating the failure probabilities of the bottom-level component nodes to the intermediate-level subsystem nodes, a spatial adjacency effect cascade determination is first performed. When a high-energy physics node is determined to be failed, the failure probability of its physical adjacent nodes is synchronously increased according to the preset spatial decay function. High-energy physics nodes are components with secondary damage effects, including battery nodes. The spatial adjacency effect cascade determination is performed according to the following formula:
[0115] ;
[0116] in, Let be the original failure probability of the physically adjacent node j. Let j be the failure probability of the physically adjacent node after adjustment for spatial adjacency effect. Let be the spatial distance between high-energy physics node i and its physically adjacent node j in the target spacecraft's coordinate system. The spatial attenuation constant, The coupling coefficient is... It is a natural constant. As a preferred embodiment, The value ranges from 0.2 meters to 0.5 meters. The value ranges from 0.3 to 0.6.
[0117] After determining the cascading effect of spatial adjacency, probability aggregation is performed layer by layer from the bottom component nodes to the top system nodes, according to the pre-defined series and parallel logic gate structure within the dynamic damage tree logic mapping model. For a group of nodes connected by series logic gates, the aggregated failure probability is the OR operation result of the failure probabilities of each node in the group, calculated using the following formula:
[0118] ;
[0119] Where M is the number of nodes connected by the serial logic gates. Let be the failure probability of the m-th node in the cascade group. This represents the failure probability after the aggregation of serial logic gates.
[0120] For a group of nodes connected by parallel logic gates, its aggregate failure probability is the sum of the failure probabilities of each node in the group, calculated using the following formula:
[0121] ;
[0122] Where M is the number of nodes connected by the parallel logic gates. Let m be the failure probability of the m-th node in the parallel group. This represents the failure probability after the parallel logic gates are aggregated.
[0123] Based on the above-mentioned series and parallel logic gate aggregation operation rules, aggregation is performed from the bottom component nodes to the intermediate subsystem nodes, and then from the intermediate subsystem nodes to the top system nodes, ultimately obtaining the failure probability of the top system nodes.
[0124] Based on the aggregated failure probabilities of the top-level system nodes, the macroscopic system-level functional loss status of the target aircraft is determined and output. Specifically, the determination rules are as follows: when the aggregated failure probability of a communication subsystem node or navigation subsystem node exceeds a preset mission damage determination threshold, a mission damage status is output; when the aggregated failure probability of a propulsion subsystem node or flight control subsystem node exceeds a preset maneuver damage determination threshold, a maneuver damage status is output. As a preferred embodiment, both the mission damage determination threshold and the maneuver damage determination threshold are set to 0.5. The mission damage status and the maneuver damage status can be output simultaneously.
[0125] The output of the macroscopic system-level functional loss state serves as the input for the subsequent closed-loop execution and strategy feedback modules.
[0126] The closed-loop execution and strategy feedback module is used to transform the macro-system-level functional loss state into an adaptive endpoint feedback reward for the control algorithm, and drive the target aircraft to execute the physical degradation response or crash response corresponding to the macro-system-level functional loss state in the real-time dynamic simulation engine, forming an end-to-end closed-loop training.
[0127] Furthermore, in the closed-loop execution and policy feedback module, the steps to transform the macroscopic system-level functional failure state into the adaptive endpoint feedback reward of the control algorithm specifically include:
[0128] Receive the macroscopic system-level functional loss status output by the dynamic damage tree logic mapping module;
[0129] When the macro-system level function loss state is the task damage state, the first-order basic reward value is generated;
[0130] When the macro-system level loss of function is in the state of motor damage, a second-order higher-order reward value is generated, which is greater than the first-order basic reward value.
[0131] When the dynamic damage tree logic mapping model determines that the preset high-value weak components are covered by the reconstructed directional fragmentation field with a probability higher than the set threshold, a precision strike tactical reward value is generated on top of the first-order basic reward value or the second-order higher-order reward value.
[0132] The first-order basic reward value, the second-order higher-order reward value, or the reward value after superimposing the precision strike tactical reward value are used as the adaptive endpoint feedback reward of the control algorithm.
[0133] Furthermore, in the closed-loop execution and strategy feedback module, the steps for driving the target aircraft to execute the physical degradation response or crash response corresponding to the macroscopic system-level functional loss state in the real-time dynamic simulation engine specifically include:
[0134] Receive the macroscopic system-level functional loss status output by the dynamic damage tree logic mapping module;
[0135] When the macro-system level is in a state of functional loss, the robot operating system command communication node sends a wandering command or return command to the target aircraft's flight control node, driving the target aircraft to perform a physical degradation response in the real-time dynamic simulation engine.
[0136] When the macro-system level loss of function is a maneuver damage state, a forced power-off and propeller lock command is issued to the flight control node of the target aircraft through the robot operating system command communication node, driving the target aircraft to execute a crash response in the real-time dynamic simulation engine;
[0137] Among them, the physical degradation response or crash response and the adaptive endpoint feedback reward together constitute the end-to-end closed-loop training.
[0138] Specifically, the closed-loop execution and policy feedback module is used to transform the macroscopic system-level functional failure state into an adaptive endpoint feedback reward for the control algorithm, and drive the target aircraft to execute the physical degradation response or crash response corresponding to the macroscopic system-level functional failure state in the real-time dynamic simulation engine, forming an end-to-end closed-loop training. Its implementation process is divided into two stages: reward shaping and physical execution.
[0139] The system receives the macro-level system-level functional loss status output by the dynamic damage tree logic mapping module. When the macro-level system-level functional loss status is a mission damage status, a first-order basic reward value is generated. When the macro-level system-level functional loss status is a maneuver damage status, a second-order higher-order reward value is generated, which is greater than the first-order basic reward value. As a preferred embodiment, the first-order basic reward value is +10, and the second-order higher-order reward value is +100. Further, the damage determination results of preset high-value weak components in the dynamic damage tree logic mapping model are obtained. High-value weak components are the key subset of high-value components that, once damaged, would cause the target aircraft to lose its core combat capability. High-value weak components include the flight controller and the battery. When the dynamic damage tree logic mapping model determines that the reconstructed directional fragmentation field of a high-value weak component covers it with a probability higher than a set threshold, a precision strike tactical reward value is generated by superimposing it on the first-order basic reward value or the second-order higher-order reward value. The determination condition for probability coverage is expressed as follows:
[0140] ;
[0141] in, This represents the failure probability of high-value, weak components in the microscopic failure probability sequence. To set a threshold. As a preferred embodiment, The value is 0.7. As a preferred embodiment, the precision strike tactical bonus value is +50.
[0142] The first-order basic reward value, the second-order higher-order reward value, or the reward value obtained by superimposing the precision strike tactical reward value are used as the adaptive endpoint feedback reward of the control algorithm. The control algorithm uses this adaptive endpoint feedback reward as the target signal to optimize the offset detonation maneuver and rendezvous attitude of the interceptor aircraft. The correspondence between the above dynamic damage tree logic mapping model and the control algorithm reward mapping is as follows: Figure 3 As shown.
[0143] While generating adaptive endpoint feedback rewards, the system receives the macroscopic system-level functional loss state output by the dynamic damage tree logic mapping module. When the macroscopic system-level functional loss state is a mission damage state, a loitering command or return-to-home command is issued to the target aircraft's flight control node via the robot operating system command communication node. The robot operating system command communication node encapsulates the command into a standard message format using a topic publishing mechanism and sends it to the control command topics subscribed to by the target aircraft's flight control node. After receiving the loitering command or return-to-home command, the target aircraft's flight control node calls the corresponding flight mode switching function, driving the target aircraft to execute a physical degradation response in the real-time dynamic simulation engine. A physical degradation response refers to the response of the target aircraft, driven by the simulation engine, to perform non-aggressive flight maneuvers when its functions are impaired but its basic flight capabilities are not lost. A physical degradation response manifests as the target aircraft no longer executing the preset attack mission, entering a fixed-point hovering loitering mode, or returning to the take-off and landing point along a preset route. When the macroscopic system-level functional loss state is a maneuver damage state, a forced power-off and propeller-locking command is issued to the target aircraft's flight control node via the robot operating system command communication node. After receiving the forced power-off and propeller lock command, the target aircraft's flight control node calls the motor emergency stop function to set the pulse width modulation output of all power motors to zero. The target aircraft loses all lift in the real-time dynamic simulation engine and falls freely along the ballistic trajectory under the action of gravity, completing the crash response.
[0144] The aforementioned physical degradation response or crash response, together with the adaptive endpoint feedback reward, constitutes an end-to-end closed-loop training. The adaptive endpoint feedback reward provides the control algorithm with numerical policy optimization signals, while the physical degradation response or crash response provides the control algorithm with visual simulation environment state feedback. Together, they form a complete closed loop from perception, intersection, damage determination to policy evaluation and state update.
[0145] Example 2:
[0146] While existing explicit dynamic simulations of real-world explosive fragment penetration and fluid-structure interaction can provide high-fidelity damage assessment results, their single calculations require days or even weeks. This makes them unsuitable for closed-loop reinforcement learning training within the physical flight engines of UAVs, which require real-time operating frequencies above 100Hz. Consequently, high-fidelity damage physics simulation and real-time flight control simulation are incompatible on a time scale. To address these issues, this invention provides an anti-UAV simulation method that integrates a warhead damage model, the structure of which is as follows: Figures 1-3 As shown. The specific implementation process of this method is as follows:
[0147] In the real-time dynamic simulation engine, the multimodal rendezvous state characteristics of the interceptor aircraft and the target aircraft are obtained;
[0148] When the interceptor and the target aircraft meet the preset virtual proximity fuse triggering conditions, the six-degree-of-freedom pose parameters of the detonation transient are extracted, and the static fragmentation velocity of the warhead and the terminal velocity of the interceptor are vector-superimposed to reconstruct the dynamic spatial kill zone.
[0149] At least the six-degree-of-freedom pose parameters of the detonation transient are used as transient intersection features and input into a pre-trained lightweight multi-damage surrogate model to output the micro-failure probability sequence of each key component inside the target aircraft.
[0150] The micro-failure probability sequence is input into the embedded dynamic damage tree logic mapping model, and the macro-system-level functional loss state of the target aircraft is output through multi-level spatial and logic gate aggregation operations.
[0151] The macroscopic system-level loss of function state is transformed into an adaptive endpoint feedback reward for the control algorithm, and the target aircraft is driven to execute the physical degradation response or crash response corresponding to the macroscopic system-level loss of function state in the real-time dynamic simulation engine, forming an end-to-end closed-loop training.
[0152] Specifically, in the real-time dynamic simulation engine, the multimodal rendezvous state characteristics of the interceptor and target aircraft are acquired. When the interceptor and target aircraft meet the preset virtual proximity fuse triggering conditions, the six-degree-of-freedom pose parameters of the detonation transient are extracted, and the static fragmentation velocity of the warhead and the terminal velocity of the interceptor are vector-superimposed to reconstruct the dynamic spatial kill zone. At least the six-degree-of-freedom pose parameters of the detonation transient are used as transient rendezvous features and input into a pre-trained lightweight multi-damage element proxy model to output the micro-failure probability sequence of each key component inside the target aircraft. The micro-failure probability sequence is input into the embedded dynamic damage tree logic mapping model, and through multi-level spatial and logic gate aggregation operations, the macro-system-level functional loss state of the target aircraft is output. The macro-system-level functional loss state is transformed into an adaptive endpoint feedback reward for the control algorithm and drives the target aircraft to execute the physical degradation response or crash response corresponding to the macro-system-level functional loss state in the real-time dynamic simulation engine, forming an end-to-end closed-loop training.
[0153] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A counter-drone simulation system integrating a warhead damage model, characterized in that, include: The multimodal feature acquisition and dimensionality reduction module is used to acquire the multimodal rendezvous state features of the interceptor and target aircraft in the real-time dynamic simulation engine. The fuze determination and rendezvous transient extraction module is used to extract the six-degree-of-freedom pose parameters of the detonation transient when the interceptor and the target aircraft meet the preset virtual proximity fuze triggering conditions, and to vector-superimpose the static fragmentation velocity of the warhead with the terminal velocity of the interceptor to reconstruct the dynamic spatial kill zone. The multi-destruction element surrogate model inference module is used to input at least the six-degree-of-freedom pose parameters of the detonation transient as transient intersection features into the pre-trained lightweight multi-destruction element surrogate model, and output the micro-failure probability sequence of each key component inside the target aircraft. The dynamic damage tree logic mapping module is used to input the micro-failure probability sequence into the embedded dynamic damage tree logic mapping model, and output the macro-system-level functional loss state of the target aircraft through multi-level spatial and logic gate aggregation operations. The closed-loop execution and strategy feedback module is used to transform the macro-system-level functional loss state into an adaptive endpoint feedback reward for the control algorithm, and drive the target aircraft to execute the physical degradation response or crash response corresponding to the macro-system-level functional loss state in the real-time dynamic simulation engine, forming an end-to-end closed-loop training.
2. The anti-drone simulation system integrating a warhead damage model according to claim 1, characterized in that, In the multimodal feature acquisition and dimensionality reduction module, the steps for acquiring the multimodal rendezvous state features of the interceptor and target aircraft specifically include: At long range, the three-dimensional coordinates of the target aircraft with Gaussian white noise are acquired as radar guidance status features to guide the interceptor aircraft to align its course. During the close encounter phase, the visual sensor node is activated, and a lightweight target detection algorithm is called to extract the two-dimensional bounding box parameters of the target aircraft from the simulated camera image. The offset vector of the high-value components inside the target aircraft relative to the center of the two-dimensional bounding box is semantically segmented as a visual dimensionality reduction feature. The radar guidance state features and the visual dimensionality reduction features are fused into a one-dimensional state vector, which serves as the multimodal intersection state features.
3. The anti-drone simulation system integrating a warhead damage model according to claim 1, characterized in that, In the fuse determination and rendezvous transient extraction module, the steps for extracting the six-degree-of-freedom pose parameters of the detonation transient specifically include: The relative state between the interceptor aircraft and the target aircraft is monitored in the real-time dynamic simulation engine. When the relative state satisfies the preset detonation distance and azimuth geometric threshold of the directional warhead carried by the interceptor aircraft, the detonation signal is triggered; Based on the detonation signal, extract the six-degree-of-freedom pose parameters of the detonation transient.
4. The anti-drone simulation system integrating a warhead damage model according to claim 1, characterized in that, In the fuse determination and rendezvous transient extraction module, the step of reconstructing the dynamic spatial kill zone specifically includes: Obtain the terminal velocity of the interceptor vehicle from the six-degree-of-freedom pose parameters of the detonation transient; The static fragmentation velocity of the warhead is vector-superimposed with the terminal velocity of the interceptor vehicle to obtain the dynamic fragmentation velocity field. The dynamic spatial kill zone is reconstructed based on the dynamic fragment dispersion velocity field.
5. The anti-drone simulation system integrating a warhead damage model according to claim 1, characterized in that, In the multi-damage surrogate model inference module, the step of outputting the microscopic failure probability sequence of each key component inside the target aircraft specifically includes: The six-degree-of-freedom pose parameters of the detonation transient are used as the transient intersection features and input into the pre-trained lightweight multi-damage surrogate model. The lightweight multi-damage surrogate model performs forward inference based on the transient intersection features and outputs the microscopic failure probability sequence of each key component inside the target aircraft.
6. The anti-drone simulation system based on a fusion warhead damage model according to claim 5, characterized in that, The step of outputting the microscopic failure probability sequence of each key component inside the target aircraft by performing forward inference based on the transient rendezvous features using the lightweight multi-damage surrogate model specifically includes: The lightweight multi-damage proxy model receives the transient intersection features as input; The lightweight multi-damage element proxy model, based on the transient rendezvous characteristics, calculates the kinetic energy failure probability of each key component inside the target aircraft under the action of fragment armor-piercing damage elements, and the overload failure probability under the action of shock wave damage elements. The kinetic failure probability and the overload failure probability are combined for probability calculation, and the microscopic failure probability sequence of each key component inside the target aircraft is output.
7. The anti-drone simulation system integrating a warhead damage model according to claim 1, characterized in that, In the dynamic damage tree logic mapping module, the step of outputting the macroscopic system-level functional loss state of the target aircraft specifically includes: Receive the microscopic failure probability sequence; The failure probabilities of each key component in the micro-failure probability sequence are input into the dynamic damage tree logic mapping model. According to the preset spatial adjacency effect cascade judgment mechanism and series-parallel logic gate structure in the dynamic damage tree logic mapping model, probability aggregation calculation is performed layer by layer from the bottom component node to the top system node. Based on the result of the probability aggregation operation, the macroscopic system-level functional loss state of the target aircraft is output, which includes mission damage state and maneuver damage state.
8. The anti-drone simulation system integrating a warhead damage model according to claim 7, characterized in that, In the closed-loop execution and policy feedback module, the step of transforming the macroscopic system-level functional loss state into an adaptive endpoint feedback reward for the control algorithm specifically includes: Receive the macroscopic system-level functional loss state output by the dynamic damage tree logic mapping module; When the macro-system-level functional loss state is a task damage state, a first-order basic reward value is generated. When the macro-system-level functional loss state is a motor damage state, a second-order higher-order reward value is generated, which is greater than the first-order basic reward value. When the dynamic damage tree logic mapping model determines that the preset high-value weak component is covered by the reconstructed directional fragmentation field with a probability higher than the set threshold, a precision strike tactical reward value is generated by superimposing it on the first-order basic reward value or the second-order higher-order reward value. The first-order basic reward value, the second-order higher-order reward value, or the reward value obtained by superimposing the precision strike tactical reward value are used as the adaptive endpoint feedback reward of the control algorithm.
9. The anti-drone simulation system integrating a warhead damage model according to claim 7, characterized in that, In the closed-loop execution and strategy feedback module, the steps of driving the target aircraft to execute the physical degradation response or crash response corresponding to the macroscopic system-level functional loss state in the real-time dynamic simulation engine specifically include: Receive the macroscopic system-level functional loss state output by the dynamic damage tree logic mapping module; When the macro-system level function loss state is a mission damage state, the robot operating system command communication node sends a wandering command or return command to the flight control node of the target aircraft, driving the target aircraft to perform a physical degradation response in the real-time dynamic simulation engine; When the macro-system level loss state is a maneuver damage state, a forced power-off and propeller lock command is issued to the flight control node of the target aircraft through the robot operating system command communication node, driving the target aircraft to execute a crash response in the real-time dynamic simulation engine; The physical degradation response or the crash response, together with the adaptive endpoint feedback reward, constitute the end-to-end closed-loop training.
10. A method for simulating anti-drone damage by integrating a warhead destruction model, characterized in that, The method of the anti-drone simulation system for a fused warhead damage model according to any one of claims 1-9 includes: In the real-time dynamic simulation engine, the multimodal rendezvous state characteristics of the interceptor aircraft and the target aircraft are obtained; When the interceptor and the target aircraft meet the preset virtual proximity fuse triggering conditions, the six-degree-of-freedom pose parameters of the detonation transient are extracted, and the static fragmentation velocity of the warhead is vector-superimposed with the terminal velocity of the interceptor to reconstruct the dynamic spatial kill zone. At least the six-degree-of-freedom pose parameters of the detonation transient are used as transient intersection features and input into a pre-trained lightweight multi-damage surrogate model to output the microscopic failure probability sequence of each key component inside the target aircraft. The micro-failure probability sequence is input into the embedded dynamic damage tree logic mapping model, and the macro-system-level functional loss state of the target aircraft is output through multi-level spatial and logic gate aggregation operations. The macroscopic system-level functional loss state is transformed into an adaptive endpoint feedback reward for the control algorithm, and the target aircraft is driven to execute a physical degradation response or crash response corresponding to the macroscopic system-level functional loss state in the real-time dynamic simulation engine, forming an end-to-end closed-loop training.