Spaceflight equipment object-scene data fusion evaluation method based on reinforcement learning

By establishing a joint evolution model and conducting alternating training of intelligent agents through a reinforcement learning-based aerospace equipment testing method, the problems of high cost, low coverage, and simulation credibility of traditional testing methods are solved, and efficient and automated aerospace equipment testing and evaluation are achieved.

CN121956581BActive Publication Date: 2026-06-09TIANMUSHAN LABORATORY +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANMUSHAN LABORATORY
Filing Date
2026-03-25
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Traditional aerospace equipment testing methods suffer from high costs, insufficient coverage, low simulation credibility, and a lack of automated adversarial scenario generation, resulting in low testing efficiency and an inability to fully explore performance boundaries and vulnerabilities.

Method used

A joint evolution model of aerospace equipment and environment is established using a reinforcement learning-based approach. The difference between data and reality is calibrated through a residual network, a high-fidelity test environment is constructed, and extreme fault conditions are automatically generated through iterative training of scenario-generating agents and object-response agents to achieve closed-loop testing and evaluation.

Benefits of technology

It effectively eliminates the deviation between numerical and actual dynamics, improves the coverage and relevance of test scenarios, and enables precise positioning of the performance vulnerabilities of aerospace equipment and dynamic safety margin assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121956581B_ABST
    Figure CN121956581B_ABST
Patent Text Reader

Abstract

This invention discloses a reinforcement learning-based object-scenario data-real fusion evaluation method for aerospace equipment, belonging to the field of aerospace equipment testing and evaluation technology. The method includes: constructing a joint evolution model containing a dynamic model of the aerospace equipment and a parameterized environmental model; defining coupling operators and a joint state space; using physical measurement data, calculating deviations and updating coupling operator parameters through a residual network; instantiating dual agents for alternating iterative training, where the scenario generation agent aims to minimize the object's safety margin; integrating the control law under test, dynamically adjusting the scenario parameter search range based on the task success rate; calculating the dynamic safety margin, and outputting a quantitative report including extreme scenario parameter distribution and safety margin statistics. This invention solves the problems of traditional testing scenarios relying on manual intervention, poor data-real consistency, and subjective evaluation, achieving full coverage of extreme operating conditions in testing scenarios and quantitative evaluation of safety boundaries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of aerospace equipment testing and evaluation technology, specifically to a reinforcement learning-based method for evaluating aerospace equipment object-scenario data-real fusion. Background Technology

[0002] With the increasing intelligence and complexity of aerospace equipment (such as drones and satellites), the need to test their autonomous adaptability, robustness, and mission completion capabilities in extreme environments is becoming increasingly urgent. Traditional testing methods have the following limitations:

[0003] Field testing is costly and has insufficient coverage: Field testing is costly, time-consuming, and carries safety risks. It is difficult to exhaustively cover all extreme, low-probability, or high-risk coupled failure conditions (such as the combination of strong gusts and control surface failure at a specific frequency), resulting in insufficient test coverage.

[0004] Purely digital simulation environments suffer from low credibility: traditional simulation models rely on human experience to construct scenarios, lacking randomness and diversity. More importantly, due to model parameters and simplification assumptions, there are significant differences between the pure simulation model and the physical entity (digital model vs. physical entity), leading to insufficient credibility of the simulation results.

[0005] Lack of automated adversarial scenario generation mechanisms: Existing testing methods mainly rely on one-way, fixed-parameter environment input, which cannot achieve proactive and adaptive attacks on the weaknesses of the tested object. This results in low testing efficiency and an inability to fully explore the performance boundaries and vulnerabilities of the tested object.

[0006] Therefore, there is an urgent need for a new testing method that can solve the problem of data consistency, achieve adaptive generation of test scenarios and object strategies, and have a closed-loop, highly reliable quantification and evaluation system. Summary of the Invention

[0007] To address the aforementioned technical issues of inconsistency between data and reality, lack of adversarial scenarios in scene generation, and incomplete evaluation systems, this invention provides a data-real fusion evaluation method for aerospace equipment objects and scenarios based on reinforcement learning.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] In a first aspect, the present invention provides a data-real fusion evaluation method for aerospace equipment objects and scenarios based on reinforcement learning, comprising the following steps:

[0010] S1. Establish the dynamic model of aerospace equipment and the parameterized model of the environmental scenario, construct a joint state space containing object state vectors and scenario state vectors, and define coupling operators to describe the interaction between object state vectors and scenario state vectors, thereby forming a joint evolution model of aerospace equipment and environmental scenario.

[0011] S2. Obtain the physical measurement data of aerospace equipment, calculate the deviation between the physical measurement data and the output data of the joint evolution model through the residual network, update the parameters of the coupling operator using the deviation, and generate a calibrated data-real fusion test environment.

[0012] S3. Instantiate the scene generation agent and the object response agent; the scene generation agent generates scene parameters with the goal of minimizing the object safety margin, and the object response agent generates control instructions with the goal of maximizing the object safety margin. The two are trained alternately and iteratively in the data-real fusion test environment.

[0013] S4. Connect the control law of the aerospace equipment under test to the data-real fusion test environment, dynamically adjust the parameter search range of the scenario generation agent according to the task success rate of the control law of the aerospace equipment under test, generate test scenario sequence and conduct simulation exercise.

[0014] S5. Record the state trajectory in the simulation exercise, calculate the dynamic safety margin of the state trajectory relative to the safety boundary, and output a test report containing the parameter distribution of the extreme scenario and the statistical values ​​of the safety margin.

[0015] In a second aspect, the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned reinforcement learning-based aerospace equipment object-scene data-real fusion evaluation method.

[0016] Thirdly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned reinforcement learning-based aerospace equipment object-scene data-real fusion evaluation method.

[0017] Beneficial effects:

[0018] 1. This invention uses a residual network correction mechanism to perform online calibration of the nonlinear coupling operators in the joint evolution model with a small amount of physical measurement data, effectively eliminating the dynamic deviation between the digital model and the physical entity, and constructing a high-fidelity digital-physical fusion test model.

[0019] 2. This invention constructs a scenario-object adversarial game framework, which utilizes reinforcement learning agents to actively search for lethal boundaries of the tested objects. It can automatically generate extreme fault conditions caused by the coupling of multiple physical fields, which are difficult to preset in traditional manual testing, and greatly improves the coverage and relevance of test scenarios.

[0020] 3. This invention establishes a dynamic safety margin index, transforming the traditional "pass / fail" binary evaluation into a continuous and quantitative assessment of the safety margin of the flight trajectory within the full envelope, which can accurately locate the performance vulnerabilities of aerospace equipment. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating the data-real fusion evaluation method for aerospace equipment objects and scenarios based on reinforcement learning, as proposed in this invention. Detailed Implementation

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific high-maneuverability flight test cases of fixed-wing unmanned aerial vehicles.

[0024] like Figure 1 As shown, the reinforcement learning-based data-real fusion evaluation method for aerospace equipment objects and scenarios of the present invention includes:

[0025] S1. Establish a joint evolution model of aerospace equipment objects and environmental scenarios: Establish a dynamic model of aerospace equipment and a parameterized scenario model of environmental scenarios, construct a joint state space containing object state vectors and scenario state vectors, and define coupling operators that describe the interaction between the two.

[0026] S2. Parameter correction of the joint evolution model based on measured data: Obtain physical measured data of aerospace equipment, calculate the deviation between the physical measured data and the output data of the joint evolution model through the residual network, update the parameters of the coupling operator using the deviation, and generate a calibrated data-real fusion test environment.

[0027] S3. Generate test scenarios and response strategies based on adversarial mechanisms: Instantiate scenario generation agents and object response agents; The scenario generation agent generates scenario parameters with the goal of minimizing the safety margin, and the object response agent generates control commands with the goal of maximizing the safety margin. The two are trained alternately and iteratively in a data-real fusion test environment.

[0028] S4. Perform dynamic closed-loop test: Connect the control law of the aerospace equipment under test to the data-real fusion test environment, dynamically adjust the parameter search range of the scenario generation agent according to the task success rate of the control law of the aerospace equipment under test, generate test scenario sequence and conduct simulation exercise.

[0029] S5. Generate test evaluation results: Record the state trajectory in the S4 simulation exercise, calculate the dynamic safety margin of the trajectory relative to the safety boundary, and output a test report containing the parameter distribution of the extreme scenario and the statistical values ​​of the safety margin.

[0030] Specifically, in S1, the state evolution equation of the joint evolution model is expressed as:

[0031] ;

[0032] in, Represents object state With scene state The combined state vector; For decoupling dynamic functions; For the learnable parameter matrix The defined nonlinear coupling operator is used to characterize the interaction between the object state and the scene state; The control input vector generated for the object's responsive agent; The term represents random noise, with the superscript T indicating the transpose of the matrix and t representing time.

[0033] Specifically, in S1, the joint state space is a space composed of object state variables and scene state variables. The object state vector includes: three-dimensional position, three-dimensional velocity, Euler angles (roll angle, pitch angle, yaw angle), and three-axis angular velocity; the scene state vector includes: local wind speed vector, atmospheric turbulence intensity coefficient, electromagnetic interference signal-to-noise ratio, and target relative azimuth angle.

[0034] Specifically, in S2, updating the parameters of the coupling operator is achieved by minimizing the distribution of the physical measured data. Distribution of model-generated data This is achieved through the 2-Wasserstein distance between them:

[0035] ;

[0036] in, Indicates the 2-Wasserstein distance; The weight parameters are those of the residual network, which is used to compensate for unmodeled dynamic errors in the joint evolution model; the learnable parameter matrix is... Parameters including coupling operators; through joint optimization and This makes the distribution of data generated by the model approximate the distribution of physically measured data.

[0037] The specific structure of the residual network is as follows:

[0038] The input to the residual network is the joint state vector. The core layer consists of three stacked LSTM (Long Short-Term Memory) layers with 128 hidden layers, used to capture the temporal lag features of the aerodynamic response. The backend is followed by two fully connected layers. The output layer uses the Tanh activation function to restrict the output value to the range [-1, 1], and is further scaled by a learnable scaling matrix. Correction increment mapped to aerodynamic coefficient The network parameters are initialized using orthogonal initialization.

[0039] This invention uses a high-fidelity model of JSBSim as the basic research model. The data flow closed-loop interaction mechanism between the residual network and the JSBSim coupling operator is described below. Since JSBSim is a non-differentiable rigid body dynamics engine, this invention designs an external parameter injection mechanism, specifically including:

[0040] Define coupling operators: In the JSBSim XML configuration file, through... <function>Label defines aerodynamic coefficient Calculation formula: External_Input is the interface for the coupling operator. The basic aerodynamic coefficient.

[0041] Forward propagation data flow: In each simulation step, the residual network calculates the correction increment based on the current state. The incremental correction will be achieved through shared memory. The property tree nodes of JSBSim are written in real time; when the JSBSim engine calculates the dynamic equations for the next frame, it automatically calls the XML... <function>The data of the attribute tree node is read and superimposed on the aerodynamic solution to achieve physical simulation of nonlinear coupling.

[0042] The output of the residual network is the correction increment. It is mainly used to compensate for unmodeled high-order dynamics, time delay effects or sensor characteristics, and is the key to improving model fidelity.

[0043] Based on this, end-to-end joint training is conducted, including:

[0044] Given a set of control command sequences and initial states, run a joint evolution model and residual network in the simulation to generate a simulation trajectory, and then distribute the state of the generated simulation trajectory. Corresponding physical measured trajectory state distribution The comparison is performed, and the 2-Wasserstein distance is calculated as the loss; the weight parameters of the residual network are updated simultaneously using gradient descent (Adam). With learnable parameter matrix .

[0045] Specifically, the alternating iterative training of S3 is performed in discrete time step k, solving the following minimax optimization problem in the form of a zero-sum game:

[0046] ;

[0047] in, The strategy for generating intelligent agents for a scene aims to minimize the safety margin of objects by generating harsh scene parameters; The policy for responding to objects in an intelligent agent aims to maximize the object's safety margin through control instructions; For evolutionary trajectory; This represents the expectation of the evolutionary trajectory; For dynamic safety margin function; Let be the joint state vector at step k; This is a scene policy entropy regularization term used to incentivize the diversity of generated scenes; is the regularization coefficient, with an initial value range of [0.01, 0.2]; K is the total duration of the task.

[0048] Specifically, in S3, the scene parameters generated by the scene-generating agent include wind field vector sequence, atmospheric density deviation, and target maneuver trajectory.

[0049] Specifically, in S4, the dynamic adjustment of the parameter search range for the scene-generating agent includes:

[0050] The difficulty index D of the test scenario is defined as the weighted sum of normalized environmental stress parameters:

[0051] ;

[0052] in, To test the ambient wind speed, This refers to the maximum wind speed. For turbulence intensity, For normalized electromagnetic interference signal-to-noise ratio, The weighting coefficients can take values ​​of (0.5, 0.2, 0.3). This represents the L2 norm.

[0053] Define the mission success rate P of the control law of the aerospace equipment under test. Mission success requires the simultaneous fulfillment of multiple conditions, including a dynamic safety margin throughout the entire process. This means that no safety boundaries were touched; the terminal navigation error was less than 50m; and the flight process did not involve stalling or crashing. The threshold range for mission success rate was set as follows: , These are the lower threshold and the upper threshold, respectively; when At the same time, maintain or narrow the parameter search range of the scene-generating agent; when At that time, the parameter search range of the scene-generating agent is expanded to increase the scene difficulty index D; only when the conditions are met... Sampling test scenario within the area. These represent the minimum and maximum values ​​of the threshold, respectively, with typical values ​​being... , .

[0054] Specifically, in S5, the dynamic safety margin The calculation formula is:

[0055] ;

[0056] in, Let be the joint state vector at time t; The Euclidean distance between the joint state vector and the preset safety boundary at time t; The time decay factor typically ranges from [0.5, 3.0], with a typical value of 1.0, and is used to adjust the weight of the influence of historical states on the current margin. Let be the integral variable, representing a historical moment; For a moment The rate of change of distance, which is negative when the object is close to the safety boundary.

[0057] The safety boundary is defined as the minimum normalized margin under multiphysics constraints. The mathematical expression is:

[0058] ;

[0059] in, Stall angle of attack limit (set to) ); The infinite norm of the normalized control surface deflection (value range [0, 1], where 1 indicates the control surface is fully saturated). This is the stall angle of attack. This formula ensures that the safety boundary value returns to zero or becomes negative as soon as any physical constraint is touched.

[0060] Example:

[0061] This embodiment selects a certain type of fixed-wing UAV as a specific example of the aerospace equipment. In the following text, "UAV" and "test object" refer to the aerospace equipment. This model has a wingspan of 4.2m and a maximum takeoff weight of 35kg. Due to its large aspect ratio, it is prone to flexible wing deformation under strong airflow disturbances, making it suitable for verifying the data-real fusion and coupled evolution effect of this invention. A rigid body model of the UAV is constructed using the C++-based six-degree-of-freedom flight dynamics calculation engine JSBSim, with the simulation frequency set to 60Hz. The Python PyTorch deep learning framework and Gymnasium reinforcement learning interface environment are used. Parallel evolution training is performed using an NVIDIA RTX 4090 GPU. During the model verification and data acquisition phases, semi-physical communication is conducted with the physical flight control computer via UDP (User Datagram Protocol). For each step of the simulation engine's derivation (1 / 60th of a second), a status packet is sent and the engine suspends to wait. Upon receiving the packet, the flight control computer calculates the control law and immediately sends back a command packet. Upon receiving the command, the simulation engine unblocks and executes the next dynamics calculation, thus strictly ensuring the alignment of the data-real time axis. To address the UDP packet loss characteristic, a 20ms timeout threshold is set. If the simulation terminal does not receive flight control commands within the timeout period, the commands from the previous frame are maintained unchanged, and the frame loss event is recorded; if more than 5 consecutive frames are lost, the communication link is reset.

[0062] The reinforcement learning-based evaluation method for aerospace equipment objects and scenarios, provided in this embodiment, includes the following steps:

[0063] S1: Construction of the Joint Evolution Model. In JSBSim, in addition to the standard flight dynamics equations, this embodiment introduces a parameterized Dryden atmospheric turbulence model as the scenario model. The aerodynamic data table of the UAV under test is loaded using the JSBSim engine. This data table includes lift coefficient, drag coefficient, side force coefficient, and triaxial moment coefficient.

[0064] The joint state vector Z is defined as a 28-dimensional vector, which includes altitude, latitude and longitude, three-axis velocity, three-axis attitude angle, three-axis angular velocity (12-dimensional object state), as well as three-axis wind speed component, turbulence scale factor, gust intensity (5-dimensional scene state), etc.

[0065] S2: Data-Real Fusion Correction. To ensure the physical realism of the simulation environment, telemetry data collected by the UAV in previous real flight tests is used for model correction. A 3-layer LSTM (Long Short-Term Memory) network is constructed as the residual network. It receives the original output of the simulation model under the same control commands and the current wind speed state. Each layer of the LSTM contains 128 neurons to capture the time-series lag features of the aerodynamic response. Finally, through a fully connected layer, the LSTM outputs the state correction.

[0066] In the JSBSim XML configuration file, via <function>The tag defines a coupling operator. This coupling operator models the effect of wing bending and torsional deformation under strong winds on aerodynamic coefficients as a function of the scene state. The parameters of the coupling operator are updated by minimizing the distribution of physically measured data. Distribution of model-generated data It is achieved by the 2-Wasserstein distance between them.

[0067] The simulation model was run under the same initial state and control sequence as the flight test log to generate a simulated trajectory. After training and convergence, the corrected model was validated under a strong crosswind of 4 m / s. The test results showed that the UAV's roll angle response error decreased from 15% before correction to less than 3%. This indicates that the environment has extremely high physical fidelity, constituting a high-fidelity data-real fusion environment, and the subsequently generated test scenarios are physically reliable.

[0068] S3: Alternating Iterative Training. In the aforementioned high-fidelity real-data fusion environment, "red and blue team" intelligent agents are instantiated for adversarial training, achieving automated generation of test scenarios and simultaneous improvement of the test object's response capabilities.

[0069] The blue team represents the object response agent, used to symbolize the flight control system of the test object. The object response agent is trained using the PPO (Proximal Policy Optimization) algorithm, which uses its pruning objective function to limit the policy update step size and prevent policy collapse during training.

[0070] The input space of the object-response agent is the object state part of the joint state vector Z. and navigation error The output of the object-responsive agent is a four-dimensional continuous control command. , These represent the elevator, aileron, rudder, and throttle, respectively.

[0071] The policy network (Actor) of PPO consists of three fully connected layers (256 neurons per layer, with Tanh activation function); the value network (Critic) has the same structure as the Actor and is used to evaluate the value of the current state.

[0072] The overall optimization objective of the object-response agent (blue side) is to maximize the following joint loss function. :

[0073] ;

[0074] in, To cut the proxy objective function, the parameters are truncated. (Value 0.2) Limits the ratio of the old and new strategies to prevent the update step size from being too large; This represents the empirical expectation operator on a finite batch of sampled data; The mean squared error loss of the value network. The value loss coefficient is set to 0.5. For the policy entropy term, is the entropy regularization coefficient, set to 0.01 to maintain exploratory behavior.

[0075] Advantage function calculation: The advantage value is calculated using generalized advantage estimation (GAE). Set discount factor GAE smoothing parameters .

[0076] The Adam optimizer is used, with a learning rate set to... Training employs an On-Policy mode, executing [the following] after collecting trajectory data for every T=2048 steps. Mini-batch stochastic gradient descent updates every epoch.

[0077] Single-step reward function of object response agent Designed as a multi-objective weighted approach, it aims to balance mission performance, flight safety, and control energy consumption.

[0078] ;

[0079] in , , These are the weighting coefficients for task tracking rewards, safety margin rewards, and control smoothness rewards, with typical weight values ​​being [values ​​to be filled in]. ; As a crash penalty, a large negative value is applied when the drone hits the ground or breaks apart; The reward is for mission tracking and is used to guide the drone to maintain a preset cruise state.

[0080]

[0081] in, These are the current speed and altitude, respectively. For target values ​​(such as 30m / s and 500m).

[0082] It is a safety margin reward, which will dynamically adjust the safety margin. The change in the value of the value is used as a reward to incentivize the agent to actively move away from the lethal boundary:

[0083] ;

[0084] in, The penalty amplification factor (set to 10) is used to impose a severe penalty when a drone approaches or touches the safety boundary (i.e., the safety margin is negative).

[0085] To prevent high-frequency oscillations from affecting flight, a penalty is applied to the rate of change of control commands. :

[0086] ;

[0087] Let be the four-dimensional control command vector at the current time t. This is the control command vector from the previous time step t-1;

[0088] When the simulation termination condition is triggered, such as a crash, a one-time crash penalty is imposed. (-1000).

[0089] The scene-generating agent, acting as the red team, generates various complex environments to achieve collaborative adversarial interaction and capability evolution with the object-response agent (blue team). The scene-generating agent employs the Flexible Actor-Commentator (SAC) algorithm.

[0090] ;

[0091] in, This indicates the scene generation strategy. Evolutionary trajectory generated by sampling The expectation operator on; The reward function is set to the negative of the blue team's dynamic safety margin to incentivize it to generate the most destructive scenarios. Discount factor; This is a temperature coefficient used to automatically adjust the exploration level. This is the strategy entropy term. Introducing this maximum entropy term can explicitly encourage the red side to explore unknown attack patterns, preventing it from falling into a single attack trap, thereby covering a wider range of scenarios.

[0092] The training process is as follows: The red team adopts an off-policy training mode. In each round of competition, the red team samples and generates a set of scene parameters (such as wind speed vector and turbulence intensity) based on the currently observed state of the blue team and injects them into the simulation environment; after the blue team's exercise ends, the tuples are... The data is stored in the experience replay pool, and parameters are updated by randomly sampling from the pool. Here, s represents the state observed by the red team at the current moment; a represents the action generated by the red team based on the current strategy. This refers to the state transitioned to at the next moment after the action is performed.

[0093] To prevent one side from becoming too dominant and preventing the other from learning, a course co-evolution strategy is adopted.

[0094] Phase 1 (Blue Team Pre-training): Freeze the Red Team's parameters and fix the scenario as a standard atmosphere. Train the Blue Team's responsive agent for 1e6 steps in a simple deterministic environment to enable it to master basic flight capabilities.

[0095] Phase Two (Adversarial Iteration): (1) In the inner loop, the red team's strategy remains unchanged, and the red team generates a batch of test scenarios based on the current strategy. The blue team updates its parameters in these scenarios using the PPO algorithm until the success rate of the task tends to stabilize due to the increase in environmental difficulty. (2) In the outer loop, the blue team's strategy remains unchanged. The red team updates its parameters based on the blue team's current weak points (i.e., the state area where the blue team's reward is the lowest) using the SAC algorithm, generating a more destructive scenario parameter distribution.

[0096] Training ends when the difficulty of the scenario generated by the red team reaches the physical limit (such as a wind speed of 15 m / s) and the blue team can still maintain a survival rate of over 90%, or when the entropy values ​​of both sides' strategies converge.

[0097] S4: Perform dynamic closed-loop testing. Integrate the actual control law to be verified into the high-fidelity real-data fusion environment evolved from S3, and conduct fully automated stress testing. Compile the UAV's onboard flight control code into a dynamic link library (.so or .dll file) using C++. Develop an adapter using Python's CTypes interface to construct a "virtual-entity" mapping interface, replacing the output port of the Blue Party object responding to the intelligent agent in S3 with the function call interface of this dynamic link library.

[0098] Therefore, the joint state vector Z of the data-real environment is directly input to the real control law, and the control surface commands calculated by the real control law are fed back to the simulation environment, realizing a closed-loop testing architecture for data-real integration. In each round of testing, the red team (the scenario-generating agent) directly invokes the attack strategy learned in the S3 phase adversarial training within the currently allowed parameter search range, actively generating the scenario sequence most likely to cause the control law under test to fail (e.g., a sudden vertical gust of wind is injected when the UAV is making a high-maneuver maneuver). The system automatically executes 5000 accelerated simulations and records the joint state data of each frame in real time.

[0099] S5: Generate test evaluation results. Perform quantitative analysis on massive amounts of simulation data and output visualized evaluation reports.

[0100] Control barrier functions (CBFs) are often used for formal verification of system safety. Therefore, based on CBF theory, the dynamic safety margin throughout the entire process is calculated. Based on the above calculation results, the system automatically generates a test report containing the following content:

[0101] Plotting the parameter distribution of the extreme scenario: Plotting using kernel density estimation (KDE) leads to control law failure. The distribution of scene parameters intuitively displays the lethal boundary of the system, that is, the physical limit region that the control law cannot cope with.

[0102] Draw a safety margin topology map: Construct a flight envelope view with Mach number and altitude as coordinate axes, and use color depth to map the average dynamic safety margin under this condition to identify weak links within the flight envelope.

[0103] Typical fault reproduction data: Extract the top N test cases with the lowest dynamic safety margin, record their complete six-degree-of-freedom state trajectories and control command sequences, and provide accurate digital evidence for subsequent fault backtracking and reproduction.

[0104] Key Failure Mode Analysis: Outputs the attribution analysis results of the dominant physical factors that lead to a sharp decline in dynamic safety margin.

[0105] Based on the identified failure modes and sensitivity analysis results, the system automatically recommends the direction of adjustment for control parameters (such as PID gain) and generates a complete improvement proposal.

[0106] The control barrier function (CBF) and kernel density estimation (KDE) are both existing technologies and will not be elaborated upon here.

[0107] The above implementation methods demonstrate that, through the detailed implementation steps described above, the present invention not only achieves automated generation of test scenarios, but also ensures the physical credibility of the generated content and the scientific nature of the evaluation results through a rigorous mathematical model. It effectively solves the problems of single test scenario construction, asynchronous data and reality, and subjective evaluation in the prior art, and has significant engineering application value.

[0108] In a second aspect, the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned reinforcement learning-based aerospace equipment object-scene data-real fusion evaluation method.

[0109] Thirdly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned reinforcement learning-based aerospace equipment object-scene data-real fusion evaluation method.

[0110] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0111] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0112] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0113] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0114] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.

[0115] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.< / function> < / function> < / function>

Claims

1. A data-real fusion evaluation method for aerospace equipment objects and scenarios based on reinforcement learning, characterized in that, Includes the following steps: S1. Establish the dynamic model of aerospace equipment and the parameterized model of the environmental scenario, construct a joint state space containing object state vectors and scenario state vectors, and define coupling operators to describe the interaction between object state vectors and scenario state vectors, thereby forming a joint evolution model of aerospace equipment and environmental scenario. S2. Obtain the physical measurement data of aerospace equipment, calculate the deviation between the physical measurement data and the output data of the joint evolution model through the residual network, update the parameters of the coupling operator using the deviation, and generate a calibrated data-real fusion test environment. S3. Instantiate the scene generation agent and the object response agent; the scene generation agent generates scene parameters with the goal of minimizing the object safety margin, and the object response agent generates control instructions with the goal of maximizing the object safety margin. The two are trained alternately and iteratively in the data-real fusion test environment. S4. Connect the control law of the aerospace equipment under test to the data-real fusion test environment, dynamically adjust the parameter search range of the scenario generation agent according to the task success rate of the control law of the aerospace equipment under test, generate test scenario sequence and conduct simulation exercise. S5. Record the state trajectory in the simulation exercise, calculate the dynamic safety margin of the state trajectory relative to the safety boundary, and output a test report containing the parameter distribution of the extreme scenario and the statistical values ​​of the safety margin.

2. The reinforcement learning-based object-scene data-real fusion evaluation method for aerospace equipment as described in claim 1, characterized in that, The state evolution equation of the joint evolution model is expressed as: ; in, Represents object state With scene state The combined state vector; For decoupling dynamic functions; For the learnable parameter matrix The defined nonlinear coupling operator is used to characterize the interaction between the object state and the scene state; The control input vector generated for the object's responsive agent; The term represents random noise, with the superscript T indicating the transpose of the matrix and t representing consecutive time intervals.

3. The reinforcement learning-based object-scene data-real fusion evaluation method for aerospace equipment as described in claim 1, characterized in that, In S1, the object state vector includes three-dimensional position, three-dimensional velocity, Euler angles, and three-axis angular velocity; the scene state vector includes local wind speed vector, atmospheric turbulence intensity coefficient, electromagnetic interference signal-to-noise ratio, and target relative azimuth.

4. The reinforcement learning-based object-scene data-real fusion evaluation method for aerospace equipment as described in claim 2, characterized in that, In S2, the parameters of the coupling operator are updated by minimizing the distribution of the physical measured data. Distribution of model-generated data This is achieved through the 2-Wasserstein distance between them: ; in, Indicates the 2-Wasserstein distance; The weight parameters are those of the residual network, which is used to compensate for unmodeled dynamic errors in the joint evolution model. The learnable parameter matrix contains the parameters of the coupling operators; the weight parameters are jointly optimized. With learnable parameter matrix This makes the model generate data distribution Approximating the distribution of physical measured data ;2-Wasserstein distance is the square bulldozer distance.

5. The reinforcement learning-based object-scene data-real fusion evaluation method for aerospace equipment as described in claim 1, characterized in that, S3's alternating iterative training includes solving a minimax optimization problem of the form of a zero-sum game at discrete time step k: ; in, The strategy for generating intelligent agents for a scene is to minimize the safety margin of objects by generating harsh scene parameters; The policy of responding to an object with an agent is used to maximize the object's safety margin through control instructions; For evolutionary trajectory; This represents the expectation of the evolutionary trajectory; For dynamic safety margin; Let be the joint state vector at step k; This is a scene policy entropy regularization term used to incentivize the diversity of generated scenes; is the regularization coefficient; K is the total duration of the task.

6. The data-real fusion evaluation method for aerospace equipment objects and scenarios based on reinforcement learning according to claim 1, characterized in that, In S4, the parameter search range for dynamically adjusting the scene-generating agent includes: Define the difficulty index D of the test scenario and the mission success rate P of the control law of the aerospace equipment under test; set the mission success rate threshold range. , These are the lower threshold and the upper threshold, respectively; when At the same time, maintain or narrow the parameter search range of the scene-generating agent; when At that time, the parameter search range of the scene-generating agent is expanded to increase the scene difficulty index D; only when the conditions are met... Sampling test scenario within the area.

7. The reinforcement learning-based object-scene data-real fusion evaluation method for aerospace equipment as described in claim 1, characterized in that, In S5, the formula for calculating the dynamic safety margin is: ; in, For dynamic safety margin, Let be the joint state vector at time t; The Euclidean distance between the joint state vector and the preset safety boundary at consecutive time t; This is a time decay factor used to adjust the weight of the influence of historical states on the current margin; Let be the integral variable, representing a historical moment; For a historic moment The rate of change of distance when the object approaches the safety boundary, at historical moments. The value is negative.

8. The reinforcement learning-based object-scene data-real fusion evaluation method for aerospace equipment according to claim 1, characterized in that, In S3, the scene parameters generated by the scene generation agent include wind field vector sequence, atmospheric density deviation, and target maneuver trajectory.

9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the reinforcement learning-based aerospace equipment object-scene data-real fusion evaluation method as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, enable the processor to implement the reinforcement learning-based aerospace equipment object-scene data-real fusion evaluation method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • MBSE spaceflight equipment model-based margin analysis method

    CN119885602A

  • Mobile terminal automatic measurement method and system based on YOLO key point detection and AR platform

    CN121353372A