An aircraft envelope testing method, device, storage medium and electronic device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BOYUAN ZHITONG TECHNOLOGY CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-08-07
AI Technical Summary
然而,该方式存在显著局限:首先,测试过程本身具有较高风险,尤其在接近失速、大迎角、大过载等临界区域时,易引发飞行事故;其次,试飞成本高昂,可执行架次有限,难以对包线进行充分、稠密的探索;此外,测试结果受试飞员主观经验、心理状态及生理耐受度影响,存在人为不确定性
[0016]This technical solution constructs a complete closed-loop system capable of autonomously, safely, and efficiently exploring and verifying flight envelopes by deploying a reinforcement learning agent trained with a reward function as the objective in a flight simulation environment. By transforming the test pilot's experience and decision-making ability into an iteratively optimized artificial intelligence agent, it can tirelessly execute envelope testing tasks, including those under extreme conditions, in a virtual environment, fundamentally avoiding the high risks and costs of human test flights. Guided by the reward function, the agent's exploration behavior is not random or conservative, but precisely shaped to actively and steadily approach various performance boundaries while ensuring flight safety. This allows it to systematically discover critical states located at the edge of the flight envelope that might be overlooked by traditional preset test procedures. Furthermore, by synchronously recording the test process and intelligently analyzing the simulation logs based on design limit parameters and gradient analysis, it can automatically and objectively identify candidate points for the flight envelope boundaries and control saturation regions, and accurately locate performance critical points representing abrupt changes in handling and stability characteristics, improving the accuracy and efficiency of boundary identification. Ultimately, the identification results of the envelope boundary and critical state are fed back to optimize the agent's reward function, enabling the agent to continuously evolve in iterations and significantly improving the automation and intelligence level of the entire envelope identification process.
Smart Images

Figure CN121778183B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of aircraft flight testing and intelligent control technology, and particularly relates to an aircraft envelope testing method, apparatus, storage medium and electronic equipment. Background Technology
[0002] Aircraft envelope testing refers to the process of determining the safe flight boundaries of an aircraft under various state parameters such as speed, altitude, overload, and attitude through experiments. It is a key step in verifying its flight performance and safety.
[0003] Traditional methods primarily rely on human test pilots to gradually approach various performance limits during actual flights to identify boundary states such as stall, flutter, and control failure. However, this approach has significant limitations: First, the testing process itself is highly risky, especially when approaching critical regions such as stall, high angle of attack, and high G-forces, which can easily lead to flight accidents; second, test flights are costly, the number of sorties that can be performed is limited, and it is difficult to conduct a thorough and dense exploration of the flight envelope; in addition, test results are affected by the test pilot's subjective experience, psychological state, and physiological tolerance, resulting in human uncertainty. Summary of the Invention
[0004] This application provides a method, apparatus, medium, and electronic equipment for testing the flight envelope of an aircraft. It can autonomously discover flight performance boundaries and accurately identify critical states. At the same time, it can achieve system self-evolution through iterative optimization, thereby reducing test flight risks and costs.
[0005] According to a first aspect of this application, a method for testing the envelope of an aircraft is provided, the method comprising:
[0006] In a flight simulation environment, based on the envelope test target, historical flight states, historical flight environment, and historical control actions, a current observation vector is generated for the agent and the current observation vector is input into the agent; wherein, the agent is pre-trained with the goal of maximizing the cumulative reward defined by the reward function;
[0007] The agent generates a current control action based on the current observation vector and uses the current control action to manipulate the aircraft to perform envelope testing.
[0008] During the envelope test, flight state parameters, flight environment parameters, and control action parameters are collected synchronously to generate a flight simulation log. Based on the flight simulation log and the design limit parameters of the aircraft, the envelope boundary and critical state of the aircraft are determined. The envelope boundary and critical state of the aircraft are used to optimize the reward function of the agent to optimize the agent.
[0009] According to a second aspect of this application, an aircraft envelope testing apparatus is provided, the apparatus comprising:
[0010] The vector generation module is used to generate a current observation vector for the agent in a flight simulation environment based on the envelope test target, historical flight state, historical flight environment, and historical control actions, and input the current observation vector into the agent; wherein, the agent is pre-trained with the goal of maximizing the cumulative reward defined by the reward function;
[0011] The envelope testing module is used to generate a current control action based on the current observation vector through the intelligent agent, and to use the current control action to manipulate the aircraft to perform envelope testing;
[0012] The result determination module is used to synchronously collect flight state parameters, flight environment parameters, and control action parameters during the envelope test to generate a flight simulation log, and to determine the envelope boundary and critical state of the aircraft based on the flight simulation log and the design limit parameters of the aircraft; wherein, the envelope boundary and critical state of the aircraft are used to optimize the reward function of the agent to optimize the agent.
[0013] According to a third aspect of the present invention, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the aircraft envelope testing method as described in embodiments of this application.
[0014] According to a fourth aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aircraft envelope testing method as described in the embodiments of the present application.
[0015] According to a fifth aspect of this application, an embodiment of this application provides a computer program product, including a computer program that, when executed by a processor, implements the aircraft envelope testing method as described in the embodiment of this application.
[0016] This technical solution constructs a complete closed-loop system capable of autonomously, safely, and efficiently exploring and verifying flight envelopes by deploying a reinforcement learning agent trained with a reward function as the objective in a flight simulation environment. By transforming the test pilot's experience and decision-making ability into an iteratively optimized artificial intelligence agent, it can tirelessly execute envelope testing tasks, including those under extreme conditions, in a virtual environment, fundamentally avoiding the high risks and costs of human test flights. Guided by the reward function, the agent's exploration behavior is not random or conservative, but precisely shaped to actively and steadily approach various performance boundaries while ensuring flight safety. This allows it to systematically discover critical states located at the edge of the flight envelope that might be overlooked by traditional preset test procedures. Furthermore, by synchronously recording the test process and intelligently analyzing the simulation logs based on design limit parameters and gradient analysis, it can automatically and objectively identify candidate points for the flight envelope boundaries and control saturation regions, and accurately locate performance critical points representing abrupt changes in handling and stability characteristics, improving the accuracy and efficiency of boundary identification. Ultimately, the identification results of the envelope boundary and critical state are fed back to optimize the agent's reward function, enabling the agent to continuously evolve in iterations and significantly improving the automation and intelligence level of the entire envelope identification process.
[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of the aircraft envelope testing method provided in Embodiment 1;
[0020] Figure 2 This is a flowchart of the aircraft envelope testing method provided in Example 2;
[0021] Figure 3 This is a schematic diagram of the structure of the aircraft envelope testing device provided in Embodiment 3 of this application;
[0022] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of this application. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0024] It should be noted that the terms "first," "second," "target," and "candidate," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] Example 1
[0026] Figure 1 This is a flowchart of the aircraft envelope testing method provided in Embodiment 1. This embodiment is applicable to various aircraft research and development testing, flight control system verification, and pilot simulation training. The method can be executed by an aircraft envelope testing device, which is implemented in hardware and / or software and can be integrated into the electronic equipment running this system.
[0027] like Figure 1 As shown, the method includes:
[0028] S110. In a flight simulation environment, based on the envelope test target, historical flight status, historical flight environment, and historical control actions, a current observation vector is generated for the agent and the current observation vector is input into the agent; wherein, the agent is pre-trained with the goal of maximizing the cumulative reward defined by the reward function.
[0029] S120. The agent generates a current control action based on the current observation vector and uses the current control action to manipulate the aircraft to perform envelope testing.
[0030] S130. During the envelope test, flight state parameters, flight environment parameters, and control action parameters are collected synchronously to generate a flight simulation log. Based on the flight simulation log and the design limit parameters of the aircraft, the envelope boundary and critical state of the aircraft are determined. The envelope boundary and critical state of the aircraft are used to optimize the reward function of the agent to optimize the agent.
[0031] The flight simulation environment refers to a virtual digital space constructed through computer programs that simulates the dynamic characteristics of an aircraft and the external atmospheric environment. The envelope test target refers to the specific performance boundaries that this test mission intends to explore, such as maximum level flight speed or maximum operational overload. Historical flight state refers to the set of physical states of the aircraft at the previous simulation moment, including position, velocity, attitude, and angular velocity. Historical flight environment refers to the atmospheric conditions at the previous moment, including disturbance information such as wind speed and turbulence. Historical control actions refer to the control commands issued by the agent at the previous moment and executed by the flight control system, such as control surface deflection and throttle position. Historical flight state, historical flight environment, and historical control actions constitute historical interactive information.
[0032] An agent is a trained reinforcement learning model, similar in role to a virtual test pilot making autonomous decisions. The current observation vector, formed by fusing and processing the aforementioned historical information, represents the agent's perception of the complete environment and its own state at the current moment, serving as the input for the agent's decisions. The agent is pre-trained with the goal of maximizing the cumulative reward defined by the reward function. The reward function is a pre-defined mathematical function used to score each state-action pair of the agent, quantifying the quality of its behavior. The cumulative reward refers to the discounted sum of all single-step rewards obtained by the agent in a complete task; the ultimate goal of agent training is to maximize this expected cumulative reward.
[0033] First, based on the test objective and historical interaction information, a current observation vector is constructed for the agent to comprehensively reflect the current situation. This ensures that the agent's decisions have the necessary contextual information. Next, the policy network within the agent calculates the current control action based on this current observation vector, outputting the current control action—that is, the continuous control commands to be executed at this moment, such as joystick inputs. The agent then manipulates the aircraft through a simulated flight control interface. Here, the aircraft manipulated by the agent is a complete virtual aircraft entity operating in a flight simulation environment, thus actively conducting envelope exploration tests. This step transforms the static policy model into dynamic, interactive test behavior.
[0034] During test execution, all key data is simultaneously recorded to generate a flight simulation log, which is a complete sequence of parameters for flight states, flight environment, and control actions, recorded in timestamps. Subsequently, this flight simulation log is used in conjunction with the aircraft's design limit parameters for offline analysis to determine the actual envelope boundaries and critical states. The design limit parameters are the theoretical limits of aircraft performance determined by aerodynamic, structural, and other factors. The envelope boundaries refer to the limits of state parameters for safe and controllable flight discovered through testing. The critical state refers to the transition point where the aircraft's handling and stability characteristics undergo a nonlinear abrupt change. This operation is necessary because quantitative boundaries and critical points cannot be directly and accurately obtained through simulation interaction alone; specialized analysis of the massive amounts of data generated is essential. Finally, the identified boundary and critical state information is fed back to optimize the design of the reward function, forming a closed loop from testing to analysis to optimization. This allows the agent to explore performance boundaries more accurately and safely in subsequent training and testing, enabling continuous iteration and capability improvement of the autonomous testing system.
[0035] This technical solution constructs a complete closed-loop system capable of autonomously, safely, and efficiently exploring and verifying flight envelopes by deploying a reinforcement learning agent trained with a reward function as the objective in a flight simulation environment. By transforming the test pilot's experience and decision-making ability into an iteratively optimized artificial intelligence agent, it can tirelessly execute envelope testing tasks, including those under extreme conditions, in a virtual environment, fundamentally avoiding the high risks and costs of human test flights. Guided by the reward function, the agent's exploration behavior is not random or conservative, but precisely shaped to actively and steadily approach various performance boundaries while ensuring flight safety. This allows it to systematically discover critical states located at the edge of the flight envelope that might be overlooked by traditional preset test procedures. Furthermore, by synchronously recording the test process and intelligently analyzing the simulation logs based on design limit parameters and gradient analysis, it can automatically and objectively identify candidate points for the flight envelope boundaries and control saturation regions, and accurately locate performance critical points representing abrupt changes in handling and stability characteristics, improving the accuracy and efficiency of boundary identification. Ultimately, the identification results of the envelope boundary and critical state are fed back to optimize the agent's reward function, enabling the agent to continuously evolve in iterations and significantly improving the automation and intelligence level of the entire envelope identification process.
[0036] In an optional embodiment, an aircraft dynamics model is deployed in the flight simulation environment. The aircraft dynamics model encapsulates a flight control system interface, an aerodynamic model, and a propulsion system model. The method further includes: determining the actual control surface commands that the aircraft can execute based on the current control action and control surface dynamic characteristics through the flight control system interface; determining the current aerodynamic parameters based on the current flight environment and the historical flight state, and inputting the current aerodynamic parameters into the aerodynamic model to determine the current aerodynamic force and aerodynamic torque through the aerodynamic model; determining the current thrust based on the current control action through the propulsion system model; substituting the current thrust and the current aerodynamic force and aerodynamic torque into the rigid body six-degree-of-freedom motion equations of the aircraft, and performing numerical integration using the historical flight state as the initial condition to solve for the next flight state.
[0037] The aircraft dynamics model is a core computational module integrating multiple subsystems. It calculates the aircraft's motion under forces in real time based on physical laws. Key components encapsulated within this model include: the flight control system interface, the aerodynamic model, and the propulsion system model. The flight control system interface is the software module responsible for receiving and processing control commands and considering the physical limitations of the actuators. The aerodynamic model is a mathematical model used to calculate the aerodynamic forces and torques acting on the aircraft, based on data obtained beforehand through computational fluid dynamics simulations or wind tunnel tests. The propulsion system model is a mathematical model used to simulate engine operating characteristics and calculate thrust.
[0038] First, the actual control surface commands are determined through the flight control system interface based on the current control actions and the dynamic characteristics of the control surfaces. These actual control surface commands refer to the rate and travel limits of the control surface deflection. Next, the current aerodynamic parameters are determined based on the current flight environment and historical flight conditions, and then input into the aerodynamic model to calculate the precise current aerodynamic forces and moments.
[0039] The current flight environment refers to the real-time simulated atmospheric conditions such as wind speed and turbulence. The historical flight state refers to the aircraft's speed and attitude at the previous simulation point. The current aerodynamic parameters mainly refer to key aerodynamic states such as angle of attack, sideslip angle, and Mach number. Optionally, the aerodynamic model is a multi-dimensional lookup table-based aerodynamic model. Current aerodynamic forces and moments are obtained through table lookup and interpolation calculations.
[0040] Simultaneously, the propulsion system model calculates the current thrust based on the throttle command in the current control action, providing another major force. Finally, the current thrust, current aerodynamic force and torque, as well as gravity and all other forces and torques, are substituted into the rigid body six-degree-of-freedom equations of motion of the aircraft. Using the historical flight state as initial conditions, numerical integration is performed to calculate the next flight state of the aircraft, i.e., the predicted flight state of the aircraft one small time step in the future. The rigid body six-degree-of-freedom equations of motion are a set of differential equations describing the translational and rotational laws of a rigid body in three-dimensional space.
[0041] The aforementioned technical solution, through modular integration of the flight control interface, aerodynamic model, and propulsion system, and based on rigorous physical laws for numerical solutions, constructs a high-fidelity, real-time response simulation kernel. Specifically, the flight control system interface translates the agent's control commands into actual control surface commands that conform to the physical limits of the actuators, ensuring the realism of the control link; it accurately calculates aerodynamic parameters based on the real-time environment and historical states, and determines the current aerodynamic torque through a high-fidelity aerodynamic model, making the aerodynamic response precise and reliable; the propulsion system model synchronously provides thrust input. Finally, all forces and torques are substituted into the rigid body's six-degree-of-freedom motion equations for numerical integration, using the previous moment's state as initial values to calculate the next flight state. This complete closed loop ensures that the simulation environment's response to any control action of the agent strictly follows physical laws, thus providing a stable, realistic, and reliable interactive environment for reinforcement learning training. This is the fundamental basis for the agent to learn effective strategies and safely explore performance boundaries, and also generates a high-quality, high-reliability physical state sequence for subsequent envelope data analysis.
[0042] In an optional embodiment, the method further includes: during envelope testing, monitoring the flight status of the aircraft by means of a flight control law deployed in the flight simulation environment; if a key flight parameter in the next flight state enters the warning margin range defined by a preset critical value, triggering a protection action and recording the flight status of the aircraft at the time of triggering.
[0043] The envelope testing process refers to the complete activity cycle of exploring the performance boundaries of an aircraft by utilizing an intelligent agent to interact with the simulation environment. The flight control law deployed in the flight simulation environment is a pre-defined set of automated control algorithms used to stabilize and control the attitude and trajectory of the aircraft; its functionality is extended for safety monitoring. Optionally, the flight control law parameters include angle-of-attack limiters, overload limiters, and attitude control gains.
[0044] Flight status refers to the set of all physical state parameters of an aircraft at a given moment. Monitoring it refers to the continuous and automated process by which the flight control law checks and analyzes the flight status data.
[0045] The next flight state of an aircraft is the object of monitoring and judgment. The next flight state is the predicted state that will come after the current control action is taken, calculated by the dynamic model, rather than the current or historical state.
[0046] Critical flight parameters are the basis for judgment. By determining whether critical flight parameters have entered the warning margin range defined by preset critical values, it can be determined whether the aircraft has exceeded the safety threshold. Critical flight parameters refer to core state quantities directly related to flight safety, such as angle of attack, airspeed, and overload. Preset near-term values refer to safety thresholds determined based on design limits. The warning margin range refers to a buffer zone pre-defined below the safety threshold.
[0047] If critical flight parameters in the next flight state enter the warning margin range, it indicates that the aircraft is approaching but has not yet breached the safety boundary. At this point, a protective action is triggered. Its purpose is to automatically correct the situation before the physical state becomes truly out of control, thereby stabilizing the aircraft within a safe area. The protective action refers to control commands automatically executed by the flight control law to remove the aircraft from a dangerous state, such as forced stick pushes and control surface limitations. Optionally, the protective mechanisms corresponding to the protective actions include angle-of-attack limiters, overload limiters, pitch / roll angle limiters, and control surface saturation protection.
[0048] While triggering the protection action, the flight status at the time of triggering is recorded, that is, a complete state snapshot at the moment of triggering protection is saved, which provides valuable accident symptoms or data reference near the boundary for subsequent analysis.
[0049] The aforementioned technical solution constructs a recoverable, intelligent safety defense by proactively monitoring and triggering protection based on predictions of the next flight state during simulation. By shifting the timing of safety intervention from passive remediation after exceeding limits to active intervention at the point of exceeding limits, it ensures the continuity and robustness of the testing process, avoids test interruptions due to unrecoverable states, and enables continuous long-term automated intensive testing. Simultaneously, by recording the complete flight state at the time of protection triggering, it automatically captures a large number of valuable critical boundary state samples, providing a high-quality data source for accurate analysis of performance critical points and quantification envelope boundaries. Ultimately, this mechanism achieves a delicate balance between encouraging exploration and ensuring safety, allowing the agent to more boldly explore performance limits in a monitored environment, thereby significantly improving the overall efficiency, data value, and engineering practicality of the autonomous testing system.
[0050] In an optional embodiment, the aircraft dynamics model is associated with an external disturbance source, which is used to introduce external disturbances into the aircraft dynamics model; the external disturbances include at least: atmospheric turbulence, gusts, and sensor noise; the atmospheric turbulence is used to simulate small-scale random velocity fluctuations in the atmosphere; the gusts are used to simulate instantaneous strong wind disturbances; the sensor noise includes at least Gaussian noise, bias drift, and sampling delay of the speedometer, altimeter, and inertial navigation system.
[0051] The external disturbance source associated with the aircraft dynamics model is a software module specifically designed to generate various types of disturbance signals. External disturbance sources are used to introduce external disturbances into the aircraft dynamics model; these are uncontrolled factors originating from outside the aircraft that affect its flight state or state measurements. Introducing these external disturbances aims to make the flight simulation environment more closely resemble the complexity and uncertainty of the real world, thereby improving the robustness of agents trained on this environment.
[0052] Optionally, external disturbances mainly include: atmospheric turbulence, gusts, and sensor noise. Atmospheric turbulence refers to a mathematical model used to simulate random, small-scale velocity fluctuations in the atmosphere, its function being to reproduce the continuous turbulence disturbances present in real flight. Optionally, the Dryden model is used to generate atmospheric turbulence. Gusts are superimposed on mean wind, turbulence, and wind shear to simulate instantaneous strong wind disturbances for testing the aircraft's response capability to sudden wind disturbances; sensor noise refers to errors superimposed on ideal measurement signals, used to simulate the non-ideal characteristics of real sensors. Sensor noise includes: Gaussian noise, bias drift, and sampling delay of the speedometer, altimeter, and inertial navigation system. Among them, the speedometer is an instrument for measuring airspeed, the altimeter is an instrument for measuring altitude, and the inertial navigation system is an inertial navigation system for measuring attitude and acceleration. Gaussian noise is a random error that follows a normal distribution. Bias drift is a systematic error in which the measurement reference value changes slowly over time. Sampling delay is the time lag between signal measurement and output.
[0053] The aforementioned technical solution, by systematically integrating external disturbance sources such as atmospheric turbulence, gusts, and sensor noise into the dynamic model, significantly improves the realism of the simulation environment and the engineering applicability of the test conclusions. By incorporating the continuous random airflow disturbances and sudden gusts present in real flight into the force environment, the agent is forced to learn robust control under dynamic uncertainty conditions. Simultaneously, the high-fidelity sensor noise model, including Gaussian noise, bias drift, and sampling delay, ensures that the quality of the observation information received by the agent is consistent with real flight, training its ability to cope with incomplete and inaccurate data. Ultimately, the envelope boundaries and critical states obtained based on this environment test inherently contain the safety margin required to cope with real disturbances, thereby greatly enhancing the reliability and direct value of transferring autonomous testing strategies and boundary data to real engineering applications.
[0054] In an optional implementation, determining the envelope boundary and critical state of the aircraft based on the flight simulation log and the aircraft's design limit parameters includes: determining state quantity thresholds and their corresponding warning margin ranges based on the design limit parameters, and detecting the state quantity sequence within a specified time window in the simulation log based on the state quantity thresholds; if any state quantity in the state quantity sequence exceeds its corresponding state quantity threshold or enters the warning margin range, the state quantity is identified as a critical state quantity, and the flight state corresponding to the time of the critical state quantity is marked as a candidate point of the envelope boundary; if the control quantity in the simulation log continuously reaches the physical limit value of its actuator within the specified time window, the control saturation zone of the aircraft is determined; gradient analysis is performed on the critical state quantity, and the performance critical point is determined based on the obtained gradient analysis results; the envelope boundary and critical state of the aircraft are determined based on the candidate points of the envelope boundary, the control saturation zone, and the performance critical point.
[0055] Flight simulation logs refer to the complete data sequence of flight state parameters, flight environment parameters, and control action parameters, recorded synchronously and arranged in chronological order during envelope testing. Aircraft design limit parameters refer to the theoretical performance limits determined by the aerodynamic, structural, and system design of the aircraft, such as maximum permissible angle of attack and maximum overload.
[0056] First, based on the design limit parameters, state quantity thresholds and their corresponding warning margin ranges are determined. The theoretical limits are transformed into quantitative standards that can be used for data comparison, and a buffer warning zone is set for states approaching the limits. The state quantity threshold refers to the absolute safety limit that each key flight state parameter, such as angle of attack or overload, cannot be exceeded. The warning margin range is a pre-defined area below the state quantity threshold, used to warn that the aircraft is approaching its limits.
[0057] Subsequently, the state quantity sequence within a specified time window in the simulation log is detected based on the state quantity threshold. The specified time window refers to a continuous simulation time segment extracted for local trend analysis. The state quantity sequence refers to the data curve of a certain state parameter changing over time within this time window. The purpose of the detection is to systematically scan the entire flight simulation log and locate all abnormal or near-abnormal states. When the detection finds that any state quantity in the state quantity sequence exceeds its corresponding state quantity threshold or enters the warning margin range, two operations are performed: first, the state quantity is identified as a critical state quantity, i.e., the parameter currently constituting a safety risk or potential risk is identified; second, the flight state corresponding to the critical state quantity at that time is marked as a candidate point for the envelope boundary, i.e., a complete state snapshot of the aircraft at the time the risk occurs is recorded as the raw material for subsequent determination of the final boundary, in order to avoid missing any possible boundary points.
[0058] Meanwhile, if the control quantity in the simulation log continuously reaches the physical limit value of its actuator within a specified time window, the control saturation zone of the aircraft is determined. The control quantity refers to the signal issued by the flight control system that instructs the actuators such as control surfaces or throttle to move. The physical limit value of the actuator refers to the inherent action boundary of the hardware, such as the maximum deflection angle of the control surface or the maximum travel of the throttle. The control saturation zone refers to the flight phase where the control quantity is continuously maintained near its limit value. Identifying this zone is crucial because control saturation often means that the aircraft has exhausted all its control capabilities to maintain its state or attempt to recover, which is itself an important indirect indicator of approaching the envelope boundary.
[0059] To further identify the continuous variation characteristics of flight quality, gradient analysis is performed on key state variables. Gradient analysis refers to mathematical operations that involve taking the first or even second derivative of the state variable sequence. The performance critical point is determined based on the gradient analysis results. This critical point is defined as the moment when the gradient curve shows an inflection point, abrupt change, or a significant shift in trend. This point characterizes the transitional state in which the aircraft's handling and stability characteristics, such as stability and control efficiency, undergo a fundamental change. It may occur earlier than the parameters reaching their absolute thresholds and is crucial for understanding the performance evolution within the flight envelope.
[0060] Finally, based on candidate points of the envelope boundary, the control saturation region, and the performance critical point, the envelope boundary and critical state of the aircraft are determined. The reason for integrating these three types of information is that a single threshold exceedance criterion may be too absolute and ignore the dynamic process. Combining control saturation and gradient change characteristics allows for cross-validation and supplementation from three dimensions: state exceedance, control capability depletion, and abrupt changes in dynamic characteristics. This results in a more comprehensive, accurate, and engineering-guided description of the envelope boundary and set of critical states.
[0061] Optionally, the aircraft envelope boundary and critical state generate Vn diagrams, angle-of-attack-lift curves, and critical attitude parameter reports.
[0062] The aforementioned technical solution sets threshold values and warning margin ranges for state variables based on design limit parameters. By detecting state variable sequences, it automatically identifies critical state variables that exceed the thresholds or enter the warning margin range, thereby marking candidate points for the envelope boundary. This ensures that all possible boundary risk points are captured by the system, avoiding omissions. Simultaneously, by monitoring whether control variables continuously reach the physical limits of the actuators, the control saturation zone is determined, providing crucial evidence for boundary identification from the perspective of depleted maneuverability. Furthermore, gradient analysis is performed on critical state variables to determine performance critical points, enabling the system to keenly capture transitional states where fundamental changes in handling and stability characteristics occur. Finally, a comprehensive determination based on candidate points for the envelope boundary, the control saturation zone, and the performance critical point ensures that the final envelope boundary and critical state integrate three core types of evidence: state exceeding limits, maneuverability saturation, and abrupt changes in dynamic characteristics. This significantly improves the objectivity, accuracy, and engineering guidance value of the analysis results, providing insights far exceeding traditional over-limit criteria for aircraft design and control law optimization.
[0063] Example 2
[0064] Figure 2 This is a flowchart of the aircraft envelope testing method provided in Embodiment 2. This embodiment further optimizes the above embodiments. Specifically, the training process of the intelligent agent is further defined.
[0065] like Figure 2 As shown, the method includes:
[0066] S210. The observation vector received by the agent in a single simulation interaction is taken as the current observation vector, and the current observation vector is input into the agent's policy network, so as to output the current control action corresponding to the current observation vector through the policy network.
[0067] In this context, a single simulation interaction refers to a basic interaction unit from receiving observations and making decisions to receiving results from the environment. The current observation vector is the observation vector received by the agent in a single simulation interaction. The current observation vector represents the data set of the current perception state, typically including features such as flight status, environmental information, and historical actions.
[0068] The policy network is one of the core components of an intelligent agent. It is a parameterized deep neural network whose function is to map the input observation vector to a specific control action or action distribution. The current control action refers to the manipulation instruction to be executed, output by the policy network based on the current observation vector.
[0069] S220. The expected reward is estimated by using the agent's evaluation network to evaluate the state-action pair consisting of the current observation vector and the current control action, so as to determine the cumulative reward corresponding to the state-action pair.
[0070] The evaluation network is another core component of the agent. It is also a parameterized deep neural network used to evaluate the long-term value of a given state-action pair. Its output is an expected reward estimate, which is a prediction of the cumulative reward that the state-action pair formed by the current observation vector and the current control action can obtain in the future. Here, the cumulative reward is the discounted sum of all future single-step rewards.
[0071] Optionally, to enhance the agent's ability to model the temporal continuity of flight dynamics, a Transformer sequence modeling module is embedded in the policy network and evaluation network to express multi-step state and action dependencies.
[0072] S230. Calculate the next flight state and the next observation vector corresponding to the current control action, and use the reward function to determine the single-step reward for the current control action based on the next flight state.
[0073] The aircraft dynamics model deployed in the flight simulation environment is invoked to calculate the next flight state corresponding to the current control action. Then, based on this newly calculated ground truth state, the flight simulation environment processes and encapsulates the external disturbance sources associated with the aircraft dynamics model to generate the next observation vector for the agent to perceive in the next step.
[0074] The reward function is a predefined evaluation criterion. It is not based on the agent's intent (the current control action) or its state at the time of decision (the current observation vector), but rather strictly on the objective physical result produced after the action is executed—the next flight state. For example, the reward function checks whether the angle of attack in the next flight state is close to stall, whether the airspeed is too low, or whether the overload is excessive. Based on these specific, objective physical conditions, it outputs a scalar value, the single-step reward. The single-step reward quantifies the quality of the causal relationship between executing the current control action and leading to the next flight state.
[0075] S240. Write the current observation vector, the current control action, the single-step reward, and the next observation vector into the experience replay pool.
[0076] The experience replay pool is a circular data buffer that stores historical experiences. The quadruple consisting of the current observation vector, current control action, single-step reward, and next observation vector represents the interaction data for a single simulation interaction. Writing the interaction data of a single simulation interaction into the experience replay pool allows for the continuous generation of realistic interaction samples for agent training.
[0077] S250. Calculate the parameter gradients of the evaluation network and the policy network based on the sample interaction data obtained from the experience replay pool, and update the network parameters of the evaluation network and the policy network according to the parameter gradients.
[0078] The parameter gradients of the evaluation network and the policy network are calculated based on sample interaction data obtained from the experience replay pool. That is, a batch of sample interaction data is randomly selected from the experience replay pool. Each sample interaction data includes the current observation vector, the current control action, the single-step reward, and the next observation vector.
[0079] For the evaluation network, its parameter gradient is calculated by minimizing the error between its own prediction of the value of the state-action pair and a more stable target value, which is composed of the single-step reward in the sample interaction data and the future value estimated based on the next observation vector.
[0080] For the policy network, its parameter gradients are calculated by optimizing the current control action of its output to maximize the expected reward estimate of the evaluation network for the corresponding state-action pair. This process also encourages the policy to maintain a certain degree of stochasticity to promote exploration. Subsequently, the network parameters of the evaluation network and the policy network are updated based on the parameter gradients. That is, the calculated gradients are applied to the optimizers of their respective networks, such as gradient descent or ascent algorithms, to adjust their internal weights, thereby completing one iteration of optimization. By repeatedly executing this process, the network parameters of the two networks are updated collaboratively, driving the agent's policy to continuously evolve towards maximizing the cumulative reward.
[0081] This application's technical solution constructs a rigorous, experience-replay-based data-driven closed loop, achieving efficient and stable policy optimization. Using a single simulation interaction as the basic unit, the policy network generates the current control action based on the current observation vector, ensuring the directness of the decision. The evaluation network then estimates the expected return for this state-action pair, establishing a foundation for accurately assessing the long-term decision value. Subsequently, the next flight state and the next observation vector corresponding to the current control action are calculated through the simulation environment, and the single-step reward is determined based on the next flight state using the reward function. This allows the reward signal to be directly and objectively associated with the specific state evolution caused by the control action. Crucially, by storing interaction data containing complete causal chains in an experience replay pool and sampling batches of historical data from the pool to calculate the network parameter gradient, not only is data utilization efficiency greatly improved, but the stability of the training process is effectively guaranteed by breaking the temporal correlation between data, avoiding policy oscillations. Ultimately, the evaluation network and policy network perform collaborative iterative updates based on the calculated parameter gradients, enabling the agent to continuously learn from massive simulation experience and gradually converge to an optimal policy that maximizes cumulative rewards while maintaining necessary exploratory activity in complex flight environments. This lays a solid foundation for its eventual deployment in autonomous envelope testing missions.
[0082] In an optional embodiment, the reward function comprises an envelope approximation reward, a flight safety reward, and an accident penalty and stability reward. The envelope approximation reward determines the reward value based on the normalized distance between real-time measurements of one or more key flight parameters and their corresponding preset thresholds. The calculation logic of the envelope approximation reward is configured to: provide a basic reward when the real-time measurement is within a safe range but far from the preset threshold; provide an incremental reward when the real-time measurement is close to but does not exceed the preset threshold; and decrease the reward value when the measurement enters a warning margin range defined by the preset threshold. The flight safety reward is used to determine the reward value based on the aircraft's current airspeed and a preset stall value. The relative magnitude relationship between velocities determines the reward value; the calculation logic of the flight safety reward item is configured to provide a positive reward when the current airspeed is higher than the stall speed, and the reward value decreases monotonically as the current airspeed decreases; the accident penalty and stability reward items include an accident penalty sub-item and a flight quality reward sub-item; the accident penalty sub-item outputs a preset negative constant as a penalty when a collision is detected; the flight quality reward sub-item is used to determine the reward value based on one or more of the aircraft's angular velocity component, angle of attack change rate, and control surface saturation degree, and the calculation logic of the flight quality reward sub-item is configured to provide a positive reward for flight states with small angular velocity components, gradual angle of attack change, and low control surface saturation.
[0083] The reward function is a mathematical function composed of weighted sub-items with explicit engineering semantics. Its core principle lies in precisely shaping and guiding the agent's behavioral strategy through a structured multi-objective reward mechanism to simultaneously meet multiple task requirements such as envelope exploration, flight safety, and handling quality. The reward function consists of envelope approximation reward items, flight safety reward items, and accident penalty and stability reward items.
[0084] The envelope approximation reward term is specifically designed to incentivize the agent to actively and stably explore the flight performance boundaries. It calculates the reward value based on the normalized distance between the real-time measurements of one or more key flight parameters and their corresponding preset thresholds. Its calculation logic is finely configured as follows: a basic reward is provided within a safe range to maintain exploration motivation; an incremental reward is provided when approaching the preset threshold to reinforce boundary exploration behavior; and the reward value is decayed when entering a warning margin range to generate a proximity penalty effect, thereby intelligently and gently preventing the agent from exceeding the safe boundaries. Key flight parameters include state variables directly related to performance limits, such as angle of attack and overload.
[0085] Optional, adopt This indicates that the envelope is approaching the bonus item. Among them, , .in, As a preset critical value, For real-time measurements, For normalization, For margin warning range, These are the coefficients of the corresponding parameters.
[0086] The flight safety reward focuses on ensuring basic flight safety. It determines the reward value based on the relative relationship between the aircraft's current airspeed and the preset stall speed. Its logic is configured to provide a positive reward when the airspeed is higher than the stall speed, and this value decreases monotonically as the airspeed decreases, thus establishing a continuous safety gradient with respect to speed. This incentivizes the agent to actively avoid a stall state. The preset stall speed refers to the minimum speed at which the aircraft can maintain stable flight.
[0087] Optional, adopt This represents a flight safety award. In the formula, For coefficients, Stall speed, This represents the current speed.
[0088] The accident penalty and stability reward are a composite item, comprising an accident penalty sub-item and a flight quality reward sub-item. The accident penalty sub-item outputs a very large preset negative constant as a penalty when a catastrophic event such as a collision is detected. Its function is to set an absolute, untouchable line for the agent's behavior. The flight quality reward sub-item is used to fine-tune the control quality during flight. It determines the reward value based on one or more of the aircraft's angular velocity components, such as roll, pitch, and yaw angular velocities, rate of change of angle of attack, and control surface saturation. Its logic clearly rewards stable flight states with low angular velocities, gradual changes in angle of attack, and low control surface saturation. Control surface saturation refers to the proportion of control surface deflection angles that are close to their physical limits.
[0089] Optionally, the accident penalty and stability bonus items are represented by the following formula:
[0090] ;
[0091] In the formula, The coefficient at which a collision occurs. For coefficients, For angular velocity components, For normalization, For the maximum allowable rate of change of angle of attack, It controls the saturation ratio. It is the corresponding coefficient.
[0092] A single reward objective cannot meet the needs of complex envelope testing tasks. The above-mentioned technical solution, by integrating envelope approximation rewards aimed at encouraging exploration, flight safety rewards aimed at ensuring a baseline, and accident penalty and stability rewards aimed at preventing disasters and encouraging smooth handling, provides the agent with a comprehensive, balanced, and conflict-free learning objective. This allows the agent to automatically weigh performance, safety, and quality during exploration, thereby autonomously evolving an envelope testing strategy that is both bold and stable, efficient and safe.
[0093] Example 3
[0094] Figure 3 This is a schematic diagram of the structure of the aircraft envelope testing device provided in Embodiment 3 of this application. This embodiment can be applied to various situations such as research and development testing of aircraft, flight control system verification, and pilot simulation training. The method can be executed by the aircraft envelope testing device, which is implemented in hardware and / or software and can be integrated into the electronic equipment running this system.
[0095] like Figure 3 As shown, the device may include:
[0096] The vector generation module 310 is used to generate a current observation vector for the agent in a flight simulation environment based on the envelope test target, historical flight state, historical flight environment and historical control actions, and input the current observation vector into the agent; wherein the agent is pre-trained with the goal of maximizing the cumulative reward defined by the reward function;
[0097] The envelope testing module 320 is used to generate a current control action based on the current observation vector through the intelligent agent, and to use the current control action to manipulate the aircraft to perform envelope testing;
[0098] The result determination module 330 is used to synchronously collect the generated flight state parameters, flight environment parameters, and control action parameters during the envelope test to generate a flight simulation log, and to determine the envelope boundary and critical state of the aircraft based on the flight simulation log and the design limit parameters of the aircraft; wherein, the envelope boundary and critical state of the aircraft are used to optimize the reward function of the agent to optimize the agent.
[0099] This technical solution constructs a complete closed-loop system capable of autonomously, safely, and efficiently exploring and verifying flight envelopes by deploying a reinforcement learning agent trained with a reward function as the objective in a flight simulation environment. By transforming the test pilot's experience and decision-making ability into an iteratively optimized artificial intelligence agent, it can tirelessly execute envelope testing tasks, including those under extreme conditions, in a virtual environment, fundamentally avoiding the high risks and costs of human test flights. Guided by the reward function, the agent's exploration behavior is not random or conservative, but precisely shaped to actively and steadily approach various performance boundaries while ensuring flight safety. This allows it to systematically discover critical states located at the edge of the flight envelope that might be overlooked by traditional preset test procedures. Furthermore, by synchronously recording the test process and intelligently analyzing the simulation logs based on design limit parameters and gradient analysis, it can automatically and objectively identify candidate points for the flight envelope boundaries and control saturation regions, and accurately locate performance critical points representing abrupt changes in handling and stability characteristics, improving the accuracy and efficiency of boundary identification. Ultimately, the identification results of the envelope boundary and critical state are fed back to optimize the agent's reward function, enabling the agent to continuously evolve in iterations and significantly improving the automation and intelligence level of the entire envelope identification process.
[0100] Optionally, the flight simulation environment deploys an aircraft dynamics model, which internally encapsulates a flight control system interface, an aerodynamic model, and a propulsion system model. The device further includes: a control surface command determination module, used to determine the actual control surface commands executable by the aircraft based on the current control action and control surface dynamic characteristics through the flight control system interface; an aerodynamic parameter determination module, used to determine the current aerodynamic parameters based on the current flight environment and the historical flight state, and input the current aerodynamic parameters into the aerodynamic model to determine the current aerodynamic force and aerodynamic torque through the aerodynamic model; a current thrust determination module, used to determine the current thrust based on the current control action through the propulsion system model; and a flight state calculation module, used to substitute the current thrust and the current aerodynamic force and aerodynamic torque into the rigid body six-degree-of-freedom motion equations of the aircraft, and perform numerical integration using the historical flight state as initial conditions to calculate the next flight state.
[0101] Optionally, the device further includes: a flight status monitoring module, used to monitor the flight status of the aircraft through a flight control law deployed in the flight simulation environment during envelope testing; and a protection action triggering module, used to trigger a protection action and record the flight status of the aircraft at the time of triggering if a key flight parameter in the next flight status enters a warning margin range defined by a preset threshold.
[0102] Optionally, the aircraft dynamics model is associated with an external disturbance source, which is used to introduce external disturbances into the aircraft dynamics model; the external disturbances include at least: atmospheric turbulence, gusts, and sensor noise; the atmospheric turbulence is used to simulate small-scale random velocity fluctuations in the atmosphere; the gusts are used to simulate instantaneous strong wind disturbances; the sensor noise includes at least Gaussian noise, bias drift, and sampling delay of the speedometer, altimeter, and inertial navigation system.
[0103] Optionally, the agent is trained in the following manner: the observation vector received by the agent in a single simulation interaction is used as the current observation vector, and the current observation vector is input into the agent's policy network to output the current control action corresponding to the current observation vector through the policy network; the expected reward is estimated by the agent's evaluation network based on the state-action pair consisting of the current observation vector and the current control action to determine the cumulative reward corresponding to the state-action pair; the next flight state and the next observation vector corresponding to the current control action are calculated, and the reward function is used to determine the single-step reward for the current control action based on the next flight state; the current observation vector, the current control action, the single-step reward, and the next observation vector are written into the experience replay pool; the parameter gradients of the evaluation network and the policy network are calculated based on the sample interaction data sampled in the experience replay pool, and the network parameters of the evaluation network and the policy network are updated according to the parameter gradients.
[0104] Optionally, the reward function consists of an envelope approximation reward, a flight safety reward, and an accident penalty and stability reward. The envelope approximation reward determines the reward value based on the normalized distance between the real-time measured values of one or more key flight parameters and their corresponding preset thresholds. The calculation logic of the envelope approximation reward is configured as follows: a basic reward is provided when the real-time measured value is within a safe range but far from the preset threshold; an incremental reward is provided when the real-time measured value is close to but does not exceed the preset threshold; and the reward value decays when the measured value enters a warning margin range defined by the preset threshold. The flight safety reward is used to determine the reward value based on the difference between the aircraft's current airspeed and a preset stall speed. The relative magnitude relationship between the values determines the reward value; the calculation logic of the flight safety reward item is configured to provide a positive reward when the current airspeed is higher than the stall speed, and the reward value decreases monotonically as the current airspeed decreases; the accident penalty and stability reward items include an accident penalty sub-item and a flight quality reward sub-item; the accident penalty sub-item outputs a preset negative constant as a penalty when a collision is detected; the flight quality reward sub-item is used to determine the reward value based on one or more of the aircraft's angular velocity component, angle of attack change rate, and control surface saturation degree, and the calculation logic of the flight quality reward sub-item is configured to provide a positive reward for flight states with small angular velocity components, gradual angle of attack change, and low control surface saturation.
[0105] Optionally, the result determination module 330 includes: a threshold determination submodule, used to determine the state quantity threshold and its corresponding warning margin range based on the design limit parameters, and to detect the state quantity sequence within a specified time window in the simulation log based on the state quantity threshold; a candidate point determination submodule, used to determine the state quantity as a critical state quantity if any state quantity in the state quantity sequence exceeds its corresponding state quantity threshold or enters the warning margin range, and to mark the flight state corresponding to the time of the critical state quantity as a candidate point of the envelope boundary; a saturation region determination submodule, used to determine the control saturation region of the aircraft if the control quantity in the simulation log continuously reaches the physical limit value of its actuator within the specified time window; a critical point determination submodule, used to perform gradient analysis on the critical state quantity and determine the performance critical point based on the obtained gradient analysis results; and a result determination submodule, used to determine the envelope boundary and critical state of the aircraft based on the candidate points of the envelope boundary, the control saturation region, and the performance critical point.
[0106] The aircraft envelope testing apparatus provided in the embodiments of the invention can execute the aircraft envelope testing method provided in any embodiment of this application, and has the corresponding performance modules and beneficial effects for executing the aircraft envelope testing method.
[0107] In the technical solution of this application, the user data involved in the aircraft envelope test is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0108] Example 4
[0109] According to embodiments of this application, this application also provides an electronic device, a readable storage medium, and a computer program product.
[0110] Figure 4A schematic diagram of an electronic device 410, which can be implemented using an embodiment, is shown. The electronic device 410 includes at least one processor 411 and a memory, such as a read-only memory (ROM) 412, a random access memory (RAM) 413, etc., communicatively connected to the at least one processor 411. The memory stores computer programs executable by the at least one processor. The processor 411 can perform various appropriate actions and processes based on the computer program stored in the ROM 412 or loaded from storage unit 418 into the RAM 413. The RAM 413 may also store various programs and data required for the operation of the electronic device 410. The processor 411, ROM 412, and RAM 413 are interconnected via a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.
[0111] Multiple components in electronic device 410 are connected to I / O interface 415, including: input unit 416, such as keyboard, mouse, etc.; output unit 417, such as various types of displays, speakers, etc.; storage unit 418, such as disk, optical disk, etc.; and communication unit 419, such as network card, modem, wireless transceiver, etc. Communication unit 419 allows electronic device 410 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0112] Processor 411 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 411 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 411 performs the various methods and processes described above, such as aircraft envelope testing methods.
[0113] In some embodiments, the aircraft envelope testing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 418. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 410 via ROM 412 and / or communication unit 419. When the computer program is loaded into RAM 413 and executed by processor 411, one or more steps of the aircraft envelope testing method described above may be performed. Alternatively, in other embodiments, processor 411 may be configured to perform the aircraft envelope testing method by any other suitable means (e.g., by means of firmware).
[0114] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0115] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable aircraft envelope testing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0116] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0118] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as an aircraft envelope test server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0119] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0120] This application also discloses a computer program product, which includes a computer program that, when executed by a processor, implements the aircraft envelope testing method provided in any embodiment of this application. This program product shares the same inventive concept as the aircraft envelope testing methods disclosed in the embodiments of this application, and therefore will not be described in detail here.
[0121] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0122] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for testing the envelope of an aircraft, characterized in that, The method includes: In a flight simulation environment, based on the envelope test target, historical flight states, historical flight environment, and historical control actions, a current observation vector is generated for the agent and the current observation vector is input into the agent; wherein, the agent is pre-trained with the goal of maximizing the cumulative reward defined by the reward function; The agent generates a current control action based on the current observation vector and uses the current control action to manipulate the aircraft to perform envelope testing. During the envelope test, flight state parameters, flight environment parameters, and control action parameters are collected synchronously to generate a flight simulation log. Based on the flight simulation log and the aircraft's design limit parameters, the aircraft's envelope boundary and critical state are determined. The aircraft's envelope boundary and critical state are used to optimize the agent's reward function to optimize the agent. The agent is trained as follows: the observation vector received by the agent in a single simulation interaction is used as the current observation vector, and the current observation vector is input into the agent's policy network to output the current control action corresponding to the current observation vector; the expected reward is estimated by the agent's evaluation network based on the state-action pair formed by the current observation vector and the current control action to determine the cumulative reward corresponding to the state-action pair; the next flight state and the next observation vector corresponding to the current control action are calculated, and the reward function is used to determine the single-step reward for the current control action based on the next flight state; the current observation vector, the current control action, the single-step reward, and the next observation vector are written into the experience replay pool; the parameter gradients of the evaluation network and the policy network are calculated based on the sample interaction data sampled in the experience replay pool, and the network parameters of the evaluation network and the policy network are updated according to the parameter gradients.
2. The method according to claim 1, characterized in that, The flight simulation environment deploys an aircraft dynamics model, which internally encapsulates a flight control system interface, an aerodynamic model, and a propulsion system model; the method further includes: The flight control system interface determines the actual control surface commands that the aircraft can execute based on the current control action and the dynamic characteristics of the control surfaces. The current aerodynamic parameters are determined based on the current flight environment and the historical flight status, and the current aerodynamic parameters are input into the aerodynamic model to determine the current aerodynamic force and aerodynamic torque through the aerodynamic model; The current thrust is determined by the propulsion system model based on the current control action; Substitute the current thrust, current aerodynamic force, and aerodynamic torque into the rigid body six-degree-of-freedom motion equations of the aircraft, and use the historical flight state as the initial condition to perform numerical integration to solve for the next flight state.
3. The method according to claim 2, characterized in that, The method further includes: During the envelope test, the flight status of the aircraft is monitored by the flight control law deployed in the flight simulation environment; If the critical flight parameters in the next flight state enter the warning margin range defined by the preset threshold, a protection action is triggered and the flight state of the aircraft at the time of triggering is recorded.
4. The method according to claim 2, characterized in that, The aircraft dynamics model is associated with an external disturbance source, which is used to introduce external disturbances into the aircraft dynamics model. The external disturbances include at least: atmospheric turbulence, gusts, and sensor noise. The atmospheric turbulence is used to simulate small-scale random velocity fluctuations in the atmosphere. The gusts are used to simulate instantaneous strong wind disturbances. The sensor noise includes at least Gaussian noise, bias drift, and sampling delay from the speedometer, altimeter, and inertial navigation system.
5. The method according to claim 1, characterized in that, The reward function consists of envelope approximation reward, flight safety reward, and accident penalty and stability reward. The envelope approximation reward is used to determine the reward value based on the normalized distance between the real-time measurement value of one or more key flight parameters and their corresponding preset threshold value. The calculation logic of the envelope approximation reward is configured as follows: a basic reward is provided when the real-time measurement value is within a safe range but far from the preset threshold value; an incremental reward is provided when the real-time measurement value is close to but does not exceed the preset threshold value; and the reward value decays when the measurement value enters the warning margin range defined by the preset threshold value. The flight safety reward item is used to determine the reward value based on the relative magnitude between the current airspeed and the preset stall speed of the aircraft; the calculation logic of the flight safety reward item is configured to provide a positive reward when the current airspeed is higher than the stall speed, and the reward value decreases monotonically as the current airspeed decreases. The accident penalty and stability reward items include an accident penalty sub-item and a flight quality reward sub-item; the accident penalty sub-item outputs a preset negative constant as a penalty when a collision is detected in the aircraft; the flight quality reward sub-item is used to determine the reward value based on one or more of the aircraft's angular velocity component, angle of attack change rate, and control surface saturation degree, and the calculation logic of the flight quality reward sub-item is configured to provide a positive reward for flight states with small angular velocity components, gentle angle of attack change, and low control surface saturation.
6. The method according to claim 1, characterized in that, The determination of the aircraft's envelope boundaries and critical states based on the flight simulation logs and the aircraft's design limit parameters includes: Based on the design limit parameters, the state quantity threshold and its corresponding warning margin range are determined, and the state quantity sequence within a specified time window in the simulation log is detected based on the state quantity threshold. If any state quantity in the state quantity sequence exceeds its corresponding state quantity threshold or enters the warning margin range, the state quantity is determined as a critical state quantity, and the flight state corresponding to the time of the critical state quantity is marked as a candidate point of the envelope boundary. If the control quantity in the simulation log continuously reaches the physical limit value of its actuator within the specified time window, the control saturation zone of the aircraft is determined. Perform gradient analysis on the key state variables and determine the performance critical point based on the obtained gradient analysis results; Based on the candidate points of the envelope boundary, the control saturation region, and the performance critical point, the envelope boundary and critical state of the aircraft are determined.
7. An aircraft envelope testing device, characterized in that, The device includes: The vector generation module is used to generate a current observation vector for the agent in a flight simulation environment based on the envelope test target, historical flight state, historical flight environment, and historical control actions, and input the current observation vector into the agent; wherein, the agent is pre-trained with the goal of maximizing the cumulative reward defined by the reward function; The envelope testing module is used to generate a current control action based on the current observation vector through the intelligent agent, and to use the current control action to manipulate the aircraft to perform envelope testing; The result determination module is used to synchronously collect flight state parameters, flight environment parameters, and control action parameters during the envelope test to generate a flight simulation log, and to determine the envelope boundary and critical state of the aircraft based on the flight simulation log and the design limit parameters of the aircraft; wherein, the envelope boundary and critical state of the aircraft are used to optimize the reward function of the agent to optimize the agent; The agent is trained as follows: the observation vector received by the agent in a single simulation interaction is used as the current observation vector, and the current observation vector is input into the agent's policy network to output the current control action corresponding to the current observation vector; the expected reward is estimated by the agent's evaluation network based on the state-action pair formed by the current observation vector and the current control action to determine the cumulative reward corresponding to the state-action pair; the next flight state and the next observation vector corresponding to the current control action are calculated, and the reward function is used to determine the single-step reward for the current control action based on the next flight state; the current observation vector, the current control action, the single-step reward, and the next observation vector are written into the experience replay pool; the parameter gradients of the evaluation network and the policy network are calculated based on the sample interaction data sampled in the experience replay pool, and the network parameters of the evaluation network and the policy network are updated according to the parameter gradients.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the aircraft envelope testing method as described in any one of claims 1-6.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the aircraft envelope testing method as described in any one of claims 1-6.
Citation Information
Patent Citations
Unmanned aerial vehicle flight envelope protection method
CN115562318A
Deformable aircraft maneuvering trajectory design method based on SAC algorithm, electronic equipment and storage medium
CN119828724A