Aircraft deformation method and device based on safety reinforcement learning, equipment and medium
Through the method based on safety reinforcement learning, flight missions and data are obtained, flight status and environmental status are determined, deformation control strategies are formulated, and the autonomous deformation of the aircraft is achieved, which solves the problem of autonomous deformation of the aircraft in the existing technology and improves adaptability and combat capabilities.
Patent Information
- Application Number
- CN202510088022.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to effectively realize the autonomous deformation of the aircraft, which limits the adaptability and combat capabilities of the aircraft in different missions and environments.
Using a method based on safety reinforcement learning, the current flight status and atmospheric environment are determined by obtaining flight missions and flight data, deformation control strategies are formulated, and the autonomous deformation of the aircraft is realized through the control module.
The aircraft is automatically deformed under different missions and environments, which improves the aircraft's adaptability and combat capabilities, and ensures the safety of the deformation process.
Smart Images

Figure CN119953557A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of aircraft deformation, and in particular, to a method, device, equipment and medium for aircraft deformation based on safety reinforcement learning. Background Art
[0002] Most traditional aircraft adopt a fixed-wing structure, and usually complete flight missions such as ascent, descent and roll by operating the rudder. The flight quality is largely restricted by the aerodynamic shape of the aircraft. The single aerodynamic layout greatly limits the aircraft from meeting the needs of multiple flight missions at the same time, and the ceiling of combat performance improvement has begun to emerge. Compared with the traditional fixed shape design, the intelligent deformable aircraft has become a hot topic in the aerospace field with its good adaptability and wide applicability. Concepts such as flexible wings and variable sweep wings have also been applied.
[0003] The use of intelligent transformable aircraft, on the one hand, can effectively control the manufacturing and maintenance costs of equipment, adapt to various battlefield needs, achieve multiple domains with one aircraft or multiple uses with one missile, simplify and merge the relevant supporting systems of various aircraft, and achieve storage and maintenance in one form and use in multiple forms, so as to achieve the effect of reducing costs and increasing efficiency. On the other hand, it can increase the number of effective participating units in the future battlefield and improve combat capabilities. Through online dynamic reconstruction, various aircraft platforms present the form that best suits the needs of the current battle situation, form a large-scale combat system in a short time, and achieve instant and powerful confrontation and attack on the enemy.
[0004] In recent years, autonomous shape-changing decision-making has been one of the hot research topics of intelligent shape-changing aircraft. It mainly refers to the ability of aircraft to autonomously change their shape according to different flight missions and flight environments to achieve the best flight state. To truly realize the intelligence and autonomy of shape-changing aircraft, the aircraft needs to adaptively change to the optimal shape according to complex environments and tasks, which is ultimately a decision-making problem.
[0005] However, the existing methods have certain limitations in the process of realizing autonomous deformation of aircraft, and it is difficult to effectively realize the deformation of aircraft. Summary of the invention
[0006] The embodiments described herein provide a method, apparatus, device, and medium for aircraft deformation based on safety reinforcement learning, which overcome the above-mentioned problems.
[0007] In a first aspect, according to the content of the present disclosure, a method for aircraft deformation based on safety reinforcement learning is provided, comprising:
[0008] Obtain the flight mission of the target aircraft;
[0009] Under the flight mission of the target aircraft, determining the current flight state of the target aircraft based on the current flight data of the target aircraft;
[0010] Determine a deformation control strategy corresponding to the target aircraft based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, wherein the deformation control strategy corresponding to the target aircraft is used to describe the wing deformation data of the target aircraft;
[0011] Based on the deformation control strategy corresponding to the target aircraft, the target aircraft is controlled to perform autonomous deformation.
[0012] Optionally, under the flight mission of the target aircraft, determining the current flight state of the target aircraft based on the current flight data of the target aircraft includes:
[0013] Based on the current flight data of the target aircraft, generating a time series flight parameter of the target aircraft, wherein the time series flight parameter of the target aircraft is used to describe the historical flight speed and historical flight altitude of the target aircraft in different time series during the current flight;
[0014] Acquire a state parameter library corresponding to the flight mission of the target aircraft, wherein the state parameter library is used to store the correspondence between the time-series flight parameters of other aircraft and the state identification parameters, and the state identification parameters are used to describe the historical flight speed and historical flight altitude of the other aircraft in different time series during the historical flight process;
[0015] The time sequence flight parameters of the target aircraft are matched with the state identification parameters in the state parameter library to obtain the current flight state of the target aircraft.
[0016] Optionally, determining the deformation control strategy corresponding to the target aircraft based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located includes:
[0017] According to the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, a deformation type corresponding to the target aircraft is obtained, where the deformation type corresponding to the target aircraft includes: symmetrical deformation and asymmetrical deformation;
[0018] For the deformation type corresponding to the target aircraft, a deformation control strategy corresponding to the target aircraft is determined based on preset trajectory data of the target aircraft.
[0019] Optionally, obtaining the deformation type corresponding to the target aircraft according to the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located includes:
[0020] Performing a flight safety analysis on the current flight state of the target aircraft according to the atmospheric environment state in which the target aircraft is located, and obtaining flight constraint conditions corresponding to the target aircraft, wherein the flight constraint conditions are used to constrain the flight stability of the target aircraft;
[0021] Based on the flight constraint condition corresponding to the target aircraft, a deformation type corresponding to the target aircraft is determined.
[0022] Optionally, the deformation type corresponding to the target aircraft, based on preset trajectory data of the target aircraft, determines the deformation control strategy corresponding to the target aircraft, including:
[0023] If the deformation type corresponding to the target aircraft is the symmetrical deformation, then based on the preset trajectory data of the target aircraft, determining the deformation angles and deformation directions of the wings on both sides of the target aircraft;
[0024] If the deformation type corresponding to the target aircraft is the asymmetric deformation, the deformation angle and deformation direction of the wings on both sides of the target aircraft are adjusted according to the current aerodynamic parameters of the target aircraft to obtain the deformation angle and deformation direction of the wing on one side of the target aircraft and the deformation angle and deformation direction of the wing on the other side of the target aircraft.
[0025] Optionally, controlling the target aircraft to perform autonomous deformation based on the deformation control strategy corresponding to the target aircraft includes:
[0026] estimating safe deformation data of the target aircraft based on current aerodynamic parameters of the target aircraft;
[0027] If the wing deformation data described by the deformation control strategy corresponding to the target aircraft does not satisfy the safety deformation data, controlling the target aircraft to perform autonomous deformation based on the safety deformation data;
[0028] If the wing deformation data described by the deformation control strategy corresponding to the target aircraft meets the safety deformation data, the target aircraft is controlled to perform autonomous deformation based on the wing deformation data described by the deformation control strategy.
[0029] Optionally, also include:
[0030] If the current flight trajectory of the target aircraft deviates from the preset flight trajectory of the target aircraft after the target aircraft performs autonomous deformation, obtaining the current flight data of the target aircraft;
[0031] Based on the current aerodynamic parameters of the target aircraft, the current flight data of the target aircraft is adjusted so that the target aircraft flies based on a preset flight trajectory.
[0032] In a second aspect, according to the content of the present disclosure, there is provided an aircraft deformation device based on safety reinforcement learning, comprising:
[0033] An acquisition module, used to acquire the flight mission of the target aircraft;
[0034] A first determination module is used to determine the current flight state of the target aircraft based on the current flight data of the target aircraft under the flight mission of the target aircraft;
[0035] A second determination module is used to determine a deformation control strategy corresponding to the target aircraft based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, wherein the deformation control strategy corresponding to the target aircraft is used to describe the wing deformation data of the target aircraft;
[0036] A control module is used to control the target aircraft to perform autonomous deformation based on the deformation control strategy corresponding to the target aircraft.
[0037] In a third aspect, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the aircraft deformation method based on safety reinforcement learning in any of the above embodiments are implemented.
[0038] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the aircraft deformation method based on safety reinforcement learning in any of the above embodiments are implemented.
[0039] The aircraft deformation method based on safety reinforcement learning provided in the embodiment of the present application obtains the flight mission of the target aircraft; under the flight mission of the target aircraft, the current flight state of the target aircraft is determined based on the current flight data of the target aircraft; based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, the deformation control strategy corresponding to the target aircraft is determined, and the deformation control strategy corresponding to the target aircraft is used to describe the wing deformation data of the target aircraft; based on the deformation control strategy corresponding to the target aircraft, the target aircraft is controlled to perform autonomous deformation. In this way, different deformation control strategies are formulated for different tasks of the aircraft to control the aircraft to perform autonomous deformation, and autonomous deformation of aircraft in different states under different tasks is effectively realized.
[0040] The above description is only an overview of the technical solution of the embodiment of the present application. In order to more clearly understand the technical means of the embodiment of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiment of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly described below. It should be noted that the drawings described below only relate to some embodiments of the present disclosure, but are not intended to limit the present disclosure, wherein:
[0042] Figure 1 It is a flow chart of an aircraft deformation method based on safety reinforcement learning provided by the present invention.
[0043] Figure 2 It is a structural schematic diagram of an aircraft deformation device based on safety reinforcement learning provided by the present invention.
[0044] Figure 3 It is a structural schematic diagram of a computer device provided by the present disclosure.
[0045] It should be noted that the elements in the drawings are schematic and not drawn to scale. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative work also fall within the scope of protection of the present disclosure.
[0047] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by a person skilled in the art to which the subject matter of the present disclosure belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the specification and the relevant art, and will not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. As used herein, a statement that two or more parts are "connected" or "coupled" together shall mean that the parts are joined together directly or through one or more intermediate components.
[0048] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase "embodiments" in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0049] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists, A and B exist at the same time, and B exists. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. Terms such as "first" and "second" are only used to distinguish one component (or a part of a component) from another component (or another part of a component).
[0050] In the description of the present application, unless otherwise specified, "plurality" means more than two (including two), and similarly, "multiple groups" means more than two groups (including two).
[0051] As a new type of combat platform, the autonomous shape-changing decision of the intelligent shape-changing aircraft has been a research hotspot in recent years. This technology can autonomously change its shape according to different flight missions and flight environments to achieve the best flight state. In the field of artificial intelligence, reinforcement learning provides a learning method that interacts with the environment. The intelligent agent maximizes the reward learning through "trial and error" to obtain the optimal strategy, improve the action plan to adapt to the external environment, and thus complete the mission objectives.
[0052] Based on the advantages of intelligent deformable aircraft and the current status of autonomous deformation control, and taking into account the continuity of state space and action space, the deep deterministic policy gradient (LSTM-DDPG) algorithm with task classifier is used to control the intelligent deformable aircraft to autonomously change its shape according to different flight conditions and tasks; for the safety issues such as stall, instability, and roll that may be faced during the shape deformation of the aircraft, the Lagrange multiplier is introduced into the reinforcement learning objective function and the Lagrange multiplier is used to balance the original objective function and the constraints, so as to realize the constraints on safety requirements in the learning process.
[0053] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0054] Figure 1 is a flow chart of a method for aircraft deformation based on safety reinforcement learning provided by an embodiment of the present disclosure, such as Figure 1As shown, the specific process of the aircraft deformation method based on safety reinforcement learning includes:
[0055] S110: Obtain the flight mission of the target aircraft.
[0056] The flight missions of the target aircraft may include but are not limited to: regional inspection, target tracking, regional detection and fixed-point cruising.
[0057] S120 . Under the flight mission of the target aircraft, determine the current flight state of the target aircraft based on the current flight data of the target aircraft.
[0058] In this embodiment, an LSTM (Long Short-Term Memory) network is used to determine the flight state. LSTM is used to process the continuity problem of state space and action space, can remember values of indefinite time length, and is suitable for processing time series data.
[0059] The LSTM network that meets the cruise missile flight mission refers to an LSTM network model designed specifically for the flight mission of aircraft such as cruise missiles. This network model can learn and predict based on the flight characteristics of the cruise missile, such as flight speed, altitude, heading, etc. Specifically, the LSTM network processes and memorizes time series data through its unique gating mechanism (including input gate, forget gate, and output gate), so that the flight trajectory of the cruise missile can be effectively predicted and controlled. This network model has significant advantages in processing continuous states and action spaces related to flight missions, and can improve the flight performance and mission execution efficiency of cruise missiles.
[0060] The current flight data of the target aircraft can be used to describe the flight altitude and flight speed of the target aircraft at each collection time point from the start of takeoff to the current moment, that is, the flight data that has been generated during this flight.
[0061] The current flight status of the target aircraft may include, but is not limited to, take-off, climb, cruise, hover, descent, approach and landing.
[0062] In some embodiments, under the flight mission of the target aircraft, determining the current flight state of the target aircraft based on the current flight data of the target aircraft includes:
[0063] Based on the current flight data of the target aircraft, the time-series flight parameters of the target aircraft are generated; the state parameter library corresponding to the flight mission of the target aircraft is obtained, and the state parameter library is used to store the corresponding relationship between the time-series flight parameters and the state identification parameters of other aircraft, and the state identification parameters are used to describe the historical flight speeds and historical flight altitudes of other aircraft in different time series during the historical flight process; the time-series flight parameters of the target aircraft are matched with the state identification parameters in the state parameter library to obtain the current flight state of the target aircraft.
[0064] Among them, the time series flight parameters of the target aircraft are used to describe the historical flight speed and historical flight altitude of the target aircraft in different time series during this flight. The time series flight parameters of the target aircraft can be obtained by arranging each group of flight speeds and flight altitudes collected at each collection time point from front to back according to the order of the collection time points.
[0065] Matching the target aircraft's sequential flight parameters with the state identification parameters in the state parameter library to obtain the target aircraft's current flight state may include: matching each historical flight altitude in the target aircraft's sequential flight parameters with each historical flight altitude in the state parameter library to obtain multiple altitude matching differences; matching each historical flight speed in the target aircraft's sequential flight parameters with each historical flight speed in the state parameter library to obtain multiple speed matching differences; based on the sum of the altitude matching differences and the speed matching differences, determining a group of state identification parameters with the smallest error in the state parameter library, and determining the flight state associated with this group of state identification parameters as the target aircraft's current flight state. Thus, the target aircraft's current flight state is effectively determined.
[0066] S130: Determine a deformation control strategy corresponding to the target aircraft based on the current flight state of the target aircraft and the state of the atmospheric environment in which the target aircraft is located.
[0067] Among them, Lag-LSTM-DDPG is adopted, and on the basis of the LSTM network, a task classifier is integrated to adapt to different flight mission requirements. The task classifier can classify the flight status according to the specific mission of the aircraft and provide customized deformation strategies for each task.
[0068] The DDPG algorithm is a model-free deep reinforcement learning algorithm that uses an actor-critic architecture to approximate the policy network and value function through a deep neural network. The algorithm uses the stochastic gradient method to train the parameters in the policy network and value network models, and uses a dual neural network architecture (i.e., the online network and the target network) to improve the stability and convergence speed of the learning process.
[0069] The deformation control strategy corresponding to the target aircraft is used to describe the wing deformation data of the target aircraft, such as the scaling angle of the wing in a certain direction.
[0070] In some embodiments, based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, determining the deformation control strategy corresponding to the target aircraft includes:
[0071] According to the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, the deformation type corresponding to the target aircraft is obtained; for the deformation type corresponding to the target aircraft, based on the preset trajectory data of the target aircraft, the deformation control strategy corresponding to the target aircraft is determined.
[0072] The deformation types corresponding to the target aircraft include: symmetrical deformation and asymmetrical deformation. Symmetric deformation means that the deformation directions and scaling angles of the two wings of the target aircraft are the same, while asymmetrical deformation means that the deformation directions and scaling angles of the two wings of the target aircraft are different.
[0073] Therefore, by determining the deformation type of the target aircraft, it is convenient to accurately determine the deformation control strategy corresponding to the target aircraft according to the deformation type in combination with the preset trajectory data of the target aircraft.
[0074] In some embodiments, the deformation type corresponding to the target aircraft is obtained according to the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, including:
[0075] A flight safety analysis is performed on the current flight state of the target aircraft according to the atmospheric environment state in which the target aircraft is located, and the flight constraint conditions corresponding to the target aircraft are obtained. The flight constraint conditions are used to constrain the flight stability of the target aircraft. Based on the flight constraint conditions corresponding to the target aircraft, the deformation type corresponding to the target aircraft is determined.
[0076] Among them, the atmospheric environment state of the target aircraft can be used to describe the strength of the atmospheric environment data in the area where the target aircraft is currently located, such as temperature, humidity, wind speed, air pressure and precipitation.
[0077] Based on the flight constraint conditions corresponding to the target aircraft, determining the deformation type corresponding to the target aircraft may include: if the flight constraint conditions corresponding to the target aircraft are pitch, then determining the deformation type corresponding to the target aircraft is symmetrical deformation; if the flight constraint conditions corresponding to the target aircraft are roll, then determining the deformation type corresponding to the target aircraft is asymmetrical deformation.
[0078] In addition, the flight constraint conditions corresponding to the target aircraft may also include: yaw. When the flight constraint condition corresponding to the target aircraft is yaw, the deformation type corresponding to the target aircraft can be determined to be asymmetric deformation, so as to ensure the flight stability of the target aircraft when adjusting the route.
[0079] Therefore, the current flight state of the target aircraft is analyzed for flight safety through the atmospheric environment state in which the target aircraft is located, and the flight constraint conditions of the target aircraft are determined to effectively measure the corresponding deformation type of the target aircraft.
[0080] In other embodiments, the deformation type corresponding to the target aircraft is obtained based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, and it may also include: obtaining the current air situation of the target aircraft; if the current air situation of the target aircraft shows that there are other aircraft in the vicinity of the target aircraft, and the other aircraft have a flight impact on the current flight state of the target aircraft, then the deformation type corresponding to the target aircraft is determined according to the atmospheric environment state in which the target aircraft is located, for example, when the wind speed is too large, the deformation type corresponding to the target aircraft is determined to be asymmetric.
[0081] In some embodiments, for the deformation type corresponding to the target aircraft, based on the preset trajectory data of the target aircraft, determining the deformation control strategy corresponding to the target aircraft includes:
[0082] If the deformation type corresponding to the target aircraft is symmetrical deformation, the deformation angle and deformation direction of the wings on both sides of the target aircraft are determined based on the preset trajectory data of the target aircraft; if the deformation type corresponding to the target aircraft is asymmetrical deformation, the deformation angle and deformation direction of the wings on both sides of the target aircraft are adjusted according to the current aerodynamic parameters of the target aircraft to obtain the deformation angle and deformation direction of the wing on one side of the target aircraft and the deformation angle and deformation direction of the wing on the other side of the target aircraft.
[0083] The current aerodynamic parameters of the target aircraft refer to the aerodynamic parameters that affect the target aircraft during flight, such as aerodynamic force, aerodynamic moment, wind resistance, lift, stall speed, etc.
[0084] According to the current aerodynamic parameters of the target aircraft, the deformation angle and deformation direction of the wings on both sides of the target aircraft are adjusted, which may include: if the current aerodynamic parameters of the target aircraft show that the current wind resistance is too large, the deformation angle of the wing on the side of the target aircraft along the wind direction is reduced, and the deformation direction angle is reduced, and the deformation angle of the wing on the side of the target aircraft against the wind direction is increased, and the deformation direction angle is increased.
[0085] Therefore, when it is determined that the deformation type corresponding to the target aircraft is symmetrical deformation, the deformation angle and deformation direction of the wings on both sides of the target aircraft are determined through the preset trajectory data of the target aircraft, so that the target aircraft can navigate according to the preset trajectory data. When it is determined that the deformation type corresponding to the target aircraft is asymmetrical deformation, the deformation angle and deformation direction of the wings on both sides of the target aircraft are adjusted in combination with the current aerodynamic parameters of the target aircraft, so that the target aircraft can adapt to the atmospheric environment during the deformation process and reduce the difficulty of deformation.
[0086] S140: Based on the deformation control strategy corresponding to the target aircraft, control the target aircraft to perform autonomous deformation.
[0087] In order to ensure the safety of the aircraft during the deformation process, the Lagrange multiplier method is introduced to balance the original objective function and the constraints. The Lagrange multiplier is used to introduce the constraints of safety requirements into the reinforcement learning objective function, thereby ensuring flight safety while optimizing the deformation strategy.
[0088] In some embodiments, based on the deformation control strategy corresponding to the target aircraft, controlling the target aircraft to perform autonomous deformation includes:
[0089] Based on the current aerodynamic parameters of the target aircraft, the safe deformation data of the target aircraft is estimated; if the wing deformation data described by the deformation control strategy corresponding to the target aircraft does not meet the safe deformation data, the target aircraft is controlled to perform autonomous deformation based on the safe deformation data; if the wing deformation data described by the deformation control strategy corresponding to the target aircraft meets the safe deformation data, the target aircraft is controlled to perform autonomous deformation based on the wing deformation data described by the deformation control strategy.
[0090] Among them, the safety deformation data of the target aircraft can be used as a deformation constraint to limit the target aircraft, in order to ensure the safety of the target aircraft during the deformation process.
[0091] The safe deformation data of the target aircraft can be used to describe the safe deformation direction and safe deformation angle of the wing of the target aircraft.
[0092] Therefore, the deformation process of the target aircraft is safely constrained by the safe deformation data of the target aircraft, so as to avoid the problem of reduced flight safety due to the influence of current aerodynamic parameters during the deformation process of the target aircraft.
[0093] In this embodiment, the flight mission of the target aircraft is obtained; under the flight mission of the target aircraft, the current flight state of the target aircraft is determined based on the current flight data of the target aircraft; based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, the deformation control strategy corresponding to the target aircraft is determined, and the deformation control strategy corresponding to the target aircraft is used to describe the wing deformation data of the target aircraft; based on the deformation control strategy corresponding to the target aircraft, the target aircraft is controlled to perform autonomous deformation. In this way, different deformation control strategies are formulated for different tasks of the aircraft to control the aircraft to perform autonomous deformation, and autonomous deformation of the aircraft in different states under different tasks is effectively realized.
[0094] In some embodiments, the method of this embodiment further includes:
[0095] If the current flight trajectory of the target aircraft deviates from the preset flight trajectory of the target aircraft after the target aircraft undergoes autonomous deformation, the current flight data of the target aircraft is obtained; based on the current aerodynamic parameters of the target aircraft, the current flight data of the target aircraft is adjusted so that the target aircraft flies based on the preset flight trajectory.
[0096] Among them, the current flight data of the target aircraft is adjusted through the current aerodynamic parameters of the target aircraft, which can ensure that the target aircraft flies in accordance with the preset flight trajectory while avoiding problems such as turbulence or unstable flight of the aircraft due to the influence of aerodynamic parameters during data adjustment.
[0097] In summary, in this embodiment, considering the safety requirements, the Lagrange multiplier method is introduced on the basis of the DDPG algorithm to deal with safety constraints. This method can ensure that the deformation process meets the preset safety standards while optimizing the deformation strategy of the aircraft, such as avoiding dangerous situations such as stalling, instability and rolling. In order to adapt to the deformation characteristics of the aircraft, the Lag-DDPG algorithm takes into account the dynamic changes of the aircraft during the deformation process, and adapts to these changes by adjusting the parameters of the strategy network and the value network, which enables the algorithm to provide the aircraft with the optimal control strategy in different flight stages and different deformation states. In practical applications, the Lag-DDPG algorithm designs a deformation strategy that adapts to symmetric and asymmetric deformation conditions by learning the aerodynamic data and deformation dynamics equations of the aircraft. Simulations show that the algorithm can converge quickly and keep the deformation error within 3%, significantly improving the adaptability of the aircraft to different flight missions.
[0098] By introducing the task classifier into the LSTM-DDPG algorithm with task classifier, the LSTM network can better handle the continuous state space and action space, and improve the accuracy and efficiency of decision-making. By introducing the Lagrange multiplier into the reinforcement learning objective function, the original objective function and the constraint conditions are effectively balanced, the constraints on safety requirements are realized, and the stability and safety of the aircraft during the deformation process are guaranteed. The intelligent deformable aircraft can autonomously change its shape according to different flight conditions and tasks, improving its combat capability and adaptability.
[0099] Figure 2 A structural schematic diagram of an aircraft deformation device based on safety reinforcement learning is provided for this embodiment. The aircraft deformation device based on safety reinforcement learning may include: an acquisition module 210, a first determination module 220, a second determination module 230 and a control module 240.
[0100] The acquisition module 210 is used to acquire the flight mission of the target aircraft.
[0101] The first determination module 220 is used to determine the current flight state of the target aircraft based on the current flight data of the target aircraft under the flight mission of the target aircraft.
[0102] The second determination module 230 is used to determine the deformation control strategy corresponding to the target aircraft based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located. The deformation control strategy corresponding to the target aircraft is used to describe the wing deformation data of the target aircraft.
[0103] The control module 240 is used to control the target aircraft to perform autonomous deformation based on the deformation control strategy corresponding to the target aircraft.
[0104] In this embodiment, optionally, the first determining module 220 is specifically configured to:
[0105] Based on the current flight data of the target aircraft, the target aircraft's time-series flight parameters are generated, and the target aircraft's time-series flight parameters are used to describe the target aircraft's historical flight speed and historical flight altitude in different time series during this flight; the state parameter library corresponding to the target aircraft's flight mission is obtained, and the state parameter library is used to store the corresponding relationship between the time-series flight parameters and state identification parameters of other aircraft, and the state identification parameters are used to describe the historical flight speed and historical flight altitude of other aircraft in different time series during historical flights; the target aircraft's time-series flight parameters are matched with the state identification parameters in the state parameter library to obtain the target aircraft's current flight state.
[0106] In this embodiment, optionally, the second determining module 230 includes: a first determining unit and a second determining unit.
[0107] The first determining unit is used to obtain a deformation type corresponding to the target aircraft according to a current flight state of the target aircraft and an atmospheric environment state where the target aircraft is located. The deformation type corresponding to the target aircraft includes: symmetrical deformation and asymmetrical deformation.
[0108] The second determination unit is used to determine a deformation control strategy corresponding to the target aircraft based on preset trajectory data of the target aircraft according to the deformation type corresponding to the target aircraft.
[0109] In this embodiment, optionally, the first determining unit is specifically configured to:
[0110] A flight safety analysis is performed on the current flight state of the target aircraft according to the atmospheric environment state in which the target aircraft is located, and the flight constraint conditions corresponding to the target aircraft are obtained. The flight constraint conditions are used to constrain the flight stability of the target aircraft. Based on the flight constraint conditions corresponding to the target aircraft, the deformation type corresponding to the target aircraft is determined.
[0111] In this embodiment, optionally, the second determining unit is specifically configured to:
[0112] If the deformation type corresponding to the target aircraft is symmetrical deformation, the deformation angle and deformation direction of the wings on both sides of the target aircraft are determined based on the preset trajectory data of the target aircraft; if the deformation type corresponding to the target aircraft is asymmetrical deformation, the deformation angle and deformation direction of the wings on both sides of the target aircraft are adjusted according to the current aerodynamic parameters of the target aircraft to obtain the deformation angle and deformation direction of the wing on one side of the target aircraft and the deformation angle and deformation direction of the wing on the other side of the target aircraft.
[0113] In this embodiment, optionally, the control module 240 is specifically configured to:
[0114] Based on the current aerodynamic parameters of the target aircraft, the safe deformation data of the target aircraft is estimated; if the wing deformation data described by the deformation control strategy corresponding to the target aircraft does not meet the safe deformation data, the target aircraft is controlled to perform autonomous deformation based on the safe deformation data; if the wing deformation data described by the deformation control strategy corresponding to the target aircraft meets the safe deformation data, the target aircraft is controlled to perform autonomous deformation based on the wing deformation data described by the deformation control strategy.
[0115] In this embodiment, optionally, it also includes: an adjustment module.
[0116] The acquisition module 210 is further configured to acquire current flight data of the target aircraft if the current flight trajectory of the target aircraft deviates from the preset flight trajectory of the target aircraft after the target aircraft undergoes autonomous deformation.
[0117] The adjustment module is used to adjust the current flight data of the target aircraft based on the current aerodynamic parameters of the target aircraft, so that the target aircraft flies based on a preset flight trajectory.
[0118] The aircraft deformation device based on safety reinforcement learning provided by the present disclosure can execute the above method embodiments. Its specific implementation principles and technical effects can be found in the above method embodiments, and the present disclosure will not repeat them here.
[0119] The present application also provides a computer device. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.
[0120] The computer device includes a memory 310 and a processor 320 that are connected to each other through a system bus. It should be noted that the figure only shows a computer device with a memory 310 and a processor 320, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.
[0121] Computer devices can be computing devices such as desktop computers, notebooks, PDAs, and cloud servers. Computer devices can interact with users through keyboards, mice, remote controls, touch pads, or voice control devices.
[0122] The memory 310 includes at least one type of readable storage medium, and the readable storage medium includes a non-volatile memory or a volatile memory, for example, a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc., and the RAM may include a static RAM or a dynamic RAM. In some embodiments, the memory 310 may be an internal storage unit of a computer device, for example, a hard disk or a memory of the computer device. In other embodiments, the memory 310 may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card or a flash card (FlashCard) equipped on the computer device. Of course, the memory 310 may also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the memory 310 is generally used to store the operating system and various application software installed on the computer device, such as the program code of the above method. In addition, the memory 310 may also be used to temporarily store various data that have been output or are to be output.
[0123] The processor 320 is generally used to perform the overall operation of the computer device. In this embodiment, the memory 310 is used to store program codes or instructions, the program code includes computer operation instructions, and the processor 320 is used to execute the program codes or instructions stored in the memory 310 or process data, such as running the program code of the above method.
[0124] In this article, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus system can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0125] Another embodiment of the present application also provides a computer-readable medium, which may be a computer-readable signal medium or a computer-readable medium. A processor in a computer reads a computer-readable program code stored in the computer-readable medium, so that the processor can execute the functional actions specified in each step or a combination of steps in the above method; and generate a device for implementing the functional actions specified in each block or a combination of blocks in the block diagram.
[0126] Computer-readable media include but are not limited to electronic, magnetic, optical, electromagnetic, infrared memory or semiconductor systems, devices or apparatuses, or any appropriate combination of the foregoing, the memory is used to store program codes or instructions, the program codes include computer operating instructions, and the processor is used to execute the program codes or instructions of the above methods stored in the memory.
[0127] For the definitions of memory and processor, please refer to the description of the aforementioned computer device embodiment and will not be repeated here.
[0128] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0129] Each functional unit or module in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0130] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program code.
[0131] In the claims, any reference symbols placed between brackets shall not be construed as limiting the claims. The word "comprising" described in the present application does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented with the aid of hardware comprising several different elements and with the aid of a suitably programmed computer. In a unit claim that lists a number of devices, several units of these devices may be embodied by the same hardware item. The use of first, second, and third, etc. does not indicate any order, and these words may be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be understood as limitations on the order of execution.
[0132] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for aircraft deformation based on safety reinforcement learning, characterized in that: include: Obtain the flight mission of the target aircraft; Under the flight mission of the target aircraft, determining the current flight state of the target aircraft based on the current flight data of the target aircraft; Determine a deformation control strategy corresponding to the target aircraft based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, wherein the deformation control strategy corresponding to the target aircraft is used to describe the wing deformation data of the target aircraft; Based on the deformation control strategy corresponding to the target aircraft, the target aircraft is controlled to perform autonomous deformation.
2. The method according to claim 1, characterized in that Under the flight mission of the target aircraft, determining the current flight state of the target aircraft based on the current flight data of the target aircraft includes: Based on the current flight data of the target aircraft, generating a time series flight parameter of the target aircraft, wherein the time series flight parameter of the target aircraft is used to describe the historical flight speed and historical flight altitude of the target aircraft in different time series during the current flight; Acquire a state parameter library corresponding to the flight mission of the target aircraft, wherein the state parameter library is used to store the correspondence between the time-series flight parameters of other aircraft and the state identification parameters, and the state identification parameters are used to describe the historical flight speed and historical flight altitude of the other aircraft in different time series during the historical flight process; The time sequence flight parameters of the target aircraft are matched with the state identification parameters in the state parameter library to obtain the current flight state of the target aircraft.
3. The method according to claim 2, characterized in that The step of determining a deformation control strategy corresponding to the target aircraft based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located comprises: According to the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, a deformation type corresponding to the target aircraft is obtained, where the deformation type corresponding to the target aircraft includes: symmetrical deformation and asymmetrical deformation; For the deformation type corresponding to the target aircraft, a deformation control strategy corresponding to the target aircraft is determined based on preset trajectory data of the target aircraft.
4. The method according to claim 3, characterized in that: The step of obtaining the deformation type corresponding to the target aircraft according to the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located comprises: Performing a flight safety analysis on the current flight state of the target aircraft according to the atmospheric environment state in which the target aircraft is located, and obtaining flight constraint conditions corresponding to the target aircraft, wherein the flight constraint conditions are used to constrain the flight stability of the target aircraft; Based on the flight constraint condition corresponding to the target aircraft, a deformation type corresponding to the target aircraft is determined.
5. The method according to claim 3, characterized in that: The method of determining a deformation control strategy corresponding to the target aircraft based on preset trajectory data of the target aircraft for the deformation type corresponding to the target aircraft includes: If the deformation type corresponding to the target aircraft is the symmetrical deformation, then based on the preset trajectory data of the target aircraft, determining the deformation angles and deformation directions of the wings on both sides of the target aircraft; If the deformation type corresponding to the target aircraft is the asymmetric deformation, the deformation angle and deformation direction of the wings on both sides of the target aircraft are adjusted according to the current aerodynamic parameters of the target aircraft to obtain the deformation angle and deformation direction of the wing on one side of the target aircraft and the deformation angle and deformation direction of the wing on the other side of the target aircraft.
6. The method according to claim 5, characterized in that The step of controlling the target aircraft to perform autonomous deformation based on the deformation control strategy corresponding to the target aircraft includes: estimating safe deformation data of the target aircraft based on current aerodynamic parameters of the target aircraft; If the wing deformation data described by the deformation control strategy corresponding to the target aircraft does not satisfy the safety deformation data, controlling the target aircraft to perform autonomous deformation based on the safety deformation data; If the wing deformation data described by the deformation control strategy corresponding to the target aircraft meets the safety deformation data, the target aircraft is controlled to perform autonomous deformation based on the wing deformation data described by the deformation control strategy.
7. The method according to claim 6, characterized in that Also includes: If the current flight trajectory of the target aircraft deviates from the preset flight trajectory of the target aircraft after the target aircraft performs autonomous deformation, obtaining the current flight data of the target aircraft; Based on the current aerodynamic parameters of the target aircraft, the current flight data of the target aircraft is adjusted so that the target aircraft flies based on a preset flight trajectory.
8. An aircraft deformation device based on safety reinforcement learning, characterized in that: include: An acquisition module, used to acquire the flight mission of the target aircraft; A first determination module is used to determine the current flight state of the target aircraft based on the current flight data of the target aircraft under the flight mission of the target aircraft; A second determination module is used to determine a deformation control strategy corresponding to the target aircraft based on the current flight state of the target aircraft and the atmospheric environment state in which the target aircraft is located, wherein the deformation control strategy corresponding to the target aircraft is used to describe the wing deformation data of the target aircraft; A control module is used to control the target aircraft to perform autonomous deformation based on the deformation control strategy corresponding to the target aircraft.
9. A computer device, characterized in that: The invention comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, an aircraft deformation method based on safety reinforcement learning as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the aircraft deformation method based on safety reinforcement learning as described in any one of claims 1 to 7 is implemented.