Aircraft recovery scheduling method, device, equipment, medium and product
Through the dual-critic architecture and generalized advantage estimation function optimization strategy network, the problems of active prevention and dynamic adjustment of safety risks in aircraft recovery scheduling are solved, the optimal scheduling under safety constraints is achieved, the violation rate is reduced and efficiency is improved.
Patent Information
- Application Number
- CN202511204947.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing aircraft recovery scheduling methods have problems such as being unable to proactively prevent safety risks in safety-critical scenarios, relying on manual experience for penalty coefficients, and having high variance in Monte Carlo return estimation, which leads to delayed constraint processing and constraint failure.
A dual-critic architecture is adopted, combined with the generalized advantage estimation function, shearing mechanism and Lagrangian safety penalty term. The performance evaluator network and the safety critic network are used to optimize the policy network to achieve accurate quantification and dynamic adjustment of safety risks and construct the optimal scheduling under safety constraints.
It achieves optimal scheduling of aircraft recovery under safety constraints, reduces safety violation rates, improves mission completion efficiency, and dynamically adapts to complex environmental changes.
Smart Images

Figure CN120706847A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of aircraft recovery scheduling, and in particular to an aircraft recovery scheduling method, device, equipment, medium and product. Background Art
[0002] Current aircraft recovery scheduling primarily relies on deep reinforcement learning algorithms such as Proximal Policy Optimization (PPO). However, these algorithms have significant limitations in safety-critical scenarios. First, the traditional single-critic architecture cannot independently quantify safety risks and only passively imposes penalties when constraints are violated. Second, the safety penalty coefficient relies on manual experience, making it difficult to dynamically adapt to complex environmental changes. Third, the Monte Carlo reward estimation suffers from high variance, leading to oscillations in policy updates. In aircraft recovery scenarios, multiple safety constraints, such as fuel consumption rate, landing gear stress peak, and runway deviation distance, are coupled. Existing methods are prone to constraint failure due to lags in constraint processing, making it impossible to achieve safety constraints. Therefore, a method is needed that can proactively prevent aircraft recovery failures and achieve optimal scheduling within safety constraints. Summary of the Invention
[0003] The purpose of this application is to provide an aircraft recovery scheduling method, device, equipment, medium and product, which can achieve optimal scheduling of aircraft recovery under safety constraints.
[0004] To achieve the above objectives, this application provides the following solutions: In a first aspect, the present application provides an aircraft recovery scheduling method, comprising: Get real-time environment status; According to the real-time environmental state, a trained policy network is used to schedule aircraft recovery to obtain the optimal action for aircraft recovery; the trained policy network is obtained by optimizing the network parameters of the initial policy network based on a historical policy network, a performance evaluator network, and a safety critic network, combined with a generalized advantage estimation function, a shearing mechanism, and a Lagrangian safety penalty term.
[0005] In one embodiment, the optimization process of the policy network specifically includes: According to the decision stage Environmental status Using the historical strategy network to interact with the environment, obtain a first action and control the environment to perform the first action; Obtaining an immediate reward and a safety cost of the environment performing the first action and generating a state-action trajectory dataset based on the immediate reward and the safety cost; Calculating a generalized advantage estimation function based on the state-action trajectory dataset based on the initial policy network and the performance evaluator network; Updating the network parameters of the initial policy network based on the historical policy network using a clipping mechanism according to the generalized advantage estimation function; Optimizing the network parameters of the performance critic network and the safety critic network respectively according to the state-action trajectory dataset; Determine the Lagrangian safety penalty term using the optimized safety critic network; According to the optimized performance critic network, the optimized safety critic network and the Lagrangian safety penalty term, the updated initial policy network is optimized using a weighted multi-objective function to obtain a trained policy network.
[0006] In one embodiment, in the decision stage Environmental status Before using the historical strategy network to interact with the environment, obtain a first action, and control the environment to execute the first action, the method further includes: Initialize the initial policy network, performance critic network, and safety critic network.
[0007] In one embodiment, the expression of the generalized advantage estimation function is: ; in: Indicates the decision stage The action advantage estimate, represents the time index of the current decision stage, Indicates the future decision stage relative to The offset index of is the reward discount factor, is the bias-variance trade-off coefficient, is the time series difference residual, is the maximum total number of decision stages in a single iteration.
[0008] In one embodiment, the expression for updating the network parameters of the initial policy network based on the historical policy network using the clipping mechanism according to the generalized advantage estimation function is: ; in: represents the clipping loss function of the policy network, are the trainable parameters of the policy network, is the importance sampling weight, Decision-making stage The generalized advantage estimate of Indicates that Clip to interval , is the clipping threshold, is the expectation operator.
[0009] In one embodiment, the Lagrangian safety penalty term is determined using the optimized safety critic network, specifically including: Determine the expected cumulative security cost based on the optimized security critic network; A Lagrangian safety penalty term is determined using a Lagrangian multiplier according to the expected cumulative safety cost.
[0010] In a second aspect, the present application provides an aircraft recovery scheduling device, comprising: Acquisition module, used to obtain real-time environment status; The scheduling module is used to use a trained policy network to perform aircraft recovery scheduling according to the real-time environmental state to obtain the optimal action for aircraft recovery; the trained policy network is obtained by optimizing the network parameters of the initial policy network based on the historical policy network, the performance evaluator network, and the safety critic network in combination with the generalized advantage estimation function, the shear mechanism, and the Lagrangian safety penalty term.
[0011] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aircraft recovery scheduling method.
[0012] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the aircraft recovery scheduling method when executed by a processor.
[0013] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the aircraft recovery scheduling method when executed by a processor.
[0014] According to the specific embodiments provided in this application, this application discloses the following technical effects: The present application provides an aircraft recovery scheduling method, apparatus, equipment, medium, and product. During the optimization of the initial policy network, accurate quantification of safety risks is achieved based on a performance evaluator network and a safety critic network. The accuracy and stability of long-term benefit estimation are balanced through a generalized advantage estimation function. The Lagrangian safety penalty term is used to balance performance and safety in real time, thereby achieving optimal scheduling of aircraft recovery under safety constraints. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 A flowchart of an aircraft recovery scheduling method provided in one embodiment of the present application.
[0017] Figure 2 Diagram of the secure reinforcement learning framework for the dual-critic architecture.
[0018] Figure 3 Schematic diagram of optimization in aircraft recovery scheduling method.
[0019] Figure 4 A schematic diagram of the functional modules of an aircraft recovery scheduling device provided in one embodiment of the present application.
[0020] Figure 5 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0021] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0022] Traditional proximal policy optimization (PPO) algorithms have significant drawbacks for aircraft recovery scheduling: First, they can only passively adjust policies after constraint violations and are unable to proactively prevent safety violations. Second, constraint penalty coefficients rely on manual experience and are unable to adapt to dynamic environmental changes. Third, a single value assessment network struggles to balance performance incentives and safety cost assessments, leading to conflicting optimization objectives. Existing improved methods attempt to address constraints through fixed penalty functions, but adjusting penalty weights requires trial and error, and static coefficients are unable to cope with complex state changes. Furthermore, the high variance of Monte Carlo methods and the estimation bias of temporal difference methods further exacerbate policy volatility, seriously threatening recovery safety. Although some studies have attempted to incorporate Lagrangian relaxation mechanisms, these have not addressed the critical issues of accurately predicting and proactively preventing safety costs. Therefore, an integrated framework integrating safety cost quantification, proactive constraint prevention, and adaptive parameter adjustment is urgently needed.
[0023] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0024] In an exemplary embodiment, Figure 1 As shown, an aircraft recovery scheduling method is provided. The method is executed by a computer device. Specifically, it can be executed by a computer device such as a terminal or a server alone, or it can be executed by a terminal and a server together. In the embodiment of the present application, the method is applied to a server as an example for explanation, and includes the following steps.
[0025] Step 101: Obtain real-time environment status.
[0026] Step 102: Using the trained policy network to perform aircraft recovery scheduling according to the real-time environmental state to obtain the optimal action for aircraft recovery; the trained policy network is obtained by optimizing the network parameters of the initial policy network based on the historical policy network, the performance evaluator network, and the safety critic network, in combination with the generalized advantage estimation function, the shearing mechanism, and the Lagrangian safety penalty term.
[0027] In the process of optimizing the initial policy network, accurate quantification of safety risks is achieved based on the performance evaluator network and the safety critic network. The accuracy and stability of long-term benefit estimation are balanced through the generalized advantage estimation function, and the Lagrangian safety penalty term is used to balance performance and safety in real time, thereby achieving optimal scheduling of aircraft recovery under safety constraints.
[0028] In an exemplary embodiment, the optimization process of the policy network specifically includes: Environmental status A historical policy network is used to interact with the environment to obtain a first action and control the environment to perform the first action; an immediate reward and a safety cost of the environment performing the first action are obtained and a state-action trajectory dataset is generated based on the immediate reward and the safety cost; a generalized advantage estimation function is calculated based on the state-action trajectory dataset based on the initial policy network and the performance evaluator network; the network parameters of the initial policy network are updated based on the historical policy network using a clipping mechanism according to the generalized advantage estimation function; the network parameters of the performance critic network and the safety critic network are optimized respectively according to the state-action trajectory dataset; the Lagrangian safety penalty term is determined using the optimized safety critic network; the updated initial policy network is optimized using a weighted multi-objective function based on the optimized performance critic network, the optimized safety critic network and the Lagrangian safety penalty term to obtain a trained policy network.
[0029] In an exemplary embodiment, in accordance with the decision stage Environmental status Before using the historical policy network to interact with the environment, obtain a first action, and control the environment to execute the first action, the method further includes: initializing the initial policy network, the performance critic network, and the safety critic network.
[0030] In an exemplary embodiment, determining a Lagrangian safety penalty term using the optimized safety critic network specifically includes: determining an expected cumulative safety cost based on the optimized safety critic network; and determining a Lagrangian safety penalty term based on the expected cumulative safety cost using a Lagrangian multiplier.
[0031] To address the challenges of traditional proximal policy optimization algorithms in constrained optimization, such as passively responding to constraint violations and relying on manual penalty coefficients, this application proposes a dual-critic architecture for safety reinforcement learning. The method includes: initializing a policy network, a historical policy network, a performance critic network, and a safety critic network; collecting trajectory data through interaction between the historical policy network and the environment; calculating a generalized advantage estimation function to evaluate the long-term benefits of actions; constraining the policy update amplitude through a clipping mechanism; updating the parameters of the performance critic network and the safety critic network in parallel; constructing an augmented Lagrangian safety penalty term and dynamically adjusting the multiplier; and optimizing network parameters by integrating multiple objective functions. The system comprises a dual-critic module, a policy execution module, a Lagrangian optimizer, and a trajectory storage. This system implements proactive, preventive optimization under safety constraints in aircraft recovery scheduling, significantly improving fuel safety during landing.
[0032] like Figure 3 As shown, a specific optimization process of the aircraft recovery scheduling method in practical application is also provided, including the following steps: Step S1: Initialize the policy network , historical strategy network , Performance Critic Network and Security Critics Network , historical strategy network The first initialization is set to the policy network The same parameters, initialized policy network =Historical Strategy Network Among them: the network is used for aircraft recovery scheduling decision: strategy network According to the environmental status Generate recycling actions (Choose which plane to land), Performance Critics Network Evaluate Action Performance rewards, security critic network Evaluate Action security risks; inputs are all environmental states (including aircraft position, speed, and fuel quantity), the outputs are the action probability distribution , state value estimation and security value estimation ; Physical quantities include: are the trainable parameters of the policy network, are the trainable parameters of the performance critic network, are the trainable parameters of the security critic network, is the discount factor; For the environmental status.
[0033] Step S2: Through the historical policy network Interacting with the environment, in the aircraft recovery simulation environment, the input is the decision stage Environmental status (provided by sensors and simulation models), the output is action ;Environmental execution After that, feedback is immediately rewarded (based on recycling efficiency indicators) and safety costs (Based on safety indicators), thereby generating a state-action trajectory dataset ,The dataset is collected through multiple interactions and stored in the trajectory memory ,in: is the decision stage index, is the maximum number of decision stages in a single iteration, for The environmental status of the stage, for The actions performed by the stage, for Stage instant reward signal, for Stage safety cost signal.
[0034] Step S3: Based on the output of steps S1 and S2, i.e. the network parameters initialized in step S1 and the trajectory dataset collected in step S2 , computes the generalized advantage estimation function.
[0035] In an exemplary embodiment, the expression of the generalized advantage estimation function is: .
[0036] in, Indicates the decision stage The action advantage estimate, represents the index of the current decision stage, Indicates the future decision stage relative to The offset index of is the bias-variance trade-off coefficient, is the time series difference residual, for The environmental status of the stage, for The environmental status of the stage, Indexing for future decision-making stages, is the total number of decision stages in a single iteration. is the bias-variance trade-off coefficient, is the reward discount factor. For performance critic networks.
[0037] Step S4: Based on the output of S3, update the trainable parameters of the initial policy network through the clipping mechanism .
[0038] In an exemplary embodiment, the expression for updating the network parameters of the initial policy network based on the historical policy network using the clipping mechanism according to the generalized advantage estimation function is: .
[0039] in, represents the clipping loss function of the policy network, are the trainable parameters of the policy network, is the importance sampling weight, Decision-making stage The generalized advantage estimate of Indicates that Clip to interval , is the clipping threshold, is the expectation operator.
[0040] Step S5: Update performance critic network parameters : Use trace memory Trajectory data in Calculate the cumulative discount reward , and optimize based on the loss function .
[0041] .
[0042] .
[0043] in, Accumulate discount rewards; Estimate the loss function for value; To express The gradient operator.
[0044] Step S6: Update security critic network parameters : Use trace memory Trajectory data in Calculating the cumulative security cost , and optimize based on the loss function .
[0045] .
[0046] .
[0047] in, For the cumulative cost of security; To express Gradient operator is the learning rate. Estimate the loss function for security.
[0048] Step S7: Based on the output of S6, construct the updated security critic network The output security value is estimated and the Lagrange multiplier security penalty term is constructed.
[0049] Specific use Evaluated expected cumulative security costs , and update the Lagrange multiplier : .
[0050] .
[0051] in, For strategy The expected cumulative security cost under is the safety threshold, is the Lagrange multiplier learning rate, is the lower limit of the multiplier.
[0052] is the Lagrangian safety penalty term; Express Gradient operator of ; is the multiplier value of the current decision stage; is the updated multiplier value; is the multiplier learning rate.
[0053] Step S8: Integrate the outputs of S4-S7, update the policy network parameters, and optimize the parameters through the weighted multi-objective function: specifically using gradient renew . Indicates based on right Find the gradient operation, is the total loss function.
[0054] .
[0055] in, for The learning rate, , is the weight coefficient and .
[0056] Step S9: Loop steps S3-S8 until parameter update is reached Round; after the cycle ends, apply the trained policy network Conduct aircraft recovery scheduling: input real-time environmental status (such as aircraft position and speed), output the optimal action (specify the recovery of a certain aircraft) to achieve safe recovery.
[0057] In practical applications, the policy network in step S1 The output action distribution follows ,in is the state-dependent mean function, is the diagonal covariance matrix, is the trainable standard deviation parameter.
[0058] In practical applications, the number of trajectory acquisition cycles in step S2 is , parameter update rounds of steps S3-S8 .
[0059] In practical applications, the generalized advantage estimation function in step S3 satisfies: The above conditions are used to illustrate the mathematical properties of the function and are not necessarily processed; specifically: when Time degenerates into single-step timing difference error ,when is equivalent to the Monte Carlo return .
[0060] In practical applications, the security cost in step S2 Fuel consumption rate during aircraft recovery , landing gear stress peak , runway deviation distance satisfy: .
[0061] in, is the weight coefficient. 、 、 The environmental state is measured in real time by the aircraft sensors and the simulation model in the environmental interaction in step S2. A portion of the input network.
[0062] Trajectory memory of step S2 Adopting the priority experience replay mechanism, sampling probability satisfy: .
[0063] in, is the time series difference residual, is the smoothing constant.
[0064] In practical applications, step S8 synchronously updates the historical strategy network parameters after each iteration: .
[0065] like Figure 2 As shown, the dual critic architecture module in this application includes a parallel performance critic network and a safety critic network; a policy execution module, including a policy network and a historical policy network; an augmented Lagrangian optimizer, dynamically updating Parameters; trajectory memory, storage quad . Among them, the performance critic network and the safety critic network: share a state feature extraction layer, and have independent reward valuation fully connected layers and cost valuation fully connected layers.
[0066] The dual critic network described in this application uses a multi-layer perceptron (MLP) as its basic architecture in its specific implementation. Its overall design is a "shared input and feature extraction, separate value estimation output" structure. Specifically, the network receives a 5-dimensional state vector consisting of relative fuel quantity, relative integrity, relative mission priority, landing success rate, and aircraft status indicator as input. The input vector first enters a state feature extraction layer that is completely shared by the performance critic network and the safety critic network. The shared part consists of two hidden fully connected layers: the first layer maps the 5-dimensional input to a 256-dimensional vector, and the second layer further processes the 256-dimensional vector into a 128-dimensional state feature vector. Both layers use rectified linear units (ReLU) as activation functions. The 128-dimensional state feature vector is used as a unified intermediate representation and is then simultaneously fed into two value valuation networks with symmetrical structures but independent parameters. The performance critic network estimates the value function related to long-term performance returns through a hidden layer with 64 neurons (with ReLU activation function) and a single-neuron output layer with a linear activation function; at the same time, the security critic network uses the same network layer and activation function settings, but uses its independent weight parameters to process the same state feature vector and ultimately output a cost-value function related to long-term security risks. , Refers to the state Input into the performance critic network and Security Critics Network middle.
[0067] 1. Parallel Critic Network: Performance Critic Network Security Critic Network The shared state feature extraction layer outputs value estimates through independent reward valuation fully connected layers and cost valuation fully connected layers respectively.
[0068] 2. Policy execution module: including policy network and historical strategy networks .
[0069] 3. Augmented Lagrangian Optimizer: Dynamic Update parameter.
[0070] 4.Trace memory : Store quad-tuples .
[0071] This application lies at the intersection of deep reinforcement learning and flight scheduling. Specifically, it involves a dual-critic reinforcement learning architecture incorporating an augmented Lagrangian mechanism. This architecture is applicable to safety-constrained scenarios during aircraft recovery, particularly high-risk control tasks such as safety reinforcement learning under multiple constraints, including fuel consumption, landing gear stress, and runway deviation. By transforming safety constraints into differentiable optimization objectives, this approach addresses the technical limitations of traditional approaches in proactively preventing safety violations in dynamic environments.
[0072] The core of the project lies in constructing a secure reinforcement learning framework with a dual-critic architecture, addressing the aforementioned shortcomings through four innovative approaches: 1) A decoupled dual-critic design is employed, with the performance critic network focusing on predicting cumulative rewards and the safety critic network independently assessing cumulative safety costs, enabling precise quantification of safety risks. 2) A generalized advantage estimation function is introduced to fuse the temporal difference residuals from multiple decision stages, balancing the accuracy and stability of long-term reward estimates through a bias-variance trade-off. 3) An augmented Lagrangian optimizer is designed to transform hard safety constraints into a differentiable objective function, balancing performance and safety in real time through dynamically adjusted Lagrangian multipliers. 4) A prioritized experience replay mechanism is constructed to adaptively adjust data sampling weights based on temporal difference residuals. During the training phase, trajectories are collected through the historical policy network, and a clipping mechanism is used to constrain the policy update amplitude, ultimately resulting in an optimal scheduling policy under safety constraints.
[0073] The implementation process first initializes the policy network, historical policy network, performance critic network, and safety critic network, setting a discount factor γ = 0.95 and a clipping threshold ε = 0.2. In an aircraft recovery simulation, the historical policy network collects 500-2000 trajectory data. Each trajectory contains states, actions, rewards, and safety costs for 100-1000 decision stages. The safety cost is calculated in real time as a weighted sum of fuel consumption rate, landing gear stress peak, and runway deviation, with the weight coefficients set based on aircraft model parameters. The collected trajectories are stored in a trajectory memory, and sampling priorities are assigned based on the absolute value of the time-series difference residuals. The parameter update phase performs 3-10 iterations: When calculating the generalized advantage estimate, the bias-variance coefficient χ is set to 0.9 to balance the rewards of multiple decision stages. The performance critic network updates parameters by minimizing the prediction error of the cumulative discounted reward. The safety critic network optimizes the prediction accuracy of the cumulative safety cost. The Lagrange multiplier is dynamically adjusted based on the deviation of the expected cumulative safety cost from a threshold of 0.3. Finally, the policy network is updated by integrating the policy loss, critic loss, and safety penalty term, and then synchronized with the historical policy network. During the real-world deployment phase, pitch angle and throttle commands are output based on real-time status during landing. During the braking phase, braking torque is adjusted based on brake temperature. During the load distribution phase, hydraulic parameters are optimized based on landing gear pressure distribution. Experiments have shown that this method reduces safety violation rates by 58% and improves mission completion efficiency by 32%.
[0074] Based on the same inventive concept, embodiments of the present application also provide an aircraft recovery scheduling device for implementing the aforementioned aircraft recovery scheduling method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more aircraft recovery scheduling device embodiments provided below can be found in the above-described limitations of the aircraft recovery scheduling method and will not be further elaborated here.
[0075] In an exemplary embodiment, Figure 4 As shown, an aircraft recovery scheduling device is provided, including: an acquisition module for acquiring real-time environmental status; a scheduling module for performing aircraft recovery scheduling based on the real-time environmental status using a trained policy network to obtain the optimal action for aircraft recovery; the trained policy network is obtained by optimizing the network parameters of the initial policy network based on a historical policy network, a performance evaluator network, and a safety critic network, combined with a generalized advantage estimation function, a shearing mechanism, and a Lagrangian safety penalty term.
[0076] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, memory and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store aircraft recovery scheduling data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an aircraft recovery scheduling method is implemented.
[0077] Those skilled in the art will understand that Figure 5 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned method embodiments when executing the computer program.
[0078] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the above-mentioned method embodiments when executed by a processor.
[0079] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.
[0080] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0081] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0082] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.
[0083] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0084] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. An aircraft recovery scheduling method, characterized in that: The aircraft recovery scheduling method includes: Get real-time environment status; According to the real-time environmental state, a trained policy network is used to schedule aircraft recovery to obtain the optimal action for aircraft recovery; the trained policy network is obtained by optimizing the network parameters of the initial policy network based on a historical policy network, a performance evaluator network, and a safety critic network, combined with a generalized advantage estimation function, a shearing mechanism, and a Lagrangian safety penalty term.
2. The aircraft recovery scheduling method according to claim 1, characterized in that: The optimization process of the policy network specifically includes: According to the decision stage Environmental status Using the historical strategy network to interact with the environment, obtain a first action and control the environment to perform the first action; Obtaining an immediate reward and a safety cost of the environment performing the first action and generating a state-action trajectory dataset based on the immediate reward and the safety cost; Calculating a generalized advantage estimation function based on the state-action trajectory dataset based on the initial policy network and the performance evaluator network; Updating the network parameters of the initial policy network based on the historical policy network using a clipping mechanism according to the generalized advantage estimation function; Optimizing the network parameters of the performance critic network and the safety critic network respectively according to the state-action trajectory dataset; Determine the Lagrangian safety penalty term using the optimized safety critic network; According to the optimized performance critic network, the optimized safety critic network and the Lagrangian safety penalty term, the updated initial policy network is optimized using a weighted multi-objective function to obtain a trained policy network.
3. The aircraft recovery scheduling method according to claim 2, characterized in that: In the decision-making stage Environmental status Before using the historical strategy network to interact with the environment, obtain a first action, and control the environment to execute the first action, the method further includes: Initialize the initial policy network, performance critic network, and safety critic network.
4. The aircraft recovery scheduling method according to claim 1, characterized in that: The expression of the generalized advantage estimation function is: ; in: Indicates the decision stage The action advantage estimate, represents the time index of the current decision stage, Indicates the future decision stage relative to The offset index of is the reward discount factor, is the bias-variance trade-off coefficient, is the time series difference residual, is the maximum total number of decision stages in a single iteration.
5. The aircraft recovery scheduling method according to claim 2, characterized in that: The expression for updating the network parameters of the initial policy network based on the historical policy network using the clipping mechanism according to the generalized advantage estimation function is: ; in: represents the clipping loss function of the policy network, are the trainable parameters of the policy network, is the importance sampling weight, Decision-making stage The generalized advantage estimate of Indicates that Clip to interval , is the clipping threshold, is the expectation operator.
6. The aircraft recovery scheduling method according to claim 2, characterized in that: The optimized safety critic network is used to determine the Lagrangian safety penalty term, which includes: Determine the expected cumulative security cost based on the optimized security critic network; A Lagrangian safety penalty term is determined using a Lagrangian multiplier according to the expected cumulative safety cost.
7. An aircraft recovery scheduling device, characterized in that: The aircraft recovery scheduling device includes: Acquisition module, used to obtain real-time environment status; The scheduling module is used to use a trained policy network to perform aircraft recovery scheduling according to the real-time environmental state to obtain the optimal action for aircraft recovery; the trained policy network is obtained by optimizing the network parameters of the initial policy network based on the historical policy network, the performance evaluator network, and the safety critic network in combination with the generalized advantage estimation function, the shear mechanism, and the Lagrangian safety penalty term.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aircraft recovery scheduling method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the aircraft recovery scheduling method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the aircraft recovery scheduling method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Reinforcement learning model optimization method and device, storage medium and electronic equipment
CN113435606A
Model-based near-end strategy optimization method
CN113947022A
Track planning method and device for data collection of unmanned aerial vehicle, equipment and medium
CN114840021A
Intelligent fish flow field simulation control method, system and device and storage medium
CN116050304A
Constraint reinforcement learning-based communication perception joint optimization method and system
CN116367337A