Aircarrier path online planning method for air-land combined transportation task

Through the online planning method of carrier paths for land-based joint transportation tasks, the problem of accuracy and efficiency of carrier path planning is solved by using reinforcement learning networks and multi-agent algorithms, and efficient planning of multi-target delivery tasks is achieved.

CN120252728AActive Publication Date: 2025-07-04HARBIN INST OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510409353.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing carrier path planning methods have problems with low planning accuracy and efficiency, especially when facing multi-objective tasks, it is difficult to plan online.

Method used

The online planning method of carrier airplane paths for land-based joint transportation tasks is adopted. By obtaining the status of the current waypoint and inputting it into the optimized reinforcement learning network, the decision variables are determined using classification agents and sorting agents, combining air control, distance and range factors, an optimization model of carrier airplane delivery point is established, and a multi-agent reinforcement learning algorithm with maximum entropy is improved for path planning.

Benefits of technology

It improves the accuracy and efficiency of carrier route planning, can perform multiple delivery tasks at once, taking into account the advantages of air transportation and automobile transportation, avoids local optimal solutions, and achieves the optimal planned route.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120252728A_ABST
    Figure CN120252728A_ABST
Patent Text Reader

Abstract

The invention discloses an aerial carrier path online planning method for an air-land combined transportation task, and relates to the technical field of path planning. The invention aims to solve the problem that an existing aerial carrier path planning method is low in planning accuracy and efficiency. The method comprises the steps of obtaining a state of a current waypoint, inputting the state of the current waypoint into an optimized reinforcement learning network, and obtaining a decision variable of the current waypoint; and determining an aerial carrier planning path according to the decision variables of all the waypoints. According to the method, the waypoints are graded, and the advantages of air transportation box automobile transportation are considered, so that the transportation efficiency and the accuracy of optimal route planning are improved. According to the invention, the aerial carrier can execute a plurality of delivery tasks at one time, and the transportation efficiency can be improved. The method is used for planning the aerial carrier transportation route.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of route planning, and particularly to an on-line path planning method for an aircraft carrier facing the air-land combined transportation task. Background Art

[0002] With the rapid development of the low-altitude economy and intelligent Internet of Things, the timeliness and economy indicators of cargo transportation have become increasingly important. Against this background, an unmanned aerial transporter that is launched by an aircraft carrier and adopts a general design of continuous power, large lift-to-drag ratio, and large internal cabin has emerged as the times require. Compared with traditional fixed-wing unmanned aerial vehicles, this kind of aircraft is faster, flies farther, and has a larger load. If an aircraft carrier can carry multiple aerial transporters at the same time, it can complete multiple transportation tasks after one takeoff. If further combined with land transportation means, it can realize the complementarity of air-land transportation, transport more types of goods, and significantly reduce costs and improve efficiency.

[0003] At present, for the multi-objective launch mission planning problem of an aircraft carrier, the existing technology first gives the effectiveness index corresponding to the launch plan through effectiveness evaluation, including the analytic hierarchy process, fuzzy evaluation method, etc. However, this process is greatly affected by subjective factors of people, and the solution efficiency of this method is limited, resulting in low accuracy and efficiency of path planning, and it is not suitable for on-line planning. Then, a non-linear programming algorithm is used to solve the optimal launch strategy, while the linear programming algorithm is insufficient in dealing with actual task interference and emergencies, such as launch interference, sudden weather, detection interference, etc., resulting in low accuracy of obtaining the optimal planned route, and the neural network model is prone to falling into local optimum when solving such problems, making it difficult to obtain the optimal planned path. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem that the existing aircraft carrier path planning method still has low planning accuracy and efficiency, and an on-line path planning method for an aircraft carrier facing the air-land combined transportation task is proposed.

[0005] The on-line path planning method for an aircraft carrier facing the air-land combined transportation task is specifically as follows:

[0006] S1. Obtain the state of the current waypoint, input the state of the current waypoint into the optimized reinforcement learning network, and obtain the decision variable of the current waypoint;

[0007] The state of the waypoint includes: the label of the current waypoint, the remaining number of aerial transporters, the air traffic control restriction of the current waypoint, the launch difficulty of the current waypoint, the meteorological condition of the current waypoint, the communication condition of the current waypoint, the remaining flight range, and the historical launch waypoint;

[0008] The reinforcement learning network includes: a classification agent and a sorting agent;

[0009] The classification agent is used to set waypoints as secondary waypoints or tertiary waypoints, and at the same time determine whether to go from waypoint a at level n to waypoint b at level n+1;

[0010] The sorting intelligent machine is used to determine the order in which the carrier aircraft passes through the secondary waypoints;

[0011] The decision variable of the waypoint is 0 or 1; a decision variable of 0 means not choosing to go from waypoint a to waypoint b; the decision variable of 1 means choosing to go from waypoint a to waypoint b;

[0012] where n takes 1 or 2;

[0013] Among them, the level 1 waypoint is the starting point of the carrier aircraft, the level 2 waypoint is the target point for the carrier aircraft to drop, and the level 3 waypoint is the waypoint that the carrier aircraft does not pass by;

[0014] S2. Determine the planned path of the carrier aircraft according to the decision variables of all waypoints.

[0015] Furthermore, the optimized reinforcement learning network is obtained in the following way:

[0016] Step 1. Establish an optimization model for the carrier aircraft drop point, specifically:

[0017] min J = J1 + J2 + J3

[0018] where J1 is the air traffic control received by the carrier aircraft at the waypoint, J2 is the driving distance of the vehicle, J3 is the total flight range of the carrier aircraft, and J is the total optimization index.

[0019] Step 2. Train the multi-agent reinforcement learning network based on the optimization model for the carrier aircraft drop point until the multi-agent reinforcement learning network converges to obtain the optimized reinforcement learning network.

[0020] Furthermore, the air traffic control J1 received by the carrier aircraft at the waypoint is as follows:

[0021]

[0022] where n is the level label of the waypoint, N is the total number of waypoint levels, M n is the number of waypoints at level n, F n ∈[0,1], F n is the restriction received by the carrier aircraft at the waypoint at level n;

[0023] The value of F n is the scoring value of the current restriction; the current restriction is the flight time interval, flight altitude, flight range or flight speed;

[0024] When the restriction is the flight time interval, F nThe larger the value, the smaller the flight time interval; when restricted to flight altitude, F n The larger the value, the smaller the flight altitude range; when restricted to flight range, F n The larger the value, the smaller the flight range; when restricted to flight speed, F n The larger the value, the smaller the maximum flight speed value; if subject to multiple restrictions, take the average of the scoring values for each restriction.

[0025] Further, the driving distance J2 of the vehicle is as follows:

[0026]

[0027] where i and j are the labels of waypoints, M n is the number of n-level waypoints, P n i,j ∈[0,1], P n i,j is the vehicle transportation efficiency from n-level waypoint i to n+1-level waypoint j, is the distance from n-level waypoint i to n+1-level waypoint j, is a decision variable;

[0028] The vehicle transportation efficiency is the transportation speed or the load capacity.

[0029] Further, the total flight range J3 of the carrier aircraft is as follows:

[0030]

[0031] where U n+1,i ∈[0,1], U n+1,i is the delivery difficulty of n-level waypoint i, represents the remaining number of aircraft transporters of the carrier aircraft;

[0032] The delivery difficulty is the scoring value for the wind force, and the greater the wind force, the greater the delivery difficulty value.

[0033] Further, in step two, training the multi-agent reinforcement learning network based on the carrier aircraft delivery point optimization model until the multi-agent reinforcement learning network converges to obtain the optimized reinforcement learning network, specifically:

[0034] Step 2-1: Initialize the agents and randomly initialize the reinforcement learning network for each agent;

[0035] Step 2-2: Set the initial state for the agents;

[0036] The state set S includes: the label of the current waypoint, the remaining number of aircraft, the air traffic control restrictions at the current waypoint, the delivery difficulty at the current waypoint, the communication conditions at the current waypoint, the remaining flight range, and the historical delivery waypoints;

[0037] Step 23: Select and execute an action for each agent, and then calculate the state-action value function Q(s,a);

[0038] Step 24: The agent trains the reinforcement learning network according to the KL-divergence based on the state, action, reward, and new state until the reinforcement learning network converges, obtains the optimal reinforcement learning network parameters and the optimal policy, and obtains the trained reinforcement learning network.

[0039] Furthermore, the state-action value function Q(s j' ,a j' ) is specifically:

[0040] Q(s j' ,a j' ) = Ε[R(s j' ,a j' )|S t = s j' ,A t = a j' )

[0041] where Q(s j' ,a j' ) is the state-action value function of action a j' under state s j' , R(s j' ,a j' ) is the reward for executing action a j' under state s j' , s j' is the joint state at the j'-th iteration, a j' is the action at the j'-th iteration, S t is the state of the agent at time t, and A t is the action of the agent at time t.

[0042] Furthermore, the agent in Step 24 trains the reinforcement learning network according to the KL-divergence based on the state, action, reward, and new state, and uses the following objective function:

[0043]

[0044] R(s j' ,a j' ) = J

[0045] where K His a hyperparameter, N' is the total number of agents, H represents entropy, and π i' is the current policy of agent i', and H(π i' (·|s j' )) is the policy entropy when the agent's state is s j' . ζ j' is the discount factor for the j'-th iteration, π is the set of agent policies, and J(π(·|s j' )) is the objective function, and A is the set of agent actions.

[0046] Furthermore, the policy of the agent is updated using the following formula:

[0047]

[0048] where D KL is the KL-divergence, π j'+1 is the policy for the (j'+1)-th step, and i' p is the specific serial number of agent i' in the agent sorting sequence. Q represents the state-action value function, Z is the partition function, is the policy of the agent with serial number i' j' under state s p , is the set of all other actions except action , is the action of the agent with serial number i' p .

[0049] Furthermore, the conditions for the convergence of the reinforcement learning network are as follows:

[0050]

[0051] where is the optimal policy of all other agents except agent i', and π i' is the optimal policy of agent i', is the joint objective function when agent i' adopts policy π i' and other agents adopt the corresponding optimal policies. a -i' is the set of all other actions in the action space except action a i' , and b i' is a hyperparameter, and A i' is the set of actions of agent i'.

[0052] The beneficial effects of the present invention are:

[0053] The present invention proposes an on-line path planning method for an aircraft carrier facing the delivery task of an air transporter. The present invention establishes an optimization model for the aircraft carrier delivery point for the delivery task of multiple air transporters, and combines the multi-agent reinforcement learning algorithm improved by the maximum entropy based on the optimization model of the aircraft carrier delivery point to enable the aircraft carrier to start from any airport, deliver air transporters to all ground targets, and complete the on-line planning of the optimal delivery point. When establishing the optimization model of the aircraft carrier delivery point, the present invention considers air traffic control factors, distance factors and range factors, and divides the targets into secondary waypoints and tertiary waypoints, taking into account the combined transportation advantages of air transporters and automobiles, enabling the aircraft carrier to perform multiple delivery tasks at one time, thereby improving the transportation efficiency and the accuracy of the optimal planned route. The present invention realizes the path planning of the aircraft carrier based on the multi-agent reinforcement learning algorithm improved by the maximum entropy, improves the planning efficiency, and at the same time avoids falling into local optimal solutions, so as to obtain the optimal planned route. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic diagram of the optimization model for the aircraft carrier delivery point of air-land combined transportation;

[0055] Figure 2 It is a flowchart of the maximum entropy agent reinforcement learning algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The first specific embodiment: The specific process of the on-line path planning method for the aircraft carrier facing the air-land combined transportation task in this embodiment is as follows:

[0057] S1. Obtain the state of the current waypoint, input the state of the current waypoint into the optimized reinforcement learning network, and obtain the decision variable of the current waypoint;

[0058] The state of the waypoint includes: the label of the current waypoint, the remaining number of air transporters, the air traffic control limit of the current waypoint, the delivery difficulty of the current waypoint, the meteorological conditions of the current waypoint, the communication conditions of the current waypoint, the remaining range, and the historical delivery waypoint;

[0059] The reinforcement learning network includes: a classification agent and a sorting agent;

[0060] The classification agent is used to set the waypoint as a secondary waypoint or a tertiary waypoint, and at the same time determine whether to go from waypoint a at level n to waypoint b at level n+1;

[0061] The sorting intelligent machine is used to determine the order of the aircraft carrier passing through the secondary waypoints;

[0062] The decision variable being 0 means not choosing to go from waypoint a to waypoint b;

[0063] Wherein, waypoint a is a waypoint at level n, and waypoint b is a waypoint at level n+1;

[0064] The decision variable being 1 indicates choosing to go from waypoint a to waypoint b;

[0065] The optimized reinforcement learning network is obtained in the following manner:

[0066] Step 1, as Figure 1 shown, establish an optimization model for the carrier release point:

[0067] min J = J1 + J2 + J3

[0068]

[0069] where J1 is the air traffic control received by the carrier at the waypoint, J2 is the driving distance of the vehicle, J3 is the total flight range of the carrier, J is the total optimization index, n is the level label of the waypoint, N is the total number of waypoint levels, M n is the number of waypoints at level n, F n is the restriction received by the carrier at the waypoint at level n, i, j are the labels of the waypoints, P n i,j is the vehicle transportation efficiency from waypoint i at level n to waypoint j at level (n + 1), is the distance from waypoint i at level n to waypoint j at level (n + 1), is the decision variable, taking 0 or 1; when waypoint i at level n reaches waypoint j at level (n + 1), Y n i,j = 1, otherwise Y n i,j = 0; U n+1,i is the release difficulty of waypoint i at level n, represents the remaining number of aircraft transporters of the carrier;

[0070] The vehicle transportation efficiency P n i,j is the normalized transportation speed or cargo capacity; P n i,j ∈[0, 1];

[0071] The release difficulty U n+1,i is the score given by current technicians for the wind force; U n+1,i ∈[0, 1]; U n+1,i The larger the value of U, the stronger the wind; for example: winds of level 1 - 3 are relatively weak, with a first - grade score, winds of level 4 - 6 are slightly stronger, with a first - grade score, and winds of level 7 - 10 are very strong, with a first - grade score.

[0072] where M1 = 1 and F1 = 0; the carrier is a waypoint at level 1, and the restriction F received by the carrier nF is the score of the flight time interval, flight altitude, flight range or flight speed; n ∈[0,1], when restricted to the flight time interval, F n The larger the value, the smaller the flight time interval; when the flight altitude is limited, F n The larger the value, the smaller the flight altitude range. When limited to the flight range, F n The larger the value, the smaller the flight range. When limited to flight speed, F n The larger the value, the smaller the maximum flight speed. If there are multiple restrictions, the average of the multiple restriction scores is taken.

[0073] Automobile transportation efficiency mainly refers to transportation speed or cargo capacity, which is a relative value and dimensionless. If normalized to 0-1, this value may be 0.11 or 0.98.

[0074] In this step, only n is added to the formula J2 to take the value in the range of 1 to N-1, and the case of n=N is not considered;

[0075] In this step, the rules for air transport aircraft placement are set as follows:

[0076] 1) The carrier aircraft departs from an airport and drops an air transport vehicle to a target along the way;

[0077] 2) The carrier aircraft is the starting point of the entire track, called the first-level waypoint;

[0078] 3) Divide the track into secondary and tertiary waypoints according to the target;

[0079] 4) The aircraft only passes through the secondary waypoint;

[0080] 5) The carrier aircraft drops an air transport vehicle to a target near a secondary waypoint, and the transport from the secondary waypoint to the tertiary waypoint is completed by a vehicle, so the carrier aircraft does not directly pass by the tertiary waypoint;

[0081] 6) Each secondary waypoint may have zero or more tertiary waypoints.

[0082] Step 2: Figure 2 As shown, based on the aircraft drop point optimization model, the multi-agent reinforcement learning network based on maximum entropy improvement is trained until the multi-agent reinforcement learning network converges to obtain the optimized reinforcement learning network:

[0083] Step 21: Initialize the classification agent and sorting agent in the reinforcement learning network:

[0084] The classification agent is used to set the waypoint as a secondary waypoint or a tertiary waypoint, and decide whether to go from the n-level waypoint a to the n+1-level waypoint b;

[0085] The sorting intelligent machine is used to determine the order in which the carrier aircraft passes through the secondary waypoints;

[0086] The parameters of the reinforcement learning network include: joint policy, hyperparameters, discount factor, number of agents;

[0087] Step 22: Set the initial state for the agent:

[0088] The state set S includes: the label of the current waypoint, the remaining number of aircraft, the air traffic control restrictions at the current waypoint, the delivery difficulty at the current waypoint, the meteorological conditions at the current waypoint, the communication conditions at the current waypoint, the remaining flight distance, and the historical delivery waypoints;

[0089] Step 23: Select and execute an action for each agent, and then calculate the state-action value function Q(s j' ,a j' ):

[0090] Q(s j' ,a j' ) = Ε[R(s j' ,a j' )|S t = s j' ,A t = a j' )

[0091] where Q(s j' ,a j' ) is the state-action value function of action a j' in state s j' , R(s j' ,a j' ) is the reward for executing action a j' in state s j' , s j' is the joint state at the j'-th iteration, a j' is the action at the j'-th iteration, S t is the state of the agent at time t, and A t is the action of the agent at time t;

[0092] The action is a decision variable

[0093] Step 24: The agent trains the reinforcement learning network according to the KL-divergence based on the state, action, reward, and new state until the reinforcement learning network converges, obtains the optimal reinforcement learning network parameters and the optimal policy, and obtains the trained reinforcement learning network;

[0094] The multi-agent training model is specifically:

[0095]

[0096] Among them, J(π(·|s j' )) is the objective function, j' is the iteration number label, s j' is the joint state in the j'-th step of iterative training, R is the joint reward function, a j' is the joint action in the j'-th step of iterative training, π is the joint policy, A is the finite joint action space, S is the finite joint state space, R(s j' ,a j' ) is the reward when the agent state is s j' and the action a j' is executed, ζ j' is the discount factor for the j'-th step of iteration, and P is the state transition matrix;

[0097] The difficulty of multi-agent training lies in coordinating the single-agent and multi-agent optimization processes to make the objective function monotonically increasing and reach the global optimum. For example, a single agent may "sacrifice" itself for this goal, gather towards other agents and fall into a local optimum. Therefore, a maximum entropy term is introduced to improve the training objective function, and each agent is made to approximate the optimal policy according to the KL-divergence to achieve joint policy optimization:

[0098]

[0099] R(s j' ,a j' ) = J

[0100] where K H is a hyperparameter, N' is the total number of agents, H represents entropy, π i' is the current policy of agent i', H(π i' (·|s j' )) is the policy entropy when the agent state is s j' , and A is the set of agent actions;

[0101] The policy of the agent is updated in the following way:

[0102] Let each agent continuously approximate the optimal policy according to the KL-divergence. Denote the sequence i' 1:n as the sorting sequence of multi-agents, and i' p (p = 1, 2…, n) as the specific serial number of agent i'. Then the new policy of the agent in the next step, i.e., the (j'+1)-th step, is:

[0103]

[0104] where D KL is the KL-divergence, π j'+1 is the policy in the (j'+1)-th step, and i' pis the specific serial number of agent i' in the agent sorting sequence, Q represents the state-action value function, and Z is the partition function that normalizes the data distribution. is the state s j' with serial number i' p of the agent, is all other actions except the action ; is the action of the agent with serial number i' p .

[0105] The convergence conditions of the reinforcement learning network are as follows:

[0106] If all agents cannot improve the joint goal by adjusting their own strategies, it is considered that the joint strategy reaches a response equilibrium, denoted as π * :

[0107]

[0108] where is the optimal strategy of other agents except agent i', π is the set of agent strategies, and π i' is the optimal strategy of agent i';

[0109] The optimal strategy of agent i' is:

[0110]

[0111] where a -i' is all other action sets in the action space except the action a i' , b i' is a hyperparameter, and A i' is the action set of agent i'.

[0112] S2. Determine the carrier aircraft planned path according to the decision variables of all waypoints;

[0113] The carrier aircraft planned path is the path where the decision variables of all waypoints in the path are 1.

[0114] The optimization variables of the present invention can be classified into two categories. One is the classification problem, that is, dividing the secondary and tertiary waypoints according to the target. The other is the sorting problem, that is, determining the order of flying through each secondary waypoint. For these two types of problems, the present invention uses two reinforcement learning agents to solve them respectively, forming a multi-agent reinforcement learning architecture. The multi-agent executes joint actions according to the joint policy and state, obtains the discounted joint reward, and iteratively optimizes the joint policy with the maximum expected cumulative return as the goal until all agents can no longer improve the objective function. The present invention uses different reinforcement learning agents to solve the multi-class non-linear programming problems involved in the optimization of secondary / tertiary waypoints respectively, forming a multi-agent reinforcement learning architecture. At the same time, a maximum entropy multi-agent learning improvement strategy is proposed to increase the randomness of training and avoid falling into local optima.

Claims

1. An on-line path planning method for the carrier aircraft facing the air-land combined transportation mission, characterized in that The specific process of the method is as follows: S1. Obtain the status of the current waypoint, input the status of the current waypoint into the optimized reinforcement learning network, and obtain the decision variable of the current waypoint; The status of the waypoint includes: the label of the current waypoint, the remaining number of aircraft, the air traffic control restriction of the current waypoint, the delivery difficulty of the current waypoint, the meteorological conditions of the current waypoint, the communication conditions of the current waypoint, the remaining voyage, and the historical delivery waypoints; The reinforcement learning network includes: a classification agent and a sorting agent; The classification agent is used to set the waypoint as a secondary waypoint or a tertiary waypoint, and at the same time determine whether to go from waypoint a at level n to waypoint b at level n + 1, and output a decision variable; The sorting intelligent machine is used to determine the order in which the carrier aircraft passes through the secondary waypoints; The decision variable of the waypoint is 0 or 1; the decision variable of 0 means not to choose to go from waypoint a to waypoint b; the decision variable of 1 means to choose to go from waypoint a to waypoint b; Wherein, n takes 1 or 2; the first-level waypoint is the starting point of the carrier aircraft, the second-level waypoint is the delivery target point of the carrier aircraft, and the third-level waypoint is the waypoint that the carrier aircraft does not pass by; S2. Determine the planned path of the carrier aircraft according to the decision variables of all waypoints.

2. The online path planning method for the carrier aircraft for air-land combined transportation tasks according to claim 1, wherein: The optimized reinforcement learning network is obtained through the following method: Step 1. Establish an optimization model for the carrier aircraft delivery point, specifically: min J = J1 + J2 + J3 Wherein, J1 is the air traffic control suffered by the carrier aircraft at the waypoint, J2 is the driving distance of the vehicle, J3 is the total voyage of the carrier aircraft, and J is the total optimization index. Step 2. Train the multi-agent reinforcement learning network based on the optimization model of the carrier aircraft delivery point until the multi-agent reinforcement learning network converges to obtain the optimized reinforcement learning network.

3. The online path planning method for the carrier aircraft for air-land combined transportation tasks according to claim 2, wherein: The air traffic control J1 suffered by the carrier aircraft at the waypoint is as follows: Among them, n is the level label of the waypoint, N is the total number of waypoint levels, and M n is the number of waypoints at level n, and F n ∈[0,1], and F n is the restriction imposed on the carrier aircraft at the waypoint of level n; The described F n value is the scoring value of the currently imposed restriction; the currently imposed restriction is the flight time interval, flight altitude, flight range, or flight speed; When restricted to a flight time interval, the larger the value of F n the smaller the flight time interval; when restricted to a flight altitude, the larger the value of F n the smaller the flight altitude range; when restricted to a flight range, the larger the value of F n the smaller the flight range; when restricted to a flight speed, the larger the value of F n the smaller the maximum flight speed value; if there are multiple restrictions, the average of the scoring values for each restriction is taken.

4. The online path planning method for the carrier aircraft for air-land combined transportation tasks according to claim 3, characterized in that: The driving distance J2 of the vehicle is as follows: where i and j are the labels of waypoints, and M n is the number of waypoints of level n, is the vehicle transportation efficiency from waypoint i of level n to waypoint j of level n + 1, is the distance from waypoint i of level n to waypoint j of level n + 1, is a decision variable; The transportation efficiency of the vehicle is the transportation speed or the cargo capacity.

5. The online path planning method for the carrier aircraft for the air-land combined transportation mission according to claim 4, wherein: The total voyage J3 of the carrier aircraft is as follows: Among them, U n+1,i ∈[0,1], U n+1,i is the delivery difficulty of the nth waypoint i, represents the number of remaining aircraft carriers; The delivery difficulty is the scoring value for the wind force, and the greater the wind force, the greater the delivery difficulty value.

6. The online path planning method for the carrier aircraft for air-land combined transportation tasks according to claim 5, wherein: In step 2, training the multi-agent reinforcement learning network based on the optimization model of the carrier aircraft delivery point until the multi-agent reinforcement learning network converges to obtain the optimized reinforcement learning network is specifically as follows: Step 21. Initialize the classification agent and the sorting agent in the reinforcement learning network; Step 22. Set the initial state for the agent; The state set S includes: the label of the current waypoint, the remaining number of aircraft, the air traffic control restriction of the current waypoint, the delivery difficulty of the current waypoint, the delivery difficulty of the current waypoint, the communication conditions of the current waypoint, the remaining voyage, and the historical delivery waypoints; Step 23. Select and execute an action for each agent, and then calculate the state-action value function Q(s, a); Step 24. The agent trains the reinforcement learning network according to the KL-divergence based on the state, action, reward, and new state until the reinforcement learning network converges, obtains the optimal reinforcement learning network parameters and the optimal policy, and obtains the trained reinforcement learning network.

7. The online path planning method for the carrier aircraft for the air-land combined transportation mission according to claim 6, characterized in that: The state-action value function Q(s j' , a j' ) is specifically as follows: Q(s j' , a j' ) = E[R(s j' , a j' ) | S t = s j' , A t = a j' ​ where Q(s j' , a j' ) is the state - action value function for action a j' in state s j' , R(s j' , a j' ) is the reward for executing action a j' in state s j' , s j' is the joint state at the j'-th iteration, a j' is the action at the j'-th iteration, S t is the state of the agent at time t, A t is the action of the agent at time t.

8. The online path planning method for the carrier aircraft for the air-land combined transportation mission according to claim 7, wherein: The agent in step 24 trains the reinforcement learning network according to the KL-divergence based on the state, action, reward, and new state, using the following objective function: Among them, K H is a hyperparameter, N' is the total number of agents, H represents entropy, and π i' is the current policy of agent i', and H(π i' (·|s j' )) is the policy entropy when the agent's state is s j' . ζ j ' is the discount factor for the j'-th iteration, π is the set of agent policies, and J(π(·|s j' )) is the objective function, and A is the set of agent actions.

9. The online path planning method for the carrier aircraft for the air-land combined transportation mission according to claim 8, wherein: The policy of the agent is updated using the following formula: Among them, D KL is the KL-divergence, π j'+1 is the policy at the (j'+1)-th step, and i' p is the specific serial number of agent i' in the agent sorting sequence. Q represents the state-action value function, and Z is the partition function. is the state s j' at which the serial number is i' p of the agent, is the set of all other actions except the action , is the action of the agent with serial number i' p .

10. The online path planning method for the carrier aircraft facing the air-land combined transportation mission according to claim 9, characterized in that: The conditions for the convergence of the reinforcement learning network are as follows: Among them, is the optimal strategy of other agents except agent i', π i' is the optimal strategy of agent i', is the situation where agent i' adopts the strategy π i' The joint objective function when other agents adopt the corresponding optimal strategies, a -i' is all other action sets in the action space except the action a i' outside, b i' is a hyperparameter, A i' is the action set of agent i'.

Citation Information

Patent Citations

  • Air-ground joint trajectory optimization and resource allocation method based on reinforcement learning

    CN114819785A

  • National park unmanned aerial vehicle patrol path optimization method based on reinforcement learning

    CN115574826A

  • Joint cache decision and trajectory optimization method under unmanned aerial vehicle assisted Internet of Vehicles

    CN116847293A

  • Multi-agent route planning method and device based on artificial potential field and PPO

    CN118670400A

  • Method for optimizing collaborative delivery path of heterogeneous system under road network and energy consumption constraints

    CN119047673A