Unmanned airship intelligent flight planning method, device, equipment and medium
By converting the trajectory planning of the stratosphere airship into the Markov decision-making process and using deep neural networks for reinforcement learning, the problems of low accuracy and non-smooth trajectory in traditional planning methods are solved, and high-precision and smooth trajectory planning are achieved.
Patent Information
- Application Number
- CN202510208036.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional unmanned airship flight planning methods have problems of low trajectory planning accuracy and unsmooth trajectory. Especially in the operating environment of stratospheric airships, changes in wind fields and clouds affect the flight safety of airships.
The trajectory planning process of the stratosphere airship is converted into the Markov decision-making process, the state space, action space and reward functions are defined, and based on the iterative update of the policy network and the value network, the strategy network and the value network are constructed using deep neural networks to carry out policy-based reinforcement learning algorithms for intelligent training.
High-precision trajectory planning is realized, and the generated trajectory is relatively smooth, which can effectively solve the problems of low accuracy and non-smooth trajectory in traditional methods, and perform data acquisition and agent training in simulation environment, improving data utilization.
Smart Images

Figure CN120066077A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of stratospheric airship flight trajectory planning, and particularly to an intelligent flight planning method, device, equipment and medium for an unmanned airship based on wind field-temperature field data. Background Art
[0002] The operating environment of a stratospheric airship has a horizontal wind field with randomly varying wind direction and speed, and due to its own dynamic characteristics, the airship is easily affected by the wind field. In addition, high-altitude cold clouds will also affect the thermodynamic equilibrium of the gas inside the airship and affect the flight safety of the airship. The mechanism is as follows: at night, the cloud layer reduces the thermal radiation received by the airship through the shielding effect on the surface radiation, and the temperature of the gas inside the airship decreases; during the day, since the albedo of the cloud layer is higher than that of the surface, the airship receives more thermal radiation, and the temperature of the gas inside the airship increases. Therefore, the trajectory planning of a stratospheric airship needs to comprehensively consider the wind field and cloud distribution.
[0003] Traditional unmanned airship flight planning methods include: the combination algorithm of A* algorithm and Rapidly-exploring Random Trees (RRT), reinforcement learning algorithm and planning method based on supervised learning. However, traditional unmanned airship flight planning methods have the problems of low accuracy of trajectory planning and insufficient smoothness of the planned trajectory. Summary of the Invention
[0004] The purpose of the present application is to provide an intelligent flight planning method, device, equipment and medium for an unmanned airship, which can improve the accuracy of trajectory planning and generate a smooth trajectory of the stratospheric airship.
[0005] To achieve the above purpose, the present application provides the following solutions:
[0006] In the first aspect, the present application provides an intelligent flight planning method for an unmanned airship, including:
[0007] Converting the trajectory planning process of the stratospheric airship into a Markov decision process, and defining a state space, an action space and a reward function; the state space includes two-dimensional wind field data, two-dimensional temperature field data, local time, target area position, current position coordinates, speed and remaining energy of the stratospheric airship; the action space includes the speed increments of the stratospheric airship in the meridional and zonal directions.
[0008] Based on the Markov decision process, training a policy network and a value network, and iteratively updating the flight state of the stratospheric airship by using the policy network and the value network until the stratospheric airship reaches the target area position, so as to obtain the trajectory planning result of the stratospheric airship from the starting position to the target area position; the policy network and the value network are deep neural networks.
[0009] In the second aspect, the present application provides an intelligent flight planning device for an unmanned airship, including:
[0010] A trajectory planning process conversion module is used to convert the trajectory planning process of a stratospheric airship into a Markov decision process, and define a state space, an action space and a reward function; the state space includes two-dimensional wind field data, two-dimensional temperature field data, local time, target area position, current position coordinates, speed and remaining energy of the stratospheric airship; the action space includes the speed increments of the stratospheric airship in the meridional and latitudinal directions.
[0011] An airship trajectory planning module is used to train a policy network and a value network based on the Markov decision process, and iteratively update the flight state of the stratospheric airship by using the policy network and the value network until the stratospheric airship reaches the target area position, so as to obtain the trajectory planning result of the stratospheric airship from the starting position to the target area position; the policy network and the value network are deep neural networks.
[0012] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the above-mentioned intelligent flight planning method for an unmanned airship.
[0013] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned intelligent flight planning method for an unmanned airship is implemented.
[0014] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0015] The present application provides an intelligent flight planning method, device, equipment and medium for an unmanned airship, which converts the trajectory planning process of a stratospheric airship into a Markov decision process, and defines a state space, an action space and a reward function; the state space includes two-dimensional wind field data, two-dimensional temperature field data, local time, target area position, current position coordinates, speed and remaining energy of the stratospheric airship; the action space includes the speed increments of the stratospheric airship in the meridional and latitudinal directions; based on the Markov decision process, a policy network and a value network are trained, and the flight state of the stratospheric airship is iteratively updated by using the policy network and the value network until the stratospheric airship reaches the target area position, so as to obtain the trajectory planning result of the stratospheric airship from the starting position to the target area position. The present application constructs a policy network and a value network through deep neural networks, and uses a policy-based reinforcement learning algorithm for agent training, which has a continuous state space, can achieve a higher-precision trajectory planning, and the trajectory is relatively smooth, solving the problems of low precision and non-smooth trajectory of traditional planning methods. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0017] Figure 1 It is an application environment diagram of an intelligent flight planning method for an unmanned airship in an embodiment of the present application;
[0018] Figure 2 It is a flowchart of an intelligent flight planning method for an unmanned airship provided in an embodiment of the present application;
[0019] Figure 3 It is a flowchart of the intelligent agent training process provided in an embodiment of the present application;
[0020] Figure 4 It is a schematic diagram of the value network structure provided in an embodiment of the present application;
[0021] Figure 5 It is a schematic diagram of the policy network structure provided in an embodiment of the present application;
[0022] Figure 6 It is a test effect diagram of trajectory planning under headwind conditions provided in an embodiment of the present application;
[0023] Figure 7 It is a test effect diagram of the stratospheric airship planning to avoid cold clouds provided in an embodiment of the present application;
[0024] Figure 8 It is an update relationship diagram of the policy network, value network and target value network provided in an embodiment of the present application;
[0025] Figure 9 It is a schematic diagram of the functional modules of an intelligent flight planning device for an unmanned airship provided in an embodiment of the present application;
[0026] Figure 10 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Detailed implementation manners
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0028] A method for stratospheric airship route planning in a dynamic environment, which combines the A* algorithm and the Rapidly-exploring Random Tree (RRT) algorithm, realizes the trajectory planning and obstacle avoidance of the airship in a dynamic environment, and optimizes the wind speed superposition value and the trajectory length of waypoints. However, it does not consider the influence of the wind field on the movement of the airship, and the airspace must be gridified, which affects the accuracy of trajectory planning. A method for intelligent flight planning of an unmanned airship based on meteorological data uses a reinforcement learning algorithm to design a deep neural network model. The agent learns the optimal strategy in the interaction with the wind field environment based on meteorological data, realizes the trajectory planning in a dynamic wind field, and optimizes the energy consumption during the flight process. However, it does not consider the influence of clouds, and the reinforcement learning method based on the value function it uses has a discrete state space, and the planned trajectory is not smooth enough. A method for intelligent trajectory planning of a stratospheric airship based on a large model uses a supervised learning planning method. It uses one-dimensional data fused from wind field data, temperature field data, and airship state data, as well as historical data of decision-making actions for supervised learning to train the large model to implement intelligent trajectory planning of the stratospheric airship. Although it considers the wind field and clouds, on the one hand, its supervised learning training method requires a large amount of historical data and has a high training cost. On the other hand, the learning goal is to fit historical actions and cannot achieve optimal trajectory planning.
[0029] In view of the above problems, the present application proposes a method, device, equipment and medium for intelligent flight planning of an unmanned airship.
[0030] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] The method for intelligent flight planning of an unmanned airship provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can send the airship flight planning request to be processed to the server 104. After receiving the airship flight planning request, for the airship flight planning request, the server 104 converts the trajectory planning process of the stratospheric airship into a Markov decision process, defines the state space, action space and reward function, trains the policy network and value network based on the Markov decision process, and uses the policy network and value network to iteratively update the flight state of the stratospheric airship until the stratospheric airship reaches the target area position, so as to obtain the trajectory planning result of the stratospheric airship from the starting position to the target area position. The server 104 can feedback the obtained trajectory planning result for the airship flight planning request to the terminal 102. In addition, in some embodiments, the intelligent flight planning method for the unmanned airship can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly perform trajectory planning processing for the airship flight planning request, or the server 104 can obtain the airship flight planning request from the data storage system and perform trajectory planning processing for the airship flight planning request.
[0032] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0033] In an exemplary embodiment, as Figure 2 shown, a method for intelligent flight planning of an unmanned airship is provided. This method is executed by a computer device, and can be specifically executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in it as an example for illustration, it includes the following steps 201 to step 202. Among them:
[0034] Step 201, convert the trajectory planning process of the stratospheric airship into a Markov decision process (MDP), and define the state space, action space and reward function.
[0035] Step 202: Based on the Markov decision process, train the policy network and the value network, and use the policy network and the value network to iteratively update the flight state of the stratospheric airship until the stratospheric airship reaches the target area position, so as to obtain the trajectory planning result of the stratospheric airship from the starting position to the target area position; the policy network and the value network are deep neural networks.
[0036] By implementing the above Steps 201 to 202, a deep neural network is constructed. With the target arrival rate, energy consumption, and cold cloud distribution passed by the trajectory as the optimization indicators, a policy-based reinforcement learning algorithm is used for agent training, and the policy function is directly obtained through policy gradient descent. This method has a continuous state space, can achieve trajectory planning with higher accuracy, and the trajectory is relatively smooth; since the training of the agent is carried out in a simulation environment, the movement of the stratospheric airship under wind field disturbances can be simulated, and the data collection and agent training adopt the off-policy method, so it has a high data utilization rate; this method reflects the optimization objectives in each item of the reward function, and through reasonable allocation of the weights of each item of the reward function, a balance is achieved among multiple optimization objectives, and the optimal policy that maximizes the total reward is obtained through training.
[0037] In another exemplary embodiment of the present application, the trajectory planning process of the stratospheric airship is converted into a sequential decision-making process, and a Markov decision model is constructed to describe the trajectory planning process, where the state space S t includes two-dimensional state data and one-dimensional state data; the two-dimensional state data includes two-dimensional wind field data and two-dimensional temperature field data, that is, the components of the temperature field and the wind field in the meridional and zonal directions; the one-dimensional state data includes local time, the coordinates of the target area position (target point position), the speed of the stratospheric airship, the current position coordinates, and the remaining energy (i.e., the remaining battery power); the action space a t includes the speed increments of the stratospheric airship in the meridional and zonal directions; according to the task requirements, the states S t and S t+1 are designed with a reward function as the optimization objective of this method.
[0038] In Step 201, in order to optimize multiple task objectives and finally obtain the optimal policy that meets the task requirements, each item of the designed reward function comprehensively considers factors such as the cumulative time spent on the task, the distance of the stratospheric airship to the target area position, the current energy consumption power of the stratospheric airship, whether the stratospheric airship reaches the target area position, the temperature distribution at the current position of the stratospheric airship, and whether the stratospheric airship goes out of bounds. The specific form of the reward function is as follows:
[0039] R = R dst + R eng + R bdr+R cld +R tim +R tgt +R out (1).
[0040] R dst = k d *(D t - D t+1 ) (2).
[0041] R eng = -k e *|W - W d | (3).
[0042]
[0043] Among them, R represents the reward function; R dst is the distance reward term, which is related to the distance between the stratospheric airship and the target area position at two adjacent time steps. When the distance in the later time step is shorter than that in the previous time step, the reward value is positive. The distance reward term is expressed in the form shown in Equation (2), and k d is the weight coefficient corresponding to the distance reward term, D t is the distance between the stratospheric airship and the target area position at time step t, and D t+1 is the distance between the stratospheric airship and the target area position at time step t + 1. R eng is the energy consumption penalty term, which is related to the energy consumption W of the stratospheric airship at the current time step. When the energy consumption deviates from the expected value W d more, the absolute value of the negative reward is larger. The energy consumption penalty term is expressed in the form shown in Equation (3), and k e is the weight coefficient corresponding to the energy consumption penalty term. R bdr is the boundary constraint term, which is related to the distance between the stratospheric airship and the boundary of the rectangular airspace. When the distance D b between the stratospheric airship and the boundary is less than the distance threshold D b0 , a negative reward value is received, and the closer to the boundary, the larger the absolute value of the negative reward. When the distance D b between the stratospheric airship and the boundary is not less than the distance threshold D b0 , the boundary constraint term is 0. The boundary constraint term is expressed in the form shown in Equation (4), and k b is the weight coefficient corresponding to the boundary constraint term. R cld is the temperature reward term, which is related to the temperature of the airspace adjacent to the stratospheric airship at two adjacent time steps. When the temperature in the later time step is higher than that in the previous time step, the reward value is positive, and k c is the weight coefficient corresponding to the temperature reward term, T t is the temperature of the airspace adjacent to the stratospheric airship at time step t, and T t+1is the temperature in the adjacent airspace of the stratospheric airship at time step t+1. R tim 、R tgt 、R out are the task termination reward terms, R tim 、R tgt 、R out are all constants, R tim represents the negative reward received when the stratospheric airship fails to reach the target area position after exceeding the maximum time limit, R tgt represents the reward received when the stratospheric airship reaches the target area position, R out represents the negative reward received when the stratospheric airship crosses the boundary.
[0044] According to the meteorological prediction data structure applied in the intelligent trajectory planning, a policy network and a value network are designed. Among them, the value network is composed of a convolutional neural network, a fully connected neural network (multi-layer perceptron), and a long short-term memory network, and its structure is as Figure 4 shown. The structure of the policy network is as Figure 5 shown, which is similar to the structure of the value network, but replaces the long short-term memory network with a fully connected neural network, that is, the policy network is composed of a convolutional neural network and a fully connected neural network. The role of the policy network is to output the corresponding policy according to the input state; the role of the value network is to output the action value function value Q(S t ,a t ) of the state-action combination according to the input state-action combination, and evaluate the reward effect of executing the action in this state. Build a simulation environment for the stratospheric airship trajectory planning task to realize the reading of wind field and temperature field data, the interaction between the stratospheric airship and the environment, and state transition, etc.
[0045] In step 202, based on the Markov decision process, the policy network and the value network are trained, specifically including: adopting the Soft Actor-Critic (SAC) algorithm, based on the Markov decision process, training the policy network and the value network.
[0046] In the above simulation environment, an intelligent agent training algorithm based on the Soft Actor-Critic algorithm is used to train the intelligent agent to find the optimal policy. This algorithm adopts the Actor-Critic framework and the off-policy update strategy to train a policy network that outputs actions, a value network that outputs the action value function, and a target value network. Among them, the value network is optimized by the method of gradient descent on the first loss function L q , and the policy network is optimized by the method of gradient descent on the second loss function L π . The calculation methods of L q and L π are as follows:
[0047]
[0048] where N is the time length, R i represents the reward obtained at the i-th step; γ represents the discount factor, with a value between 0 and 1, used to measure the proportion of delayed rewards. The closer it is to 1, the more importance is attached to long-term rewards; Q j (S,a) is the action-value function, used to evaluate the expected reward after taking action a in state S; the term αlogπ(a|S) represents the policy entropy. Its introduction can improve policy exploration and help the training converge to the global optimum. αlogπ(a i+1 |S i+1 ) represents the policy entropy at the (i + 1)-th step, and αlogπ(a i |S i ) represents the policy entropy at the i-th step. Its weight α is called the temperature coefficient and can be adaptively adjusted during training to control the update progress of the policy function. The loss function L used for updating the weight α α As shown in Equation (8), where H 0 represents the target policy entropy, which is a user-specified constant. represents the mathematical expectation symbol, a t ~π(*|s t ) means that under the condition of determining state s t , the distribution of the random variable a t obeys the function π. The agent consists of a policy network, a first value network, a second value network, a first target value network, and a second target value network. The update relationships of each network are as Figure 8 shown. The target value network has the same structure as the value network, both as Figure 4 shown. minQ j=1,2 (a i ,S i ) means selecting the smaller output value from the output values of the two value networks as the estimated value of the action-value function at the i-th step. Q 1 (a i ,S i ) represents the estimated value of the action-value function output by the first value network. Q 2 (a i ,S i ) represents the estimated value of the action-value function output by the second value network; min j=1,2 Q′ j (S i+1 ,a i+1 ) means selecting the smaller output value from the output values Q' j (S i+1 ,a i+1Among them, select the output value with a smaller value as the estimated value of the action value function at time t+1 to solve the problem of the overestimated estimated value of the action value function, Q' 1 (S i+1 ,a i+1 ) represents the estimated value of the action value function output by the first target value network, Q' 2 (S i+1 ,a i+1 ) represents the estimated value of the action value function output by the second target value network. The model parameters of the target value network are updated through the soft update mechanism shown in Equation (9), where represents the parameters of the target value network, is the model parameter of the first target value network, is the model parameter of the second target value network, θ j represents the value network parameters, θ 1 is the model parameter of the first value network, θ 2 is the model parameter of the second value network, and the parameter τ is a constant specified manually.
[0049] The forward propagation processes of the value network and the policy network are as follows:
[0050] (1) The forward propagation process of the value network is as follows: Read the state data of B batches and T consecutive time steps from the replay pool, organize the two-dimensional temperature field data, meridional and zonal wind speed field data with dimensions of length×width into a three-dimensional array of (B×T)×3×length×width, input it into the convolutional neural network of the value network for feature extraction to obtain the first convolutional feature. Organize the extracted first convolutional feature (the extracted feature map) into a vector with dimensions of B×T×N, and splice it with other one-dimensional state data (i.e., local time, target area position coordinates, stratospheric airship speed, current position coordinates, and remaining energy) and the action vector in the last dimension to form the first spliced array, with dimensions of B×T×M. Input the first spliced array into the first network composed of a multi-layer fully connected neural network and a long short-term memory network. The final output value of the first network is the action value function value.
[0051] (2) The difference in the forward propagation process of the policy network from that of the value network is that: only need to splice other one-dimensional state data with the output of the convolutional neural network, without inputting the action. Therefore, the spliced array is B×T×(M - 2), and the final output value is the mean and standard deviation of the Gaussian distribution of the action a t and the decision-making action a t is sampled from the mean and standard deviation.
[0052] The forward propagation process of the policy network is as follows: Read the state data of B batches and T consecutive time steps from the replay pool, and organize the two-dimensional temperature field data, meridional and zonal wind speed field data with dimensions of length × width into a three-dimensional array of (B×T)×3×length×width, and input it into the convolutional neural network of the policy network for feature extraction to obtain the second convolutional feature. Concatenate the other one-dimensional state data with the second convolutional feature output by the convolutional neural network to obtain the second concatenated array with a size of B×T×(M - 2). Input the second concatenated array into the fully connected neural network to obtain the decision-making action a t 。
[0053] The process of the agent processing the input information and making decisions includes the following steps 301 to 303.
[0054] Step 301: Read the wind field, temperature field, and the state information of the stratospheric airship itself, and concatenate them into the state S t 。
[0055] Step 302: According to the forward propagation process of the policy network, input the state S t into the policy network to obtain the output decision-making action a t 。During the decision-making process, both the batch B and the consecutive time steps T are 1.
[0056] Step 303: Apply the action a t to the stratospheric airship to cause the state of the stratospheric airship itself to transfer, and read the environmental state (including two-dimensional state data and one-dimensional state data) at the next moment, and concatenate them into the state S t+1 。Obtain the reward R of the agent at the current time step t at the t time step through the comparison between the state S t+1 at the next time step and the state S t at the current time step. t 。
[0057] The training process of the policy network and the value network includes the following steps 401 to 405.
[0058] Step 401: Obtain the wind field - temperature field data sequence.
[0059] Step 402: In each time step of each round of training, use the policy network to determine the action at the current time step, use the value network to interact with the environment to transfer the state at the current time step to the state at the next time step, and calculate the reward value at the current time step.
[0060] Step 403: Determine whether the stratospheric airship reaches the task termination condition. The task termination condition is that the stratospheric airship arrives within the target area, goes out of bounds, or the cumulative time step exceeds the preset maximum time step.
[0061] Step 404: When the stratospheric airship reaches the mission termination condition, the state of the current time step, the action of the current time step, the reward value of the current time step and the state of the next time step are put into the data memory playback area as an interactive data.
[0062] Step 405: Select a number of interaction data from the data memory playback area to form a training set, and use the soft actor-critic algorithm to update the network parameters of the policy network and the value network according to the training set in a gradient descent manner to obtain a trained policy network and a trained value network.
[0063] like Figure 3 As shown in Figure 1, the training process of the agent specifically includes the following steps:
[0064] Read several continuous historical wind field and temperature field observation data sequences (wind field-temperature field data sequences, also Figure 3 The data sequence in the data), the read wind field-temperature field data is stored in the buffer area, and the data in the buffer area is continuously updated as the number of training iterations increases;
[0065] Initialization of agent neural network parameters;
[0066] Before each round of agent training begins, a wind field-temperature field sequence is randomly read from the cache, and the stratospheric airship initialization state is randomly generated and spliced into the initial state S 0 ;
[0067] At each time step in each training round, the agent makes a decision a t , interact with the environment to make the state S t Transfer to S t+1 , and receive reward R t If the stratospheric airship arrives in the target area, goes out of the boundary, or the cumulative time step exceeds the preset maximum time step, the MDP process terminates. t ,a t ,R t ,S t+1 ) is stored in the data memory playback area;
[0068] Randomly read some interaction data from the data memory playback area, calculate the loss function of each network according to the neural network update rule of the SAC algorithm, and update the network parameters of the neural network by back propagation in the gradient descent method.
[0069] like Figure 8 As shown, the updating process of the strategy network, the first value network, the second value network, the first target value network and the second target value network is as follows: input the interaction data set (S t ,a t,R t ,S t+1 ), the policy network outputs a decision-making action a based on the state S t and updates the weight α of the policy entropy of the policy function using Equation (8) to obtain the updated weight α. Based on the updated weight α and Equation (7), the model parameters of the policy function are updated. Based on Equation (6), the first value network Q is updated using the interaction data group t to obtain the updated first value network. Based on the soft update mechanism shown in Equation (9) and the model parameters of the updated first value network, the first target value network is updated to obtain the updated first target value network; based on Equation (6), the second value network Q 1 is updated using the interaction data group to obtain the updated second value network. Based on the soft update mechanism shown in Equation (9) and the model parameters of the updated second value network, the second target value network is updated to obtain the updated second target value network. The above process is repeated to obtain the trained policy network, the trained first value network, the trained second value network, the trained first target value network, and the trained second target value network. 2 After training is completed, the trajectory planning effect of this method is simulated and tested. The test effect of trajectory planning under the headwind condition is as
[0070] shown, and the test effect of the stratospheric airship planning to avoid cold clouds is as Figure 6 shown. The trajectory planning effect includes the flight state of the stratospheric airship at each planning moment. The input size of the wind field - temperature field used in the test is 12 * 21, where the area with a temperature lower than minus 30 degrees Celsius is regarded as the cold cloud area. The simulation step length is 7.5 minutes, and the maximum number of simulation steps is 272 steps (equivalent to 34 hours). The initial energy of the stratospheric airship battery is 100%, the initial speed of the stratospheric airship is 0 m / s, and the starting position and the target area position are randomly selected. In 200 tests, the agent achieved an 86.5% target arrival rate, the average minimum battery power in the task was 52%, the average number of steps in the cold cloud area was 9.38 steps (equivalent to 1.17 hours), and the average number of steps consumed to reach the target area was 60.45 steps (equivalent to 7.56 hours). Figure 7 Therefore, this application realizes the trajectory planning of the stratospheric airship to any specified area in the time-varying wind field, and optimizes the energy consumption of the stratospheric airship and its ability to avoid cold cloud areas.
[0071]
[0072] The advantages of the unmanned airship intelligent flight planning method provided in this application are as follows: (1) The trajectory planning accuracy is relatively high, and a smooth stratospheric airship trajectory can be generated; (2) The method of unsupervised learning is adopted, and the intelligent agent training process does not rely on prior knowledge and does not require data annotation; (3) When the algorithm makes a decision, various state data with different structures such as the wind field, temperature field, stratospheric airship energy, and time are comprehensively considered; (4) The intelligent agent always maintains a certain degree of exploration during the training process, so a global optimal solution with the highest total reward can be obtained.
[0073] This application also provides an application scenario that applies the above-mentioned unmanned airship intelligent flight planning method. Specifically: The unmanned airship intelligent flight planning method provided in this embodiment can be applied to the flight control scenario of stratospheric airships. The flight control scenario of stratospheric airships includes a request generation link, a flight planning link, and a control link; the flight planning request enters the flight planning link from the request generation link, obtains the corresponding trajectory planning result, and enters the downstream control link. The unmanned airship intelligent flight planning method provided in this embodiment belongs to the flight planning link. Specifically, during the flight planning link process for the flight planning request, the trajectory planning process of the stratospheric airship can be converted into a Markov decision process, the state space, action space, and reward function are defined, and based on the Markov decision process, the policy network and value network are trained, and the flight state of the stratospheric airship is iteratively updated using the policy network and value network until the stratospheric airship reaches the target area position to obtain the trajectory planning result of the stratospheric airship from the starting position to the target area position.
[0074] Based on the same inventive concept, the embodiment of this application also provides an unmanned airship intelligent flight planning device for implementing the above-mentioned unmanned airship intelligent flight planning method. The implementation solutions provided by this device to solve problems are similar to those described in the above method. Therefore, the specific limitations in one or more embodiments of the unmanned airship intelligent flight planning device provided below can refer to the limitations on the unmanned airship intelligent flight planning method in the above text and will not be elaborated here.
[0075] In an exemplary embodiment, as Figure 9 shown, an unmanned airship intelligent flight planning device is provided, which includes the following modules:
[0076] A trajectory planning process conversion module T1, configured to convert the trajectory planning process of the stratospheric airship into a Markov decision process, and define the state space, action space, and reward function; the state space includes two-dimensional wind field data, two-dimensional temperature field data, local time, target area position, current position coordinates, speed, and remaining energy of the stratospheric airship; the action space includes the speed increments of the stratospheric airship in the meridional and latitudinal directions;
[0077] The airship trajectory planning module T2 is used to train the policy network and the value network based on the Markov decision process, and iteratively update the flight state of the stratospheric airship using the policy network and the value network until the stratospheric airship reaches the target area position, so as to obtain the trajectory planning result of the stratospheric airship from the starting position to the target area position; the policy network and the value network are deep neural networks.
[0078] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store trajectory planning data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an intelligent flight planning method for an unmanned airship.
[0079] Those skilled in the art can understand that Figure 10 the structure shown in
[0080] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0081] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0082] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0083] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0084] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.
[0085] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0086] In this article, specific examples are used to illustrate the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An unmanned airship intelligent flight planning method, characterized in that: The unmanned airship intelligent flight planning method comprises: The trajectory planning process of the stratospheric airship is converted into a Markov decision process, and a state space, an action space and a reward function are defined; the state space includes two-dimensional wind field data, two-dimensional temperature field data, local time, target area position, current position coordinates of the stratospheric airship, speed and residual energy; the action space includes the speed increments of the stratospheric airship in the longitudinal and latitudinal directions; Based on the Markov decision process, the policy network and the value network are trained, and the flight state of the stratospheric airship is iteratively updated using the policy network and the value network until the stratospheric airship reaches the target area position, so as to obtain the trajectory planning result of the stratospheric airship from the starting position to the target area position; the policy network and the value network are deep neural networks.
2. The unmanned airship intelligent flight planning method according to claim 1, characterized in that: The reward function is: R=R dst +R eng +R bdr +R cld +R tim +R tgt +R out ; Among them, R represents the reward function, R dst is the distance reward item, which is related to the distance between the stratospheric airship and the target area in two adjacent time steps; R eng is the energy penalty term, which is related to the energy consumption of the stratospheric airship at the current time step; R bdr is the boundary constraint term, which is related to the distance from the stratospheric airship to the boundary of the rectangular airspace; R cld is the temperature bonus item, which is related to the temperature of the airspace adjacent to the stratospheric airship in two consecutive time steps; R tim , R tgt , R out is the task termination reward, R tim represents the negative reward received when the stratospheric airship exceeds the maximum time limit and still fails to reach the target area, R tgt Represents the reward received when the stratospheric airship reaches the target area, R out represents the negative reward received when the stratospheric airship crosses the boundary.
3. The unmanned airship intelligent flight planning method according to claim 1, characterized in that: The strategy network is composed of a convolutional neural network and a fully connected neural network.
4. The unmanned airship intelligent flight planning method according to claim 1, characterized in that: The value network is composed of a convolutional neural network, a fully connected neural network and a long short-term memory network.
5. The unmanned airship intelligent flight planning method according to claim 1, characterized in that: Based on the Markov decision process, the policy network and value network are trained, including: The policy network and value network are trained based on the soft actor-critic algorithm and Markov decision process.
6. The unmanned airship intelligent flight planning method according to claim 5, characterized in that: The training process of the policy network and value network includes: Obtain wind field-temperature field data sequence; In each time step of each round of training, the policy network is used to determine the action of the current time step, and the value network is used to interact with the environment to transfer the state of the current time step to the state of the next time step, and calculate the reward value of the current time step; Determining whether the stratospheric airship has reached a mission termination condition; When the stratospheric airship reaches the mission termination condition, the state of the current time step, the action of the current time step, the reward value of the current time step and the state of the next time step are taken as an interactive data and put into the data memory playback area; A number of interaction data are selected from the data memory playback area to form a training set, and the network parameters of the policy network and the value network are updated according to the training set in a gradient descent manner using a soft actor-critic algorithm to obtain a trained policy network and a trained value network.
7. The unmanned airship intelligent flight planning method according to claim 6, characterized in that: The mission termination condition is that the stratospheric airship arrives in the target area, goes out of bounds, or the cumulative time step exceeds the preset maximum time step value.
8. An unmanned airship intelligent flight planning device, characterized in that: The unmanned airship intelligent flight planning device comprises: A trajectory planning process conversion module is used to convert the trajectory planning process of the stratospheric airship into a Markov decision process, and define a state space, an action space, and a reward function; the state space includes two-dimensional wind field data, two-dimensional temperature field data, local time, target area position, current position coordinates of the stratospheric airship, speed, and residual energy; the action space includes the speed increments of the stratospheric airship in the longitudinal and latitudinal directions; The airship trajectory planning module is used to train the policy network and the value network based on the Markov decision process, and iteratively update the flight state of the stratospheric airship using the policy network and the value network until the stratospheric airship reaches the target area position, so as to obtain the trajectory planning result of the stratospheric airship from the starting position to the target area position; the policy network and the value network are deep neural networks.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the unmanned airship intelligent flight planning method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the unmanned airship intelligent flight planning method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Intelligent control method for horizontal trajectory of stratospheric airship
CN111538241A
Track planning method and device for data collection of unmanned aerial vehicle, equipment and medium
CN114840021A
Unmanned surface vehicle path tracking method based on deep reinforcement learning
CN115016496A
Multi-unmanned ship deep reinforcement learning collaborative navigation method based on near-end strategy optimization
CN117168468A
Intelligent trajectory planning method for stratospheric airship based on large model
CN117494564A
Cited By
Aircraft recovery scheduling method, device, equipment, medium and product
CN120706847A
Multi-target track collaborative planning method and system based on hierarchical reinforcement learning
CN120869165A