Airship path planning method and system based on deep reinforcement learning
By applying a deep reinforcement learning method in airship path planning, combining dynamic wind farm information and airship's own state information to generate the propulsion speed and heading angle of Gaussian distribution, the problems of low path planning accuracy and neglect of wind farm uncertainty in the existing technology are solved, and more efficient and more accurate path planning is achieved.
Patent Information
- Application Number
- CN202510654203.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing airship path planning method has low computational efficiency and is prone to local optimality when dealing with high-dimensional and dynamic environments, and ignores the uncertainty of the forecast wind field, resulting in low path planning accuracy and low energy efficiency.
The airship path planning method based on deep reinforcement learning is adopted. By inputting dynamic wind farm information and airship's own state information into the trained strategy network, Gaussian distribution of airship propulsion speed and heading angle is generated, and training is carried out through the near-end strategy optimization method. The method includes a wind farm fusion module, a wind farm feature extraction module and a fully connected network module, which can perceive dynamically changing wind farm information and provide a reasonable path planning scheme.
It improves the accuracy and efficiency of airship path planning, can better adapt to the dynamic changes of complex wind fields, reduces energy consumption, and improves the airship's arrival and mission capabilities.
Smart Images

Figure CN120176684A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of airship path planning, and specifically relates to an airship path planning method and system based on deep reinforcement learning. Background Art
[0002] As an aircraft relying on light gases (such as helium) to generate buoyancy, the airship demonstrates broad application prospects in both civilian and military fields due to its long endurance, low energy consumption, large payload capacity, and flight stability. Compared with traditional airplanes and satellites, the stratospheric airship occupies a unique position in modern technological applications due to its lower operating costs and longer hovering time. Designing a suitable path planning algorithm for the stratospheric airship is a very important research content. At the same time, due to the complex flight environment of the airship, multi-dimensional factors such as wind fields, altitude, and solar radiation need to be considered, and the airship's own energy is limited, which increases the difficulty of path planning.
[0003] Traditional methods mainly use heuristic algorithms (such as the A* algorithm, Dijkstra algorithm) and intelligent optimization algorithms (such as genetic algorithms, particle swarm optimization algorithms) to solve path planning problems. To a certain extent, these methods can generate feasible paths, but they often face problems such as low computational efficiency and being easily trapped in local optima when dealing with high-dimensional and dynamic environments. Especially for a system like the stratospheric airship that needs to operate in a complex wind field for a long time, traditional methods are difficult to effectively cope with the dynamic changes of the environment. With the development of artificial intelligence technology, the emergence of deep learning and reinforcement learning provides new possibilities for solving path planning problems. Reinforcement learning (RL), as a machine learning method that learns optimal strategies by interacting with the environment, has high intelligence, flexibility, and adaptability. In the path planning of airships and other tasks, researchers have proposed reinforcement learning-based methods such as D3RQN, DQN, etc. These methods gradually optimize strategies through trial-and-error learning by reasonably designing the state space and reward function, providing new solutions for the path planning of airships in complex environments.
[0004] Although reinforcement learning methods have many advantages, they still face some problems in the practical application of airship path planning. On the one hand, existing methods mostly adopt discrete action spaces. Although this design simplifies the calculation of the airship dynamics model, it severely restricts the movement ability of the airship. This limitation not only reduces the accuracy of path planning but also may cause the airship to be unable to effectively utilize the characteristics of the wind field, thereby increasing energy consumption. In addition, the performance of reinforcement learning highly depends on the design of the state space, and most existing studies only use the global forecast wind field or the local detected wind field as the state input in the processing of the state space. This simplified environmental representation is difficult to comprehensively reflect the real flight environment in which the airship is located, resulting in the airship being difficult to accurately perceive the dynamic changes in the environment, thus affecting the quality of decision-making. On the other hand, current research usually uses the forecast wind field as the environmental input, and the uncertainty of the forecast wind field makes it difficult for the airship to make decisions. Although prediction methods such as time series methods, Kalman filters, and neural networks can obtain relatively accurate forecast wind fields, due to the complexity and variability of the stratospheric wind field, there are often significant errors between the forecast wind field and the actual wind field, and this error is uncertain. However, existing research generally ignores the uncertainty of the forecast wind field, resulting in significant deviations that may occur in the application of the path planning results in the actual wind field. Especially in long-term and large-scale flight missions, this uncertainty will seriously affect the arrival ability and energy efficiency of the airship, further restricting the actual mission ability of the airship. Summary of the Invention
[0005] The purpose of this application is to overcome the defects of low path planning accuracy and ignoring the uncertainty of the forecast wind field in the prior art.
[0006] To achieve the above purpose, this application proposes an airship path planning method based on deep reinforcement learning, including:
[0007] Input the dynamic wind field information and the airship's own state information into the trained policy network to obtain the Gaussian distribution of the planned airship propulsion speed and heading angle, and sample the Gaussian distribution to obtain the planned airship propulsion speed and heading angle;
[0008] The policy network is trained by the proximal policy optimization method.
[0009] As an improvement of the above method, the dynamic wind field information includes the measured wind speed, forecast wind field data, and uncertainty at the current position of the airship.
[0010] As an improvement of the above method, the forecast wind field data is obtained through a sliding window centered on the airship.
[0011] As an improvement of the above method, the uncertainty is:
[0012]
[0013] Among them, represents the basic uncertainty of the predicted wind field itself; represents the influence function:
[0014]
[0015] Among them, represents the Euclidean distance from each position in the sliding window to the position of the airship; represents the parameter for controlling the influence range; represents the difference between the measured wind speed and the predicted wind speed at the same position at the current position of the airship.
[0016] As an improvement of the above method, the predicted wind field is updated according to the error between the current actual wind field and the original predicted wind field:
[0017]
[0018] Among them, represents the updated predicted wind field; represents the original predicted wind field.
[0019] As an improvement of the above method, the state information of the airship itself includes the speed, navigation, position, and energy information of the airship;
[0020] Within the time interval the energy change of the airship battery is expressed as:
[0021]
[0022] Among them, represents the power consumption of the avionics equipment; represents the solar power generation:
[0023]
[0024] Among them, represents the solar irradiance; represents the solar energy conversion efficiency; represents time of the solar energy effective factor:
[0025]
[0026] Among them, represents the daylight duration of 12 h, represents the sunrise time;
[0027] Indicate Motor power consumption at a moment:
[0028]
[0029] Wherein, Indicates the coefficient of the airship motor doing work; Indicates the airship at the moment Actual speed at the position;
[0030] Energy state of the battery at the moment Is:
[0031]
[0032]
[0033] Wherein, Indicates the energy state of the battery at time t; Indicates the maximum capacity of the battery.
[0034] As an improvement of the above method, the policy network includes:
[0035] A wind field fusion module, configured to fuse the measured wind speed and the predicted wind field data at the current position of the airship to obtain a fused wind field matrix;
[0036] A wind field feature extraction module, including two layers of convolutional networks, configured to extract wind field features from the fused wind field matrix and the uncertainty;
[0037] A fully connected network module, configured to calculate the comprehensive state vector after fusing the wind field features and the airship's own state, and output the Gaussian distribution of the planned airship propulsion speed and heading angle.
[0038] As an improvement of the above method, the reward function during the training of the proximal policy optimization method Is:
[0039]
[0040] Wherein, , , And Indicate the weight coefficients of each reward function; Indicates the proximity to the target reward function:
[0041]
[0042] Wherein: The distance Denotes the distance from the position of the airship at time t to the target position;
[0043] Denotes the energy consumption penalty function:
[0044]
[0045] Wherein, Denotes the battery energy state at time t; Denotes the maximum capacity of the battery;
[0046] Denotes the arrival target reward and boundary penalty:
[0047]
[0048] Denotes the stability reward function:
[0049]
[0050] Wherein, and Are the weight coefficients of the direction change penalty and the speed change penalty; Is the change amount of the propulsion direction between the current step and the previous step; Is the change amount of the propulsion speed between the current step and the previous step.
[0051] This application also provides an airship path planning system based on deep reinforcement learning, which is implemented based on the above method. The system includes:
[0052] An environment simulation module, which is used to provide an airship simulation operation environment and generate dynamic wind field information and airship own state information;
[0053] A policy network, which is used to plan the propulsion speed and heading angle of the airship;
[0054] A path planning module, which is used to input the dynamic wind field information and the airship own state information into the trained policy network to obtain the Gaussian distribution of the planned propulsion speed and heading angle of the airship, and obtain the planned propulsion speed and heading angle of the airship through sampling the Gaussian distribution;
[0055] A training module, which is used to train the policy network by the proximal policy optimization method.
[0056] Compared with the prior art, the advantages of this application are:
[0057] 1. This application proposes the application of the Proximal Policy Optimization (PPO) method with a continuous action space in the path planning of stratospheric airships in an uncertain wind field. This method can sense the dynamically changing wind field information and give a reasonable path planning scheme.
[0058] 2. This application proposes a Gaussian space enhanced wind field estimation method. By fusing airship sounding data and forecast wind fields, it provides more accurate wind field data and wind field uncertainty information for the path planning of airships.
[0059] 3. Through the analysis of the energy characteristics of airships and wind field data, this application proposes a state space that combines a wind field information window and uncertainty. Through a carefully designed reward function, the practicability and reliability of the model are improved. Through experiments, the method we proposed has higher efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 The figure shows a flowchart of an airship path planning method based on deep reinforcement learning. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The technical solutions of this application will be described in detail below with reference to the accompanying drawings.
[0062] The present invention proposes an uncertain wind field path planning method based on deep reinforcement learning, that is, an Improved Windfield Information-based Proximal Policy Optimization algorithm (IWI-PPO), to solve the problem of airship path planning under uncertain wind fields.
[0063] As Figure 1As shown in the figure, the method for uncertain wind field path planning based on deep reinforcement learning provided by this application uses a policy network to generate the path planning of the airship. The policy network includes three core capabilities: wind field fusion, wind field feature extraction, and a fully connected network. The policy network takes heterogeneous data as input, including dynamic wind field information and the airship's own state information. Specifically, the airship's own state information includes key parameters such as the airship's position coordinates, speed, heading angle, current time, and remaining energy. The dynamic wind field information is obtained by cropping the surrounding wind field matrix window centered on the airship's current position. The wind field data within this window undergoes wind field fusion processing to generate a fused wind field matrix and simultaneously quantify the uncertainty of the wind field. The fused wind field matrix and the uncertainty information are further subjected to wind field feature extraction, and two-layer convolutional neural networks are used to extract effective spatial features from the complex wind field data. These spatial features are then converted into one-dimensional vectors and fused with the airship's own state information to form a comprehensive state vector. This fusion mechanism of heterogeneous data ensures that the agent can simultaneously consider the real-time state of the airship and dynamic environmental changes, thereby enhancing its adaptability to complex environments. Finally, the comprehensive state vector is calculated through a fully connected network, and Gaussian distributions of the airship's propulsion speed and heading angle are output. The Gaussian distribution is determined by two parameters, variance and expectation. Based on these outputs, the policy network obtains the propulsion speed and heading angle of the airship's action through sampling.
[0064] The input information of the policy network, wind field fusion, and reinforcement learning algorithm design will be introduced separately below.
[0065] 1. Airship's own state information
[0066] The airship's own state information includes key parameters such as the airship's position coordinates, speed, heading angle, current time, and remaining energy.
[0067] An environment simulation provides a flight environment close to reality for the airship flying in the stratosphere. The flight of the airship in the stratosphere is mainly affected by the wind field and is also subject to energy constraints. The environment simulation includes kinematic simulation and airship energy simulation.
[0068] (1) Kinematic simulation
[0069] The kinematic simulation environment mainly consists of a wind field and an airship dynamics model. The position of the airship can be expressed as , where is the longitude, is the latitude.
[0070] The wind field is a collection of wind speed and wind direction at each position in space. The wind speed at position at time t can be expressed as:
[0071] (1)
[0072] Among them, are respectively the wind speed components in the longitude and latitude directions of the position at a certain moment.
[0073] The propulsion speed of the airship is the speed relative to the air, denoted as , the magnitude of the propulsion speed and the course are obtained through the output of the policy network. Since the airship is affected by the stratospheric wind during flight, the influence of the wind field on the movement of the airship must be considered. The actual speed of the airship relative to the ground is the propulsion speed and the wind speed vector sum of:
[0074] (2)
[0075] (3)
[0076] The position of the airship changes with time, is the initial position of the airship, and the position of the airship at time is obtained by integrating the actual speed:
[0077] (4)
[0078] At the discrete time step , the position at the next moment is updated as:
[0079] (5)
[0080] is the actual speed of the position of the airship at time is . In this way, we obtain the transfer formula of the airship position.
[0081] (2) Energy simulation
[0082] The path planning feature that differentiates the airship from other missions lies in energy management. The flight missions of airships generally last for a long time, possibly exceeding 7 days or even longer, and the energy reserve of the airship's battery is not sufficient to handle such a long-time power propulsion. At the same time, on-board electronic devices on the airship, such as environmental monitoring devices and controllers, consume energy throughout the day. Therefore, the airship usually needs to have its own charging function. Airships usually rely on on-board solar panels to provide electrical energy, but there is a phenomenon that the charging power is different due to different solar radiation angles within different time ranges of a day. Therefore, we have established an airship energy model that includes airship energy consumption and solar charging.
[0083] When the airship is flying in the air, in order to maintain a constant speed, it needs to overcome air resistance, and the power required by the motor is proportional to the cube of the speed. Motor power consumption at a certain moment It can be calculated by the following formula:
[0084] (6)
[0085] Among them, is the coefficient of the motor work of the airship. Avionics equipment (such as navigation, communication, sensors, etc.) continuously consumes energy during flight and can be regarded as constant power consumption .
[0086] Solar energy is one of the main energy sources of the airship. The solar charging ability depends on solar irradiance , conversion efficiency and time (day-night cycle). The solar radiation at the altitude where the airship is located is not blocked by clouds.
[0087] (7)
[0088] Among them, represents time of the solar energy effective factor, which represents the proportion of the actually available solar energy. To simulate the day-night cycle, W can be expressed by the following function:
[0089] (8)
[0090] Among them, is the daylight duration of 12 h, and are the sunrise time and the current time respectively, in hours. During the time interval , the energy change of the battery is determined by the input (solar power generation ) and the output (motor power consumption and avionics equipment power consumption ).
[0091] (9)
[0092] Among them, is the maximum capacity of the battery. The battery power at the current moment is , and the new energy state of the battery at the next moment is:
[0093]
[0094] (10)
[0095] When occurs, the airship cannot continue to operate.
[0096] 2. Gaussian Space Enhanced Wind Field Fusion
[0097] This application proposes a Gaussian space enhanced wind field fusion method, which fuses the accurate wind measurement values at the current position of the airship with the predicted wind field data, and quantifies the wind field uncertainty at each position to obtain the uncertainty, providing a new perspective and basis for decision-making optimization in path planning. The core idea is to adjust the predicted wind field based on the difference between the measured value and the predicted value at the current position of the airship, and use a decaying influence function to propagate this adjustment in space. The specific algorithm steps are as follows:
[0098] First, calculate the measured wind speed at the current position of the airship and the predicted wind speed at the same position The difference between them provides a reference difference value for subsequent wind field adjustment. The reference difference value can reflect the error scale and overall situation between the current actual wind field and the predicted wind field.
[0099] (11)
[0100] For each position in the window , the algorithm calculates the Euclidean distance to the central position , providing the necessary spatial distance information for the calculation of the influence function.
[0101] (12)
[0102] (13)
[0103] Use the Gaussian function as the influence function , which decreases as the distance from the center increases under the condition of error determination, reflecting the influence of the error value on the wind field adjustment. Among them, is the parameter controlling the influence range. Using the influence function and the difference , the algorithm adjusts the predicted wind field to obtain the updated wind field .
[0104] (14)
[0105] Furthermore, calculate the uncertainty at each position.
[0106] (15)
[0107] (16)
[0108] An uncertainty is represented by a matrix of the same size as the wind field. The uncertainty at each position represents the uncertainty of the error between the predicted value of the wind field at the current position and the true wind field. The uncertainty is determined by the basic uncertainty of the predicted wind field itself and the uncertainty calculated from the influence function based on the error and distance . The uncertainty increases with the distance and the error , reflecting the reduction in the degree of trust in the predicted wind field at positions with a larger distance from the measurement value and a larger error. The uncertainty represents the degree of credibility of the wind field at the current position. In this method, the uncertainty matrix is used as part of the environmental state as the input to the policy network.
[0109] 3. Reinforcement Learning Algorithm Design
[0110] In this application, the proximal policy optimization method is used to train the policy network.
[0111] (1) State Space
[0112] For the wind field information around the airship, in this application, a sliding window centered on the airship is used to obtain it, which can better obtain the wind field information required by the airship. The size of the sliding window has a certain impact on the path planning performance. If the window size is too large, it may lead to too low computational efficiency and difficult training convergence. The wind field information at a distance too large from the current position is not very helpful for the path planning of the wind field. If the window size is too small, it may cause the path planner to not obtain enough wind field data, resulting in the inability to obtain the optimal path. The minimum size of the sliding window is 1×1, and the maximum is the size of the entire map. In actual tests, the size range of the sliding window is 1×1~64×64.
[0113] While obtaining the prediction wind field information window at the current position, the airship can detect the wind field information at the current position in real time. After fusing the two, it is used as the wind field information window, and the obtained uncertainty matrix is used as part of the airship state information.
[0114] (2) Action Space
[0115] To enable the airship to complete a variety of different tasks, we adopt a continuous action space. According to the maneuvering characteristics of the airship, we select the flight speed and heading as the action space, and use : the magnitude of the propulsion speed relative to the air and : Heading angle of advancement representation. The policy network can output continuous values of speed and heading, enabling precise control of the airship. Due to the relatively weak self-power of the airship and its large volume and weight, the acceleration ability and turning ability of the airship are limited to a certain extent. Therefore, we have imposed constraints on the action space of the airship.
[0116] (17)
[0117] (18)
[0118] Among them, is the change in heading angle, is the change in speed. This application adds two constraints, (19) and (20):
[0119] (19)
[0120] (20)
[0121] (3) Reward function design
[0122] The reward function is mainly divided into the following parts:
[0123] 1) Proximity-to-goal reward
[0124] (21)
[0125] Among them:
[0126] (22)
[0127] Among them, the distance represents the distance from the current position to the target position. After the airship takes an action in a certain state, the distance from the current position to the target position is recalculated. The reward setting can encourage the airship to move towards the target at each step, effectively solving the problem of sparse rewards in the path planning task.
[0128] 2) Energy consumption penalty
[0129] (23)
[0130] Among them, the current battery energy state is , and as the airship consumes electrical energy and charges with solar energy during flight. it will keep changing. To ensure that the airship has enough electrical energy to guarantee the normal operation of the avionics equipment and the ability to handle emergencies, it is necessary to ensure that the airship has a certain energy reserve. When the battery power is lower than a certain threshold, the airship is regarded as an unsafe state, and a penalty should be imposed on this state.
[0131] 3) Reaching the target reward and boundary penalty
[0132] (24)
[0133] Among them, is the position of the target point. After the airship reaches the target point area, the airship is given a large positive reward, which can encourage the airship to still have a long-term goal of reaching the target position even when the energy is insufficient. When the airship flies out of the current wind field area, a large penalty is given, and a sign of environmental termination is returned to prevent the airship from flying out of the boundary quickly to obtain the minimum penalty.
[0134] 4) Stability reward
[0135] By reducing the frequent speed and direction changes of the airship, the flight stability is maintained.
[0136] (25)
[0137] Among them, and are the weight coefficients of the direction change penalty and the speed change penalty, which are 0.8 and 1 respectively. is the change amount of the propulsion direction between the current step and the previous step, is the change amount of the propulsion speed between the current step and the previous step, punishing the frequent speed and direction adjustments. Encourage the airship to plan a smooth flight path, reducing energy waste and mechanical wear.
[0138] Combine different rewards according to different parameters:
[0139] (26)
[0140] Among them, = 1, = 1, = 1, = 5.
[0141] Using the uncertain wind field path planning method based on deep reinforcement learning provided by this application for simulation calculation, the results are as follows:
[0142] The proposed method is implemented using the PyTorch framework. Training is carried out in a simulation environment with real wind fields. The wind fields in the East Asian region from 2021 to 2023 are used for training, and the wind fields within the range of 0 - 40 degrees north latitude and 100 - 140 degrees east longitude are intercepted in different seasons. The real wind field data is sourced from ERA5 reanalysis data, and the forecast data is sourced from the wind field forecast data of the National Space Science Center. A discrete 160*160 grid is used to represent the wind field, where each grid represents 0.25 degrees of longitude and latitude, and the time resolution is 1 hour. The size of the wind field window in the state space is 32*32, and the time step is 15 minutes.
[0143] Table 1 summarizes the parameters of the learning algorithm. The deployed policy network is trained using a GPU for 30 hours with a total of 5 million simulations, and data is collected using 10 parallel simulation environments. Training is carried out on Windows 10 with a training environment of 64GB RAM and an RTX4090 GPU.
[0144] Table 1 Training Parameters
[0145]
[0146] To verify the performance and effectiveness of the proposed method, two different methods are used as comparisons for testing, D3RQN and the heuristic algorithm. The D3RQN algorithm adopts a discrete action space and uses the global forecast wind field information as the state input. The heuristic algorithm (Heuristic Algorithm, HA) is representative of traditional methods. By comparing the method of the present invention with D3RQN and the heuristic algorithm, the performance advantages of the algorithm in aspects such as path planning efficiency, energy state, and environmental adaptability can be comprehensively evaluated.
[0147] This application selects a typical test scenario in June 2024 for visual display. The starting point is located at 133.25 degrees east longitude and 8.5 degrees north latitude, and the ending point is located at 134 degrees east longitude and 34 degrees north latitude. In this scenario, this application comparatively analyzes the path planning performances of three algorithms: the method provided by this application (IWI - PPO), HE, and D3RQN. The arrival times at the ending point are 85h, 96.5h, and 92.5h respectively. The path planned by the method of this application can reach the ending point earlier and the path is smoother. The reason is that the algorithm proposed in this application better grasps the characteristics of the wind field and moves along the direction of the wind field, thus effectively saving time and energy.
[0148] This application also provides an airship path planning system based on deep reinforcement learning, implemented based on the above method. The system includes:
[0149] An environmental simulation module for providing an airship simulation operation environment and generating dynamic wind field information and airship own state information;
[0150] A policy network for planning the propulsion speed and heading angle of the airship;
[0151] A path planning module for inputting the dynamic wind field information and the airship own state information into the trained policy network to obtain the planned propulsion speed and heading angle of the airship;
[0152] A training module for training the policy network by the proximal policy optimization method.
[0153] The present application can also provide a computer device, including: at least one processor, a memory, at least one network interface, and a user interface. Each component in the device is coupled together through a bus system. It can be understood that the bus system is used to realize the connection and communication between these components. The bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.
[0154] Among them, the user interface may include a display, a keyboard, or a pointing device. For example, a mouse, a trackball, a touchpad, or a touch screen, etc.
[0155] It can be understood that the memory in the disclosed embodiments of the present application can be a volatile memory, a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The memories described herein are intended to include but not be limited to these and any other suitable types of memories.
[0156] In some embodiments, the memory stores the following elements, executable modules, or data structures, or subsets thereof, or extended sets thereof: an operating system and application programs.
[0157] Among them, the operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., and is used to implement various basic services and handle hardware-based tasks. The application programs include various application programs, such as a media player and a browser, etc., and are used to implement various application services. The program for implementing the method of the disclosed embodiments of the present application can be included in the application programs.
[0158] In the above-mentioned embodiments, by calling the programs or instructions stored in the memory, specifically, the programs or instructions stored in the application programs, the processor is used to:
[0159] Execute the steps of the above method.
[0160] The above method can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed above. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Combining the steps of the above-disclosed method can be directly embodied as being completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0161] It can be understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or a combination thereof.
[0162] For software implementation, the technology of this application can be implemented by executing the functional modules of this application (such as procedures, functions, etc.). The software code can be stored in the memory and executed by the processor. The memory can be implemented inside or outside the processor.
[0163] The present application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, the various steps in the above method embodiments can be implemented.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present application does not depart from the spirit and scope of the technical solutions of the present application, and they should all be covered within the scope of the claims of the present application.
Claims
1. A method for airship path planning based on deep reinforcement learning, comprising: The dynamic wind field information and the airship's own state information are input into the trained strategy network to obtain the planned airship propulsion speed and heading angle Gaussian distribution, and the planned airship propulsion speed and heading angle are obtained by sampling the Gaussian distribution; The policy network is trained via a proximal policy optimization method.
2. The airship path planning method based on deep reinforcement learning according to claim 1, characterized in that: The dynamic wind field information includes the measured wind speed, forecast wind field data and uncertainty at the current position of the airship.
3. The airship path planning method based on deep reinforcement learning according to claim 2, characterized in that: The forecast wind field data is obtained through a sliding window centered on the airship.
4. The airship path planning method based on deep reinforcement learning according to claim 3, characterized in that: The uncertainty stated for: ; in, It represents the basic uncertainty of the forecast wind field itself; Represents the influence function: ; in, Represents each position in the sliding window Euclidean distance to the airship's position; Parameters that represent the scope of control influence; Represents the difference between the measured wind speed at the airship's current location and the forecast wind speed at the same location.
5. The airship path planning method based on deep reinforcement learning according to claim 4, characterized in that: The forecast wind field is updated according to the error between the current actual wind field and the original forecast wind field: ; in, represents the updated forecast wind field; Represents the original forecast wind field.
6. The airship path planning method based on deep reinforcement learning according to claim 1, characterized in that: The airship's own state information includes the airship's speed, navigation, position and energy information; In time interval Energy changes of the airship battery It is expressed as: ; in, represents the power consumption of avionics equipment; Indicates solar power generation power: ; in, represents solar irradiance; It represents the solar energy conversion efficiency; Indicates time Solar Effective Factor: ; in, The daytime length is 12 hours. Indicates sunrise time; express Motor power consumption at each moment: ; in, The coefficient that represents the work done by the airship motor; Indicates that the airship is at time The actual speed of the position; Battery energy status at all times for: ; ; in, Indicates the energy state of the battery at time t; Indicates the maximum capacity of the battery.
7. The airship path planning method based on deep reinforcement learning according to claim 1, characterized in that: The policy network includes: The wind field fusion module is used to fuse the measured wind speed at the current position of the airship with the forecast wind field data to obtain a fused wind field matrix; The wind field feature extraction module includes a two-layer convolutional network, which is used to extract wind field features from the fused wind field matrix and uncertainty; The fully connected network module is used to calculate the comprehensive state vector after the fusion of wind field characteristics and the airship's own state, and output the Gaussian distribution of the planned airship propulsion speed and heading angle.
8. The airship path planning method based on deep reinforcement learning according to claim 1, characterized in that: The reward function when the proximal policy optimization method is trained for: ; in, , , and Represents the weight coefficient of each reward function; Represents a reward function close to the target: ; Where: Distance represents the distance from the airship position to the target position at time t; Represents the energy consumption penalty function: ; in, Indicates the battery energy status at time t; Indicates the maximum capacity of the battery; Represents the target reward and boundary penalty: ; Denote the stability reward function: ; in, and The weight coefficient of the penalty for direction change and the weight coefficient of the penalty for speed change; is the change in the advancement direction between the current step and the previous step; is the change in advancement speed between the current step and the previous step.
9. An airship path planning system based on deep reinforcement learning, implemented based on the method described in any one of claims 1 to 8, characterized in that: The system comprises: The environmental simulation module is used to provide the airship simulation operation environment and generate dynamic wind field information and the airship's own status information; The policy network is used to plan the propulsion speed and heading angle of the airship; A path planning module is used to input dynamic wind field information and the airship's own state information into the trained strategy network to obtain the planned airship propulsion speed and heading angle; and The training module is used to train the policy network through the proximal policy optimization method.
Citation Information
Patent Citations
Unmanned ship weather adaptive obstacle avoidance method based on deep reinforcement learning
CN113176776A
Stratospheric airship cluster area coverage control method and system under uncertain wind field
CN116339388A
Unmanned airship intelligent flight planning method based on meteorological data
CN116522802A
Intelligent trajectory planning method for stratospheric airship based on large model
CN117494564A
Modeling and trajectory tracking method for stratospheric airship in unknown wind environment
CN118550197A