A method and system for airship path planning based on deep reinforcement learning

Through the deep reinforcement learning method, combined with dynamic wind farm information and airship status, the propulsion speed and heading angle of Gaussian distribution are generated, which solves the problems of low computing efficiency and wind farm uncertainty in airship path planning, and achieves more efficient and accurate path planning.

CN120176684BActive Publication Date: 2025-08-26NAT SPACE SCI CENT CAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510654203.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-26
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing airship path planning method has low computational efficiency and is prone to local optimality when dealing with high-dimensional and dynamic environments, and fails to effectively utilize the wind field characteristics, resulting in increased energy consumption and low path planning accuracy, neglecting the uncertainty of the forecast wind field and affecting the quality of decision-making.

Method used

Using a method based on deep reinforcement learning, by inputting dynamic wind farm information and airship's own state information into the policy network, a Gaussian distribution of airship propulsion speed and heading angle is generated, combined with near-end strategy optimization methods and Gaussian space enhanced wind farm estimation, integrating airship sounding data and forecast wind farm, designing continuous action space and careful reward functions to improve the accuracy and adaptability of path planning.

Benefits of technology

It improves the accuracy and energy efficiency of airships in uncertain wind farms, can better perceive dynamic environmental changes, provide more efficient path planning solutions, and reduce energy consumption and time costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120176684B_ABST
    Figure CN120176684B_ABST
Patent Text Reader

Abstract

The present application provides an airship path planning method and system based on deep reinforcement learning, which includes: inputting dynamic wind field information and airship state information into a trained policy network to obtain the planned airship propulsion speed and heading angle; the policy network is trained by a proximal policy optimization method; the dynamic wind field information is to use a moving window to intercept the forecast wind field, and the forecast wind field and the detected wind speed information are Gaussian enhanced wind field fusion and wind field uncertainty and other wind field information; the dynamic wind field information is extracted through the convolutional neural network of the policy network, and after the wind field information is fused with the airship state information, the airship action is calculated through the fully connected network of the policy network. The advantages of this application are: by fusing airship sounding data and forecast wind field, more accurate wind field data and wind field uncertainty information are provided for airship path planning; and through a carefully designed reward function, the practicality and reliability of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of airship path planning, and specifically to an airship path planning method and system based on deep reinforcement learning. Background Art

[0002] Airships, as aircraft that rely on lightweight gases (such as helium) for buoyancy, offer broad potential for civil and military applications due to their long endurance, low energy consumption, large payload capacity, and stable flight. Compared to traditional aircraft and satellites, stratospheric airships hold a unique position in modern technological applications due to their lower operating costs and longer flight times. Designing appropriate path planning algorithms for stratospheric airships is a crucial research topic. However, the complex flight environment of airships requires consideration of multiple factors, such as wind speed, altitude, and solar radiation, as well as the limited energy resources of the airships themselves, making path planning more challenging.

[0003] Traditional approaches to path planning primarily rely on heuristic algorithms (such as the A* algorithm and the Dijkstra algorithm) and intelligent optimization algorithms (such as genetic algorithms and particle swarm optimization) to solve the problem. While these methods can generate feasible paths to a certain extent, they often suffer from low computational efficiency and a tendency to get stuck in local optima when dealing with high-dimensional, dynamic environments. This is particularly true for systems like stratospheric airships, which must operate for extended periods in complex wind fields. Traditional approaches struggle to effectively cope with the dynamic nature of the environment. The advancement of artificial intelligence (AI) technologies, including deep learning and reinforcement learning, has provided new possibilities for solving path planning problems. Reinforcement learning (RL), a machine learning method that learns optimal policies through interaction with the environment, offers high intelligence, flexibility, and adaptability. Researchers have proposed RL-based methods for path planning in airships and other missions, such as D3RQN and DQN. These methods utilize trial-and-error learning to gradually optimize policies by rationally designing state spaces and reward functions, offering new solutions for airship path planning in complex environments.

[0004] Despite the numerous advantages of reinforcement learning, practical applications of airship path planning still face several challenges. For one thing, existing methods often employ a discrete action space. While this design simplifies the computation of the airship's dynamic model, it severely limits the airship's maneuverability. This limitation not only reduces path planning accuracy but can also prevent the airship from effectively utilizing wind field characteristics, thereby increasing energy consumption. Furthermore, the performance of reinforcement learning is highly dependent on the design of the state space. Existing research often uses only global forecast wind fields or local detected wind fields as state inputs. This simplified representation of the environment fails to fully reflect the airship's actual flight environment, making it difficult for the airship to accurately perceive dynamic environmental changes, thus compromising decision-making quality. Furthermore, current research typically uses forecast wind fields as input, but the uncertainty of these forecast wind fields complicates airship decision-making. Although prediction methods such as time series methods, Kalman filters, and neural networks can generate relatively accurate wind forecasts, the complexity and variability of stratospheric wind fields often lead to significant errors between the forecast and actual wind fields, and these errors are often uncertain. However, existing research generally ignores the uncertainty of wind forecasts, which can lead to significant deviations in the application of path planning results to actual wind fields. This uncertainty can severely impact the airship's reach and energy efficiency, particularly during long, wide-area missions, further limiting its mission capabilities. Summary of the Invention

[0005] The purpose of this application is to overcome the defects of the existing technology such as low path planning accuracy and neglect of the uncertainty of the forecast wind field.

[0006] To achieve the above objectives, this application proposes an airship path planning method based on deep reinforcement learning, including:

[0007] The dynamic wind field information and the airship's own state information are input into the trained strategy network to obtain the planned Gaussian distribution of the airship's propulsion speed and heading angle. The planned propulsion speed and heading angle of the airship are obtained by sampling the Gaussian distribution.

[0008] The policy network is trained via a proximal policy optimization method.

[0009] As an improvement to the above method, the dynamic wind field information includes the measured wind speed, forecast wind field data and uncertainty at the current position of the airship.

[0010] As an improvement to the above method, the forecast wind field data is obtained through a sliding window centered on the airship.

[0011] As an improvement to the above method, the uncertainty for:

[0012]

[0013] in, It represents the basic uncertainty of the forecast wind field itself; Represents the influence function:

[0014]

[0015] in, Represents each position in the sliding window Euclidean distance to the airship's location; Parameters that represent the scope of control influence; Represents the difference between the measured wind speed at the airship's current location and the forecast wind speed at the same location.

[0016] As an improvement to the above method, the forecast wind field is updated based on the error between the current actual wind field and the original forecast wind field:

[0017]

[0018] in, represents the updated forecast wind field; Represents the original forecast wind field.

[0019] As an improvement to the above method, the airship state information includes the airship's speed, navigation, position and energy information;

[0020] In time interval Energy changes in the airship battery Expressed as:

[0021]

[0022] in, represents the power consumption of avionics equipment; Indicates solar power generation power:

[0023]

[0024] in, represents solar irradiance; represents the solar energy conversion efficiency; Indicates time Solar efficiency factor:

[0025]

[0026] in, Indicates that the daytime is 12 hours long. Indicates sunrise time;

[0027] express Motor power consumption at each moment:

[0028]

[0029] in, The coefficient that represents the work done by the airship motor; Indicates that the airship is at time The actual speed of the position;

[0030] Battery energy status at all times for:

[0031]

[0032]

[0033] in, Indicates the energy state of the battery at time t; Indicates the maximum capacity of the battery.

[0034] As an improvement to the above method, the policy network includes:

[0035] The wind field fusion module is used to fuse the measured wind speed at the current position of the airship with the forecast wind field data to obtain the fused wind field matrix;

[0036] The wind field feature extraction module includes a two-layer convolutional network, which is used to extract wind field features from the fused wind field matrix and uncertainty;

[0037] The fully connected network module is used to calculate the comprehensive state vector after integrating the wind field characteristics and the airship's own state, and output the Gaussian distribution of the planned airship propulsion speed and heading angle.

[0038] As an improvement to the above method, the reward function of the proximal policy optimization method during training is for:

[0039]

[0040] in, 、 、 and Represents the weight coefficient of each reward function; Represents a reward function close to the target:

[0041]

[0042] Where: Distance represents the distance from the airship position to the target position at time t;

[0043] Represents the energy consumption penalty function:

[0044]

[0045] in, Indicates the battery energy status at time t; Indicates the maximum capacity of the battery;

[0046] Represents the goal reward and boundary penalty:

[0047]

[0048] Denote the stability reward function:

[0049]

[0050] in, and The weight coefficient of the direction change penalty and the weight coefficient of the speed change penalty; is the change in advancement direction between the current step and the previous step; is the change in advancement speed between the current step and the previous step.

[0051] The present application also provides an airship path planning system based on deep reinforcement learning, which is implemented based on the above method. The system includes:

[0052] The environmental simulation module is used to provide the airship simulation operation environment and generate dynamic wind field information and airship status information;

[0053] A policy network, used to plan the propulsion speed and heading angle of the airship;

[0054] The path planning module is used to input dynamic wind field information and the airship's own state information into the trained strategy network to obtain the planned Gaussian distribution of the airship's propulsion speed and heading angle. The planned propulsion speed and heading angle of the airship are obtained by sampling the Gaussian distribution.

[0055] The training module is used to train the policy network through the proximal policy optimization method.

[0056] Compared with the prior art, the advantages of this application are:

[0057] 1. This application proposes the application of the proximal policy optimization (PPO) method in a continuous action space for path planning of a stratospheric airship in an uncertain wind field. This method can perceive the dynamically changing wind field information and provide a reasonable path planning solution.

[0058] 2. This application proposes a Gaussian spatially enhanced wind field estimation method, which provides more accurate wind field data and wind field uncertainty information for airship path planning by fusing airship sounding data and forecast wind fields.

[0059] 3. This application analyzes airship energy characteristics and wind field data to propose a state space that combines wind field information windows and uncertainty. Through a carefully designed reward function, this improves the practicality and reliability of the model. Experiments have shown that our proposed method is more efficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 Shown is a flow chart of the airship path planning method based on deep reinforcement learning. DETAILED DESCRIPTION

[0061] The technical solution of this application is described in detail below with reference to the accompanying drawings.

[0062] This paper proposes a deep reinforcement learning-based path planning method for uncertain wind fields, namely the Proximal Policy Optimization algorithm integrating wind field information (IWI-PPO), to solve the airship path planning problem under uncertain wind fields.

[0063] like Figure 1As shown in the figure, the uncertain wind field path planning method based on deep reinforcement learning provided in this application utilizes a policy network to generate an airship path plan. The policy network includes three core capabilities: wind field fusion, wind field feature extraction, and a fully connected network. The policy network takes heterogeneous data as input, including dynamic wind field information and the airship's own state information. Specifically, the airship's own state information includes key parameters such as the airship's position coordinates, speed, heading angle, current time, and remaining energy. Dynamic wind field information is obtained by cropping a wind field matrix window centered on the airship's current position. The wind field data within this window undergoes wind field fusion processing to generate a fused wind field matrix, and simultaneously quantifies the wind field uncertainty. The fused wind field matrix and uncertainty information are further subjected to wind field feature extraction. A two-layer convolutional neural network is used to extract effective spatial features from the complex wind field data. These spatial features are then converted to one-dimensional vectors and fused with the airship's own state information to form a comprehensive state vector. This heterogeneous data fusion mechanism ensures that the intelligent agent can simultaneously consider the airship's real-time state and dynamic environmental changes, thereby enhancing its adaptability to complex environments. Finally, the integrated state vector is calculated through a fully connected network, outputting a Gaussian distribution for the airship's propulsion speed and heading angle. The Gaussian distributions are determined by two parameters: variance and expectation. Based on these outputs, the policy network samples the propulsion speed and heading angle for the airship's maneuvers.

[0064] The following introduces the input information of the strategy network, wind farm fusion and reinforcement learning algorithm design respectively.

[0065] 1. Airship status information

[0066] The airship's own status information includes key parameters such as the airship's position coordinates, speed, heading angle, current time, and remaining energy.

[0067] Environmental simulation provides a tight and realistic flight environment for airships flying in the stratosphere. Airships flying in the stratosphere are primarily affected by wind fields and are subject to energy constraints. Environmental simulation includes kinematic simulation and airship energy simulation.

[0068] (1) Kinematic simulation

[0069] The operational simulation environment mainly consists of the wind field and airship dynamics models. The position of the airship can be expressed as ,in is the longitude, is the latitude.

[0070] The wind field is the set of wind speed and direction at each location in space. The wind speed at a location can be expressed as:

[0071] (1)

[0072] in, They are Location longitude and latitude The wind speed component at the moment.

[0073] The propulsion speed of the airship is relative to the air speed, which is recorded as , the magnitude of the propulsion speed and heading The output of the policy network is obtained. Since the airship will be affected by the stratospheric wind during flight, the impact of the wind field on the airship's motion must be considered. The actual speed of the airship relative to the ground Is the propulsion speed and wind speed The vector sum of:

[0074] (2)

[0075] (3)

[0076] The position of the airship changes over time, is the initial position of the airship, and the airship is at time The position is obtained by integrating the actual velocity:

[0077] (4)

[0078] In discrete time steps Next, the position at the next moment is updated to:

[0079] (5)

[0080] For the airship at the moment The actual velocity of the position is . In this way, we get the transfer formula of the airship position.

[0081] (2) Energy simulation

[0082] What distinguishes airship path planning from other missions is energy management. Airship missions are typically long, potentially exceeding seven days or even longer. The energy reserves in the airship's batteries are insufficient to sustain such extended propulsion. Furthermore, onboard electronic equipment, such as environmental monitoring equipment and controllers, consume energy throughout the day, necessitating the airship's own charging capabilities. Airships typically rely on onboard solar panels for power, but solar charging can result in varying charging power levels at different times of the day due to varying solar radiation angles. Therefore, we developed an airship energy model that incorporates both airship energy consumption and solar charging.

[0083] When an airship flies in the air, it needs to overcome air resistance to maintain a constant speed. The power that the motor needs to provide is proportional to the cube of the speed. Motor power consumption at all times It can be calculated using the following formula:

[0084] (6)

[0085] in, The coefficient of the work done by the airship's motor. Avionics equipment (such as navigation, communication, sensors, etc.) consumes energy continuously during flight and can be regarded as constant power consumption. .

[0086] Solar energy is one of the main energy sources for airships, and solar charging capacity depends on solar irradiance. , conversion efficiency and time (day-night cycle). The solar radiation at the airship's altitude is not blocked by clouds.

[0087] (7)

[0088] in, Indicates time The solar energy effectiveness factor represents the actual proportion of solar energy available to simulate the day and night cycle, W It can be expressed as the following function:

[0089] (8)

[0090] in, The daylight hours are 12 hours. and The sunrise time and the current time are in hours, respectively. The energy of the battery changes with the input (solar power ) and output (motor power consumption and avionics power consumption )Decide.

[0091] (9)

[0092] in, is the maximum capacity of the battery, at the current moment The battery level is , the new energy state of the battery at the next moment is:

[0093]

[0094] (10)

[0095] when The airship cannot continue to operate.

[0096] 2. Gaussian spatial enhanced wind field fusion

[0097] This application proposes a Gaussian spatially enhanced wind field fusion method that combines precise wind measurements at the airship's current location with forecast wind field data. This method also quantifies the uncertainty of the wind field at each location, providing a new perspective and basis for optimizing decision-making in path planning. The core idea is to adjust the forecast wind field based on the difference between the measured and forecast values ​​at the airship's current location and propagate this adjustment spatially using a decaying influence function. The specific algorithm steps are as follows:

[0098] First calculate the measured wind speed at the current position of the airship Forecast wind speed at the same location The difference between , providing a benchmark difference value for subsequent wind farm adjustments. It can reflect the error scale and overall situation between the current actual wind field and the forecast wind field.

[0099] (11)

[0100] For each position in the window , the algorithm calculates the center position Euclidean distance , providing the necessary spatial distance information for the calculation of the influence function.

[0101] (12)

[0102] (13)

[0103] Use Gaussian function as influence function , this function decreases as the distance from the center increases when the error is determined, reflecting the impact of the error value on the wind field adjustment. Is a parameter that controls the scope of influence. Using the influence function and the difference , algorithm adjusts the forecast wind field To obtain the updated wind field .

[0104] (14)

[0105] Furthermore, the uncertainty of each position is calculated .

[0106] (15)

[0107] (16)

[0108] A matrix with the same size as the wind field is used to represent the uncertainty. The uncertainty of each location represents the uncertainty of the error between the wind field prediction value at the current location and the actual wind field. The uncertainty is determined by the basic uncertainty of the forecast wind field itself. and the uncertainty calculated based on the influence function of error and distance Determine. Uncertainty increases with distance and error The increase in the size of reflects the decrease in the trust in the predicted wind field at locations that are farther away from the measured value and have larger errors. The uncertainty represents the trustworthiness of the wind field at the current location. In this method, the uncertainty matrix is ​​used as part of the environmental state as the input of the strategy network.

[0109] 3. Reinforcement learning algorithm design

[0110] This application uses the proximal strategy optimization method to train the strategy network.

[0111] (1) State space

[0112] For the wind field information around the airship, this application uses a sliding window centered on the airship to obtain it, which can better obtain the wind field information required by the airship. The size of the sliding window has a certain impact on the path planning performance. If the window size is too large, the calculation efficiency may be too low and the training is difficult to converge. The wind field information that is too far away from the current position is not very helpful for the wind field path planning. If the window size is too small, the path planner may not be able to obtain enough wind field data, resulting in the inability to obtain the optimal path. The minimum sliding window is 1×1, and the maximum is the size of the entire map. In actual tests, the size of the sliding window ranges from 1×1 to 64×64.

[0113] The airship can detect the wind field information of the current position in real time while obtaining the forecast wind field information window of the current position. The two are fused as the window of wind field information, and the obtained uncertainty matrix is ​​used as part of the airship status information.

[0114] (2) Action Space

[0115] In order to adapt the airship to complete a variety of tasks, we use a continuous action space. According to the maneuverability characteristics of the airship, we select the flight speed and heading as the action space, respectively. : relative to the magnitude of the air propulsion velocity and : represents the propulsion heading angle. The policy network can output continuous values ​​for speed and heading, enabling precise control of the airship. Due to the airship's weak power, large size, and weight, its acceleration and steering capabilities are limited. Therefore, we constrained the airship's action space.

[0116] (17)

[0117] (18)

[0118] in, is the heading angle change, For the velocity variation, this application adds two constraints (19) and (20):

[0119] (19)

[0120] (20)

[0121] (3) Reward function design

[0122] The reward function is mainly divided into the following parts:

[0123] 1) Rewards for approaching the target

[0124] (twenty one)

[0125] in:

[0126] (twenty two)

[0127] Among them, the distance Represents the distance from the current position to the target position. After the airship takes an action in a certain state, the distance from the current position to the target position is recalculated. The reward setting can motivate the airship to move towards the target at every step, which can effectively solve the problem of sparse rewards in path planning tasks.

[0128] 2) Energy consumption penalty

[0129] (twenty three)

[0130] Among them, the current battery energy state is , as the airship consumes electricity and charges with solar energy during the flight. The power level of the airship will change constantly. To ensure that the airship has sufficient power to maintain the normal operation of the avionics equipment and the ability to respond to emergencies, it is necessary to maintain a certain energy reserve. When the power level of the airship falls below a certain threshold, it is considered unsafe and will be penalized.

[0131] 3) Rewards for reaching the target and penalties for reaching the boundary

[0132] (twenty four)

[0133] in, is the location of the target point. After the airship reaches the target area, it receives a large positive reward, motivating it to pursue the long-term goal of reaching the target location even when energy is insufficient. If the airship flies outside the current wind farm area, it is given a large penalty and returns an environmental abort flag, preventing the airship from rapidly flying out of bounds to minimize the penalty.

[0134] 4) Stability Rewards

[0135] Maintaining smooth flight by reducing the airship's frequent speed and direction changes.

[0136] (25)

[0137] in, and are the weight coefficients for direction change penalty and speed change penalty, which are 0.8 and 1 respectively. is the change in advancement direction between the current step and the previous step, The propulsion speed change between the current step and the previous step is penalized for frequent speed and direction adjustments. This encourages the airship to plan a smooth flight path, reducing energy waste and mechanical losses.

[0138] Combine different rewards according to different parameters:

[0139] (26)

[0140] in, =1, =1, =1, =5.

[0141] The uncertain wind farm path planning method based on deep reinforcement learning provided in this application was used for simulation calculations, and the results are as follows:

[0142] The proposed method is implemented using the pytorch framework. Training is performed in a simulation environment using a real wind field. The wind field in East Asia from 2021 to 2023 is used for training, and the wind fields within the range of 0-40 degrees north latitude and 100-140 degrees east longitude are intercepted in different seasons. The real wind field data comes from the ERA5 reanalysis data, and the forecast data comes from the wind field forecast data of the National Space Science Center. A discrete 160*160 grid is used to represent the wind field, each grid represents 0.25 degrees of longitude and latitude, and the time resolution is 1 hour. The size of the wind field window in the state space is 32*32, and the time step is 15 minutes.

[0143] Table 1 summarizes the learning algorithm parameters. The deployed policy network was trained using a GPU for 30 hours and 5 million simulations. Data was collected using 10 parallel simulation environments. Training was performed on Windows 10 with 64GB of RAM and an RTX 4090 GPU.

[0144] Table 1 Training parameters

[0145]

[0146] To verify the performance and effectiveness of the proposed method, two different approaches were tested for comparison: D3RQN and a heuristic algorithm. The D3RQN algorithm uses a discrete action space and takes global wind forecast information as state input. The heuristic algorithm (HA) represents a representative traditional approach. By comparing the proposed method with D3RQN and the heuristic algorithm, we can comprehensively evaluate the performance advantages of the algorithms in terms of path planning efficiency, energy status, and environmental adaptability.

[0147] This application selects a typical test scenario of 2024.6 for visual display. The starting point is at 133.25 degrees east longitude and 8.5 degrees north latitude, and the end point is at 134 degrees east longitude and 34 degrees north latitude. In this scenario, this application compares and analyzes the path planning performance of the three algorithms provided by this application (IWI-PPO), HE and D3RQN, and the arrival time at the end point is 85h, 96.5h and 92.5h respectively. The path planned by the method of this application can reach the end point earlier and the path is smoother. The reason is that the algorithm proposed in this application better grasps the characteristics of the wind field and moves in the direction of the wind field, thereby effectively saving time and energy.

[0148] The present application also provides an airship path planning system based on deep reinforcement learning, which is implemented based on the above method. The system includes:

[0149] The environmental simulation module is used to provide the airship simulation operation environment and generate dynamic wind field information and airship status information;

[0150] A policy network, used to plan the propulsion speed and heading angle of the airship;

[0151] The path planning module is used to input dynamic wind field information and the airship's own state information into the trained strategy network to obtain the planned airship propulsion speed and heading angle;

[0152] The training module is used to train the policy network through the proximal policy optimization method.

[0153] The present application may also provide a computer device comprising: at least one processor, memory, at least one network interface, and a user interface. The various components in the device are coupled together via a bus system. It will be understood that the bus system is used to enable communication between these components. In addition to a data bus, the bus system also includes a power bus, a control bus, and a status signal bus.

[0154] The user interface may include a display, a keyboard, or a pointing device, such as a mouse, a trackball, a touchpad, or a touch screen.

[0155] It is understood that the memory in the embodiments disclosed in the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memory.

[0156] In some embodiments, the memory stores the following elements, executable modules or data structures, or a subset or an extension thereof: an operating system and applications.

[0157] The operating system includes various system programs, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and handle hardware-based tasks. Application programs include various application programs, such as media players and browsers, which are used to implement various application services. The program that implements the method of the embodiment of the present disclosure can be included in the application program.

[0158] In the above embodiment, the processor may also call a program or instruction stored in the memory, specifically, a program or instruction stored in the application program, to:

[0159] Perform the steps of the above method.

[0160] The above method can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The above-disclosed methods, steps, and logic block diagrams can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the above-disclosed method can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0161] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, or other electronic units or combinations thereof for performing the functions described herein.

[0162] For software implementation, the technology of the present application can be implemented by executing the functional modules (e.g., procedures, functions, etc.) of the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0163] The present application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, each step in the above method embodiment can be implemented.

[0164] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of this application and are not intended to limit the scope of the present invention. Although this application has been described in detail with reference to the embodiments, it should be understood by those skilled in the art that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application and should be encompassed by the claims of this application.

Claims

1. A deep reinforcement learning-based airship path planning method, comprising: The dynamic wind field information and the airship's own state information are input into the trained strategy network to obtain the planned Gaussian distribution of the airship's propulsion speed and heading angle. The planned propulsion speed and heading angle of the airship are obtained by sampling the Gaussian distribution. The policy network is trained by a proximal policy optimization method; The dynamic wind field information includes the measured wind speed, forecast wind field data and uncertainty at the current position of the airship; The forecast wind field data is obtained through a sliding window centered on the airship; The uncertainty stated for: ; in, It represents the basic uncertainty of the forecast wind field itself; Represents the influence function: ; in, Represents each position in the sliding window Euclidean distance to the airship's location; Parameters that represent the scope of control influence; Represents the difference between the measured wind speed at the airship's current location and the forecast wind speed at the same location; The forecast wind field is updated based on the error between the current actual wind field and the original forecast wind field: ; in, represents the updated forecast wind field; Represents the original forecast wind field.

2. The airship path planning method based on deep reinforcement learning according to claim 1, characterized in that: The airship's own state information includes the airship's speed, navigation, position and energy information; In time interval Energy changes in the airship battery Expressed as: ; in, represents the power consumption of avionics equipment; Indicates solar power generation power: ; in, represents solar irradiance; represents the solar energy conversion efficiency; Indicates time Solar efficiency factor: ; in, Indicates that the daytime is 12 hours long. Indicates sunrise time; express Motor power consumption at each moment: ; in, The coefficient that represents the work done by the airship motor; Indicates that the airship is at time The actual speed of the position; Battery energy status at all times for: ; ; in, express t The energy status of the battery at all times; Indicates the maximum capacity of the battery.

3. The airship path planning method based on deep reinforcement learning according to claim 1, characterized in that: The policy network includes: The wind field fusion module is used to fuse the measured wind speed at the current position of the airship with the forecast wind field data to obtain the fused wind field matrix and uncertainty; The wind field feature extraction module includes a two-layer convolutional network, which is used to extract wind field features from the fused wind field matrix and uncertainty; The fully connected network module is used to calculate the comprehensive state vector after integrating the wind field characteristics and the airship's own state, and output the Gaussian distribution of the planned airship propulsion speed and heading angle.

4. The airship path planning method based on deep reinforcement learning according to claim 1, characterized in that: The reward function when the proximal policy optimization method is trained for: ; in, 、 、 and Represents the weight coefficient of each reward function; Represents a reward function close to the target: ; Where: Distance Indicates time t The distance from the airship position to the target position; Represents the energy consumption penalty function: ; in, express t Time battery energy status; Indicates the maximum capacity of the battery; Represents the goal reward and boundary penalty: ; Denote the stability reward function: ; in, and The weight coefficient of the direction change penalty and the weight coefficient of the speed change penalty; is the change in advancement direction between the current step and the previous step; is the change in advancement speed between the current step and the previous step.

5. A deep reinforcement learning-based airship path planning system, implemented based on the method of any one of claims 1-4, characterized in that: The system comprises: The environmental simulation module is used to provide the airship simulation operation environment and generate dynamic wind field information and airship status information; A policy network, used to plan the propulsion speed and heading angle of the airship; A path planning module is used to input dynamic wind field information and the airship's own state information into the trained strategy network to obtain the planned airship propulsion speed and heading angle; and The training module is used to train the policy network through the proximal policy optimization method.

Citation Information

Patent Citations

  • Unmanned airship intelligent flight planning method based on meteorological data

    CN116522802A