Evaporation waveguide channel map-oriented ship multi-waypoint path planning method
By constructing a ship multi-waypoint path planning method based on an evaporating waveguide channel map, and combining a near-end policy optimization (PPO) algorithm with offline generalization training and online real-time inference, the real-time decision-making problem in dynamic maritime data collection tasks is solved, achieving efficient and stable multi-waypoint path planning and meeting the time constraints of multi-node data collection and transmission.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-10
- Publication Date
- 2026-04-07
AI Technical Summary
Existing path planning methods struggle to achieve real-time decision-making and dynamic adaptation when faced with dynamic, multi-constraint, multi-waypoint maritime data collection tasks. Traditional convex optimization methods have high computational complexity, metaheuristic algorithms are computationally complex and time-consuming, and deep reinforcement learning methods lack training stability in complex communication-motion coupling scenarios. These methods cannot meet the needs of ships for online trajectory replanning and real-time decision-making in dynamic maritime environments.
A multi-waypoint path planning method for ships based on evaporating waveguide channel maps is adopted. By establishing a multi-segment segmented transmission motion model and communication model for ships, a channel gain map is constructed. A near-end policy optimization algorithm (PPO) that integrates offline generalization training and online real-time inference is proposed. The generalization policy model is trained offline, and the current state is input in real time online for rapid decision-making, so as to achieve efficient and stable response in multi-waypoint maritime operations.
It significantly reduces the computation time and convergence cost of traditional methods, supports real-time decision-making in dynamic environments, can quickly output optimal trajectories and transmission schemes, adapts to the addition or removal of waypoints and adjustment of data volume, and meets the time constraints of multi-node data collection and transmission.
Smart Images

Figure CN121809801A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, specifically a method for ship multi-waypoint path planning based on evaporating waveguide channel maps. Background Technology
[0002] Evaporated waveguides are a special atmospheric phenomenon that causes superrefractive effects that confine electromagnetic wave energy between the waveguide layer and the sea surface, significantly reducing path loss over long distances and thus improving communication capabilities and coverage between ships and shore-based base stations. Channel knowledge graphs (CKM) and their derived channel gain graphs (CGM), as key technologies for environmental cognitive communication, can integrate historical data with physical models (such as evaporated waveguide models and electromagnetic propagation models) to construct a priori maps reflecting the characteristics of the marine wireless environment. This makes it possible to predict channel conditions and optimize ship trajectories before communication begins.
[0003] Existing path planning methods generally face challenges in real-time decision-making and dynamic adaptation when dealing with dynamic, multi-waypoint maritime data collection tasks with multiple constraints. These challenges are mainly reflected in the following aspects:
[0004] Convex optimization methods heavily rely on the convex structure of the problem. For complex non-convex problems such as ship motion, channel nonlinearity, and multi-objective trade-offs, they are usually difficult to apply directly or require extensive simplification, which may result in a loss of optimality of the solution. Furthermore, the computational delay of online replanning is relatively large.
[0005] Metaheuristic algorithms: While these algorithms do not rely on problem convexity and can handle complex models, they are essentially offline global search methods for specific problem instances. Whenever the task scenario changes (e.g., waypoint additions or subtractions, data volume adjustments, or fluctuations in evaporation waveguide height), the population needs to be reinitialized and a complete iterative optimization performed. This process is computationally complex and has significant convergence time, making it difficult to support millisecond- or second-level online trajectory replanning and real-time decision-making for ships in dynamic maritime environments.
[0006] Deep reinforcement learning (DRL) methods offer a new paradigm for solving decision-making problems, theoretically possessing the ability to learn online and adapt to dynamic changes through interaction with the environment. However, some early DRL algorithms suffered from shortcomings in training stability and sample efficiency, affecting their reliable application in complex communication-motion coupling scenarios.
[0007] Document CN 120074715 A discloses a multi-objective ship communication path planning method based on marine waveguide channel maps, comprising the following steps: establishing a single-input single-output ship motion model and communication model for long-distance transmission at sea; establishing a CGM based on existing gas phase data, evaporation waveguide models, and electromagnetic wave propagation models; proposing an optimization problem model; proposing the NSGA-II-PSO-VS algorithm; comparing the running results of the NSGA-II algorithm and the NSGA-II-PSO-VS algorithm at the same time; and conducting simulation experiments at different evaporation waveguide heights to prove the rationality and applicability of the proposed algorithm. Essentially, it is an "offline population iterative search"—requiring re-initialization of the population, crossover mutation, and particle swarm optimization for a single scenario, relying on multiple iterations to obtain the optimal solution. The method in this document focuses on "single path planning," supporting only the optimization of a single trajectory from "start point → end point," without constraints related to transfer points, and is only applicable to simple direct communication tasks. Moreover, the offline iterative mode of the patent results in a long decision-making time (re-planning is required for scenario changes), making it unable to cope with dynamic scenarios such as the addition or reduction of waypoints and adjustments in data volume.
[0008] Therefore, a new technical solution is needed to solve the above-mentioned technical problems. Summary of the Invention
[0009] To address the aforementioned issues, this invention discloses a ship multi-waypoint path planning method for evaporating waveguide channel maps. This method can handle complex non-convex models like a metaheuristic algorithm and achieve efficient and stable training and online real-time decision-making like an ideal DRL method. It can more effectively respond to multi-waypoint maritime operation scenarios, simultaneously optimize data transmission time and ship sailing time, and focus on solving key problems such as multi-node data collection and transmission time constraints for each segment.
[0010] The technical solution of this invention is: a ship multi-waypoint path planning method based on evaporating waveguide channel maps, comprising the following steps:
[0011] S1: In the scenario of multi-node data collection at sea, establish a single-input, single-output segmented transmission motion model and communication model for multi-segment ship navigation.
[0012] S2: Construct a channel gain map using an electromagnetic wave propagation model based on the parabolic equation;
[0013] S3: Construct a multi-objective optimization problem model with multi-segment constraints;
[0014] S4: A near-end policy optimization algorithm that integrates offline generalization training and online real-time inference is proposed;
[0015] S5: Visualize and verify the multi-waypoint trajectory output results of the PPO algorithm to demonstrate its adaptability and segmented transmission reliability in multi-node transmission scenarios.
[0016] The specific steps of S1 are as follows:
[0017] For a multi-waypoint segmented data transmission scenario, a maritime communication system consisting of ship users and shore-based base stations has N relay points at sea, denoted as... The ship from its starting position Starting from there, sail in sequence through [location]. point, Point until the destination is reached. A point refers to any position a ship is at during navigation. Collected size The data is collected and transmitted to the base station, and then transmitted to the next location. The data transmission will be completed before the process is finished;
[0018] To simplify the problem and focus on ship trajectory optimization, it is assumed that both the base station and the ship user are equipped with a single antenna for communication, and that the ship is at any waypoint Collecting data here does not take time;
[0019] The goal is to minimize data transmission time and SU's travel time by optimizing the ship's trajectory and utilizing information from the marine evaporative waveguide environment. Next, the coordinate system, ship motion model, and signal model are first described, and then an optimization problem is constructed for the data transmission task under consideration.
[0020] Assuming the height of the transmitting antenna is The height of the receiving antenna is ,starting point and the end point The coordinates are defined as follows: and The shore-based base station is located at coordinates At the location, the ship was at The location of each time slot is:
[0021] ;
[0022] in Indicates the maximum permissible navigation time slot, the ship in the... The velocity of each time slot is denoted as The angle between its velocity vector and the x-axis is denoted as . Therefore, the velocity components of the ship in the x and y directions are as follows:
[0023] ;
[0024] in , and They are respectively The lower and upper limits are constrained by physical conditions, and the ship's maneuverability is limited. For the velocity direction between any two adjacent time slots, the steering angle must satisfy the following constraints:
[0025] ;
[0026] In the In each time slot, the received signal at the base station is:
[0027] ;
[0028] in This is the channel between ship users and base stations. In order to transmit signals, This represents additive white Gaussian noise. Furthermore, it is assumed that the transmit power is constant, i.e.:
[0029] ;
[0030] Received signal power It can be represented as:
[0031] ;
[0032] in This indicates the path loss between the ship user and the base station. Indicates the transmit antenna gain. Let represent the receiving antenna gain. In this case, we express the maximum achievable transmission rate as:
[0033] ;
[0034] in, For channel bandwidth, This represents the noise power spectral density.
[0035] The specific steps of S2 are as follows:
[0036] The electromagnetic wave propagation in the evaporating waveguide environment is modeled using the parabolic equation method. For any point in the environment, the large-scale channel gain relative to the base station is calculated using PE. This gain mainly characterizes the path loss.
[0037] The three-dimensional space is divided into many grids, each grid having a width of [missing value]. The vertical grid height is Within this grid, the channel state remains unchanged. Assuming the base station is located at the origin, the maximum effective radius of the environment in the horizontal direction is... ,in The integer represents the characteristic height of the evaporation waveguide in the vertical direction. ,in It is also an integer, thus each grid node can be uniquely indexed, and its coordinates are:
[0038] ;
[0039] In the above formula, , and same, = , will any position The channel gain at point is defined as Because the position of the ship in any time slot is Therefore, based on the grid to which the ship's position belongs at this moment, the large-scale channel gain between the base station and the ship is obtained. .
[0040] In S3, the objective is to optimize the ship's trajectory to minimize its transmission and arrival times, assuming the total sailing time slots are... The total transmission time slots are The transmission time slot for each trajectory segment of the ship is The sailing time slot for each segment is ,in This serves as an index for the current segment of the ship's trajectory. The ship's trajectory is The multi-objective optimization problem is expressed as:
[0041] ;
[0042] ;
[0043] ;
[0044] ;
[0045] ;
[0046] ;
[0047] ;
[0048] ;
[0049] ;
[0050] ;
[0051] ;
[0052] ;
[0053] ;
[0054] In the above formula, This means that the ship must transmit all data before reaching its destination, and the total transmission time is less than the total sailing time.
[0055] This indicates that the total sailing time is less than the maximum allowed sailing time;
[0056] Indicates the first The transmission time slot of the ship in the segment trajectory is less than or equal to the navigation time slot;
[0057] This means that the sum of the transmission time slots of all segments equals the total transmission time slots;
[0058] This means that the sum of all the segmented time slots equals the total time slot.
[0059] Denotes the evolution of the queue representing data transmission constraints, in the... The amount of data to be transmitted in the next time slot of a segment trajectory is the amount of data to be transmitted in the current time slot minus the amount of data transmitted in that time slot, until the amount of data to be transmitted in that segment trajectory is... ,in and They represent the first time. The first segment of the trajectory The and the first Data to be transmitted in each time slot For the first Data transmission rate per time slot The length of the time slot, That is, the first The amount of data transmitted per time slot;
[0060] This indicates the size of the data to be transmitted in each segment, where For the first The amount of data to be transmitted in the segment trajectory;
[0061] Indicates the first The amount of data to be transmitted at the end of each flight segment is ;
[0062] Indicates the starting position constraint of the vessel;
[0063] Indicates the constraints on the location of the ship's transfer point;
[0064] Indicates the endpoint position constraint of the vessel;
[0065] This indicates the ship's steering angle constraint.
[0066] The specific steps of S4 are as follows:
[0067] The proposed online real-time decision-making algorithm based on near-end policy optimization mainly consists of two stages:
[0068] Offline training phase: The agent, i.e. the ship, in a simulated evaporating waveguide channel environment, learns the complex relationship between channel fluctuations, motion constraints and multi-segment transmission tasks through a large number of trial and error explorations, and trains a unified policy model with strong generalization ability. This phase is completed in one go and does not require real-time performance.
[0069] Online inference phase: During actual navigation, the ship deploys the trained strategy model on the shipboard computing unit. In each decision time slot, the system inputs the current state in real time, and the strategy model instantly performs forward inference and outputs the optimal course adjustment action. This process has extremely low computational cost, realizing a fundamental shift from "scenario replanning" to "real-time state response", meeting the real-time decision-making needs in dynamic environments. In addition, the online phase inputs different scenario parameters than the previous model.
[0070] By adopting the above technical solution, the algorithm can quickly output the optimal trajectory and transmission scheme through strategy reasoning, which significantly reduces the computation time and convergence cost of traditional methods in multi-point tasks, and provides technical support for real-time decision-making.
[0071] In particular, step S4 models the problem as a Markov decision process: ;
[0072] 1) State Space :
[0073] ;
[0074] Normalized coordinates, range: ;
[0075] : The normalized distance to the current target point, ranging from ;
[0076] The current course of the vessel, normalized range is: ;
[0077] The angle error between the current heading and the direction pointing to the current target, with a normalized range of [missing information]. ;
[0078] The current flight segment's "remaining data percentage / total percentage" is normalized to a range of [missing information]. ;
[0079] : Time progress, which is "time sailed / maximum allowed sailing time", with a range of ;
[0080] : Segment index normalization, when the ship is arrive During the flight segment, its normalized index is , range ;
[0081] 2) Action Space Because the ship's turning angle is constrained, its maneuver space is also limited. Therefore, a distribution with a range of values, such as a Gaussian distribution, is adopted. When dealing with continuous probability distributions, action pruning is required, which can cause problems in subsequent logarithmic probability calculations. Therefore, a range of [missing information] was adopted. The Beta distribution is first obtained through sampling. , , and Given the parameters of the Beta distribution, we then perform motion mapping to control the motion range within the ship's heading angle limits, resulting in... At this time, the range of motion is ;
[0082] 3) State transition ;
[0083] 4) Reward Function The reward function consists of three phases, including the data transmission phase. Used to balance the two parts of transportation and navigation; the transportation completion stage. Phase 1: The vessel rapidly and directly reaches the target point; Phase 2: The vessel reaches the destination. And all data has been transmitted. When the voyage is successful, or when the destination is not reached within the stipulated time. And data transmission failure upon reaching any waypoint However, if the voyage fails, corresponding penalties will be imposed.
[0084] The overall reward function expression is:
[0085] ;
[0086] in As a reward for the data transmission phase, As a reward for completing the transmission phase, As a reward for a failed voyage, As a reward for a successful voyage, , , , The judgment conditions for each stage are expressed as follows:
[0087] ;
[0088] ;
[0089] ;
[0090] ;
[0091] Reward function during data transmission phase Represented as:
[0092] ;
[0093] in , , The transmission reward, navigation progress reward, and distance and time penalties in this process are represented as follows:
[0094] ;
[0095] ;
[0096] ;
[0097] in Indicates the transmission weight parameter. The amount of data transmitted within a single timeslot. For exploration bonuses, it is represented as:
[0098] ;
[0099] Its reward value is limited to between, This represents the average transmission rate over the most recent 10 time slots. This indicates the highest transmission rate within the last 10 time slots. To avoid terms with a denominator of 0, middle For navigation weight parameters, This represents the actual distance to the target point within a single time slot. This indicates the distance from the target point before the current time slot. This indicates the distance from the target point after the current time slot has moved. Indicates the error between the ship's direction and the direction of the target point. The closer the current direction is to the target point's direction at the same angle, the better. The smaller the value, middle This represents the current distance from the destination. This represents the distance between the starting point and the target point. The number of time slots already navigable. , , Here, 400 represents the weighting factor, and 400 represents the ship's travel distance within a single time slot.
[0100] Reward function for transmission completion phase for:
[0101] ;
[0102] in , , , The navigation progress reward, heading reward, angle penalty, and time penalty in this process are represented as follows:
[0103] ;
[0104] ;
[0105] ;
[0106] ;
[0107] in The reward value has been further increased to facilitate ships' navigation to the target point. This indicates a heading reward. A positive heading reward is given when the actual distance between the ship and the target point decreases; otherwise, it is 0. This is to prevent ships from deviating from the target point in order to obtain a heading reward. Angle penalties are imposed to prevent ships from making large changes in angle. As a time penalty, it compels ships to reach the target point as quickly as possible. , , These are the weighting coefficients.
[0108] The termination reward for a failed voyage is represented as follows:
[0109] ;
[0110] in, The penalty for a failed voyage is set to a negative value as a hard constraint penalty;
[0111] The termination reward upon successful navigation is represented as follows:
[0112] ;
[0113] in To terminate slot rewards, it is represented as:
[0114] ;
[0115] in These are the weighting coefficients. For weighted rewards, in order to obtain possible non-convex Pareto fronts, the reward value here is expressed using Chebyshev decomposition as follows:
[0116] ;
[0117] To achieve the ideal sailing time slot, a smaller value is set to encourage ships to reach their destination in the shortest possible time, thus reducing the penalty. For the ideal transmission time slot, it is also set to a smaller value to help ships shorten transmission time. These are the weighting coefficients.
[0118] Preferably, the specific process of the algorithm is as follows:
[0119] (1) During the initialization phase, the parameters are configured, and the environment and preference conditions of the PPO agent are created. The actor and critic networks in the agent are both inputted with the state and preference vector [α, β].
[0120] (2) The training is divided into multiple predetermined weight stages, and the learning rate is changed or a pre-trained model is loaded as needed when switching stages;
[0121] (3) The agent samples an action using a Beta distribution based on the current environmental state and the preference vector of the current stage;
[0122] (4) The agent interacts with the environment, calculates the reward according to the reward function under the current preference weight, and moves to the next state, and stores the data of this interaction in the experience buffer;
[0123] (5) When the accumulated data reaches the preset trajectory length, the agent uses this data to calculate the advantage function through generalized advantage estimation and performs multiple rounds of mini-batch gradient updates to optimize the network parameters of the actors and critics;
[0124] (6) When the training step target of the current weight setting stage is reached, switch to the next preference weight stage and reset the parameters to continue training;
[0125] (7) Repeat this process until all weight stages have been trained.
[0126] Preferably, the trajectory of the PPO algorithm is plotted to prove the rationality and applicability of the proposed algorithm.
[0127] The advantages of this invention are as follows: 1. Compared with traditional swarm intelligence algorithms (such as NSGA-II, PSO, etc.), which require time-consuming and independent iterative searches for each specific scenario, the PPO framework proposed in this invention adopts a new paradigm of "offline training of general strategies and online real-time reasoning and decision-making". In the offline stage, the agent learns the characteristics of the evaporating waveguide channel and the rules of multi-waypoint tasks by interacting with the simulated environment, and trains a policy model with strong generalization. When applied online, only the current real-time state (position, channel quality, data margin) needs to be input into the model to achieve rapid reasoning and decision-making, which completely avoids the huge computational overhead of traditional methods that need to "re-plan" due to changes in waypoints, data volume or channel environment.
[0128] 2. This invention proposes a state-space modeling method that integrates multi-dimensional task progress (such as segment index, data completion rate, and time schedule), enabling the agent to accurately perceive task stages. The designed phased, multi-objective reward function not only balances the long-term goals of transmission and navigation, but its built-in exploration enhancement mechanism also incentivizes the agent to actively seek high-quality communication areas. This design enables the trained policy model to exhibit dynamic adaptability and autonomous trade-off capabilities when facing complex evaporating waveguide environments, achieving global collaborative optimization under segmented transmission time constraints.
[0129] 3. This invention verifies the effectiveness of the algorithm in dynamic scenarios and its real-time decision-making advantages through simulation. Experimental results show that the PPO-based strategy can converge stably under different weight preferences, the multi-waypoint trajectory is clearly visualized, and the ship can dynamically adjust its path according to real-time channel conditions, providing an efficient and reliable autonomous decision-making solution for dynamic maritime data collection tasks.
[0130] 4. This invention employs the Proximal Policy Optimization (PPO) deep reinforcement learning algorithm to construct a two-stage framework of "offline training + online inference". In the offline stage, a generalized policy model is trained through multi-weight scenarios. In the online stage, only real-time status (position, channel quality, remaining data) needs to be input to output decisions in milliseconds, completely eliminating the limitation of traditional algorithms that "require re-iteration when the scenario changes". It is specifically designed for "multi-waypoint segmented transmission scenarios", which require ships to pass through N transfer points in sequence, and each segment must complete the full transmission of data before reaching the next transfer point. It is suitable for complex multi-task scenarios such as marine monitoring and multi-node material transportation. The online inference mode can achieve "train once, reuse many times", supports real-time response to dynamic scenarios, and does not require retraining when adding waypoints or adjusting task parameters. The optimal trajectory can be quickly output through policy inference alone, thus improving decision-making efficiency. Attached Figure Description
[0131] Figure 1 This is a flowchart of an algorithm for a ship multi-waypoint path planning method based on an evaporating waveguide channel map, according to an embodiment of the present invention.
[0132] Figure 2 This is a system block diagram of a ship multi-waypoint path planning method based on an evaporating waveguide channel map according to an embodiment of the present invention.
[0133] Figure 3 This is a diagram of ship trajectories under different navigation and transmission weights in an embodiment of the present invention.
[0134] The accompanying drawings for this invention should be in color. Detailed Implementation
[0135] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0136] like Figure 1-3 As shown, a method for ship multi-waypoint path planning based on evaporating waveguide channel maps includes the following steps:
[0137] S1: In the scenario of multi-node data collection at sea, establish a single-input single-output (SISO) segmented transmission motion model and communication model for multi-segment ship navigation.
[0138] S2: Construct a channel gain map (CGM) using an electromagnetic wave propagation model based on the parabolic equation (PE).
[0139] S3: Construct a multi-objective optimization problem model with multi-segment constraints;
[0140] S4: Propose a proximal policy optimization (PPO) algorithm that integrates offline generalization training and online real-time inference;
[0141] S5: Visualize and verify the multi-waypoint trajectory output results of the PPO algorithm to demonstrate its adaptability and segmented transmission reliability in multi-node transmission scenarios.
[0142] Step S1 is as follows:
[0143] For a multi-waypoint segmented data transmission scenario, consider a maritime communication system consisting of ship users (SU) and shore-based base stations (BS). In this case, there are N relay points at sea, representing... The ship from its starting position Starting from there, sail in sequence through [location]. point, Point until the destination is reached. A point refers to any position a ship is at during navigation. Collected size The data is collected and transmitted to the base station as quickly as possible, and upon reaching the next location... Before data transmission is completed, to simplify the problem and focus on ship trajectory optimization, we assume that both the base station and the ship user are equipped with a single antenna for communication, and that the ship is at any waypoint. Collecting data here does not take time.
[0144] The goal is to minimize data transmission time and SU's travel time by optimizing the ship's trajectory and utilizing information from the marine evaporative waveguide environment. We first describe the coordinate system, ship motion model, and signal model. Then, we construct an optimization problem for the considered data transmission task.
[0145] Assuming the height of the transmitting antenna is The height of the receiving antenna is .starting point and the end point The coordinates are defined as follows: and The shore-based base station is located at coordinates At the location, the ship was at The location of each time slot is:
[0146] ;
[0147] in Indicates the maximum permissible navigation time slot. The ship is in the... The velocity of each time slot is denoted as The angle between its velocity vector and the x-axis is denoted as . Therefore, the velocity components of the ship in the x and y directions are as follows:
[0148] ;
[0149] in , and They are respectively The lower and upper limits are defined. Due to physical constraints, the ship's maneuverability is limited. For the velocity direction between any two adjacent time slots, the steering angle must satisfy the following constraints:
[0150] ;
[0151] In the In each time slot, the received signal at the base station is:
[0152] ;
[0153] in This is the channel between ship users and base stations. In order to transmit signals, This represents additive white Gaussian noise. Furthermore, it is assumed that the transmit power is constant, i.e.:
[0154] ;
[0155] Received signal power It can be represented as:
[0156] ;
[0157] in This indicates the path loss between the ship user and the base station. Indicates the transmit antenna gain. Let represent the receiving antenna gain. In this case, we express the maximum achievable transmission rate as:
[0158] ;
[0159] in, For channel bandwidth, This represents the noise power spectral density.
[0160] Step S2 is as follows:
[0161] The electromagnetic wave propagation in an evaporating waveguide environment is modeled using the parabolic equation method. For any point in the environment, the large-scale channel gain relative to the base station can be calculated using the PE, which mainly characterizes the path loss.
[0162] The three-dimensional space is divided into many grids, each grid having a width of [missing value]. The vertical grid height is Within this grid, the channel state remains unchanged. Assuming the base station is located at the origin, the maximum effective radius of the environment in the horizontal direction is... ,in The integer represents the characteristic height of the evaporation waveguide in the vertical direction. ,in It is also an integer, thus each grid node can be uniquely indexed, and its coordinates are:
[0163] ;
[0164] In the above formula, , and same, = , will any position The channel gain at point is defined as Because the position of the ship in any time slot is Therefore, based on the grid to which the ship's position belongs at this moment, the large-scale channel gain between the base station and the ship can be obtained. .
[0165] Step S3 is as follows:
[0166] We assume the total sailing time slots of the ship are The total transmission time slots are The transmission time slot for each trajectory segment of the ship is The sailing time slot for each segment is ,in This serves as an index for the current segment of the ship's trajectory. Ship tracks ;
[0167] The multi-objective optimization problem is represented as:
[0168] ;
[0169] ;
[0170] ;
[0171] ;
[0172] ;
[0173] ;
[0174] ;
[0175] ;
[0176] ;
[0177] ;
[0178] ;
[0179] ;
[0180] ;
[0181] In the above formula, This means that the ship must transmit all data before reaching its destination, and the total transmission time is less than the total sailing time.
[0182] This indicates that the total sailing time is less than the maximum allowed sailing time;
[0183] Indicates the first The transmission time slot of the ship in the segment trajectory is less than or equal to the navigation time slot;
[0184] This means that the sum of the transmission time slots of all segments equals the total transmission time slots;
[0185] This means that the sum of all the segmented time slots equals the total time slot.
[0186] Denotes the evolution of the queue representing data transmission constraints, in the... The amount of data to be transmitted in the next time slot of a segment trajectory is the amount of data to be transmitted in the current time slot minus the amount of data transmitted in that time slot, until the amount of data to be transmitted in that segment trajectory is... ,in and They represent the first time. The first segment of the trajectory The and the first Data to be transmitted in each time slot For the first Data transmission rate per time slot The length of the time slot, That is, the first The amount of data transmitted per time slot;
[0187] This indicates the size of the data to be transmitted in each segment, where For the first The amount of data to be transmitted in the segment trajectory;
[0188] Indicates the first The amount of data to be transmitted at the end of each flight segment is ;
[0189] Indicates the starting position constraint of the vessel;
[0190] Indicates the constraints on the location of the ship's transfer point;
[0191] Indicates the endpoint position constraint of the vessel;
[0192] Indicates the ship's steering angle constraints;
[0193] Step S4 is as follows:
[0194] The proposed online real-time decision-making algorithm based on near-end policy optimization mainly consists of two stages:
[0195] Offline training phase: In a simulated evaporating waveguide channel environment, the agent (ship) learns the complex relationship between channel fluctuations, motion constraints and multi-segment transmission tasks through extensive trial and error exploration, and trains a unified policy model with strong generalization ability. This phase is completed in one go and does not require real-time performance.
[0196] In the online inference phase: During actual navigation, the ship deploys the trained strategy model on its onboard computing unit. At each decision time slot, the system receives real-time input of the current state (normalized position, remaining data, channel gain, etc.), and the strategy model instantly performs forward inference, outputting the optimal heading adjustment action. This process has extremely low computational cost, achieving a fundamental shift from "scenario replanning" to "real-time state response," meeting the real-time decision-making needs in dynamic environments. Furthermore, the online phase can input scenario parameters different from those in the previous model (such as data volume and preference weights). The algorithm can quickly output the optimal trajectory and transmission scheme through strategy inference, significantly reducing the computation time and convergence cost of traditional methods in multi-point tasks, providing technical support for real-time decision-making.
[0197] The problem will then be modeled as a Markov decision process: ;
[0198] 1) State Space :
[0199] ;
[0200] Normalized coordinates, range: ;
[0201] : The normalized distance to the current target point, ranging from ;
[0202] The current course of the vessel, normalized range is: ;
[0203] The angle error between the current heading and the direction pointing to the current target, with a normalized range of [missing information]. ;
[0204] The current flight segment's "remaining data percentage / total percentage" is normalized to a range of [missing information]. ;
[0205] : Time progress, which is "time sailed / maximum allowed sailing time", with a range of ;
[0206] : Segment index normalization, when the ship is arrive During the flight segment, its normalized index is , range ;
[0207] 2) Action Space Because the ship's turning angle is constrained, its maneuver space is also limited. Therefore, a distribution with a range of values, such as a Gaussian distribution, is adopted. When dealing with continuous probability distributions, action pruning is required, which may cause problems in subsequent logarithmic probability calculations. Therefore, we adopted a range of... The Beta distribution is first obtained through sampling. , , and Given the parameters of the Beta distribution, we then employ motion mapping to control the motion range within the ship's heading angle limits, resulting in... At this time, the range of motion is ;
[0208] 3) State transition ;
[0209] 4) Reward Function The reward function consists of three phases, including the data transmission phase. Used to balance the two parts of transportation and navigation. Transportation completion phase. The ship quickly reaches the target point in a straight line; final stage: the ship reaches the destination. And all data has been transmitted. When the voyage is successful, or when the destination is not reached within the stipulated time. And data transmission failure upon reaching any waypoint However, if the voyage fails, corresponding penalties will be imposed.
[0210] The overall reward function expression is:
[0211] ;
[0212] in As a reward for the data transmission phase, As a reward for completing the transmission phase, As a reward for a failed voyage, As a reward for a successful voyage, , , , The judgment conditions for each stage are expressed as follows:
[0213] ;
[0214] ;
[0215] ;
[0216] ;
[0217] Reward function during data transmission phase Represented as
[0218]
[0219] in , , The transmission reward, navigation progress reward, and distance and time penalties in this process are represented as follows:
[0220] ;
[0221] ;
[0222] ;
[0223] in Indicates the transmission weight parameter. The amount of data transmitted within a single timeslot. For exploration bonuses, it is represented as:
[0224] ;
[0225] Its reward value is limited to between, This represents the average transmission rate over the most recent 10 time slots. This indicates the highest transmission rate within the last 10 time slots. To avoid terms with a denominator of 0. middle For navigation weight parameters, This represents the actual distance to the target point within a single time slot. This indicates the distance from the target point before the current time slot. This indicates the distance from the target point after the current time slot has moved. Indicates the error between the ship's direction and the direction of the target point. The closer the current direction is to the target point's direction at the same angle, the better. The smaller the value. middle This represents the current distance from the destination. This represents the distance between the starting point and the target point. This represents the number of time slots that have already been navigated. , , 400 represents the weighting coefficient, and 400 represents the ship's travel distance within a single time slot.
[0226] Reward function for transmission completion phase for:
[0227] ;
[0228] in , , , The navigation progress reward, heading reward, angle penalty, and time penalty in this process are represented as follows:
[0229] ;
[0230] ;
[0231] ;
[0232] ;
[0233] in The reward value has been further increased to facilitate ships' navigation to the target point. This represents a heading bonus. A positive heading bonus is awarded when the actual distance between the ship and the target point decreases; otherwise, it is zero. This prevents ships from deviating from the target point to obtain heading bonuses. Angle penalties are imposed to prevent ships from making large changes in angle. As a time penalty, it compels ships to reach the target point as quickly as possible. , , These are the weighting coefficients.
[0234] The termination reward for a failed voyage is represented as follows:
[0235] ;
[0236] in, The penalty for a failed voyage is set to a negative value as a hard constraint penalty.
[0237] The termination reward upon successful navigation is represented as follows:
[0238] ;
[0239] in To terminate slot rewards, it is represented as:
[0240] ;
[0241] in These are the weighting coefficients. For weighted rewards, in order to obtain possible non-convex Pareto fronts, the reward value here is expressed using Chebyshev decomposition as follows:
[0242] ;
[0243] For ideal navigation time slots, a small value is typically set to encourage ships to reach their destination in the shortest possible time, thus reducing the penalty. For the ideal transmission time slot, it is also set to a smaller value to help ships shorten transmission time. These are the weighting coefficients.
[0244] The specific process of the algorithm is as follows:
[0245] (1) During the initialization phase, configuration parameters are performed to create a PPO agent with environment and preference conditions. Both the actor and critic networks in the agent take the state and preference vector [α, β] as input;
[0246] (2) The training is divided into multiple predetermined weight stages, and the learning rate is changed or a pre-trained model is loaded as needed when switching stages;
[0247] (3) The agent samples an action using a Beta distribution based on the current environmental state and the preference vector of the current stage;
[0248] (4) The agent interacts with the environment, calculates the reward according to the reward function under the current preference weight, and moves to the next state, and stores the data of this interaction in the experience buffer;
[0249] (5) When the accumulated data reaches the preset trajectory length, the agent uses this data to calculate the advantage function through generalized advantage estimation and performs multiple rounds of mini-batch gradient updates to optimize the network parameters of the actors and critics;
[0250] (6) When the training step target of the current weight setting stage is reached, switch to the next preference weight stage and reset the learning rate, entropy coefficient and other parameters to continue training;
[0251] (7) Repeat this process until all weight stages have been trained.
[0252] Step S5 is as follows:
[0253] The trajectory of the PPO algorithm is plotted to prove the rationality and applicability of the proposed algorithm.
[0254] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are merely examples and do not limit the present invention; the objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been demonstrated and explained in the embodiments, and any modifications or variations of the embodiments of the present invention may be made without departing from the stated principles.
Claims
1. A method for ship multi-waypoint path planning based on evaporating waveguide channel maps, characterized in that, Includes the following steps: S1: In the scenario of multi-node data collection at sea, establish a single-input, single-output segmented transmission motion model and communication model for multi-segment ship navigation. S2: Construct a channel gain map using an electromagnetic wave propagation model based on the parabolic equation; S3: Construct a multi-objective optimization problem model with multi-segment constraints; S4: A near-end policy optimization algorithm that integrates offline generalization training and online real-time inference is proposed; S5: Visualize and verify the multi-waypoint trajectory output results of the PPO algorithm to demonstrate its adaptability and segmented transmission reliability in multi-node transmission scenarios.
2. The method for ship multi-waypoint path planning based on evaporating waveguide channel maps according to claim 1, characterized in that: The specific steps of S1 are as follows: For a multi-waypoint segmented data transmission scenario, a maritime communication system consisting of ship users and shore-based base stations has N relay points at sea, denoted as... The ship from its starting position Starting from there, sail in sequence through [location]. point, Point until reaching the destination A point refers to any position a ship is at during navigation. Collected size The data is collected and transmitted to the base station, and then transmitted to the next location. The data transmission will be completed before the process is finished; To simplify the problem and focus on ship trajectory optimization, it is assumed that both the base station and the ship user are equipped with a single antenna for communication, and that the ship is at any waypoint Collecting data here does not take time; The goal is to minimize data transmission time and SU's travel time by optimizing the ship's trajectory and utilizing information from the marine evaporative waveguide environment. Next, the coordinate system, ship motion model, and signal model are first described, and then an optimization problem is constructed for the data transmission task under consideration. Assuming the height of the transmitting antenna is The height of the receiving antenna is ,starting point and the end point The coordinates are defined as follows: and The shore-based base station is located at coordinates At the location, the ship was at The location of each time slot is: ; in Indicates the maximum permissible navigation time slot, the ship in the... The velocity of each time slot is denoted as The angle between its velocity vector and the x-axis is denoted as . Therefore, the velocity components of the ship in the x and y directions are as follows: ; in , and They are respectively The lower and upper limits are constrained by physical conditions, and the ship's maneuverability is limited. For the velocity direction between any two adjacent time slots, the steering angle must satisfy the following constraints: ; In the In each time slot, the received signal at the base station is: ; in This is the channel between ship users and base stations. In order to transmit signals, This represents additive white Gaussian noise. Furthermore, it is assumed that the transmit power is constant, i.e.: ; Received signal power It can be represented as: ; in This indicates the path loss between the ship user and the base station. Indicates the transmit antenna gain. Let represent the receiving antenna gain. In this case, we express the maximum achievable transmission rate as: ; in, For channel bandwidth, This represents the noise power spectral density.
3. The method for ship multi-waypoint path planning based on evaporating waveguide channel maps according to claim 1, characterized in that: The specific steps of S2 are as follows: The electromagnetic wave propagation in the evaporating waveguide environment is modeled using the parabolic equation method. For any point in the environment, the large-scale channel gain relative to the base station is calculated using PE. This gain mainly characterizes the path loss. The three-dimensional space is divided into many grids, each grid having a width of [missing value]. The vertical grid height is Within this grid, the channel state remains unchanged. Assuming the base station is located at the origin, the maximum effective radius of the environment in the horizontal direction is... ,in The integer represents the characteristic height of the evaporation waveguide in the vertical direction. ,in It is also an integer, thus each grid node can be uniquely indexed, and its coordinates are: ; In the above formula, , and same, = , will any position The channel gain at point is defined as Because the position of the ship in any time slot is Therefore, based on the grid to which the ship's position belongs at this moment, the large-scale channel gain between the base station and the ship is obtained. .
4. The method for ship multi-waypoint path planning based on evaporating waveguide channel maps according to claim 1, characterized in that: The objective in S3 is to optimize the ship's trajectory to minimize its transmission and arrival times, assuming the total sailing time slots are... The total transmission time slots are The transmission time slot for each trajectory segment of the ship is The sailing time slot for each segment is ,in This serves as an index for the current segment of the ship's trajectory. The ship's trajectory is The multi-objective optimization problem is expressed as: ; ; ; ; ; ; ; ; ; ; ; ; ; In the above formula, This means that the ship must transmit all data before reaching its destination, and the total transmission time is less than the total sailing time. This indicates that the total sailing time is less than the maximum allowed sailing time; Indicates the first The transmission time slot of the ship in the segment trajectory is less than or equal to the navigation time slot; This means that the sum of the transmission time slots of all segments equals the total transmission time slots; This means that the sum of all the segmented time slots equals the total time slot. Denotes the evolution of the queue representing data transmission constraints, in the... The amount of data to be transmitted in the next time slot of a segment trajectory is the amount of data to be transmitted in the current time slot minus the amount of data transmitted in that time slot, until the amount of data to be transmitted in that segment trajectory is... ,in and They represent the first time. The first segment of the trajectory The and the first Data to be transmitted in each time slot, For the first Data transmission rate per time slot The length of the time slot, That is, the first The amount of data transmitted per time slot; This indicates the size of the data to be transmitted in each segment, where For the first The amount of data to be transmitted in the segment trajectory; Indicates the first The amount of data to be transmitted at the end of each flight segment is ; Indicates the starting position constraint of the vessel; Indicates the constraints on the location of the ship's transfer point; Indicates the endpoint position constraint of the vessel; This indicates the ship's steering angle constraint.
5. The ship multi-waypoint path planning method based on evaporating waveguide channel maps according to claim 1, characterized in that: The specific steps of S4 are as follows: The proposed online real-time decision-making algorithm based on near-end policy optimization mainly consists of two stages: Offline training phase: The agent, i.e. the ship, in a simulated evaporating waveguide channel environment, learns the complex relationship between channel fluctuations, motion constraints and multi-segment transmission tasks through a large number of trial and error explorations, and trains a unified policy model with strong generalization ability. This phase is completed in one go and does not require real-time performance. Online inference phase: During actual navigation, the ship deploys the trained strategy model on the shipboard computing unit. In each decision time slot, the system inputs the current state in real time, and the strategy model instantly performs forward inference and outputs the optimal course adjustment action. This process has extremely low computational cost, realizing a fundamental shift from "scenario replanning" to "real-time state response", meeting the real-time decision-making needs in dynamic environments. In addition, the online phase inputs different scenario parameters than the previous model.
6. The method for ship multi-waypoint path planning based on evaporating waveguide channel maps according to claim 5, characterized in that: In the specific step S4, the problem is modeled as a Markov decision process: ; 1) State Space : ; Normalized coordinates, range: ; : The normalized distance to the current target point, ranging from ; The current course of the vessel, normalized range is: ; The angle error between the current heading and the direction pointing to the current target, with a normalized range of [missing value]. ; The current flight segment's "remaining data percentage / total percentage" is normalized to a range of [missing information]. ; : Time progress, which is "time sailed / maximum allowed sailing time", with a range of ; : Segment index normalization, when the ship is arrive During the flight segment, its normalized index is , range ; 2) Action Space Because the ship's turning angle is constrained, its maneuver space is also limited. Therefore, a distribution with a range of values, such as a Gaussian distribution, is adopted. When dealing with continuous probability distributions, action pruning is required, which can cause problems in subsequent logarithmic probability calculations. Therefore, a range of [missing information] was adopted. The Beta distribution is first obtained by sampling. , , and Given the parameters of the Beta distribution, we then perform motion mapping to control the motion range within the ship's heading angle limits, resulting in... At this time, the range of motion is ; 3) State transition ; 4) Reward Function ; The reward function consists of three phases, including the data transmission phase. Used to balance the two parts of transportation and navigation; the transportation completion stage. Phase 1: The vessel rapidly and directly reaches the target point; Phase 2: The vessel reaches the destination. And all data has been transmitted. When the voyage is successful, or when the destination is not reached within the stipulated time. And data transmission failure upon reaching any waypoint However, if the voyage fails, corresponding penalties will be imposed. The overall reward function expression is: ; in As a reward for the data transmission phase, As a reward for completing the transmission phase, As a reward for a failed voyage, As a reward for a successful voyage, , , , The judgment conditions for each stage are expressed as follows: ; ; ; ; Reward function during data transmission phase Represented as: ; in , , The transmission reward, navigation progress reward, and distance and time penalties in this process are represented as follows: ; ; ; in Indicates the transmission weight parameter. The amount of data transmitted within a single time slot. For exploration bonuses, it is represented as: ; Its reward value is limited to between, This represents the average transmission rate over the last 10 time slots. This indicates the highest transmission rate within the last 10 time slots. To avoid terms with a denominator of 0, middle For navigation weight parameters, This represents the actual distance to the target point within a single time slot. This indicates the distance from the target point before the current time slot. This indicates the distance from the target point after the current time slot has moved. Indicates the error between the ship's direction and the direction of the target point. The closer the current direction is to the target point's direction at the same angle, the better. The smaller the value, middle This represents the current distance from the destination. This represents the distance between the starting point and the target point. This represents the number of time slots already navigable. , , Here, 400 represents the weighting factor, and 400 represents the ship's travel distance within a single time slot. Reward function for transmission completion phase for: ; in , , , The navigation progress reward, heading reward, angle penalty, and time penalty in this process are represented as follows: ; ; ; ; in The reward value has been further increased to facilitate ships' navigation to the target point. This represents a heading bonus. A positive heading bonus is awarded when the actual distance between the ship and the target point decreases; otherwise, it is zero. This prevents ships from deviating from the target point to obtain heading bonuses. Angle penalties are imposed to prevent ships from making large changes in angle. As a time penalty, it compels ships to reach the target point as quickly as possible. , , These are the weighting coefficients; The termination reward for a failed voyage is represented as follows: ; in, The penalty for a failed voyage is set to a negative value as a hard constraint penalty; The termination reward upon successful navigation is represented as follows: ; in To terminate slot rewards, it is represented as: ; in These are the weighting coefficients. For weighted rewards, in order to obtain possible non-convex Pareto fronts, the reward value here is expressed using Chebyshev decomposition as follows: ; To achieve the ideal sailing time slot, a smaller value is set to encourage ships to reach their destination in the shortest possible time, thus reducing the penalty. For the ideal transmission time slot, it is also set to a smaller value to help ships shorten transmission time. These are the weighting coefficients.
7. The ship multi-waypoint path planning method based on evaporating waveguide channel maps according to claim 6, characterized in that: The specific process of the algorithm is as follows: (1) During the initialization phase, the parameters are configured, and the environment and preference conditions of the PPO agent are created. The actor and critic networks in the agent are both input with the state and preference vector [α, β]. (2) The training is divided into multiple predetermined weight stages, and the learning rate is changed or a pre-trained model is loaded as needed when switching stages; (3) The agent samples an action using a Beta distribution based on the current environmental state and the preference vector of the current stage; (4) The agent interacts with the environment, calculates the reward according to the reward function under the current preference weight, and moves to the next state, and stores the data of this interaction in the experience buffer; (5) When the accumulated data reaches the preset trajectory length, the agent uses this data to calculate the advantage function through generalized advantage estimation and performs multiple rounds of mini-batch gradient updates to optimize the network parameters of the actors and critics; (6) When the training step target of the current weight setting stage is reached, switch to the next preference weight stage and reset the parameters to continue training; (7) Repeat this process until all weight stages have been trained.
8. A ship multi-waypoint path planning method based on evaporating waveguide channel maps according to claim 7, characterized in that: The trajectory of the PPO algorithm is plotted to demonstrate the rationality and applicability of the proposed algorithm.
Citation Information
Patent Citations
Unmanned ship collision avoidance decision-making method and system fusing captain experience
CN118746979A
Multi-target ship communication path planning method based on offshore waveguide channel map
CN120074715A
Big data-based passenger-roll transport demand prediction and ship intelligent scheduling method and system
CN121072879A
Emergency rescue resource scheduling method based on Internet of Things
CN121481065A