Intelligent route planning method based on deep reinforcement learning
Through the intelligent route planning method based on deep reinforcement learning, AIS data preprocessing and DQN algorithm combined with K-means clustering, design state and action space and reward functions, the optimal route is realized independently planning by ships, solving the problem of labor-intensive and inaccurate planning in traditional manual methods, and improving navigation safety and efficiency.
Patent Information
- Application Number
- CN202510331785.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, ship path planning relies on manual means to consume a lot of manpower and the planning is not accurate enough, making it difficult to achieve efficient and safe path planning in complex marine environments.
Using an intelligent route planning method based on deep reinforcement learning, the state space, action space and reward functions are designed to achieve the ship's autonomous planning of the optimal route through AIS data preprocessing, DQN algorithm and K-means clustering.
It improves the accuracy and efficiency of path planning, reduces human resource consumption, ensures the safety and economicality of navigation, and adapts to dynamic environmental changes.
Smart Images

Figure CN120274746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent route planning, and specifically to an intelligent route planning method based on deep reinforcement learning. Background Art
[0002] With the development of economic globalization and frequent international trade. As a means of transportation, ships have become one of the cornerstones of international trade and globalization due to their low transportation costs, strong transportation capacity, low energy consumption, etc., providing reliable, efficient and low-cost transportation services for international trade and logistics. The rapid economic development has put forward higher requirements for maritime intelligent transportation. How to improve maritime transportation efficiency, reduce costs, enhance safety, etc. to promote the development of the shipping industry has become the hottest topic in the field of shipping technology. At the same time, the largest proportion of ship transportation costs is the fuel cost consumed. The longer the route, the more fuel the ship consumes accordingly, and the higher the cost of ship navigation. Therefore, planning the shortest possible route under the premise of ensuring safety has become one of the factors for ships to reduce transportation costs and improve transportation efficiency. Navigation safety is the primary consideration in ship route planning. When a ship sails in a complex marine environment, it needs to avoid collisions with other ships, obstacles, and the marine ecosystem to ensure navigation safety. In addition, a ship will encounter various dynamically changing environmental factors during navigation, such as weather, sea currents, and hydrological conditions. These factors have an important impact on ship route planning, and it is necessary to consider how to adapt to and respond to these changes to ensure the real-time nature and adaptability of the route. Route planning in a dynamic environment requires timely acquisition of environmental information, prediction of future changes, and corresponding adjustments to maintain navigation safety and efficiency. Therefore, an efficient route planning algorithm is crucial for reducing accident risks and protecting the safety of personnel and ships.
[0003] Ship route planning is also closely related to navigation efficiency. By selecting the best route, a ship can reduce navigation time, fuel consumption, and operating costs, improving transportation efficiency and economic benefits. Optimizing the navigation route can reduce unnecessary navigation distance, avoid congested areas, and reduce navigation resistance, thereby improving navigation efficiency. At the same time, a ship arriving at the port on time can ensure the timely delivery of goods, which is conducive to maintaining the smoothness of the supply chain and improving the operating efficiency of ports and shipping companies. Ports can better arrange ship docking, cargo handling, and resource scheduling, while shipping companies can better plan ship routes and resource utilization to improve transportation efficiency. Delayed arrival at the port may lead to cargo detention and supply chain interruption, causing losses to suppliers, manufacturers, and customers. The development of ocean transportation is mainly restricted by the scientific nature of ship navigation route planning. Traditional ocean route planning mainly relies on crew members to draw manually, which not only consumes a large amount of manpower but also the planned route is not accurate enough. Therefore, the development level of route planning is related to the country's economic development, and realizing the automation of ship navigation route planning has important practical significance. Summary of the invention
[0004] In order to solve the problem that the current ship path planning technology requires the crew to draw the ship's navigation route on the paper chart based on the sailing experience accumulated over the years and refer to the previous sailing routes according to different sailing tasks. This manual method has low working efficiency, consumes a lot of manpower and material resources, and has poor safety and economy, the present invention proposes an intelligent route planning method based on deep reinforcement learning, which realizes the ability of the ship to autonomously plan a route from the current position to the destination port that is closest or close to the destination port without manual intervention and under the premise of ensuring safety.
[0005] The specific plan is as follows:
[0006] An intelligent route planning method based on deep reinforcement learning,
[0007] S1, preprocessing of AIS data: receiving AIS data messages and removing abnormal stop points, abnormal drift points, and abnormal turning points;
[0008] S2, build the experimental environment: organize the processed AIS data messages into track diagrams based on the ship numbers, and select the ship's historical track as the safe track;
[0009] S3: Establish a deep reinforcement learning model: Consider the ship as an intelligent agent, use the DQN algorithm to make the intelligent agent interact with the environment and learn the optimal strategy to reach the destination; and combine the K-means algorithm clustering to narrow the scope of the state space;
[0010] S31: Design state space, action space and reward function: Design a 2-dimensional state space to represent all possible positions of the agent on the plane; discretize the action space of the agent into 8 directions; design the reward function of the agent based on the penalty term, set the reward value as a continuous penalty value, and the penalty term changes with the change of the agent state;
[0011] S32: compressing the state space based on the K-means algorithm: when the ship selects the next position, the scope of the state space is reduced; clustering the safe tracks described in S2, and only selecting the center cluster of the nearest class as the next state of the ship;
[0012] S4: Formulate route planning strategy: The current state of the ship is used as the input of the deep neural network, and the output is the expected future reward for each action, i.e., the Q value. The route planning strategy of the ship is formulated by selecting the action with the highest Q value.
[0013] Preferably, the method for determining the abnormal stopping point in S1 is:
[0014] In an AIS data packet, if the (i + 1)-th point satisfying the following formula is an abnormal stop point;
[0015]
[0016] v i is the speed of the i-th point, v i+1 is the speed of the (i + 1)-th point, (lon i , lat i ) is the coordinate of the i-th point, (lon i+1 , lat i+1 ) is the coordinate of the (i + 1)-th point.
[0017] Preferably, the method for determining the abnormal drift point described in S1 is: If the distance between a trajectory point and the previous trajectory point exceeds the distance threshold lambda, then this trajectory point is called an abnormal drift point; the calculation method of the distance threshold lambda is:
[0018] S11: Let the ship length be L, the target speed be v t , the maximum acceleration be a max , the minimum acceleration be a min , the acceleration time be t1, the deceleration time be t2, and then assume that the ship makes a uniformly accelerated or uniformly decelerated motion during navigation, then there is:
[0019]
[0020] S12: The maximum acceleration a max and the minimum acceleration a min of the ship can be obtained as,
[0021]
[0022] Let the time at the i-th point in the track sequence of the ship be t i , the speed be v i , the time at the (i + 1)-th point be t i+1 , the speed be v i+1 , then the distance between the i-th point and the (i + 1)-th point in the ship's track is
[0023] S13: The maximum travel distance lambda of the ship can be expressed as,
[0024]
[0025] Among them, assuming according to the maximum travel distance situation, starting from the i-th point, the ship should first perform maximum acceleration and then maximum deceleration at a certain moment t m ,
[0026] Make the speed reaching point i+1 exactly v i+1 ; t m Satisfy v i+1 = v i + a max (t m - t) + a min (t i+1 - t m );
[0027] S14: The abnormal drift point can be judged as
[0028] dis(i, i+1) > lambda (5)
[0029] The straight-line distance dis(i, i+1) between trajectory point i and trajectory point i+1 can be calculated according to their latitude and longitude coordinates.
[0030] Preferably, the method for determining the abnormal turning point described in S1 is: calculating the route segment based on the turning point. If the slope change between a route segment s i and the previous route segment s i-1 exceeds 90°, then the point p i is an abnormal turning point.
[0031] Preferably, the action space described in step S3 is the set of actions that the agent may take in a given state, which is discrete. The agent can take actions in 8 directions in the neighborhood.
[0032] Preferably, the closer the reward function is to the destination, the greater the reward value; the farther away from the destination, the smaller the reward value; the reward function sets the reward value to a penalty value less than 0 to change the discrete value to a continuous value.
[0033] Preferably, the setting method of the reward function is:[[]]
[0034]
[0035] Among them, d max represents the distance from the agent to the destination at the beginning, d goal represents the distance from the current position of the agent to the destination, and P is the penalty term.
[0036] Preferably, the calculation method of the penalty term P is:[[]]
[0037]
[0038] Among them, if not in feasible area means that if the agent is not on the safe ship track; d lastIndicates the distance from the agent to the destination in the previous step; else if the next position is the target, it means the next state of the agent is the target position.
[0039] Beneficial effects:
[0040] The present invention proposes an intelligent route planning method based on deep reinforcement learning. By combining AIS data preprocessing, reinforcement learning algorithms, state space design, action space design, and reward function design, an efficient, safe, and scalable intelligent route planning method is achieved.
[0041] The following are the main innovations and technical effects of this algorithm:
[0042] 1. Through AIS data preprocessing, identify and eliminate abnormal stop points to ensure the accuracy of route planning; identify and eliminate abnormal drift points to screen out track data that conforms to the actual navigation rules; identify and eliminate abnormal turning points to ensure the smoothness and rationality of the route. Through preprocessing, noise and abnormal data are eliminated, improving the reliability and accuracy of AIS data and providing a high-quality data foundation for subsequent route planning.
[0043] 2. Experimental environment construction: By sorting and screening historical route data, a safety database of ship historical tracks is constructed. By screening historical tracks, a reference basis for intelligent route planning is provided. Thus, a real, reliable, and representative experimental environment is constructed to ensure that the verification and optimization of the algorithm are based on actual navigation data.
[0044] 3. Establish a deep reinforcement learning model. By adopting the DQN algorithm, the agent interacts with the environment to learn the optimal strategy to reach the destination. By clustering and compressing the state space, only the center cluster of the nearest category is selected as the next state, improving the convergence speed and scalability of the algorithm. Combining the DQN and K-means algorithms not only improves the convergence speed of the algorithm but also enhances the scalability of the algorithm in a large-scale environment, providing an efficient learning mechanism for intelligent route planning. Through two-dimensional state space design, the problem complexity is simplified, enabling the agent to efficiently learn the optimal route planning strategy. By using K-means clustering to narrow the state space range, the exploration complexity of the agent is reduced, thereby improving the model convergence speed. Through discrete action space design, the action selection of the agent is ensured to be clear and operable, providing a clear decision-making basis for route planning. At the same time, the reward value is set as a continuous penalty value instead of the traditional positive and negative rewards, providing more refined behavior guidance. Through the design of penalty terms and continuous reward values, the agent can obtain richer feedback information, accelerating the model convergence speed, while simulating the navigation process at the cost of fuel consumption, improving the economy and practicality of route planning.
[0045] In summary, the present invention proposes an intelligent route planning method based on deep reinforcement learning, which realizes efficient, safe and scalable intelligent route planning through AIS data preprocessing, reinforcement learning algorithm, state space and action space design and innovative reward function design. The algorithm has important innovation in theory and has also demonstrated significant technical advantages in practical applications, providing strong technical support for intelligent shipping and maritime traffic management. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Flowchart of an intelligent route planning method based on deep reinforcement learning.
[0047] Figure 2 In the embodiment, the track display diagram of the AIS data set is sorted by mmsi number.
[0048] Figure 3 A diagram of the deep neural network framework constructed in the embodiment.
[0049] Figure 4 A display diagram of the position information of pixels in the environment image in the embodiment. DETAILED DESCRIPTION
[0050] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0051] The present invention uses PostgreSQL, Python, and image technology to build an experimental environment, and establishes a deep reinforcement learning model based on AIS data.
[0052] The present invention is mainly divided into the following steps:
[0053] like Figure 1 As shown in Figure 1, an intelligent route planning method based on deep reinforcement learning is
[0054] S1, AIS data preprocessing: receiving AIS data messages and removing abnormal stop points, abnormal drift points, and abnormal turning points;
[0055] S2, build the experimental environment: organize the processed AIS data messages into track diagrams based on the ship numbers, and select the ship's historical track as the safe track;
[0056] S3, establish a deep reinforcement learning model: regard the ship as an intelligent agent, use the DQN algorithm to make the intelligent agent interact with the environment, learn the optimal strategy to reach the destination, and combine the K-means algorithm clustering to narrow the scope of the state space;
[0057] S31. Design the state space, action space, and reward function: Design a 2D state space to represent all possible positions of the agent on the plane; discretize the agent's action space into 8 directions; design the agent's reward function based on penalty terms, and set the reward value as a continuous penalty value, where the penalty terms change with the change of the agent's state.
[0058] S32. Compress the state space based on the K-means algorithm: When the ship selects the next position, narrow the range of the state space; cluster the safe track described in S2, and only select the center cluster of the nearest class as the next state of the ship.
[0059] S4. Develop a route planning strategy: Take the current state of the ship as the input of the deep neural network, and the output is the expected future reward for each action, that is, the Q value. Develop the ship's route planning strategy by selecting the action with the highest Q value.
[0060] S1: Preprocessing of AIS data
[0061] AIS data messages are not always reliable. When the ship is far from the coast and the position of the ship is beyond the coverage range of the shore base station, it can only rely on satellite signals to transmit AIS data messages. Especially when affected by electromagnetic interference, AIS data messages often get lost and incorrect during the transmission process. Incorrect position data may enter the maritime or shipping management information system, which may cause the management department to misjudge the current maritime traffic status, and may even fail to predict in time the marine casualties such as collisions, groundings, or strandings, resulting in heavy losses. From a macroscopic perspective, this article identifies and eliminates abnormal stop points, abnormal drift points, and abnormal turning points in AIS.
[0062] (1) Abnormal stop point
[0063] There is a forwarding mechanism in the process of sending AIS data messages, and the receiving end may receive multiple duplicate messages of the same track point. Except for the time information, all other information in the duplicate messages is the same. If these duplicate messages are not processed, it will be misjudged that the ship is in a stopped state. The method for determining abnormal stop points defined in this patent is: In an AIS track sequence, if the speed of the i-th point is greater than 2 knots, and the coordinates (lon, lat), speed v, and AIS signal transmission time t of the ship of the (i + 1)-th point are all the same as those of the i-th point, then the (i + 1)-th point is an abnormal stop point, as shown in the following formula.
[0064]
[0065] (2) Abnormal drift point
[0066] The present invention defines that if the distance between a trajectory point and the previous trajectory point exceeds the distance threshold lambda, then this trajectory point is called an abnormal drift point. Under full load, the distance for a ship to accelerate from rest to the target speed is about 20 times the ship's length. In the case of no load, the acceleration distance is shortened to 1 / 2 - 1 / 3 of the original length. The ship's stopping stroke is affected by the ship's displacement and is generally 8 - 20 times the ship's length. According to this standard, the acceleration of a ship accelerating to its target speed at a distance of 15 times the ship's length can be used as its theoretical maximum acceleration; the acceleration of the ship decelerating from its target speed to 0 at a distance of 12 times the ship's length can be used as its theoretical minimum acceleration. If the ship's length is set as L and the target speed is v t , the maximum acceleration is a max , the minimum acceleration is a min , the acceleration time is t1, the deceleration time is t2, and assuming that the ship moves in a uniformly accelerated or uniformly decelerated motion during the navigation process, then there is
[0067]
[0068] The maximum acceleration a of the ship can be obtained max and the minimum acceleration a min are
[0069]
[0070] Let the time at the i-th point of the ship's track sequence be t i , the speed is v i , the time at the i + 1-th point is t i+1 , the speed is v i+1 , then the distance between the i-th point and the i + 1-th point of the ship's track should be During the actual navigation of the ship, according to the received AIS data message, only the speeds v i 、v i+1 and the times t i 、t i+1 of the ship at two trajectory points can be obtained, and the speed change situation between the two trajectory points of the ship cannot be known. According to the maximum travel situation, starting from the i-th point, the ship should first perform maximum acceleration, reach a certain moment t m and then perform maximum deceleration so that the speed when reaching the i + 1-th point is exactly v i+1 . At this time, the maximum travel lambda of the ship can be expressed as
[0071]
[0072] Among them, t m satisfies v i+1 =v i +a max (t m -t)+amin (t i+1 -t m )。Due to the influence of wind, waves and obstacle avoidance, the actual track of the ship between point i and point i + 1 is not a strict straight line. Therefore, the straight-line distance dis(i, i + 1) between track point i and track point i + 1 must be less than lambda. The straight-line distance dis(i, i + 1) between track point i and track point i + 1 can be calculated based on their latitude and longitude coordinates. Accordingly, the abnormal drift point can be judged as
[0073] dis(i, i + 1) > lambda (5)
[0074] In the AIS dataset, we define the same track as (mmsi, startport, T, V, P, endport), where mmsi is the ship number, startport is the starting port, endport is the ending port, T = {t1, t2,..., t n}, V = {v 1, v 2, ,..., v n}, P = {p1 = (lon1, lat1), p2 = (lon2, lat2),..., p n = (lon n , lat n )}, t i and v i respectively represent the time and speed corresponding to the ship with mmsi number at p i . Calculate the distance dis(p i and p i+1 ) between every two adjacent positions p i and p i+1 respectively. If dis(p i , p i+1 ) > lambda, it is an abnormal drift point, and the track containing this track point is filtered. Finally, the track data without abnormal drift points is selected as the alternative AIS dataset.
[0075] (3) Abnormal turning point
[0076] Theoretically, the time interval for a ship in a moving state to send AIS data packets under good communication conditions is generally 6 seconds (as shown in Table 3). However, since the transmission of AIS data packets is not always reliable, we choose to eliminate abnormal turning points from a macroscopic perspective, that is, if a route segment s i = abs(p i+1 - p i ) and the previous route segment s i-1 = abs(p i - pi-1 ) If the slope change between them exceeds 90°, the point p is called an abnormal turning point. Since the AIS dataset in our study has a very large amount of data, we can remove the track data containing abnormal turning points from the AIS dataset. i
[0077] S2: Setup of the experimental environment
[0078] Taking the Bohai Bay area as an example, after the data preprocessing process, the AIS dataset is sorted into tracks according to the mmsi number, as shown in the following representation. Figure 2
[0079] Extract the historical routes of container ships in the Bohai Bay to Tianjin Port. That is, the present invention believes that all historical ship tracks are safe tracks, and selecting the positions where the ship will sail in the historical tracks can ensure the safety of ship navigation.
[0080] S3: Selection of the reinforcement learning algorithm
[0081] In the present invention, the DQN (Deep Q-Network) algorithm is used as the basic reinforcement learning algorithm and the K-means algorithm is introduced. This algorithm is a policy learning method that outputs discrete actions. By an agent taking an action a in a certain state s and interacting with the environment, its state is changed to s′, and a reward r is obtained. The probability from state s to state s’ is called the state transition probability. The agent continuously interacts with the environment and finally learns the optimal policy to reach the destination. In the DQN algorithm, the Q-value is updated based on the current estimated Q-value, the immediate reward, and the maximum future Q-value to gradually approach the optimal Q-value. As shown in the formula:
[0082] Q(s,a) = Q(s,a) + α * (r + γ maxQ(s',a') - Q(s,a)) (8)
[0083] Where Q(s,a) is the current estimated Q-value of taking action a in state s, α is the learning rate, which is used to control the amplitude of each update. r is the immediate reward obtained after taking action a in state s. γ is the discount factor, which is used to balance the importance of the current reward and the future reward. γ is usually between 0.9 and 0.99, which means more attention is paid to obtaining more rewards in the near future rather than in the future.
[0084] The K-means algorithm is used to compress or narrow the range of the state space when the ship selects the next position, that is, after clustering the feasible paths of the ship, only the center cluster of the nearest class is selected as the next state of the ship, which improves the convergence speed of the algorithm and enhances the scalability of the algorithm.
[0085] 1) State space design
[0086] In the context of the DQN algorithm, the state space refers to the set of all possible states that an agent may encounter during its interaction with the environment. In this paper, the state space of the agent is set to two-dimensional, i.e., in the x and y directions to represent all possible positions of the agent. By considering the state space, the DQN algorithm learns to approximate the Q-value, which represents the expected future reward for taking different actions in each state. This enables the agent to make informed decisions and choose actions that maximize its cumulative reward over time.
[0087] 2) Action Space Design
[0088] In the DQN algorithm, the action space refers to the set of actions that an agent may take in a given state. It represents the available choices or decisions that the ship can make to interact with the environment. In this paper, the action space of the ship is discrete, and the action space of the ship is set to 8, i.e., 8 directions in the neighborhood. By considering the action space, the DQN algorithm learns to estimate the Q-value associated with each state-action pair. This enables the agent to select the action with the highest Q-value in a given state and interact with the environment based on this to learn and optimize its behavior.
[0089] 3) Reward Function Design
[0090] The present invention does not consider whether the agent will collide with other ships, that is, it pays more attention to how to reach the destination along a better path in the historical track. Therefore, it only needs to consider the distance from the current agent to the destination port. When calculating the distance, this paper uses the Chebyshev Distance, as shown in the formula:
[0091] d = max(|x2 - x1|, |y2 - y1|) (9)
[0092] Similar to the idea of designing the reward function in traditional reinforcement learning, the closer to the destination, the greater the reward value, and the farther from the destination, the smaller the reward value. The difference is that in traditional reinforcement learning, the agent can only obtain positive and negative reward values by reaching the destination or colliding with obstacles. However, other actions do not receive any positive or negative feedback, and most data cannot reflect its own quality. This means that the model does not receive any feedback before receiving the first reward or punishment, so it may not learn useful experience or stop learning. To solve this problem, this paper introduces a penalty term that changes with the change of the agent's state, so that the agent receives corresponding feedback for each action it takes, as shown in the formula:
[0093]
[0094] where d max represents the distance from the agent to the destination at the beginning, d goalRepresents the distance from the current position of the agent to the destination. P is the penalty term, and there is
[0095]
[0096] where d last is the distance from the agent to the destination in the previous step. Converting the discrete penalty into a continuous one can provide the agent with richer feedback and more refined guidance for its behavior. P has different values according to the current state of the agent, which can grade the state of the agent: the cost is the highest when the agent goes out of the feasible region, and the lowest when the agent moves towards the destination. When d goal is larger, it means the agent is farther from the destination and the penalty is greater. When d goal is smaller, it means the agent is closer to the destination and the penalty is smaller.
[0097] In deep reinforcement learning, the reward function plays an important role in evaluating the effectiveness of behavior decisions and the safety of obstacle avoidance. Different from the concept of designing reward mechanisms in other current methods and techniques in this field, the reward function designed in this invention sets the reward values as penalty values less than 0, which more realistically simulates the process of a ship sailing towards the destination port at the cost of fuel consumption. In addition, the reward function in this paper changes the discrete values into continuous values, which can make timely feedback on the agent's behavior to accelerate the convergence speed of the model.
[0098] S4: Construct the optimal policy
[0099] The deep neural network takes the state as input and outputs the Q-value estimation for each action. A policy is formulated by selecting the action with the highest Q-value, enabling the agent to make optimal decisions in the environment. Among them, Q-network and Target-network are two neural networks respectively, and their structures are as Figure 3 shown.
[0100] Taking the ship with Tianjin as the destination port in the Bohai Bay as an example for intelligent route planning:
[0101] 1) Input the current longitude and latitude information of the ship and the longitude and latitude information of the destination port into the deep reinforcement learning model.
[0102] 2) The model outputs a series of continuous position information of pixel points in the environmental picture, as Figure 4 shown.
[0103] 3) According to the corresponding relationship of the proportional scaling between the picture and the electronic nautical chart, convert the pixel points into longitude and latitude and draw a route from the current position of the ship to Tianjin Port.
[0104] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.
Claims
1. An intelligent route planning method based on deep reinforcement learning, characterized in that S1, preprocessing of AIS data: Receive AIS data messages and eliminate abnormal stop points, abnormal drift points, and abnormal turning points; S2, build an experimental environment: Organize the processed AIS data messages into a track chart based on ship numbers, and select the ship's historical track as the safe track; S3: Establish a deep reinforcement learning model: Regard the ship as an agent, adopt the DQN algorithm, enable the agent to interact with the environment, and learn the optimal strategy to reach the destination; and combine the K-means algorithm clustering to narrow the range of the state space; S31: Design the state space, action space, and reward function: Design a 2-dimensional state space to represent all possible positions of the agent on the plane; Discretize the action space of the agent into 8 directions; Design the reward function of the agent based on the penalty term, and set the reward value as a continuous penalty value, and the penalty term changes with the change of the agent's state; S32: Compress the state space based on the K-means algorithm: When the ship selects the next position, narrow the range of the state space; Cluster the safe track described in S2, and only select the center cluster of the nearest class as the next state of the ship; S4: Develop a route planning strategy: Use the current state of the ship as the input of the deep neural network, and the output is the expected future reward of each action, that is, the Q value. Develop the route planning strategy of the ship by selecting the action with the highest Q value.
2. The intelligent route planning method based on deep reinforcement learning according to claim 1, characterized in that, The method for determining the abnormal stop point described in S1 is: In an AIS data message, if the (i + 1)-th point that satisfies the following formula is an abnormal stop point; v i is the velocity of the i-th point, v i+1 is the velocity of the (i + 1)-th point, (lon i , lat i ) are the coordinates of the i-th point, (lon i+1 , lat i+1 ) are the coordinates of the (i + 1)-th point.
3. An intelligent route planning method based on deep reinforcement learning according to claim 1, characterized in that, The method for determining the abnormal drift point described in S1 is: If the distance between a track point and the previous track point exceeds the distance threshold lambda, then this track point is called an abnormal drift point; The calculation method of the distance threshold lambda is: S11: Let the ship length be L, the target speed be v t , the maximum acceleration be a max , the minimum acceleration be a min , the acceleration time be t1, the deceleration time be t2. Assuming that the ship moves with uniform acceleration or uniform deceleration during navigation, then we have: S12: The maximum acceleration a and the minimum acceleration a of the ship can be obtained as follows max and min for Let the time of the ship at the \(i\)-th point of the track sequence be \(t\). i and the speed be \(v\). i The time at the \((i + 1)\)-th point is \(t'\). i+1 and the speed be \(v'\). i+1 Then the distance between the \(i\)-th point and the \((i + 1)\)-th point of the ship's track is S13: The maximum travel distance lambda of the ship can be expressed as Among them, assuming according to the maximum stroke situation, starting from point i, the ship should first perform maximum acceleration and reach a certain moment t m and then perform maximum deceleration so that the speed when reaching point i + 1 is exactly v i+1 ; t m satisfies v i+1 = v i + a max (t m - t)+ a min (t i+1 - t m ); S14: The abnormal drift point can be judged as dis(i, i + 1) > lambda (5) The straight-line distance dis(i, i + 1) between track point i and track point i + 1 can be calculated according to their longitude and latitude coordinates.
4. An intelligent route planning method based on deep reinforcement learning according to claim 1, characterized in that The method for determining the abnormal turning point described in S1 is as follows: Calculate the route segment based on the turning point. If the slope change between one route segment s i and the previous route segment s i-1 exceeds 90°, then the point p i is an abnormal turning point.
5. An intelligent route planning method based on deep reinforcement learning according to claim 1, characterized in that, The action space described in step S3 is the set of actions that the agent may take in a given state, which is discrete, and the agent can take actions in 8 directions in the neighborhood.
6. An intelligent route planning method based on deep reinforcement learning according to claim 1, characterized in that, The closer the reward function described in step S3 is to the destination, the greater the reward value, and the farther away from the destination, the smaller the reward value; The reward function sets the reward value as a penalty value less than 0, and changes the discrete value to a continuous value.
7. An intelligent route planning method based on deep reinforcement learning according to claim 1, characterized in that, The setting method of the reward function is: Use the Chebyshev distance to calculate the distance from the agent to the target port; where d max represents the distance from the agent to the destination initially, and d goal represents the distance from the current position of the agent to the destination, and P is the penalty term.
8. An intelligent route planning method based on deep reinforcement learning according to claim 1, characterized in that, The calculation method of the penalty term P is: Among them, "if not in feasible area" means that if the agent is not on the safe ship track; d last represents the distance from the agent to the destination in the previous time; "else if next position is target" means that the next state of the agent is the target position.