Polar region route intelligent decision-making method based on improved near-end strategy optimization algorithm

By constructing a comprehensive potential energy field map and generating policy rewards, the problem of low exploration efficiency of intelligent agents in the sparse reward scenario of the Arctic shipping route is solved, and stable learning and efficient navigation of intelligent agents in the Arctic shipping route are realized.

CN121632142APending Publication Date: 2026-03-10CHINA UNIV OF GEOSCIENCES (WUHAN) +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing deep reinforcement learning methods struggle to establish correct causal relationships in extremely sparse reward scenarios like the Arctic shipping route, leading to problems such as low exploration efficiency, oscillating training processes, and even failure to converge.

Method used

A comprehensive potential energy field map is constructed, which integrates navigation risk and travel distance to the destination into a continuously distributed comprehensive potential energy value. The navigation strategy of the ship agent is generated through the strategy model to be trained, and the strategy reward is generated according to the comprehensive potential energy value. The experience pool is updated to improve the training model.

Benefits of technology

It significantly alleviates the sparse reward problem, enhances the agent's ability to explore the environment and its learning stability, and improves the safety and economic balance of route planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121632142A_ABST
    Figure CN121632142A_ABST
Patent Text Reader

Abstract

The invention provides a polar region route intelligent decision-making method based on an improved near-end strategy optimization algorithm, and relates to the field of channel planning. Wherein the electronic equipment constructs a comprehensive potential energy field map according to the environment information of the navigation area; generating a navigation strategy of the ship intelligent body through the to-be-trained strategy model; the ship intelligent body is controlled to execute a navigation strategy, so that the ship intelligent body navigates from the current first grid to the second grid in the comprehensive potential energy field map; obtaining a strategy reward of the navigation strategy according to the comprehensive potential energy value of the first grid and the comprehensive potential energy value of the second grid, and adding a piece of navigation experience including the strategy reward into an experience pool; and updating the to-be-trained strategy model based on the experience pool. Therefore, in the moving process of each step of the ship intelligent body, the strategy reward can be generated according to the change of the comprehensive potential energy value between the adjacent grids, so that the sparse reward problem is remarkably relieved, and the exploration ability and learning stability of the intelligent body to the environment are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of waterway planning, and more specifically, to an intelligent decision-making method for polar routes based on an improved near-end strategy optimization algorithm. Background Technology

[0002] As an emerging maritime route connecting the Atlantic and Pacific Oceans, the Arctic shipping route is gaining increasing strategic importance in the global shipping landscape due to its ability to significantly shorten the voyage between Asia and Europe. With the continuous melting of sea ice caused by climate change, the navigation window of the Arctic route is lengthening year by year, attracting more and more commercial vessels to attempt to traverse it. However, the navigation environment in this region is extremely complex, constrained year-round by sea ice cover, severe weather conditions, a lack of navigational facilities, and weak emergency rescue capabilities, posing significant safety risks to ships during navigation. Therefore, optimizing route paths while ensuring navigational safety, and balancing the relationship between voyage length and ice risk, has become the core challenge of intelligent planning for the Arctic shipping route.

[0003] Currently, deep reinforcement learning has been widely applied in recent years to autonomous navigation research for unmanned surface vessels due to its ability to handle high-dimensional state spaces and continuous decision-making processes, particularly for path planning tasks in complex environments. It should be understood that this type of method accumulates experience through interaction between the agent and the simulated environment, gradually learning a mapping strategy from environmental perception to action selection. Theoretically, it can achieve adaptive path generation under unknown or dynamically changing ice conditions.

[0004] However, research has revealed significant shortcomings in current deep reinforcement learning-based methods when faced with extremely sparse reward scenarios like the Northeast Passage in the Arctic. Specifically, because effective reward signals are only obtained near the destination or upon collision, and there is a lack of intermediate feedback during the long voyage, the agent struggles to establish correct causal relationships, leading to low exploration efficiency, oscillating training processes, and even convergence failure. Summary of the Invention

[0005] To overcome at least one deficiency in the prior art, this application provides a polar route intelligent decision-making method based on an improved near-end strategy optimization algorithm, the method comprising: Based on the environmental information of the navigation area, a comprehensive potential energy field map is constructed. The comprehensive potential energy field map includes multiple grids, and each grid is configured with a comprehensive potential energy value. The comprehensive potential energy value is obtained based on the navigation risk of the corresponding grid and the travel distance to the target destination, and represents the comprehensive advantages and disadvantages of the corresponding grid in terms of safety and economy. Generate navigation strategies for ship agents using a strategy model to be trained; Control the ship agent to execute the navigation strategy so that the ship agent navigates from the current first grid to the second grid in the integrated potential energy field map; Based on the combined potential energy value of the first grid and the combined potential energy value of the second grid, the strategy reward of the navigation strategy is obtained, and a navigation experience including the strategy reward is added to the experience pool. The training policy model is updated based on the experience pool.

[0006] Compared with the prior art, this application has the following beneficial effects: This application constructs a comprehensive potential energy field map, integrating navigation risks and travel distances to the destination into a continuously distributed comprehensive potential energy value. This gives each grid location a quantitative indicator reflecting the trade-off between safety and economy, enabling the ship agent to generate policy rewards based on the changes in comprehensive potential energy values ​​between adjacent grids during each step of movement. This significantly alleviates the sparse reward problem and enhances the agent's ability to explore the environment and its learning stability. Attached Figure Description

[0007] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 This is a schematic diagram of the method flow provided in the embodiments of this application; Figure 2 This is a schematic diagram illustrating the effect of navigation area division provided in an embodiment of this application; Figure 3 This is a schematic diagram illustrating the RIO calculation principle provided in an embodiment of this application. Figure 4 The diagram illustrates the effect of the risk potential energy field provided in the embodiments of this application. Figure 5 This is a diagram illustrating the effect of the precise distance field provided in the embodiments of this application. Figure 6 A schematic diagram illustrating the principle of the policy model to be trained provided in an embodiment of this application; Figure 7A This is one of the schematic diagrams showing the comparison of experimental results provided in the embodiments of this application; Figure 7B This is the second schematic diagram showing the comparison of experimental effects provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0009] To make the objectives, technical solutions, and advantages of the embodiments of this application (hereinafter referred to as "the embodiments") clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0010] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0011] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0012] In the description of this application, it should be noted that the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0013] Based on the above statement, as described in the background section, it reveals significant shortcomings when facing an extremely sparse reward scenario such as the Arctic Northeast Passage.

[0014] It should be noted that the defects in the solutions in the prior art are the result of practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of this application in the following text should be regarded as contributions to this application in the process of invention and creation, and should not be understood as technical content known to those skilled in the art.

[0015] Based on the discovery of the above-mentioned technical problems, this embodiment provides a polar route intelligent decision-making method based on an improved near-end strategy optimization algorithm. For example... Figure 1 As shown, the method includes: S1, construct a comprehensive potential energy field map based on the environmental information of the navigation area.

[0016] The integrated potential energy field map includes multiple grids, each with an integrated potential energy value. The integrated potential energy value is based on the navigation risk of the corresponding grid and the travel distance to the target destination, representing the overall advantages and disadvantages of the corresponding grid in terms of safety and economy.

[0017] S2 generates the navigation strategy of the ship agent through the policy model to be trained.

[0018] S3 controls the ship agent to execute a navigation strategy so that the ship agent can navigate from the current first grid to the second grid in the integrated potential energy field map.

[0019] S4. Based on the combined potential energy values ​​of the first grid and the second grid, the strategy reward for the navigation strategy is obtained, and a navigation experience including the strategy reward is added to the experience pool.

[0020] S5 updates the policy model to be trained based on the experience pool.

[0021] In this way, by constructing a comprehensive potential energy field map, navigation risks and travel distances to the destination are integrated into a continuously distributed comprehensive potential energy value. This gives each grid location a quantitative indicator that reflects the trade-off between safety and economy. As a result, the ship agent can generate policy rewards based on the changes in comprehensive potential energy values ​​between adjacent grids during each step of its movement, thereby significantly alleviating the sparse reward problem and enhancing the agent's ability to explore the environment and its learning stability.

[0022] It should be understood that the electronic device implementing this method can be, but is not limited to, a high-performance server deployed in a shore-based data center, or an embedded industrial control computer installed in the bridge of an unmanned surface vessel or merchant ship. The server can be a single server or a group of servers. The server group can be centralized or distributed (e.g., the servers can be a distributed system). In some embodiments, the server can be local or remote relative to the user terminal. In some embodiments, the server can be implemented on a cloud platform; by way of example only, the cloud platform can include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, inter-cloud, multi-cloud, etc., or any combination thereof. In some embodiments, the server can be implemented on an electronic device having one or more components.

[0023] To make the solution provided in this embodiment clearer, the following uses a server as an example, combined with... Figure 1Each step of the method is described in detail. However, it should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical contextual relationships may be reversed in order or implemented simultaneously. Furthermore, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowchart, or remove one or more operations from the flowchart. Figure 1 As shown, the method includes: S1, construct a comprehensive potential energy field map based on the environmental information of the navigation area.

[0024] The integrated potential energy field map includes multiple grids, each with an integrated potential energy value. The integrated potential energy value is based on the navigation risk of the corresponding grid and the travel distance to the target destination, representing the overall advantages and disadvantages of the corresponding grid in terms of safety and economy.

[0025] Research has found that current reinforcement learning algorithms face the problem of reward sparsity in long-distance polar voyages. The root cause lies in the fact that traditional reward functions rely solely on discrete signals indicating whether the destination has been reached, lacking effective evaluation of intermediate states. Therefore, this embodiment provides the following optional implementation methods for step S1: S1-1 divides the navigation area into multiple sub-areas.

[0026] This can be understood as the navigation area being the polar region where route planning is required. The above steps discretize the continuous navigation area into sub-regions with clear boundaries and attribute carrying capacity, so that environmental information (such as sea ice distribution, obstacle location, navigation risks, etc.) can be digitally modeled and stored in each sub-region.

[0027] For example, such as Figure 2 As shown, the navigation area is assumed to be divided into 720×720 regular sub-regions, forming a two-dimensional rasterized simulation environment. Each sub-region corresponds to a certain area in the actual geographic space and has a one-to-one mapping relationship with the latitude and longitude coordinates in satellite remote sensing data, thereby ensuring the geographic accuracy of the environmental modeling. In this process, each sub-region serves as the smallest spatial analysis unit, used to carry various attribute information such as land markers, iceberg markers, and the Risk Index of Operation (RIO) at that location.

[0028] It is worth noting that, due to the influence of the Earth's polar projection characteristics, the spatial resolution of the multiple sub-regions defined above exhibits a non-uniform distribution across different latitudes. Specifically, in the latitudinal direction, the resolution gradually increases from approximately 6.32 kilometers at the map edge to approximately 0.97 kilometers near the poles, reflecting that the actual distance represented by a unit pixel is smaller in high-latitude regions; while in the longitude direction, the resolution remains constant at 36.4 kilometers, maintaining spatial scale consistency along the meridians.

[0029] In this way, the navigable area can be transformed into a discrete spatial model consisting of 518,400 sub-regions with 720 rows and 720 columns, each sub-region being an independent grid cell.

[0030] Based on the description of the navigation area division and sub-regions in the above embodiments, step S1 further includes: S1-2, based on the environmental information of each sub-region, obtain the operational risk index of each sub-region.

[0031] It should be understood that the purpose of the above steps is to transform multi-source ice condition data in the complex polar navigation environment into quantifiable and calculable risk assessment indicators. Therefore, the above steps achieve the quantification of navigation risks in each sub-region by fusing multi-source sea ice data generated from satellite remote sensing and numerical models, and by performing weighted modeling based on ship navigation safety rules.

[0032] In practice, the server first acquires two types of key environmental data: the PIOMAS Arctic Sea Ice Volume Reanalysis dataset (hereinafter referred to as the PIOMAS dataset) and the MASAM2 Daily 4 km Arctic Sea Ice Concentration, Version 1 dataset (hereinafter referred to as the MASAM2 dataset). The PIOMAS dataset provides monthly temporal resolution sea ice volume information with a spatial resolution of approximately 1° longitude and 0.375° latitude; while the MASAM2 dataset has a daily temporal resolution and a spatial resolution of a 4 km grid.

[0033] Due to inconsistencies in spatiotemporal scales between the two datasets, spatiotemporal alignment and standardization are necessary. To address the low temporal resolution of the PIOMAS dataset, the server employs linear interpolation to extend its monthly sea ice volume data into daily time series, matching the daily observation frequency of the MASAM2 dataset and ensuring temporal synchronization between the two datasets. Simultaneously, considering the higher spatial resolution of the MASAM2 dataset and the coarser grid scale of the PIOMAS dataset, resulting in spatial differences, the server uses spatial resampling technology to downsample the high-resolution sea ice concentration data from the MASAM2 dataset to a unified grid framework consistent with the PIOMAS dataset, achieving spatial alignment and effective fusion of multi-source data.

[0034] After completing the above preprocessing, as follows Figure 3 As shown, the server further calculates the operational risk index. In this embodiment, the calculation of the operational risk index follows the assessment method in the POLARIS rules. Specifically, the server determines the corresponding ship risk value based on the ship type (e.g., PC4 class ship) and the ice conditions (including sea ice decay and sea ice type) in the sub-region. This value reflects the navigation risk level faced by ships under specific ice conditions. Then, the risk contribution under each ice condition is weighted and summed using the following expression to obtain the operational risk index:

[0035] In the formula, This represents the operational risk index, characterizing the navigation safety of this sub-region; Indicates the first Sea ice concentration for different ice conditions; This represents the ship risk value under the corresponding ice condition. The calculation process traverses all ice condition types existing in each sub-region, accumulates them based on their concentration ratio and corresponding risk value, and finally outputs a unique RIO value for that sub-region as its operational risk index.

[0036] Based on the above description of the line risk index in the embodiments, step S1 further includes: S1-3, based on the environmental information of the navigation area, obtain the shortest distance from each sub-area to the navigation destination safely.

[0037] In this embodiment, the server can obtain the shortest distance to the destination safely from each sub-region using a path planning algorithm based on the environmental information of the navigation area. This path planning algorithm can be a mature and widely used one.

[0038] The purpose of the above steps can be understood as providing key inputs for the economic dimension of constructing a comprehensive potential energy field, namely, quantifying the actual travel cost from each sub-region to the destination while considering obstacle constraints. This embodiment can pre-generate the shortest travel distance from each location to the destination across the entire map using a graph search-based global optimal path calculation mechanism.

[0039] In practice, the server models the entire navigation area as a two-dimensional mesh graph, where each sub-region corresponds to a node in the graph. The connectivity between nodes is determined by the direction of movement (e.g., four-neighbor or eight-neighbor) and the distribution of obstacles. Based on this, Dijkstra's algorithm is used as the path planning algorithm to calculate the shortest path distance from any sub-region to the navigation destination point by point. It should be understood that this algorithm is a classic single-source shortest path solution method, suitable for graph structures with non-negative edge weights, and can effectively handle the complex Arctic sea conditions in this scenario, including impassable areas such as land and icebergs.

[0040] Thus, by establishing a graphical model of the navigable area and applying Dijkstra's algorithm, the shortest reachable path length from each sub-region is calculated. This not only considers the actual geographical distance but also incorporates safety requirements for obstacle avoidance. Therefore, the resulting shortest distance is a practical measure of navigation costs that balances feasibility and economy.

[0041] Based on the shortest distance from each sub-region to the destination and the operational risk index of each sub-region as described in the above embodiments, step S1 further includes: S1-4, based on the operational risk index and shortest distance of each sub-region, obtain the comprehensive potential energy value of each sub-region.

[0042] It should be understood that the operational risk index and the shortest distance are two different assessment indicators. The operational risk index reflects navigation safety; a higher value indicates a safer environment. Its value range is typically influenced by ice conditions and ship class, and may fall within a limited range such as [-10, 30]. The shortest distance, on the other hand, represents route economy. Its value reflects the actual travel distance in geographical space, and in long-distance Arctic shipping routes, it can reach thousands of kilometers. Therefore, if the two are directly linearly weighted and merged, the distance term, because its absolute value is much larger than the risk index term, will have an overwhelming weight in the overall calculation, leading to a severe weakening or even neglect of safety factors.

[0043] Therefore, the server can normalize the operating risk index and shortest distance of each sub-region separately to obtain the normalized operating risk index and shortest distance of each sub-region; and then weight and fuse the normalized operating risk index and shortest distance of each sub-region to obtain the comprehensive potential value of each sub-region.

[0044] This can be understood as follows: the above steps eliminate the scale difference of the original data through linear normalization, and fuse the normalized indicators based on the weight coefficients to construct a comprehensive potential value that represents the overall quality.

[0045] In practice, the RIO (Risk Index of Operation) attribute of the sub-region is used as a key metric for navigation risk. A higher RIO value indicates lower navigation risk and greater safety. The expression for converting this into a normalized operational risk index (also known as a safety potential value) is as follows:

[0046] In the formula, Subregion The normalized operational risk index; This indicates the original operational risk index of the sub-region; and These represent the maximum and minimum RIO values ​​in all sub-regions within the entire navigation area, respectively. The difference between the two values ​​is normalized using the range (the difference between the maximum and minimum values) so that the result falls within the [0,1] interval.

[0047] In this process, the normalized operational risk index directly reflects the safety advantage of the location. Therefore, when a sub-region is under optimal ice conditions, its... A value close to 1 indicates the safest area; conversely, a value close to 0 indicates a high-risk area. For example... Figure 4 As shown in the figure, the distribution characteristics of the operational risk index in the main channel of the Arctic shipping route are presented. Since the normalized operational risk index is also known as the safety potential energy value, the distribution characteristics of the operational risk index in the main channel of the Arctic shipping route is also known as the risk potential energy field.

[0048] Meanwhile, after obtaining the shortest travel distance matrix D calculated by the path planning algorithm, the server needs to further eliminate the influence of absolute distance caused by different starting and ending point settings or differences in route length, and needs to perform normalization processing:

[0049] In the formula, Subregion The normalized shortest distance (also known as the distance potential energy value); This represents the shortest path distance from the sub-region to the destination. This represents the total distance from the initial starting point to the ending point, serving as the baseline value for normalization. For example... Figure 5As shown, this diagram presents the distribution characteristics of the shortest distance in the main Arctic shipping route. Since the normalized shortest distance is also known as the distance potential value, the distribution characteristic map of the shortest distance in the main Arctic shipping route is also called the precise distance field.

[0050] Thus, during the execution of the above steps, the normalized operational risk index is calculated based on the global RIO range, and the local shortest distance is proportionally scaled based on the total distance from the start point to the end point to obtain the normalized shortest distance, thereby converting both into dimensionless values, ranging between [0,1].

[0051] Based on this, the server performs weighted fusion to generate the final comprehensive potential energy value. This comprehensive potential energy value integrates the normalized operational risk index and the normalized shortest distance potential energy through a linear weighting method, expressed as:

[0052] In the formula, Subregion The overall potential energy value, and These are the weighting coefficients for safety and economy, respectively, and they satisfy... This weighting configuration allows for flexible adjustment of navigation strategy preferences based on mission requirements. For example, it can increase the weighting in severe ice conditions. Prioritizing safety, this can be improved in open waters. The goal is to find the shortest path.

[0053] Based on the comprehensive potential energy value of each sub-region obtained in the above embodiments, step S1 further includes: S1-5: Construct a comprehensive potential energy field map based on the comprehensive potential energy value of each sub-region.

[0054] From an overall conceptual perspective, the above steps establish a mapping relationship between the discretized grid cells and their corresponding comprehensive potential energy values, forming a two-dimensional potential energy distribution map covering the entire navigation area, thereby providing an environmental reward field with global guidance capabilities for subsequent reinforcement learning models.

[0055] In practice, after calculating the comprehensive potential energy value for each sub-region, a gridded simulation environment of the same size as the navigation area can be obtained. Each element corresponds to the position coordinates of a sub-region and stores the comprehensive potential energy value at that location. This value is obtained by weighted fusion of the normalized operational risk index and the normalized shortest distance, comprehensively reflecting the overall advantages and disadvantages of the location in terms of both safety and economy. That is, the higher the value, the more favorable the location is for navigation decisions.

[0056] In this process, the combined potential energy values ​​of all sub-regions are arranged according to their geographical location to form a regular two-dimensional array or tensor structure, thus obtaining a combined potential energy field map.

[0057] Based on the above description of step S1 in the embodiments, please refer to... Figure 1 Next, for Figure 1 Step S2 will be explained as follows: S2 generates the navigation strategy of the ship agent through the policy model to be trained.

[0058] This can be understood as follows: by adopting the Actor-Critic dual network architecture, the above steps enable the policy generation process and the value evaluation process to work together, which not only ensures the diversity and exploratory nature of action outputs, but also improves the stability and convergence efficiency of the training process.

[0059] In specific implementation, such as Figure 6 As shown, the policy model to be trained consists of two key neural networks: a policy network (Actor network) and a critic network (Critic network). It should be understood that this architecture is designed with reference to the basic principles of Proximal Policy Optimization (PPO), enabling efficient policy search and value estimation in high-dimensional state spaces. The policy network in the figure interacts with the partitioned grid simulation environment to generate a navigation policy. This navigation policy is further enhanced with a reward using a comprehensive potential field obtained by fusing the precise distance field and the risk potential field, ultimately outputting a navigation experience and adding it to the experience pool. The critic network then collects historical experience from the experience pool for evaluation.

[0060] based on Figure 6 The model architecture shown illustrates that when the policy network generates the navigation strategy for the ship's intelligent agent, its input is the current environmental state. This state is a tuple containing observable information across multiple dimensions, specifically including: Location-related information, including the ship's intelligent agent's current normalized position vector. ,in and These represent the maximum number of grid cells on the map along the x-axis and y-axis (both are 720 cells). and The current coordinates; Normalized position vectors of the start and end points and ; Distance metric information, including the normalized Euclidean distance from the current location to the starting point. and the normalized distance to the destination ,in This is the actual distance from the starting point to the destination, used to reflect the progress of the voyage and the remaining distance. Endpoint Normalization Angle It is used to characterize the deviation between the current course and the target direction.

[0061] In addition, this state also contains an eight-dimensional Boolean action mask vector. It is used to indicate whether eight possible movement directions (up, down, left, right, upper right, lower right, upper left, lower left) are reachable, thus preventing the agent from choosing actions that lead to obstacles such as land or icebergs.

[0062] Finally, this state also includes an agent-centric one. Mesh Local View Field This window encodes the RIO value (Operational Risk Index) and obstacle distribution information of the surrounding area, forming a window of size [size missing]. The three-dimensional tensor provides the policy network with local environment awareness capabilities.

[0063] Therefore, during the above steps, the policy network receives the multi-source fusion state input and completes feature extraction and decision output through its internal structure. Specifically, the policy network uses a fully connected layer (Multilayer Perceptron, MLP) to process scalar state features (e.g., normalized coordinates, distance, angle, etc.) and a convolutional neural network (CNN) to process image-type state features (e.g., ... The local observation field tensor is used to effectively capture local spatial patterns. After the two types of features are fused, the final output is a probability distribution in the action space, corresponding to the probability of eight movement directions being selected. For example, if a certain direction leads to open water and is heading towards the destination, its probability of being selected is high; conversely, if there are obstacles or deviation from the course, the probability decreases.

[0064] Meanwhile, the critique network is responsible for evaluating the value of the navigation strategies generated by the policy network. Its input is also the current environmental state. The output is an estimate of the expected cumulative reward in that state, i.e., the state value function. This value signal is used to measure the effectiveness of the current strategy and provides a basis for the gradient direction of strategy updates.

[0065] Based on the above description of the navigation strategy in the embodiments, the following will discuss... Figure 1 Step S3 will be explained below: S3 controls the ship agent to execute a navigation strategy so that the ship agent can navigate from the current first grid to the second grid in the integrated potential energy field map.

[0066] This can be understood as follows: by mapping the action instructions output by the policy network to the specific displacement operations of the ship agent in the discretized space, the above steps realize the transformation of decision results into actual motion behavior, thereby driving the agent to gradually explore and approach the navigation endpoint in the comprehensive potential energy field map.

[0067] During the execution of the above steps, the server randomly samples the action probability distribution output by the policy network to determine the specific movement direction to be taken in this decision; based on this action instruction, the ship agent calculates the movement direction from its current first grid cell. The second grid reached after moving in the specified direction For example, if the "top right" action is selected, the new position will be... If the "down" action is selected, then it is... The coordinates are updated according to the direction rules.

[0068] S4. Based on the combined potential energy values ​​of the first grid and the second grid, the strategy reward for the navigation strategy is obtained, and a navigation experience including the strategy reward is added to the experience pool.

[0069] Optionally, the server can obtain the comprehensive potential energy change factor based on the difference between the comprehensive potential energy value of the first grid and the comprehensive potential energy value of the second grid; based on the comprehensive potential energy change factor, the strategy reward for the navigation strategy is obtained, and the corresponding expression is:

[0070] In the formula, Indicates strategy reward, This represents the comprehensive potential energy change factor. This represents the combined potential energy value of the first grid. This represents the combined potential energy value of the second grid. This represents the adjustment coefficient.

[0071] Optionally, based on the aforementioned comprehensive potential energy change factor, the server can also obtain a travel efficiency adjustment factor based on the number of travel steps after the ship's intelligent agent executes the navigation strategy; and obtain the strategy reward for the navigation strategy based on the travel efficiency adjustment factor and the comprehensive potential energy change factor. The corresponding expression is:

[0072] In the formula, Indicates strategy reward; This represents the travel efficiency adjustment factor. This represents the basic penalty factor for each step of a ship's intelligent navigation. Indicates the number of steps taken; This represents the comprehensive potential energy change factor. This represents the combined potential energy value of the first grid. This represents the combined potential energy value of the second grid. This represents the adjustment coefficient.

[0073] Optionally, based on the aforementioned travel efficiency adjustment factor, the server can also obtain a navigation status reward / penalty factor according to the ship's intelligent state in the second grid; and obtain a strategy reward for the navigation strategy based on the navigation status reward / penalty factor, the travel efficiency adjustment factor, and the comprehensive potential energy change factor.

[0074] The relationship between strategy reward, comprehensive potential energy change factor, journey efficiency adjustment factor, and navigation status reward / penalty factor is as follows:

[0075] Indicates strategy reward; This represents the travel efficiency adjustment factor. This represents the basic penalty factor for each step of a ship's intelligent navigation. Indicates the number of steps taken; This represents the comprehensive potential energy change factor. This represents the combined potential energy value of the first grid. This represents the combined potential energy value of the second grid. Indicates the adjustment coefficient; This represents the navigation status reward / penalty factor. If the ship's intelligence collides or times out at the second grid position, the navigation status reward / penalty factor is a preset value less than 0. If the ship's intelligence reaches the navigation destination at the second grid position, the navigation status reward / penalty factor is a preset value greater than 0.

[0076] In summary, the above steps integrate three factors—dynamic potential energy change, navigation efficiency penalty, and key event reward and punishment—to form a dense and semantically rich enhanced reward signal, thereby providing continuous and effective learning guidance for the policy model to be trained.

[0077] In practice, the strategy reward It includes a trip efficiency adjustment factor, a comprehensive potential energy change factor, and a navigation state reward / penalty factor. The reward function can be expressed in the following form:

[0078] In the formula, Indicates the state Next action Later transitioned to a new state The strategy reward obtained is explained in detail below for each term in the expression: The travel efficiency adjustment factor is reflected in Item. Here, This indicates the number of navigation steps currently executed. This represents the base penalty factor for each step, typically negative (e.g., -0.1), used to impose a linear penalty on the agent's path length. This design encourages the agent to complete the navigation task in as few steps as possible, indirectly optimizing path economy. For example, if a trajectory takes 200 steps, the cumulative contribution of this factor is... This significantly affects the total reward value.

[0079] The comprehensive potential energy change factor is reflected as Item. Among them, and These represent the combined potential energy values ​​of the ship's intelligent agent in the first grid before movement and the second grid after movement, respectively. The difference between the two reflects the potential energy gain or loss resulting from this action. Parameters This is an adjustment factor used to control the weight of this item in the overall reward and can be adjusted according to the task complexity. For example, it can be appropriately increased in tasks with dense obstacles or winding routes. To enhance guidance strength.

[0080] Therefore, when a ship's intelligent agent moves towards a region of higher potential energy (i.e., it is safer or closer to the destination), it is more likely to be safer. If the target value is positive, it provides a positive incentive; conversely, if the target value enters a low-potential region, it generates a negative reward, prompting the policy to avoid such behavior. In this way, the sparse feedback problem caused by relying solely on the final reward in the original PPO algorithm can be effectively alleviated, allowing the agent to obtain meaningful learning signals at each step.

[0081] The navigation status reward / penalty factor is represented by a piecewise function:

[0082] In the formula, This indicates a severe penalty for collisions or timeouts, and is usually set to a large negative value (such as -100) to severely suppress unsafe or ineffective behavior. This represents a general reward under normal circumstances, usually a positive value (such as +10), used to reward the behavior of successfully reaching the finish line.

[0083] Therefore, during the execution of the above steps, whenever the ship's intelligent agent completes a state transition from the first grid to the second grid, the server determines its current state based on its real-time location information and determines the navigation state reward / penalty factor. If the second grid contains obstacles such as land or icebergs, it is determined as a "collision," triggering... If the agent fails to reach the destination within the preset maximum number of steps, it is considered a "timeout" and a penalty is triggered. A penalty is only triggered when the agent accurately reaches the target destination. Items that will earn you a high reward.

[0084] It is evident that this reward mechanism not only considers continuous feedback during the process (i.e., through potential energy changes and step penalties), but also retains discrete rewards and penalties for key events (i.e., reaching the endpoint and avoiding accidents), thus achieving a comprehensive evaluation of the agent's behavior.

[0085] Based on the above description of the strategy reward in the embodiments, the following will discuss... Figure 1 Step S5 will be explained below: S5 updates the policy model to be trained based on the experience pool.

[0086] In this embodiment, the server can sample multiple navigation experiences from the experience pool to form a training batch; calculate the advantage function estimate based on the training batch; perform gradient updates on the policy model to be trained based on the advantage function estimate, and use a pruning mechanism to limit the policy update magnitude.

[0087] This can be understood as follows: by introducing a key mechanism from the Proximal Policy Optimization (PPO) algorithm, the above steps effectively control the step size of policy updates while ensuring learning efficiency, thus avoiding instability in the training process or drastic performance fluctuations due to excessively large update magnitudes.

[0088] In practice, historical interaction data is first extracted from the experience pool (i.e., the trajectory pool). This experience pool stores complete information about each state transition in tuple form, specifically represented as:

[0089] In the formula, Indicates time state, The action to be performed in this state. The actual reward obtained (i.e., the enhanced strategy reward). This refers to the new state entered after performing an action. This data is gradually accumulated through the interaction between the ship's intelligent agent and the simulation environment, forming the learning sample set required for model updates.

[0090] Therefore, during the above steps, the server samples the experience pool (e.g., randomly) to obtain multiple navigation experiences, forming a training batch. The server then uses this training batch to calculate the advantage function estimate. The advantage function measures the relative merit of taking a specific action in a given state compared to average performance. In this embodiment, the Generalized Advantage Estimation (GAE) method can be used to generate this value. It should be understood that GAE balances bias and variance by exponentially weighting the temporal difference error (TD-error), providing a more stable advantage estimate in long-sequence tasks. This value is then used to guide the gradient direction of the policy network (Actor network).

[0091] Based on this, gradient updates are performed on the policy model to be trained. The update process consists of two parts: updating the policy network and updating the critic network. For the policy network, the goal is to minimize the objective function of the following pruning form:

[0092] In the formula, This represents the probability of the action output by the current policy network. The probabilities are the same as the old strategy before the update. This is the estimated value of the dominance function. A preset small positive number (e.g., 0.2) is used to define the clipping interval. This mechanism limits the range of variation in the ratio of new to old policies, preventing drastic changes in the policy during a single update and thus ensuring the stability of the training process.

[0093] For critical networks, the goal is to accurately predict the value function of the current state. This provides a benchmark for advantage calculation. The network updates independently by minimizing the value loss function, as shown below:

[0094] in, To criticize the state of the internet The value estimate, This corresponds to the true reward (usually obtained by discounting and summing the reward sequence). Supervised learning is used to continuously narrow the gap between the predicted and true values, thereby improving the accuracy of the evaluation.

[0095] Therefore, this embodiment improves sample efficiency by making full use of historical experience and effectively constrains the policy update magnitude through the unique pruning design of PPO, which significantly enhances the convergence and robustness of the model in complex and sparse reward environments.

[0096] In summary, the polar route intelligent decision-making method based on the improved near-end strategy optimization algorithm (REPPO) provided in this embodiment was compared with the PPO algorithm in tests, and the actual test results showed that it performed better. Figure 7A As shown, in the short-distance channel planning scenario, both the REPPO and PPO algorithms can achieve convergence. However, the REPPO algorithm exhibits more stable reward growth in the early stages of training, indicating that it has better initial training stability in low-complexity tasks. At the same time, the REPPO algorithm also shows a smoother and slower loss descent trajectory, indicating that it has a more comprehensive environmental exploration capability.

[0097] like Figure 7B As shown, with the task difficulty increasing to a medium-distance flight path planning scenario, the efficiency advantage of REPPO becomes even more prominent. The reward curves indicate that, compared to the PPO algorithm, REPPO's reward curve quickly escapes the negative reward range early in training and continues to rise, eventually stabilizing faster. Furthermore, the loss curves show that PPO's policy loss consistently fails to stabilize near zero, while REPPO's policy loss converges to zero much faster. Therefore, the REPPO algorithm achieves more efficient exploration and policy optimization of large state spaces through its potential field reward mechanism, and its rapid convergence of policy loss directly reflects the algorithm's training efficiency advantage in complex environments.

[0098] It should also be understood that if the above embodiments are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0099] Therefore, this embodiment also provides a storage medium, which is a computer-readable storage medium. This storage medium stores a computer program, which, when executed by a processor, implements the polar route intelligent decision-making method based on an improved near-end strategy optimization algorithm provided in this embodiment. The storage medium can be any medium capable of storing program code, such as a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0100] This embodiment provides an electronic device for implementing a polar route intelligent decision-making method based on an improved near-end strategy optimization algorithm. For example... Figure 8 The electronic device may include a processor 22 and a memory 21. The memory 21 stores a computer program, and the processor reads and executes the computer program corresponding to the above-described embodiments in the memory 21 to implement the polar route intelligent decision-making method based on the improved near-end strategy optimization algorithm provided in this embodiment.

[0101] See also Figure 8 The electronic device also includes a communication unit 23. The memory 21, processor 22 and communication unit 23 are electrically connected to each other directly or indirectly through system bus 24 to realize data transmission or interaction.

[0102] The memory 21 can be an information recording device based on any electronic, magnetic, optical, or other physical principles, used to record execution instructions, data, etc. In some embodiments, the memory 21 can be, but is not limited to, volatile memory, non-volatile memory, memory drive, etc.

[0103] In some embodiments, the volatile memory may be random access memory (RAM); in some embodiments, the non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc.; in some embodiments, the storage drive may be a disk drive, solid-state drive, any type of storage disk (such as optical disc, DVD, etc.), or similar storage media, or a combination thereof.

[0104] The communication unit 23 is used to send and receive data over a network. In some embodiments, the network may include a wired network, a wireless network, a fiber optic network, a telecommunications network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, or a near field communication (NFC) network, or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include wired or wireless network access points, such as base stations and / or network switching nodes, through which one or more components of the service request processing system can connect to the network to exchange data and / or information.

[0105] The processor 22 may be an integrated circuit chip with signal processing capabilities, and may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the processor described above may include a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), an Application Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC) computer, or a microprocessor, or any combination thereof.

[0106] Understandable. Figure 8The structure shown is for illustrative purposes only. Electronic devices may also have more advanced features. Figure 8 Showing more or fewer components, or having with Figure 8 The different configurations shown. Figure 8 The components shown can be implemented using hardware, software, or a combination thereof.

[0107] It should be understood that the apparatus and methods disclosed in the above embodiments can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0108] The above descriptions are merely various embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A polar route intelligent decision-making method based on an improved proximal policy optimization algorithm, characterized in that, The method comprises: According to the environmental information of the navigation area, a comprehensive potential field map is constructed, wherein the comprehensive potential field map comprises a plurality of grids, each of which is configured with a comprehensive potential value, the comprehensive potential value being obtained based on the navigation risk of the corresponding grid and the navigation distance to the target terminal point, and representing the comprehensive advantages and disadvantages of the corresponding grid between safety and economy; A navigation strategy of a ship agent is generated by a to-be-trained strategy model; The ship agent is controlled to execute the navigation strategy, so that the ship agent navigates from a first grid to a second grid in the comprehensive potential field map; According to the comprehensive potential value of the first grid and the comprehensive potential value of the second grid, a strategy reward of the navigation strategy is obtained, and a navigation experience including the strategy reward is added to an experience pool; The to-be-trained strategy model is updated based on the experience pool.

2. The polar route intelligent decision-making method based on the improved proximal policy optimization algorithm according to claim 1, characterized in that, According to the comprehensive potential value of the first grid and the comprehensive potential value of the second grid, a strategy reward of the navigation strategy is obtained, comprising: According to the difference between the comprehensive potential value of the first grid and the comprehensive potential value of the second grid, a comprehensive potential change factor is obtained; According to the comprehensive potential change factor, a strategy reward of the navigation strategy is obtained.

3. The polar route intelligent decision-making method based on the improved proximal policy optimization algorithm according to claim 2, characterized in that, According to the comprehensive potential value of the first grid and the comprehensive potential value of the second grid, a strategy reward of the navigation strategy is obtained, further comprising: According to the number of navigation steps after the ship agent executes the navigation strategy, a travel efficiency adjustment factor is obtained; According to the travel efficiency adjustment factor, a strategy reward of the navigation strategy is obtained.

4. The polar route intelligent decision-making method based on the improved proximal policy optimization algorithm according to claim 3, characterized in that, According to the comprehensive potential value of the first grid and the comprehensive potential value of the second grid, a strategy reward of the navigation strategy is obtained, further comprising: According to the state of the ship agent in the second grid, a navigation state reward and punishment factor is obtained; According to the navigation state reward and punishment factor, a strategy reward of the navigation strategy is obtained.

5. The polar route intelligent decision-making method based on the improved proximal policy optimization algorithm according to claim 4, characterized in that, The relationship among the strategy reward, the comprehensive potential change factor, the travel efficiency adjustment factor, and the navigation state reward and punishment factor is: represents a strategy reward; represents a trip efficiency adjustment factor, represents a base penalty factor for each step of the voyage by the ship intelligence, represents the number of steps of the voyage; represents a comprehensive potential change factor, represents a comprehensive potential value of the first grid, represents a comprehensive potential value of the second grid, represents an adjustment coefficient; represents a sailing state reward and punishment factor, if the ship intelligence collides or times out at the second grid position, the sailing state reward and punishment factor is a preset value less than 0, and if the ship intelligence reaches a sailing terminal point at the second grid position, the sailing state reward and punishment factor is a preset value greater than 0.

6. The polar route intelligent decision-making method based on the improved proximal policy optimization algorithm according to claim 1, characterized in that, According to the environmental information of the navigation area, a comprehensive potential field map is constructed, comprising: The navigation area is divided into a plurality of sub-areas; According to the environmental information of each of the sub-areas, a running risk index of each of the sub-areas is obtained; According to the environmental information of the navigation area, the shortest distance from each of the sub-areas to the navigation terminal point is obtained; According to the running risk index and the shortest distance of each of the sub-areas, a comprehensive potential value of each of the sub-areas is obtained; According to the comprehensive potential value of each of the sub-areas, the comprehensive potential field map is constructed.

7. The polar route intelligent decision-making method based on the improved proximal policy optimization algorithm according to claim 6, characterized in that, According to the running risk index and the shortest distance of each of the sub-areas, a comprehensive potential value of each of the sub-areas is obtained, comprising: The running risk index and the shortest distance of each of the sub-areas are normalized respectively to obtain the normalized running risk index and the shortest distance of each of the sub-areas; The normalized running risk index and the shortest distance of each of the sub-areas are weighted and fused respectively to obtain the comprehensive potential value of each of the sub-areas.

8. The polar route intelligent decision-making method based on the improved proximal policy optimization algorithm according to claim 6, characterized in that, According to the environmental information of the navigation area, the shortest distance from each of the sub-areas to the navigation end point is obtained, including: According to the environmental information of the navigation area, the shortest distance from each of the sub-areas to the navigation end point is obtained by a path planning algorithm.

9. The polar route intelligent decision-making method based on the improved proximal policy optimization algorithm according to claim 1, characterized in that, The strategy model to be trained includes a strategy network and a criticism network, the strategy network is used to generate a navigation strategy, and the criticism network is used to evaluate the navigation strategy generated by the strategy network.

10. The polar route intelligent decision-making method based on the improved proximal policy optimization algorithm according to claim 9, characterized in that, Based on the experience pool, the strategy model to be trained is updated, including: Randomly sampling multiple navigation experiences from the experience pool to form a training batch; According to the training batch, an advantage function estimate value is calculated; According to the advantage function estimate value, gradient update is performed on the strategy model to be trained, and a clipping mechanism is used to limit the strategy update amplitude.