Unmanned ship gridding control method and system based on reinforcement learning

By dividing the unmanned boat control tasks into three levels: global navigation, local obstacle avoidance and energy management, and using independent reinforcement learning model optimization, combining meta-learning and grid control, the task coordination and energy management problems of unmanned boats in complex marine environments are solved, and efficient and stable autonomous navigation is achieved.

CN120370934APending Publication Date: 2025-07-25ZHONGYING FUND MANAGEMENT CO LTD +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510454036.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing unmanned boat control and path planning optimization methods have problems such as a single task control structure, a lack of multi-task coordination mechanism, and difficulty in local autonomy and spatial mapping of control strategies. They are especially facing policy conflicts, inconsistent responses, and inefficient energy utilization in complex marine environments.

Method used

The control tasks of unmanned boats are divided into three levels: global navigation, local obstacle avoidance and energy management. Each level is optimized using an independent reinforcement learning model, and the decisions between each level are coordinated through meta-learning, control strategies are dynamically adjusted according to task needs, and navigation areas are divided into multiple grid units to optimize real-time environmental data-driven control decisions.

Benefits of technology

It achieves efficient response and stability of multi-objective tasks, improves the adaptability and operation stability of unmanned boats in complex environments, and improves the system's response efficiency and energy consumption control capabilities through resource optimization and space autonomous control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120370934A_ABST
    Figure CN120370934A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned surface vehicle gridding control method and system based on reinforcement learning, and relates to the technical field of unmanned surface vehicle intelligent control, and the method comprises the steps: dividing the control task of an unmanned surface vehicle into three levels of global navigation, local obstacle avoidance and energy management, and each level is optimized through an independent reinforcement learning model. Decisions among the layers are coordinated through meta-learning, and control strategies of the layers are dynamically adjusted according to task requirements. And after each level of control strategy is optimized, dividing the navigation area of the unmanned ship into a plurality of grid units, and optimizing the control decision in each grid unit according to the real-time environment data. According to the method, through a three-step control architecture design of hierarchical modeling-task coordination-space autonomy, significant progress is made in the aspects of task decomposition, control precision and energy consumption optimization, and an evolvable control foundation framework is provided for autonomous operation of the unmanned ship in a complex and changeable marine environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent control of unmanned boats, and particularly to a grid control method and system for unmanned boats based on reinforcement learning. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, intelligent decision-making systems based on deep learning and reinforcement learning have been gradually applied to the autonomous navigation control of surface unmanned platforms. As an intelligent carrier with the capabilities of autonomous navigation, remote control, and task execution, unmanned boats are playing an increasingly important role in fields such as hydrological detection, ocean patrol, and emergency rescue. Most traditional control methods rely on rule-based path planning, state feedback control, or model-based task execution strategies. However, with the increase in task complexity and environmental dynamics, the single-level and static planning control mode gradually exposes problems such as response lag, poor adaptability, and difficulty in coordination. To overcome these bottlenecks, reinforcement learning methods have become one of the current research focuses due to their ability to learn optimal strategies from interactions and adapt to complex state transitions, especially showing good development prospects in the multi-task and multi-objective dynamic control of unmanned boats.

[0003] Although some literature has proposed introducing reinforcement learning into the navigation control process of unmanned boats, the existing technologies mainly focus on the strategy training of a single functional layer, such as path optimization or obstacle avoidance processing, lacking a task hierarchical structure and a multi-strategy coordination mechanism. Specifically, traditional methods often treat tasks as a whole black box, ignoring the independence and interactivity of subtasks such as navigation, obstacle avoidance, and energy management in the control system, resulting in problems such as strategy conflicts, inconsistent responses, and low energy utilization efficiency in complex marine environments. At the same time, there is generally a lack of a local autonomy mechanism at the spatial region level in existing control methods, and the overall navigation area has not been divided into grid regions that can be independently optimized and executed, making it difficult to achieve an efficient control framework that combines task decomposition and local control. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the technical problems solved by the present invention are: the existing control and path planning optimization methods for unmanned boats have problems such as a single task control structure, lack of a multi-task coordination mechanism, difficulty in local autonomy and spatial mapping of control strategies, and the problem of how to construct a highly adaptable control system that integrates hierarchical reinforcement learning and regional grid autonomy mechanism.

[0006] To solve the above technical problems, the present invention provides the following technical solution: A grid control method for an unmanned boat based on reinforcement learning, which includes dividing the control tasks of the unmanned boat into three levels: global navigation, local obstacle avoidance, and energy management. Each level uses an independent reinforcement learning model for optimization. Meta-learning is used to coordinate the decisions between levels, and the control strategies of each level are dynamically adjusted according to the task requirements. After the control strategies of each level are optimized, the navigation area of the unmanned boat is divided into multiple grid cells, and the control decisions within each grid cell are optimized according to the real-time environmental data. Coordinating the decisions between levels by meta-learning includes dynamically adjusting the decision weights of the global navigation, local obstacle avoidance, and energy management layers based on the task priority and urgency. The global navigation task weight is calculated according to the time limit of the task objective, the length and complexity of the navigation path. The local obstacle avoidance task weight is calculated based on the number of dynamic obstacles in the environment and the urgency of the appearance of the obstacles. The energy management task weight is calculated according to the current battery level, task duration, and environmental conditions.

[0007] As a preferred solution of the grid control method for an unmanned boat based on reinforcement learning according to the present invention, wherein: dividing the control tasks of the unmanned boat into three levels of global navigation, local obstacle avoidance, and energy management includes that the global navigation layer performs path planning and optimizes the shortest distance and time consumption of the path. The local obstacle avoidance layer detects and avoids dynamic obstacles. The energy management layer optimizes the sailing speed and direction of the unmanned boat to maximize energy efficiency.

[0008] As a preferred solution of the grid control method for an unmanned boat based on reinforcement learning according to the present invention, wherein: the global navigation layer includes using a graph search algorithm to generate a preliminary path and combining reinforcement learning to optimize the global path selection. By adjusting the heading and speed in real time, the global path is kept optimal under different environmental conditions.

[0009] As a preferred solution of the grid control method for an unmanned boat based on reinforcement learning according to the present invention, wherein: the local obstacle avoidance layer includes combining lidar, sonar, and visual sensor data to detect the position of obstacles in real time, and adjusting the heading to avoid collisions through a reinforcement learning algorithm. When an obstacle enters the preset avoidance area of the unmanned boat, calculate and select the optimal obstacle avoidance path.

[0010] As a preferred solution of the grid control method for an unmanned boat based on reinforcement learning according to the present invention, wherein: the energy management layer includes optimizing the sailing strategy in real time according to the remaining battery level, current speed, and task duration of the unmanned boat, and adjusting the speed and heading of the unmanned boat. In long-term tasks, combine weather changes and environmental data to dynamically adjust the energy consumption strategy of the unmanned boat.

[0011] As a preferred solution of the grid control method for unmanned boats based on reinforcement learning according to the present invention, wherein: the dynamic adjustment of the control strategies at each level according to task requirements includes dynamically adjusting the decision weights of the global navigation, local obstacle avoidance, and energy management layers based on task priority and urgency. The global navigation task weight is calculated according to the time limit of the task objective, the length and complexity of the navigation path. The local obstacle avoidance task weight is calculated based on the number of dynamic obstacles in the environment and the urgency of the appearance of the obstacles. The energy management task weight is calculated according to the current battery level, task duration, and environmental conditions.

[0012] As a preferred solution of the grid control method for unmanned boats based on reinforcement learning according to the present invention, wherein: the optimization of the control decision in each grid cell according to real-time environmental data includes, within each grid cell, using the reinforcement learning models of the global navigation, local obstacle avoidance, and energy management layers to perform local path planning, obstacle avoidance adjustment, and energy consumption optimization, and the tasks in each grid are executed independently. The task objectives of each grid cell are dynamically updated according to real-time environmental data, and are adaptively adjusted within each grid through the reinforcement learning algorithm to cope with environmental changes and task requirements.

[0013] Another object of the present invention is to provide a grid control system for unmanned boats based on reinforcement learning, which can divide the control tasks of the unmanned boat into three levels: global navigation, local obstacle avoidance, and energy management, and each level is optimized using an independent reinforcement learning model, solving the problems of the existing unmanned boat control and path planning optimization methods, such as a single task control structure, lack of multi-task coordination mechanism, difficulty in local autonomy and spatial mapping of control strategies, and the problem of how to construct a highly adaptable control system that integrates hierarchical reinforcement learning and regional grid autonomy mechanism.

[0014] As a preferred solution of the grid control system for unmanned boats based on reinforcement learning according to the present invention, wherein: it includes an unmanned boat control task division module, a layer task optimization module, and a grid module.

[0015] The unmanned boat control task division module is used to divide the control tasks of the unmanned boat into three levels: global navigation, local obstacle avoidance, and energy management, and each level is optimized using an independent reinforcement learning model.

[0016] The layer task optimization module is used to coordinate the decisions between each level through meta-learning, and dynamically adjust the control strategies of each level according to task requirements.

[0017] The grid module is used to divide the navigation area of the unmanned boat into multiple grid cells after the control strategies at each level are optimized, and optimize the control decisions in each grid cell according to real-time environmental data.

[0018] A computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of a grid control method for an unmanned boat based on reinforcement learning.

[0019] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of a grid control method for an unmanned boat based on reinforcement learning are implemented.

[0020] Advantages of the present invention: The grid control method for an unmanned boat based on reinforcement learning provided by the present invention trains a control model through hierarchical reinforcement learning, realizes the structured modeling and independent optimization of three types of typical tasks, and further improves the response efficiency and stability of the system to multi-objective tasks. Finally, the parallel execution of heterogeneous tasks in a complex environment and the separation of control loads are achieved. It significantly enhances the interpretability and adjustability of the control system, improves the adaptive ability and operation stability of the unmanned boat in a high-dynamic environment, and breaks through the control bottleneck of the traditional reinforcement learning system of "single model for all tasks".

[0021] By introducing a meta-learning framework to dynamically adjust control weights, the system can automatically allocate computing and decision-making priorities according to task urgency and resource conditions, and then realize the real-time reconciliation of policy conflicts between multiple tasks. Finally, the resource utilization efficiency and response elasticity of the system are effectively improved. Through this coordination mechanism, the conflicts between control strategies are reconciled, computing resources are allocated on demand, and the operation load of the overall decision-making chain of the system is more balanced, significantly improving the response efficiency and energy consumption control ability of the system during the collaborative execution of multiple tasks.

[0022] By dividing the navigation area into grids and constructing an independent local decision-making mechanism within the area, the deployment of a spatially distributed autonomous control strategy is realized, and then the local adaptation ability and real-time decision-making ability of the system in the spatial dimension are enhanced. Finally, the control field accuracy and task reliability of the unmanned boat in a wide-area complex environment are significantly improved. Description of the Drawings

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is the overall flowchart of a grid control method for an unmanned boat based on reinforcement learning provided by the first embodiment of the present invention.

[0025] Figure 2System flowchart of a grid control method for unmanned boats based on reinforcement learning provided for the third embodiment of the present invention. Detailed implementation manners

[0026] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0027] Example 1, referring to Figure 1 , which is an embodiment of the present invention, provides a grid control method for unmanned boats based on reinforcement learning, including:

[0028] S1: Divide the control tasks of the unmanned boat into three levels: global navigation, local obstacle avoidance, and energy management. Each level is optimized using an independent reinforcement learning model.

[0029] Dividing the control tasks of the unmanned boat into three levels of global navigation, local obstacle avoidance, and energy management includes that the global navigation layer performs path planning and optimizes the shortest distance and time consumption of the path. The local obstacle avoidance layer detects and avoids dynamic obstacles. The energy management layer optimizes the sailing speed and direction of the unmanned boat to maximize energy efficiency. The global navigation layer uses a graph search algorithm to generate a preliminary path and combines reinforcement learning to optimize the global path selection. By adjusting the heading and speed in real time, the global path is kept optimal under different environmental conditions. The local obstacle avoidance layer combines lidar, sonar, and visual sensor data to detect the position of obstacles in real time and adjusts the heading through a reinforcement learning algorithm to avoid collisions. When an obstacle enters the preset avoidance area of the unmanned boat, the optimal obstacle avoidance path is calculated and selected. The energy management layer optimizes the sailing strategy in real time according to the remaining power, current speed, and task duration of the unmanned boat, and adjusts the speed and heading of the unmanned boat. In long-term tasks, the energy consumption strategy of the unmanned boat is dynamically adjusted in combination with weather changes and environmental data.

[0030] A preferred solution of the global navigation layer specifically includes that in the global navigation layer, the goal of path planning is to minimize the path length of the unmanned boat from the starting point to the ending point and optimize the time consumption of the path. When optimizing the path through the reinforcement learning model, environmental dynamics (such as sea current, wind speed, etc.) need to be considered, and a graph search algorithm is used to generate a preliminary path.

[0031] The preliminary path generation is generated by a graph search algorithm, and the path selection is optimized in combination with the reinforcement learning model. The goal of optimization is to reduce the sailing time and energy consumption through the shortest path time. Expressed as:

[0032]

[0033] Among them, P opt represents the optimal path selection. d i represents the Euclidean distance of the i-th path segment. f time (v i ) represents the time consumption function based on the speed v of the unmanned boat i of the unmanned boat, and this function takes into account the change of time consumption at different speeds. λ env represents the environmental factor weight, which adjusts the influence of environmental factors such as sea current and wind speed on the path. represents the environmental dynamic function, which considers the influence of environmental factors (such as sea current, wind speed, etc.), and represents the environmental influence at the position of the i-th path point of the unmanned boat. n represents the number of path segments. x i represents the position of the i-th path point.

[0034] Furthermore, the calculation formula of f time (v i ) is expressed as:

[0035]

[0036] The calculation formula is expressed as:

[0037]

[0038] Among them, γ1 represents the weight factor of the influence of wind speed on environmental disturbance. γ2 represents the weight factor of the influence of sea current on environmental disturbance. s i represents the total intensity index of environmental disturbance at position i, which is obtained by weighted accumulation after normalization and fusion of the wind speed, flow velocity and turbulence change data collected by the multi-source sensors carried by the unmanned boat. s0 represents the reference level value of the disturbance intensity. δ represents the steepness adjustment coefficient of the disturbance function.

[0039] A preferred scheme of the local obstacle avoidance layer specifically includes that the task of the local obstacle avoidance layer is to detect the position of the obstacle in real time and calculate the obstacle avoidance path through the reinforcement learning model. Sensor data (such as lidar, sonar and vision sensors) is used to detect obstacles in real time. When the obstacle enters the preset avoidance area of the unmanned boat, the local obstacle avoidance layer will perform path correction.

[0040] When it is detected that the obstacle enters the avoidance area of the unmanned boat, the local obstacle avoidance layer adjusts the heading through the reinforcement learning model and selects the optimal obstacle avoidance path. The calculation formula is as follows:

[0041]

[0042] Among them, Δθ avoid represents the adjustment of the obstacle avoidance angle. d jRepresents the distance between obstacle j and the unmanned boat. f collision (v j ) represents the collision risk function, and the value range is [0, 1], which calculates the collision risk between the unmanned boat and the obstacle. λ adjust Represents the obstacle avoidance adjustment weight. Represents the function describing the influence of the obstacle on the selection of the obstacle avoidance path, and the value range is positive real numbers. m represents the number of obstacles. o j Represents the set of obstacle parameters.

[0043] Furthermore, the formula of the collision risk function is expressed as:

[0044]

[0045] The function describing the influence of the obstacle on the selection of the obstacle avoidance path is expressed as:

[0046]

[0047] Among them, v j Represents the navigation speed of the unmanned boat relative to the j-th obstacle. v th Represents the speed threshold at which the collision risk rises significantly. In the present invention, the speed threshold is set to 2.5 m / s. k1 represents the slope adjustment coefficient of the risk function. θ j Represents the angle of the obstacle relative to the current heading. θ u Represents the current heading angle of the unmanned boat. ω1 represents the adjustment weight coefficient of the distance term, and the larger it is, the more attention is paid to the influence of nearby obstacles. ω2 represents the adjustment weight coefficient of the angle included angle term, which reflects the path interference intensity when the orientation of the obstacle is consistent with the heading.

[0048] A preferred solution of the energy management layer specifically includes that the task of the energy management layer is to optimize the navigation strategy in real time according to the remaining power of the unmanned boat, the current speed and the mission duration, and adjust the speed and heading of the unmanned boat. In a long-term mission, it is also necessary to make dynamic adjustments in combination with weather changes and environmental data to maximize energy efficiency.

[0049] The energy consumption of the unmanned boat is optimized by the following formula:

[0050]

[0051] Among them, E opt Is the optimal energy consumption strategy. P cons (v k ) is the energy consumption power at speed v k . f energy (t k ) is the energy consumption function based on time t k . is the influence function of weather factors, which describes the impact of weather changes on energy consumption. q represents the total number of time slices for energy management task partitioning. w k represents the weather state within the k-th time slice.

[0052] Furthermore, based on time t k the energy consumption function is expressed as:

[0053]

[0054] The influence function of weather factors is expressed as:

[0055]

[0056] where t k represents the length of the k-th time period in the energy management task. η represents the basic energy consumption coefficient, which controls the starting magnification of energy consumption growth over time. φ represents the power exponent coefficient of energy consumption time growth, reflecting the cumulative non-linear energy consumption of long-duration continuous tasks. T k represents the external environmental temperature within this time period. T ref represents the reference temperature of the recommended operating environment of the device. W k represents the wind speed within this time period. H k represents the humidity level within this time period. θ1 represents the weight coefficient of the impact of temperature deviation on energy consumption. θ2 represents the weight coefficient of the increase in navigation energy consumption due to the square of the wind speed. θ3 represents the weight coefficient of the impact of humidity on the operating efficiency of the device.

[0057] Furthermore, S1 divides the control tasks of the unmanned boat into three levels: global navigation, local obstacle avoidance, and energy management. Each level is optimized using an independent reinforcement learning model. This hierarchical structure can carefully handle different task objectives within each level, ensuring that tasks can be completed efficiently and accurately. The global navigation layer optimizes the path through the combination of graph search algorithms and reinforcement learning models. The local obstacle avoidance layer detects obstacles and corrects the path through real-time sensor data and reinforcement learning. The energy management layer optimizes the speed and heading through reinforcement learning, and dynamically adjusts the task execution strategy according to environmental changes and remaining battery power.

[0058] By dividing the tasks into multiple levels and independently optimizing each level, complex control problems can be transformed into multiple simple local optimization problems, thereby improving efficiency. Using reinforcement learning models can adjust strategies according to real-time environments and task requirements, enhancing the system's adaptability and ensuring that the unmanned boat can flexibly respond to different environmental and task changes. Combining global path optimization and local obstacle avoidance strategies can not only reduce the navigation time but also optimize energy consumption through the energy management layer, extending the task execution time.

[0059] Furthermore, traditional unmanned boat control systems usually rely on fixed path planning and static obstacle avoidance strategies, and are unable to flexibly cope with complex and dynamic environments. The hierarchical reinforcement learning structure designed in step S1 solves the limitations in traditional methods by adjusting the control strategies of each layer in real time, especially the comprehensive optimization problems of path planning, obstacle avoidance, and energy management in complex marine environments. By introducing reinforcement learning and dynamic adjustment strategies, the unmanned boat can maintain efficient and stable task execution in changing environments, making up for the deficiencies of existing technologies.

[0060] S2: Coordinate the decisions between layers through meta-learning, and dynamically adjust the control strategies of each layer according to task requirements. Based on task priorities and urgencies, dynamically adjust the decision weights of the global navigation, local obstacle avoidance, and energy management layers. The global navigation task weight is calculated according to the time limit of the task objective, the length and complexity of the navigation path. The local obstacle avoidance task weight is calculated based on the number of dynamic obstacles in the environment and the urgency of the appearance of obstacles. The energy management task weight is calculated according to the current battery level, task duration, and environmental conditions. In complex environments, automatically select the most appropriate decision-making strategy to ensure the balance and coordination of each task objective.

[0061] A preferred solution for the global navigation task weight specifically includes that the weight of the global navigation task needs to consider the complexity of path planning, the time limit of the task, and the length of the navigation path. The longer the path length and the higher the urgency of the task, the higher the weight needs to be allocated. The more urgent the time limit of the task, the higher the task priority and the greater the weight. The impact of environmental dynamics (such as sea currents, wind speeds, etc.) on path selection also affects the calculation of the weight. The weight calculation formula is expressed as:

[0062]

[0063] where, w global represents the calculated weight of the global navigation task. α1, α2, and α3 represent the adjustment coefficients of the global navigation task weight. L path represents the total path length. T limit represents the time limit of the task. β1 represents the environmental impact coefficient. represents the environmental dynamic impact function of each path point x path in the path.

[0064] Further, the environmental dynamic impact function of each path point x path in the path is expressed as:

[0065]

[0066] where represents the set of all path points on the complete path. w i represents the path point x iWind speed value at c i Represents waypoint x i Water flow speed value at Represents waypoint x i Rate of change of ambient temperature per unit time at T i Represents waypoint x i Ambient temperature at

[0067] A preferred scheme for the weight of the local obstacle avoidance task specifically includes that the weight of the local obstacle avoidance task is calculated according to the number of dynamic obstacles in the environment and the urgency of the obstacles. The weight of this task increases as the obstacles approach and the threat level increases. The more obstacles there are, the higher the priority of the obstacle avoidance task. The closer the distance of the obstacle, the greater the threat, and the higher the priority of the task. The weight calculation formula is expressed as:

[0068]

[0069] Where, w avoid Represents the calculated weight of the local obstacle avoidance task. α4 and λ1 represent the adjustment coefficients of the local obstacle avoidance task. f collision (v j ) represents the collision risk function, which calculates the collision risk according to the distance. Represents the dynamic change function, considering the dynamics of the obstacles. Represents the obstacle influence function, which evaluates the influence of the obstacles on the obstacle avoidance decision.

[0070] Furthermore, Is expressed as:

[0071]

[0072] Where, Represents the rate of change of the speed of the j-th obstacle within consecutive time slices. Represents the current movement direction angle of the j-th obstacle. Represents the acceleration of the j-th obstacle.

[0073] A preferred scheme for the weight of the energy management task specifically includes that the weight of the energy management task is calculated according to the remaining battery power, the duration of the task, and environmental conditions (such as weather changes and ocean currents). Energy management is a key part of long-term tasks, so it needs to be dynamically adjusted. The less remaining battery power, the higher the priority of energy management. The longer the task time, the higher the priority of energy management. Changes in the environment (such as wind speed, ocean current) will affect energy consumption, so dynamic adjustment is required. The weight calculation formula is expressed as:

[0074]

[0075] Where, wenergy Represents the calculation weight of the energy management task. α5 and λ2 represent the adjustment coefficients of the energy management task. E remaining Represents the remaining battery power. T duration Represents the duration of the task.

[0076] In the meta - learning stage, the most suitable decision - making strategy is automatically selected according to the weights of each task to ensure the balance and coordination of task objectives. Through meta - learning calculations, it is possible to adjust the decisions at each level by combining the priorities and urgencies of global navigation, local obstacle avoidance, and energy management tasks.

[0077] The decision - making coordination formula is expressed as:

[0078]

[0079] Among them, w total Represents the final decision - making weight of the global task, which combines the weights of each level. w z Represents the weights of global navigation, local obstacle avoidance, and energy management. Represents the meta - learning function, which adjusts the weight of each task based on the priorities of each level of tasks and environmental factors.

[0080] Furthermore, the meta - learning function Is expressed as:

[0081]

[0082] Among them, Represents the energy consumption change rate per unit time of the task at the z - th layer in the current cycle. ε z Represents the cumulative energy consumption value of the z - th control task layer. z represents the task layer index, with values of 1 (global navigation), 2 (local obstacle avoidance), and 3 (energy management).

[0083] It should be noted that the design concept of step S2 is to introduce a meta - learning mechanism to adjust the weights and coordinate the strategies for the global navigation, local obstacle avoidance, and energy management tasks in the multi - layer control system of the unmanned boat. Different from the traditional system that executes tasks according to preset rules, this step dynamically adjusts the proportion of the execution of each layer's strategy by real - time evaluating the priority and urgency of tasks, making the task scheduling more intelligent and refined. Its core design lies in: introducing a structured mathematical function model for each task layer, including quantitative factors such as task timeliness, path complexity, obstacle threat level, power level, and environmental dynamics, and finally summarizing them into the meta - learning framework to complete the adaptive selection of strategies. The greatest advantage of this step is that the system has the ability of adaptive scheduling, which can automatically select the optimal task execution strategy according to the actual changes of tasks, avoiding problems such as resource waste or key task lag caused by "uniform priority". Compared with the existing technology, this method effectively solves problems such as static setting of task priorities, non - adjustable conflicts in hierarchical task execution, and lag in emergency response in traditional unmanned boat control, significantly enhancing the adaptability and overall execution efficiency of the unmanned boat system to complex and changing marine environments.

[0084] S3: After the control strategies at each level are optimized, divide the navigation area of the unmanned boat into multiple grid cells, and optimize the control decisions within each grid cell according to real - time environmental data.

[0085] Within each grid cell, use the reinforcement learning models of the global navigation, local obstacle avoidance, and energy management layers to perform local path planning, obstacle avoidance adjustment, and energy consumption optimization, ensuring that the tasks within each grid can be executed efficiently and independently. The task objectives of each grid cell are dynamically updated according to real - time environmental data, and are adaptively adjusted within each grid through the reinforcement learning algorithm to cope with environmental changes and task requirements.

[0086] Based on the three - layer control strategies and their dynamic weight allocation results obtained in S1 and S2, divide the overall navigation area of the unmanned boat into several grid cells with fixed numbers. Each grid cell represents an independently controllable local navigation area. The division criteria can be based on geographical coordinates, task density, or environmental change degree. The boundary information of the grid has been loaded into the control system during the system initialization stage.

[0087] Inside each grid cell, the unmanned boat system will call the control models corresponding to the global navigation, local obstacle avoidance, and energy management layers respectively, and execute local control decisions according to the environmental state and task requirements in this local area. The three - layer models are not repeatedly trained, but are derived from the main model trained by previous reinforcement learning to generate local versions applicable to this grid, and perform fine - tuning and optimization of local objectives.

[0088] The task objectives and status awareness of each grid cell are dynamically updated. Continuously collect environmental data from the grid, establish a time-sliding window of the current grid state, and monitor the change trends at consecutive moments. When the system detects significant fluctuations in the environmental or task status within the grid, such as a sudden increase in the number of obstacles, a change in the water flow direction, or abnormal power consumption, etc., it will trigger a rapid adjustment of the local control strategy. Regenerate the adjusted control instructions and update the corresponding decision logic to adapt to the current state.

[0089] To ensure the continuity between multiple grids and the consistency of the control strategy, a state synchronization mechanism is set at the grid boundaries. When the current grid completes a task or is about to enter the next grid, it will transfer the current control state, heading trend, and remaining energy information to the next grid, so that the next grid can quickly load the decision context of the previous area, achieve seamless connection across grids, and avoid policy jumps or instruction breaks.

[0090] Furthermore, by dividing the area into multiple independent grid cells and running a three-layer control strategy trained by reinforcement learning in each cell, and combining the dynamic environmental perception and policy adaptive adjustment mechanism, the present invention realizes a grid-based intelligent control system with local independence, high efficiency in response, and continuous connection, effectively improving the stability, energy conservation, and task completion efficiency of the unmanned boat in a large-scale complex environment.

[0091] Embodiment 2, an embodiment of the present invention, provides a grid-based control method for an unmanned boat based on reinforcement learning. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0092] The selected unmanned boat sample in the experiment is equipped with a variety of sensors, including lidar, sonar, GPS, and accelerometer, etc., and has an adaptive reinforcement learning control system. The test site is selected in a simulated marine environment, including sea areas with different water flow speeds, wind speed changes, and complex obstacle distributions.

[0093] At the beginning of the experiment, the unmanned boat starts in the starting area and conducts global path planning. During the global navigation process, the unmanned boat divides the navigation area into multiple grid cells and applies a local reinforcement learning model in each grid for path adjustment, obstacle avoidance, and energy optimization. The local obstacle avoidance task adjusts the heading to avoid collisions when detecting obstacles. The energy management task then optimizes the navigation strategy in real time according to the remaining power, navigation speed, and environmental conditions to ensure the efficient use of energy.

[0094] During the entire experiment, the control strategies for all tasks were independently optimized within different grid cells to ensure the efficient execution of tasks in each area. During the experiment, the system continuously collected data, including the completion time, energy consumption, obstacle avoidance success rate, task success rate of each task, and the system's adaptability to environmental changes. All data was analyzed in real-time through the control system of the unmanned boat, and finally, the advantages of the present invention were verified by comparing with existing traditional control methods according to the experimental results. Some experimental data is shown in Table 1.

[0095] Table 1 Experimental Data Table

[0096]

[0097]

[0098] It can be clearly seen from Table 1 that the grid-based control method of the unmanned boat of the present invention demonstrates significant innovation and advantages compared with the prior art. In terms of path planning time, the global navigation and local obstacle avoidance task times in the present invention are relatively short. The path planning time of the global navigation task shows high execution efficiency in different environments. Especially compared with existing control methods, the task completion time is significantly shortened, indicating that the present invention can adapt to the environment and generate paths more quickly.

[0099] In terms of the obstacle avoidance success rate, the success rate of the local obstacle avoidance task of the present invention is stable and high. Especially in complex environments, the task execution is more accurate and timely. Compared with traditional methods, existing control strategies are prone to slow reaction or unreasonable obstacle avoidance paths when facing dynamic obstacles, while the present invention makes the obstacle avoidance task more efficient and successful through the real-time optimization and local decision-making of the reinforcement learning model.

[0100] In terms of energy consumption, the energy management layer of the present invention ensures the reasonable use of energy through the real-time optimization of multiple variables such as remaining battery power, speed, and task duration, and reduces unnecessary energy waste during task execution. By comparing with traditional control methods, the present invention consumes relatively less energy during the task, thus achieving longer endurance and higher resource utilization efficiency.

[0101] From the perspectives of task success rate and environmental adaptability, the reinforcement learning control strategy of the present invention demonstrates excellent adaptability in dynamic environments. It can adjust the control strategy to cope with various emergencies when environmental changes such as wind speed and ocean current occur. Especially in complex environments, the unmanned boat can make adaptive adjustments in a short time, and the success rate of completing tasks is much higher than that of traditional control methods.

[0102] Example 3, refer to Figure 2, which is an embodiment of the present invention, provides a grid control system for an unmanned boat based on reinforcement learning, including an unmanned boat control task division module 100, a layer task optimization module 200, and a grid module 300.

[0103] S4: The unmanned boat control task division module 100 is used to divide the control tasks of the unmanned boat into three levels: global navigation, local obstacle avoidance, and energy management. Each level is optimized using an independent reinforcement learning model.

[0104] The unmanned boat control task division module 100 includes a task stratification sub-module 101 and a model allocation sub-module 102.

[0105] Furthermore, the task stratification sub-module 101 is used to divide the overall control tasks of the unmanned boat into three levels: global navigation, local obstacle avoidance, and energy management according to the task nature and response cycle. Among them, the global navigation task is mainly responsible for the planning of medium- and long-distance targets in the path and the determination of the macro-trajectory. The local obstacle avoidance task is used to handle real-time sudden obstacle avoidance behaviors. The energy management task is responsible for scheduling the ship speed and attitude adjustment to optimize the power consumption strategy. The model allocation sub-module 102 is used to allocate an independent reinforcement learning model to each task layer, perform policy training and parameter tuning respectively, and load preset experience parameters for each layer model during the system initialization phase to improve the training efficiency and avoid the convergence fluctuations caused by cold start.

[0106] It should be noted that the task stratification sub-module 101, as the entrance of the unmanned boat control task division module 100, is directly related to the functional boundary division of the task structure and provides a reasonable model structure basis for subsequent policy optimization. The model allocation sub-module 102 matches appropriate model structures and computing resources according to the response timeliness and target complexity of different levels, so as to ensure that each sub-task has an independent and accurate policy expression ability. The unmanned boat control task division module 100, as the starting point of this system, is the basis for constructing a multi-layer reinforcement learning control system, and the rationality of its design directly affects the cooperation efficiency of subsequent optimization modules and grid control.

[0107] S5: The layer task optimization module 200 is used to coordinate the decisions between each level through meta-learning and dynamically adjust the control strategies of each level according to the task requirements.

[0108] The layer task optimization module 200 includes a weight calculation sub-module 201 and a meta-policy coordination sub-module 202.

[0109] Furthermore, the weight calculation sub-module 201 is used to calculate the weight coefficients of the three types of task layers in real time according to dynamic indicators such as the current priority, urgency, remaining resources, and environmental complexity of the task, and complete the task priority sorting according to the predefined weight update rules. The meta-policy coordination sub-module 202 is used to dynamically allocate the execution priorities of the three-layer model according to the weight results, and coordinate the intervention degrees of different policies in scenarios of multi-task conflicts or resource constraints, so as to achieve the optimal balance between task goals.

[0110] It should be noted that the weight calculation sub-module 201 directly reflects the real-time requirements and resource status in the current task scheduling, and its results are used to guide the meta-policy coordination sub-module 202 to achieve complementary regulation between policies. The meta-policy coordination sub-module 202 realizes the coupled decision-making and synchronous optimization between reinforcement learning models by constructing an information integration interface across models, so as to maintain the coherence and dynamic robustness of the system response during multi-task execution. The layer task optimization module 200, as the central module connecting the task structure and area control, is the key mechanism to realize "multi-policy fusion" and "weight-driven scheduling", and its optimization results directly determine the efficiency and response quality of the control behavior within the grid.

[0111] S6: The grid module 300 is used to divide the navigation area of the unmanned boat into multiple grid cells after the control strategies at each level are optimized, and optimize the control decisions within each grid cell according to the real-time environmental data.

[0112] The grid module 300 includes a region division sub-module 301 and a local execution sub-module 302.

[0113] Furthermore, the region division sub-module 301 is used to divide the overall navigation sea area of the unmanned boat into a two-dimensional numbered grid structure, bind fixed coordinate boundary information to each grid cell, and store the regional state parameters (such as flow velocity, obstacle density, electromagnetic interference, etc.) in the grid mapping table for the control system to index as needed. The local execution sub-module 302 is used to call the corresponding sub-strategy of the reinforcement learning model within each grid, perform local path adjustment, obstacle avoidance reaction, and energy consumption regulation in combination with the state information of the current grid, and dynamically update the local policy parameters through the state sliding window mechanism to achieve high-frequency policy adjustment and smooth boundary connection at the micro-region level.

[0114] It should be noted that the area division sub-module 301 is the basis of the grid module 300, and its grid boundary design directly determines the granularity and computational overhead of the system's spatial autonomous control. The local execution sub-module 302, as the specific executor for policy implementation, not only realizes the task closed-loop response within each grid but also achieves the cross-grid continuation of the policy at the area boundary, ensuring the spatial continuity of the overall control logic and the dynamic stability of the navigation path. The grid module 300 is the terminal control unit of this system, and its design quality determines the local adaptability of the entire system in a changing environment and the controllability and stability of executing tasks over a large range.

[0115] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0116] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0117] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which a program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0118] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application specific integrated circuit having appropriate combinational logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. An unmanned boat grid control method based on reinforcement learning, characterized in that, Including: Dividing the control tasks of the unmanned boat into three levels: global navigation, local obstacle avoidance, and energy management, and optimizing each level using independent reinforcement learning models; Coordinating the decisions between levels through meta-learning, and dynamically adjusting the control strategies of each level according to task requirements; After the control strategies of each level are optimized, dividing the navigation area of the unmanned boat into multiple grid cells, and optimizing the control decisions within each grid cell according to real-time environmental data; Coordinating the decisions between levels through meta-learning includes dynamically adjusting the decision weights of the global navigation, local obstacle avoidance, and energy management levels based on task priorities and urgencies; The global navigation task weight is calculated according to the time limit of the task objective, the length and complexity of the navigation path; the local obstacle avoidance task weight is calculated based on the number of dynamic obstacles in the environment and the urgency of the appearance of the obstacles; the energy management task weight is calculated according to the current battery level, task duration, and environmental conditions.

2. The grid control method for an unmanned boat based on reinforcement learning according to claim 1, characterized in that: The dividing the control tasks of the unmanned boat into three levels of global navigation, local obstacle avoidance, and energy management includes, The global navigation layer performs path planning and optimizes the shortest distance and time consumption of the path; The local obstacle avoidance layer detects and avoids dynamic obstacles; The energy management layer optimizes the navigation speed and direction of the unmanned boat to maximize energy efficiency.

3. The grid control method for an unmanned boat based on reinforcement learning according to claim 2, wherein: The global navigation layer includes, Using a graph search algorithm to generate a preliminary path and combining reinforcement learning to optimize the global path selection; by adjusting the heading and speed in real time, the global path is kept optimal under different environmental conditions.

4. The grid control method for an unmanned boat based on reinforcement learning according to claim 2, characterized in that: The local obstacle avoidance layer includes, By combining lidar, sonar, and visual sensor data, the position of obstacles is detected in real time, and the heading is adjusted through a reinforcement learning algorithm to avoid collisions; When an obstacle enters the preset avoidance area of the unmanned boat, calculate and select the optimal obstacle avoidance path.

5. The grid control method for an unmanned boat based on reinforcement learning according to claim 2, characterized in that: The energy management layer includes, According to the remaining battery level, current speed, and task duration of the unmanned boat, the navigation strategy is optimized in real time, and the speed and heading of the unmanned boat are adjusted; in long-term tasks, combined with weather changes and environmental data, the energy consumption strategy of the unmanned boat is dynamically adjusted.

6. The grid control method for an unmanned boat based on reinforcement learning according to claim 5, characterized in that: The dynamically adjusting the control strategies of each level according to task requirements includes, Based on task priorities and urgencies, dynamically adjusting the decision weights of the global navigation, local obstacle avoidance, and energy management levels; The global navigation task weight is calculated according to the time limit of the task objective, the length and complexity of the navigation path; The local obstacle avoidance task weight is calculated based on the number of dynamic obstacles in the environment and the urgency of the appearance of the obstacles; The energy management task weight is calculated according to the current battery level, task duration, and environmental conditions.

7. The grid control method for an unmanned boat based on reinforcement learning according to any one of claims 1, 2, 4 or 6, characterized in that: The optimizing the control decisions within each grid cell according to real-time environmental data includes, Within each grid cell, use the reinforcement learning models of the global navigation, local obstacle avoidance, and energy management layers to perform local path planning, obstacle avoidance adjustment, and energy consumption optimization, and the tasks within each grid are executed independently; The task objectives of each grid cell are dynamically updated according to real-time environmental data, and adaptive adjustment is performed within each grid through a reinforcement learning algorithm to cope with environmental changes and task requirements.

8. An unmanned boat grid control system based on reinforcement learning, characterized in that: It includes an unmanned boat control task division module (100), a layer task optimization module (200), and a grid module (300); The unmanned boat control task division module (100) is used to divide the control tasks of the unmanned boat into three levels: global navigation, local obstacle avoidance, and energy management, and each level is optimized using an independent reinforcement learning model; The layer task optimization module (200) is used to coordinate the decisions between each level through meta-learning and dynamically adjust the control strategies of each level according to the task requirements; The grid module (300) is used to divide the navigation area of the unmanned boat into multiple grid cells after the control strategies of each level are optimized, and optimize the control decisions in each grid cell according to the real-time environmental data.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the reinforcement learning-based unmanned boat grid control method described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the reinforcement learning-based unmanned boat grid control method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Ship intelligent navigation state space construction method based on double-layer state modeling

    CN121432912A

  • Urban low-altitude unmanned aerial vehicle multi-constraint path collaborative planning method and system

    CN122062707A

  • A Multi-Constraint Path Collaborative Planning Method and System for Urban Low-Altitude Unmanned Aerial Vehicles

    CN122062707B