Unmanned ship intelligent guarding and cruising method, system, equipment and medium

By establishing a two-dimensional sea space and state model on unmanned boats, integrating energy consumption and conflict judgment, and using the Q-learning method to update the path, the problem of path quality evaluation and dynamic adjustment in unmanned boat path planning is solved, and adaptive optimization of paths and efficient resource utilization is achieved.

CN120295309APending Publication Date: 2025-07-11ZHONGYING FUND MANAGEMENT CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510433900.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing intelligent cruise route planning method of unmanned boats has the problem of a single path cost model, lack of dynamic strategy optimization capabilities, and inability to conduct path quality level assessment, making it difficult to take into account energy consumption, obstacle avoidance and path conflict assessment in a dynamic environment.

Method used

Establish a two-dimensional sea space and set up an unmanned boat state model, integrate energy consumption and conflict judgment, build an objective function for optimization of unmanned boat missions, use the Q-learning method to perform path updates and optimal action selection, and use the multi-factor weighted combination function to perform path evaluation and dynamic adjustment.

Benefits of technology

It realizes the adaptability and optimization of path planning in complex dynamic environments, improves path security and resource utilization efficiency, has self-learning ability, and can perform optimal control in multi-objective collaborative tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295309A_ABST
    Figure CN120295309A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned ship intelligent guarding and cruising method, system, equipment and medium, and relates to the technical field of intelligent path planning and task guarding control of ocean unmanned ships, and the method comprises the steps: building a two-dimensional sea area space, and setting an unmanned ship state model; integrating energy consumption and conflict judgment to construct an unmanned ship task optimization objective function, and performing classification management on a current path; and path updating and optimal action selection are carried out based on a Q-learning method. According to the method, energy consumption estimation is carried out by fusing the navigational speed, the acceleration and the resistance coefficient, the multi-factor cost model is constructed by combining the path conflict and the avoidance priority, and the path safety and the resource utilization efficiency are improved. And path updating and optimal action selection are further performed by adopting a Q-learning method, so that optimal control and intelligent scheduling of a task path can be completed in a complex dynamic sea area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent path planning and mission duty control of unmanned marine boats, and in particular to an intelligent duty and cruising method, system, equipment and medium for unmanned boats. Background Art

[0002] In recent years, with the continuous development of marine intelligent equipment and autonomous control technology, unmanned boats have gained wide attention as an important platform for performing surface patrol, reconnaissance, search and rescue, and monitoring tasks. Their intelligence level has gradually improved, especially in terms of autonomous navigation and path planning capabilities in mission execution, and they have begun to integrate multi-source perception, path prediction, and intelligent decision-making algorithms. At present, some unmanned boat systems can already realize basic functions such as regional autonomous cruising and target area monitoring. Combined with real-time task scheduling, local obstacle avoidance, and energy consumption control, they have gradually formed a system framework with certain task autonomy capabilities. However, issues such as multi-boat collaboration, adaptive path adjustment, and task benefit optimization in complex mission scenarios still need further research.

[0003] However, existing unmanned boat path planning and guarding technologies generally have several key deficiencies. Most current path planning methods are based only on geometric shortest paths or static obstacle modeling, and lack the ability to model multi-objective path costs for comprehensive factors (such as energy consumption, conflict probability, and collaborative obstacle avoidance) in dynamic environments. It is difficult to balance task efficiency and energy optimization while ensuring path feasibility. Most methods use static strategies or heuristic algorithms for path planning, which do not have the ability to self-learn and adapt, and cannot dynamically adjust strategies based on historical mission experience, making it difficult to adapt to the characteristics of frequent environmental changes and flexible target adjustments in actual tasks. Existing path selection mechanisms usually do not introduce path classification or task level evaluation, and lack unified path quality standards, resulting in inconsistent path evaluation standards and non-optimal selection. Summary of the invention

[0004] In view of the above-mentioned problems, the present invention is proposed.

[0005] Therefore, the technical problem solved by the present invention is: the existing unmanned boat intelligent cruise path planning method has the problems of a single path cost model, lack of dynamic strategy optimization capability, and inability to perform path quality level evaluation; and how to construct an optimized path model that takes into account energy consumption, obstacle avoidance and path conflict evaluation, and realize path updating and optimal action selection through reinforcement learning methods.

[0006] To solve the above technical problems, the present invention provides the following technical solution: an intelligent monitoring and cruising method for an unmanned boat, including establishing a two-dimensional sea area space and setting an unmanned boat state model; integrating energy consumption and conflict judgment to construct an optimization objective function for the unmanned boat's tasks, and classifying and managing the current path; performing path update and optimal action selection based on the Q-learning method; establishing a two-dimensional sea area space including setting horizontal and vertical coordinate boundaries to limit the navigable range of the unmanned boat, and state modeling including constructing a joint vector of the current position, speed, and heading angle of the unmanned boat; the optimization objective function for the unmanned boat's tasks includes constructing a weighted combination function of multiple cost factors based on time, energy consumption, path conflict, and obstacle avoidance priority, and dynamically evaluating it within the task cycle; path update includes using Q-learning to update the Q value of the state-action pair, and selecting the current execution action by combining the immediate cost feedback and the predicted maximum benefit.

[0007] As a preferred embodiment of the intelligent monitoring and cruising method for an unmanned boat described in the present invention, wherein: establishing the two-dimensional sea area space includes modeling the sea area as a plane coordinate region, and setting horizontal and vertical boundaries to limit the navigable range of the unmanned boat; the motion state of the unmanned boat is composed of the position, speed, and sailing direction angle of the unmanned boat in the sea area, and by setting the starting coordinate and the target coordinate, and combining the current speed parameter of the unmanned boat, the time required for the unmanned boat to reach the target point is calculated.

[0008] As a preferred embodiment of the intelligent monitoring and cruising method for an unmanned boat described in the present invention, wherein: integrating energy consumption and conflict judgment includes evaluating the cumulative energy consumption of the energy consumption cost based on the motion parameters of the unmanned boat and combining the sailing resistance coefficient, and judging the path conflict cost based on the real-time spatial distance relationship between unmanned boats. If the distance between two boats is detected to be lower than a preset safety threshold, it will be marked as having a potential conflict and a conflict loss value will be imposed. At the same time, according to the priority setting of surrounding unmanned boats, a penalty cost is calculated for the unmanned boat that needs to actively avoid, guiding path adjustment.

[0009] As a preferred embodiment of the intelligent monitoring and cruising method for an unmanned boat described in the present invention, wherein: the optimization objective function for the unmanned boat's tasks includes accumulating the weighted combination of multiple cost factors, including the time required for sailing, the estimated energy consumption, the path conflict evaluation with different unmanned boats, and the avoidance cost of neighboring boats that may cause collisions; by configuring the weight coefficients of each cost factor, the path planning takes into account the execution efficiency, safety, and energy consumption, evaluates the total cost of the path within the entire task cycle, and determines the optimization objective based on the evaluation results.

[0010] As a preferred solution of the unmanned boat intelligent monitoring and cruising method of the present invention, wherein: the classification management of the current path includes that according to the value calculated by the unmanned boat task optimization objective function, the path will be divided into three types: optimal, acceptable, and high-risk. The classification basis is whether the path cost value is within the optimal range, the tolerable range, or exceeds the risk threshold.

[0011] As a preferred solution of the unmanned boat intelligent monitoring and cruising method of the present invention, wherein: the path update and optimal action selection include that during the path optimization process, the state and action of the unmanned boat are formed into a state-action pair, and the expected benefit of the executed action in the current state is learned based on the Q-learning method. The expected benefit value is updated through the execution result of each step of the task. During the update process, the immediate feedback situation and the predicted future feedback are combined to dynamically optimize the path quality.

[0012] As a preferred solution of the unmanned boat intelligent monitoring and cruising method of the present invention, wherein: the path update and optimal action selection further include that in each task time step, based on the expected benefits corresponding to all possible actions in the current state, the action with the highest benefit value is selected as the execution instruction for this time step. The action with the highest benefit value will be transmitted as a control parameter to the heading and propulsion module of the unmanned boat to generate a control command.

[0013] Another object of the present invention is to provide an unmanned boat intelligent monitoring and cruising system, which can construct an unmanned boat task optimization objective function by integrating energy consumption and conflict judgment, classify and manage the current path, and solve the problem of low efficiency in the current unmanned boat intelligent cruising path planning technology.

[0014] As a preferred solution of the unmanned boat intelligent monitoring and cruising system of the present invention, wherein: it includes a state construction module, a cost calculation module, and a path optimization module; the state construction module is used to establish a two-dimensional sea area space and set an unmanned boat state model; the cost calculation module is used to integrate energy consumption and conflict judgment to construct an unmanned boat task optimization objective function and classify and manage the current path; the path optimization module is used to perform path update and optimal action selection based on the Q-learning method.

[0015] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the unmanned boat intelligent monitoring and cruising method are implemented.

[0016] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the unmanned boat intelligent monitoring and cruising method are implemented.

[0017] Advantages of the present invention: The intelligent monitoring and cruising method for unmanned boats provided by the present invention realizes the precise abstraction of the sea area task environment into a plane coordinate system with boundary constraints by establishing a two-dimensional sea area space and setting an unmanned boat state model, and constructs a unified state vector in combination with position information, speed and heading angle, providing continuously adjustable input variables for path planning and dynamic control, and enhancing the system's mapping and time-effect modeling capabilities for the actual geographical environment. On the basis of integrating energy consumption and conflict judgment to construct an optimization objective function for unmanned boat tasks, energy consumption is estimated by fusing speed, acceleration and drag coefficient, and a multi-factor cost model is constructed by combining path conflicts and avoidance priorities, realizing the unified evaluation and risk classification management of path selection, and improving path safety and resource utilization efficiency. Further, the Q-learning method is used for path update and optimal action selection, and the current state and behavior feedback of the unmanned boat are used for Q-value iterative update, realizing the closed-loop reinforcement learning of action selection and environment feedback, enabling the path planning to have the capabilities of self-adaptation, learnability and self-optimization, so as to ensure the optimal control and intelligent scheduling of the task path in a complex and dynamic sea area. The present invention has better effects in terms of accuracy, reliability and flexibility. Brief Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0019] Figure 1 It is the overall flowchart of the intelligent monitoring and cruising method for unmanned boats provided by the first embodiment of the present invention. Detailed Embodiments

[0020] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention in conjunction with the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0021] Embodiment 1, referring to Figure 1 , which is an embodiment of the present invention, provides an intelligent monitoring and cruising method for unmanned boats, including:

[0022] S1: Establish a two-dimensional sea area space and set an unmanned boat state model.

[0023] Furthermore, a two-dimensional sea area space is established, where the sea area is modeled as a planar coordinate region. The region defines the navigable range of the unmanned boat by setting horizontal and vertical boundaries. The motion state of the unmanned boat consists of its position, speed, and navigation direction angle in the sea area. By setting the starting point coordinates and the target point coordinates, and combining the current speed parameters of the unmanned boat, the time required for the unmanned boat to reach the target point is calculated.

[0024] It should be noted that a preferred scheme for establishing a two-dimensional sea area space and setting the unmanned boat state model specifically includes mathematical modeling of the ocean cruise task of the unmanned boat. The sea area space S for the unmanned boat cruise is constructed through a two-dimensional plane and is expressed as:

[0025] S = {(X, Y)|X ∈ [X min , X max , Y ∈ [Y min , Y max}

[0026] Among them, S is the sea area space set, X is the horizontal coordinate in the sea area, Y is the vertical coordinate in the sea area, X min and X max are respectively the minimum and maximum horizontal coordinate ranges of the unmanned boat in the sea area, and Y min and Y max are respectively the minimum and maximum vertical coordinate ranges of the unmanned boat in the sea area; the state of each unmanned boat u i is expressed as:

[0027] S i = (X i , Y i , V i , Θ i )

[0028] Among them, u i represents the i-th unmanned boat in the unmanned boat cluster, S i is the state vector of the unmanned boat i, including position information, speed, and heading. X i and Y i respectively represent the current abscissa and ordinate of the unmanned boat i, V i represents the navigation speed of the unmanned boat i, and Θ i represents the navigation direction angle of the unmanned boat i; the unmanned boat moves from the starting point (X0, Y0) to the target point (X T , Y T ), and the time required to complete the task is expressed as:

[0029]

[0030] Among them, T irepresents the time required for the unmanned boat i to travel from the starting point to the ending point. X0 and Y0 respectively represent the starting horizontal coordinate and vertical coordinate of the unmanned boat, and X T and Y T respectively represent the target horizontal coordinate and vertical coordinate of the unmanned boat.

[0031] It should also be noted that by constructing a two-dimensional sea area space and abstracting the sea surface mission environment into a plane coordinate system with boundary constraints through mathematical modeling, the navigable area of the unmanned boat is limited, thereby improving the mapping ability of the path planning model to the actual geographical boundary. Then, by setting the state vector of the unmanned boat, including position, speed, and heading angle, quantifiable input state information is provided for subsequent path optimization. By setting the starting point and the ending point and estimating the mission time in combination with the current sailing speed, the ability of the model to control the mission timeliness is further enhanced, and a path time estimation mechanism based on the unified modeling of geometric distance and dynamic state is realized. This step not only constructs the environmental basis for path planning but also makes the state dynamics of the unmanned boat predictable and controllable, providing a standardized and continuous input expression for subsequent energy consumption calculation, path cost evaluation, and motion optimization, thereby achieving the beneficial effect of improving the path modeling accuracy and global planning adaptability.

[0032] S2: Integrate energy consumption and conflict judgment to construct an optimization objective function for the unmanned boat mission and classify and manage the current path.

[0033] Furthermore, integrating energy consumption and conflict judgment includes that the energy consumption cost is based on the motion parameters of the unmanned boat and combines the sailing resistance coefficient to evaluate the cumulative energy consumption. The judgment of the path conflict cost is based on the real-time spatial distance relationship between unmanned boats. If the distance between two boats is detected to be lower than the preset safety threshold, it will be marked as having a potential conflict and a conflict loss value will be imposed. At the same time, according to the priority setting of the surrounding unmanned boats, a penalty cost is calculated for the unmanned boat that needs to actively avoid, guiding path adjustment.

[0034] It should also be noted that a preferred scheme for integrating energy consumption and conflict judgment specifically includes that the energy consumption of the unmanned boat during navigation is expressed as:

[0035]

[0036]

[0037] where E i represents the total energy consumed by the unmanned boat i during navigation, a i is the acceleration of the unmanned boat i, k1 is the water resistance coefficient, ρ is the density of water, C D is the resistance coefficient of the unmanned boat; s is the stress area of the unmanned boat; detecting whether the unmanned boat i has a path intersection with the unmanned boat j is expressed as:

[0038]

[0039] Among them, D col.i is the cost of path conflict for the unmanned boat i, and δ i,j is the path intersection indication function. If there is a path intersection between i and j, then δ i,j = 1, otherwise it is 0. κ is the sensitivity parameter for conflict detection, and d i,j is the minimum Euclidean distance between the unmanned boats i and j, and d th is the safety distance threshold for path conflict; The unmanned boat needs to avoid high-priority unmanned boats during navigation. The calculation of the collision avoidance loss is expressed as:

[0040]

[0041] Among them, C adj,i is the cost paid by the unmanned boat i to avoid high-priority unmanned boats during navigation. N is the total number of unmanned boats, and β k is the priority weight of the unmanned boat k. The larger the value of β k , the more urgent the target to be avoided. X k and Y k represent the lateral coordinate and longitudinal coordinate of the unmanned boat k respectively.

[0042] It should be noted that the optimization objective function of the unmanned boat mission includes combining multiple cost factors with weights and then accumulating them, including the time required for navigation, the estimated energy consumption, the path conflict assessment with different unmanned boats, and the avoidance cost of neighboring boats that may cause collisions; By configuring the weight coefficients of each cost factor, the path planning takes into account the execution efficiency, safety, and energy consumption, evaluates the total cost of the path during the entire mission cycle, and determines the optimization objective based on the evaluation results.

[0043] It should also be noted that the final optimization objective function of the unmanned boat mission is expressed as:

[0044]

[0045] Among them, J i represents the total optimization objective of the unmanned boat i. T is the total number of time steps of the mission. α, β, γ, δ are the weight parameters of the mission time, energy consumption, path conflict, and obstacle avoidance respectively, which are used to dynamically adjust the optimization strategy.

[0046] It should also be noted that the classification management of the current path includes that according to the value calculated by the optimization objective function of the unmanned boat mission, the path will be divided into three types: optimal, acceptable, and high-risk. The classification basis is whether the path cost value is within the optimal range, the tolerable range, or exceeds the risk threshold.

[0047] It should also be noted that a preferred solution for classifying and managing the current path specifically includes if d i,j >d th , then D col,i ≈0, indicating that there is no path conflict and the optimization goal can ignore path adjustment; J i The smaller the value, the better the path planning; set the optimal threshold J opt and the risk threshold J risk ; when J i ≤J opt , the task path is optimal, with high cruising efficiency, low energy consumption, and very few path conflicts; when J opt <J i ≤J risk , the task path is acceptable, the cruising can still be optimized, and the speed or avoidance strategy needs to be adjusted; when J i >J risk , the task path has a greater risk, there may be serious path conflicts or high energy consumption, and the path needs to be forced to be optimized.

[0048] It should also be noted that by integrating energy consumption calculation, path conflict identification, and avoidance mechanism into a unified task objective function model, the path of the unmanned boat is comprehensively evaluated at the cost level, significantly enhancing the adaptability to multi-constraint scenarios in the path planning process. The energy consumption calculation part not only considers the combined influence of acceleration and speed, but also introduces the navigation resistance coefficient, which truly reflects the differences in energy consumption of different dynamic behaviors. The conflict judgment is based on the dynamic Euclidean distance and safety threshold, and quantifies and punishes the adjacent path crossing and approaching behaviors, assisting the system to effectively avoid path overlap or collision. By combining cost factors with different weights into a unified evaluation index, the unified modeling of path execution time, energy consumption, conflict, and obstacle avoidance cost is realized. Furthermore, the path quality can be classified and graded according to the total cost value, and it can be dynamically determined whether the path needs to be adjusted or re-planned. It improves the adaptive ability and intelligent scheduling level of the path planning system, enables the path evaluation to have a quantitative standard and the possibility of iterative optimization, thus achieving the beneficial effects of improving path safety, energy efficiency ratio, and dynamic management ability.

[0049] S3: Perform path update and optimal action selection based on the Q-learning method.

[0050] Furthermore, performing path update and optimal action selection includes, in the process of path optimization, forming a state-action pair with the state of the unmanned boat and the executed action, learning the expected reward of the executed action in the current state based on the Q-learning method, updating the expected reward value through the result of each task execution, and dynamically optimizing the path quality by combining the immediate feedback situation and the predicted future feedback during the update process.

[0051] It should be noted that path update and optimal action selection also include, in each task time step, based on the expected rewards corresponding to all possible actions in the current state, selecting the action with the highest reward value as the execution instruction for this time step. The action with the highest reward value will be transmitted as a control parameter to the heading and propulsion modules of the unmanned boat to generate a control command.

[0052] It should also be noted that a preferred solution for path update and optimal action selection specifically includes, to improve the path selection ability of the unmanned boat in complex environments, using reinforcement learning Q-learning for dynamic path optimization to enable it to have an adaptive adjustment ability. The value update of reinforcement learning Q-learning is expressed as:

[0053]

[0054] Where Q(S t ,A t ) is the Q value of taking action A t in state S t ; S t is the current state of the unmanned boat, including position, speed, task information, etc.; A t is the path optimization strategy selected by the current unmanned boat (such as obstacle avoidance, acceleration, etc.); R t is the immediate reward, measuring the quality of the current path selection; α' is the learning rate, determining the update speed of the Q value, with a value range of (0,1]; γ' is the discount factor, with a value of (0,1], used to measure the importance of future rewards; A' is the set of possible actions that can be taken in the next state S t+1 ; Q(S t+1 ,A′) is the Q value of the unmanned boat taking the optimal action in the next state S t+1 ; The immediate reward R t is expressed as:

[0055] R t = -(λ1D col,i + λ2C adj,i + λ3E i + λ4T i )

[0056] λ1, λ2, λ3, λ4 are the weight coefficients of different costs, used to dynamically adjust the optimization strategy. The optimal strategy selection is expressed as:

[0057]

[0058] A traverses all possible actions; if the task of the unmanned boat is urgent, the weight coefficient λ4 should be appropriately increased to make T iIf the impact is greater, the shortest path is preferentially selected; if the marine environment is complex, λ3 needs to be reduced to decrease the energy consumption optimization weight so that the unmanned boat can avoid obstacles preferentially; if the endurance is limited, then λ3 needs to be adjusted to reduce E i influence and improve the endurance.

[0059] It should also be noted that by introducing the Q-learning reinforcement learning mechanism, the state and actions of the unmanned boat are formed into a state-action pair, and the Q value is updated by means of interactive learning, realizing the transformation of path selection from empirical heuristic to data-driven policy learning. After each path execution, the system updates the Q value according to the immediate reward, and optimizes the current decision-making strategy by combining the prediction of future benefits. The immediate reward incorporates path conflict, energy consumption, time, and obstacle avoidance factors, and its dynamic weight adjustment mechanism enables the system to make differential action selections under the influence of multiple factors such as task urgency, environmental complexity, and energy state. Finally, the action with the largest Q value is selected from all optional actions through the policy selection function to achieve the optimal policy output in the local state, and this is used to drive the generation of instructions for adjusting the heading and speed of the unmanned boat. This step not only improves the online learning ability of path optimization, but also makes the path update process have task adaptability and sustainable optimality, ultimately achieving the beneficial effects of realizing autonomous learning, dynamic adjustment, and optimal path control.

[0060] Embodiment 2, an embodiment of the present invention, provides an intelligent monitoring and cruising method for an unmanned boat. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0061] First, the experimental area is simulated as a closed rectangular sea area of 1000m×800m, and the navigable space of the unmanned boat is restricted by setting the horizontal and vertical boundary coordinates. A total of 6 simulated unmanned boats with controllable parameters are deployed. The initial starting point and target point of each unmanned boat are known, and the length of the starting and ending paths is between 450m and 600m. The state vector of each boat includes the current position, speed, and heading angle, and the initial values are set by the task scheduling module. All unmanned boats receive preliminary path planning instructions before the task starts, and the navigation speed, energy consumption, conflict risk coefficient, and obstacle avoidance reaction data are recorded in real time by the path control system during the execution process.

[0062] The path optimization evaluation function comprehensively considers four indicators: sailing time, energy consumption, conflict risk, and obstacle avoidance cost. All cost items are collected through real-time sensors and path monitoring systems and transmitted back to the control terminal for cumulative statistics. At each time step, the system updates the cost values corresponding to each boat according to the current path performance and performs path classification and strategy adjustment. The system sets the optimization path threshold to 260 and the risk path threshold to 300. The control module automatically marks the task as "optimal", "acceptable", or "high risk" and triggers the corresponding path reconstruction mechanism. During the experiment, some unmanned boats adopted traditional heuristic path strategies as the comparison baseline, while the others updated their path selections based on reinforcement learning. The training iteration step size is 50 steps, and the learning rate dynamically adjusts within the range of 0.05 - 0.2. Refer to Table 1 for recording and analyzing some experimental data.

[0063] Table 1 Experimental Data Record Table

[0064]

[0065] From the experimental data in the table, it can be seen that the unmanned boats adopting the solution of the present invention (such as unmanned boats B, D, and F) generally show lower path costs, which are 256.1, 244.3, and 250.5 respectively, all within the "optimal path" range defined by the system. This indicates that these boats have achieved an overall balance in terms of mission time, energy consumption, conflict, and obstacle avoidance performance, and show superior multi-objective collaborative optimization capabilities compared to traditional strategies (such as unmanned boat C). As the control group, the path cost of unmanned boat C is as high as 308.4, exceeding the high-risk threshold. The main reason is that there are relatively high conflict risks and obstacle avoidance costs in its path, indicating the lack of an effective conflict detection and dynamic adjustment mechanism.

[0066] Further analysis reveals that during the path update process, the method of the present invention strengthens the behavior selection of low conflict and low energy consumption through the Q-learning strategy, effectively suppressing the trend of high-risk path evolution. Especially in unmanned boats D and F, even though their speeds are relatively high, the system can still adjust their action strategies in real time to control the path cost within the optimal range, demonstrating that the reinforcement learning path scheduling mechanism has good convergence and task adaptation capabilities for high-speed boats. In addition, although the speed of unmanned boat A is medium, due to the relatively high conflict risk, its path cost is slightly higher than that of B, D, and F, further confirming that the present invention gives full weight to the conflict factor in the process of constructing the cost function and improves the response ability to safety.

[0067] The experimental results fully prove that the multi-factor optimization objective function and Q-learning adaptive path update method based on the present invention have better task completion quality, path safety, and resource utilization efficiency when facing dynamic environments and multi-boat collaborative tasks, and have obvious technological innovation and practical value compared to traditional heuristic path planning methods.

[0068] Embodiment 3 is an embodiment of the present invention, which provides an intelligent monitoring and cruising system for unmanned boats, including a state construction module 100, a cost calculation module 200, and a path optimization module 300.

[0069] Among them, the state construction module 100 is used to establish a two-dimensional sea area space and set the unmanned boat state model; the cost calculation module 200 is used to integrate energy consumption and conflict judgment to construct an unmanned boat task optimization objective function and classify and manage the current path; the path optimization module 300 is used to update the path and select the optimal action based on the Q-learning method.

[0070] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0071] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0072] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0073] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. Method for intelligent monitoring and cruising of unmanned boat, characterized in that, Including: Establish a two-dimensional sea area space and set up an unmanned boat state model; Integrate energy consumption and conflict judgment to construct an optimization objective function for the unmanned boat's mission, and classify and manage the current path; Based on the Q-learning method, perform path update and optimal action selection; Establishing a two-dimensional sea area space includes setting horizontal and vertical coordinate boundaries to limit the navigable range of the unmanned boat. State modeling includes constructing a joint vector of the current position, speed, and heading angle of the unmanned boat. The optimization objective function for the unmanned boat's mission includes constructing a weighted combination function of multiple cost factors based on time, energy consumption, path conflict, and obstacle avoidance priority, and performing dynamic evaluation within the mission cycle. Path update includes updating the Q value of the state-action pair using Q-learning, and selecting the current execution action by combining the immediate cost feedback and the predicted maximum benefit.

2. The intelligent monitoring and cruising method for unmanned boats according to claim 1, wherein: The establishment of the two-dimensional sea area space includes, The sea area is modeled as a planar coordinate region, and the region limits the navigable range of the unmanned boat by setting horizontal and vertical boundaries; The motion state of the unmanned boat is composed of the position, speed, and sailing direction angle of the unmanned boat in the sea area. By setting the starting coordinate and the target coordinate, and combining the current speed parameter of the unmanned boat, calculate the time required for the unmanned boat to reach the target point expectedly.

3. The intelligent monitoring and cruising method for unmanned boats according to claim 1 or 2, characterized in that: The integration of energy consumption and conflict judgment includes, The energy consumption cost is based on the motion parameters of the unmanned boat and combines the sailing resistance coefficient to evaluate the cumulative energy consumption. The judgment of the path conflict cost is based on the real-time spatial distance relationship between unmanned boats. If the distance between two boats is detected to be lower than the preset safety threshold, it will be marked as having a potential conflict and a conflict loss value will be imposed. At the same time, according to the priority setting of the surrounding unmanned boats, a penalty cost will be calculated for the unmanned boat that needs to actively avoid, guiding path adjustment.

4. The intelligent monitoring and cruising method for unmanned boats according to claim 3, characterized in that: The optimization objective function for the unmanned boat's mission includes, Accumulate after weighting and combining multiple cost factors, including the time required for sailing, the expected energy consumption, the path conflict evaluation with different unmanned boats, and the avoidance cost of neighboring boats that may cause collisions; By configuring the weight coefficients of each cost factor, make the path planning take into account the execution efficiency, safety, and energy consumption, evaluate the total cost of the path within the entire mission cycle, and determine the optimization objective based on the evaluation results.

5. The intelligent monitoring and cruising method of the unmanned boat according to claim 1, 2 or 4, characterized in that: The classification and management of the current path includes, According to the value calculated by the optimization objective function for the unmanned boat's mission, the path will be divided into three types: optimal, acceptable, and high-risk. The classification basis is whether the path cost value is within the optimal range, the tolerable range, or exceeds the risk threshold.

6. The intelligent monitoring and cruising method for unmanned boats according to claim 5, characterized in that: The path update and optimal action selection include, During the path optimization process, form a state-action pair with the state of the unmanned boat and the executed action. Based on the Q-learning method, learn the expected benefit of the executed action in the current state, and update the expected benefit value through the result of each step of the task execution. During the update process, combine the immediate feedback situation and the predicted future feedback to dynamically optimize the path quality.

7. The intelligent monitoring and cruise method for unmanned boats according to claim 1, 2, 4 or 6, characterized in that: The path update and optimal action selection also includes, In each task time step, based on the expected rewards corresponding to all possible actions in the current state, select the action with the highest reward value as the execution instruction for this time step. The action with the highest reward value will be transmitted as a control parameter to the heading and propulsion modules of the unmanned boat to generate a control command.

8. Unmanned boat intelligent monitoring and cruising system, characterized in that: It includes a state construction module (100), a cost calculation module (200), and a path optimization module (300); The state construction module (100) is used to establish a two-dimensional sea area space and set up an unmanned boat state model; The cost calculation module (200) is used to integrate energy consumption and conflict judgment to construct an unmanned boat task optimization objective function and classify and manage the current path; The path optimization module (300) is used to update the path and select the optimal action based on the Q-learning method.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the unmanned boat intelligent monitoring and cruising method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the unmanned boat intelligent monitoring and cruising method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Robot cluster collaborative collision avoidance method and device for electric power system inspection

    CN121165716A