Efficient operation path planning method for aquaculture unmanned mowing boat

Through the improved Q-Learning algorithm, the problem that unmanned mowing boats are difficult to efficiently plan their operation paths in complex and dynamic aquaculture pond environments is solved, and efficient and low-energy mowing operation path planning is achieved.

CN120194699APending Publication Date: 2025-06-24JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510247765.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently plan the operation path of unmanned mowed boats in complex and dynamic aquaculture pond environments, resulting in low operational efficiency, high energy consumption and difficulty in avoiding dynamic obstacles.

Method used

Using the improved Q-Learning algorithm, an optimal Q value table is generated to plan the optimal collision-free operation path through eight direction exploration mechanisms, redesigned reward functions and dynamically adjusted exploration rate.

Benefits of technology

It significantly improves the operating efficiency of unmanned mowing boats, reduces energy consumption, and effectively avoids dynamic obstacles, ensuring efficient completion of mowing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120194699A_ABST
    Figure CN120194699A_ABST
Patent Text Reader

Abstract

The invention discloses a high-efficiency operation path planning method for an aquaculture unmanned mowing boat, and the method comprises the steps: scanning the overall environment of a pond, collecting the latitude and longitude coordinates of the vertexes and obstacles of the pond, and converting the latitude and longitude coordinates into plane coordinates; modeling an environment map by adopting a grid method, and establishing an obstacle matrix; selecting the position of a grass unloading wharf; the direction, the reward function and the exploration rate of a traditional Q-Learning algorithm are improved and trained, and an optimal Q value table is generated; and an optimal Q value table generated by using an improved Q-Learning algorithm is utilized to plan an optimal operation path from a return point to the grass unloading wharf in real time. And if a dynamic obstacle is encountered in the course of returning, reselecting the action and updating the Q value table to avoid the obstacle. And after completing grass unloading, the mowing boat travels according to the original path to continue to execute the mowing operation at the interruption position of the last operation until the mowing task of the whole pond is completed. The method can solve the problem of actual path planning of returning to a wharf for unloading grass for many times due to the fact that the collection box is full of grass for the aquaculture unmanned grass cutting ship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of aquaculture and path planning, and relates to an efficient operation path planning method for an unmanned grass cutting boat in aquaculture. Background Art

[0002] With the development of social economy, the demand for river crabs and crayfish is increasing day by day, and the aquaculture industry of river crabs and crayfish has developed rapidly. At present, the pond culture mode of river crabs and crayfish mainly adopts grass planting. While acting as a shelter, aquatic plants can also absorb and decompose harmful substances in water. Due to the excessive growth of aquatic plants in summer, the water quality is likely to deteriorate after decay. Therefore, the aquaculture ponds of river crabs and crayfish need to clean the aquatic plants regularly. At present, it mainly relies on manual boating to salvage aquatic plants, with high labor intensity and low operation efficiency. The unmanned grass cutting boat has changed the previous operation mode of manual salvaging of aquatic plants. The operation boat can automatically harvest aquatic plants according to the set cruising route, significantly improving the operation efficiency and reducing the labor cost. Since the current aquaculture ponds of river crabs and crayfish are relatively large in area, it is difficult for the unmanned grass cutting boat to harvest all the aquatic plants in the water area in one operation. Therefore, it is necessary to return to the unloading dock multiple times to transport the harvested aquatic plants, and then return to the grass cutting interruption point to continue the operation. Due to the presence of various obstacles such as aeration pipes, bamboo poles, and aerators in the pond, traditional point-to-point path planning algorithms are mostly applicable to maps in static and simple environments, and are prone to falling into local optima in complex environments, and it is difficult to plan a collision-free optimal path with a short path and few turns in a complex dynamic environment.

[0003] The Q-Learning algorithm is one of the commonly used algorithms for solving dynamic path planning problems. When dealing with known obstacle maps, it has better learning, adaptation, and dynamic adjustment capabilities. However, the traditional Q-Learning algorithm has problems such as slow convergence speed and difficulty in balancing the relationship between exploration and exploitation, and the planned path is difficult to meet the requirements of efficient operation of the unmanned grass cutting boat. Therefore, the present invention proposes an efficient operation path planning method for an unmanned grass cutting boat in aquaculture based on an improved Q-Learning algorithm to improve the operation efficiency of the grass cutting boat and reduce energy consumption. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides an efficient operation path planning method for an unmanned grass cutting boat in aquaculture.

[0005] The present invention is realized through the following technical solutions:

[0006] Step S1: Scan the overall environment of the pond, collect the longitude and latitude coordinates of the pond vertices and obstacles, and convert them into plane coordinates;

[0007] Step S2: Model the environmental map using the grid method to establish an obstacle matrix;

[0008] Step S3: Select the location of the unloading dock;

[0009] Step S4: Improve the direction, reward function, and exploration rate of the traditional Q-Learning algorithm, and train it to generate an optimal Q-value table;

[0010] Step S5: Use the optimal Q-value table generated by the improved Q-Learning algorithm to real-time plan the optimal operation path from the return point to the unloading dock; if a dynamic obstacle is encountered during the return journey, reselect the action and update the Q-value table to avoid this obstacle; after the unmanned mowing boat finishes unloading the grass, drive back along the original path to the interruption point of the previous operation and continue with the mowing operation until the mowing task of the entire pond is completed.

[0011] Furthermore, the specific steps of the said Step S2 are as follows:

[0012] Use the grid method to model the pond environment, and the grid size is determined by the size of the mowing boat and the size of the operation area. For some boundary obstacles, after rasterization, they only occupy a small part of the grid area, and it is impossible to directly set them as obstacle areas. Determine whether the grid on the boundary line is an obstacle area by checking whether the center point of the grid is located in the obstacle area.

[0013] Furthermore, the specific steps of the said Step S4 are as follows:

[0014] Step S41: Change the four-direction exploration of the traditional Q-Learning algorithm's environment map to eight-direction exploration, increasing the number of available step samples. Since the mowing boat is relatively large, in order to solve problems such as easy missed mowing during right-angle turning and excessive energy consumption due to too long path length during operation, the present invention expands the original four directions to eight directions for exploration, adding four exploration directions: upper left, lower left, upper right, and lower right. In the present invention, the path length of translating one unit is set as d. If translating one grid in the diagonal direction, the path unit length is

[0015] Step S42: Redesign the reward function R. The reward mechanism of the traditional Q-learning algorithm only includes rewards for normal walking, encountering obstacles, and reaching the target point. The present invention improves the reward function as shown in formula (1). A corresponding reward value is set for each step during the movement of the unmanned mowing boat, reducing the influence of misleading rewards, providing more detailed and specific feedback information to guide the unmanned mowing boat to learn, and enabling it to converge to the optimal strategy faster. The improved rewards are as follows:

[0016] (1) Give a penalty r4 for turning. In order to reduce the energy consumption of the mowing boat during operation, give a turning penalty to reduce the number of turns in the path.

[0017] (2) Impose a penalty r5 for the mower barge staying in place when hitting the map boundary. During the training process, the mower barge constantly hits the map boundary. Therefore, to accelerate the convergence speed, a penalty r5 is imposed after hitting the obstacle boundary.

[0018] (3) Impose a penalty r6 for hitting an obstacle obliquely. Due to the working width limitation of the mower barge, oblique movement is likely to hit obstacles. Therefore, a penalty is imposed to avoid such situations.

[0019]

[0020] Step S43: To balance the relationship between exploration and exploitation, the present invention designs a dynamic adjustment function ξ(n) with the iteration number n as the independent variable to dynamically adjust the greedy factor. The traditional Q-learning algorithm selects actions using the ε-greedy strategy, and the exploration rate ε ∈ (0, 1) is generally a fixed value, as shown in formula (2). A represents the action selected according to the ε-greedy strategy, a * represents the action that selects the maximum Q value, a r represents randomly selecting an action, and P represents the probability of selecting action a * or a r of.

[0021]

[0022] Regarding the problem in the traditional Q-learning algorithm that when the value of ε is too large, the mower barge has been constantly exploring the environment, resulting in a slow convergence speed; when the value of ε is too small, the mower barge will overly utilize the existing knowledge to select the action with the maximum Q value during the learning process, and may fall into a local optimum due to insufficient learning of the unknown environment, the present invention proposes a dynamic adjustment function ξ(n), and the value of the adjustment factor ξ gradually decreases as the iteration number n increases. After the algorithm iterates k times, the value of ε remains unchanged. The optimized dynamic adjustment function enables the algorithm to fully explore the environment in the early stage. As the iteration number increases, the exploration rate gradually decreases, and the mower barge is more inclined to utilize the known environmental information, accelerating the algorithm convergence speed while avoiding the algorithm falling into a local optimum.

[0023] Among them, the dynamic adjustment function ξ(n) is defined as follows, where i is an adjustable parameter and N is the maximum number of iterations.

[0024]

[0025] Step S44: Based on the above improvements, first create a Q-table to store each state-action pair, initialize the Q-table to 0, and initialize the learning rate (α), discount factor (γ), and exploration rate (ε). Among them, the Q-value function update rule is as follows:

[0026] Q(st , a t ) ← Q(s t , a t ) + α[r t+1 + γ max Q(s t+1 , a t+1 ) - Q(s t , a t )](4)

[0027] Where γ ∈ (0, 1) is the discount rate, α ∈ (0, 1) is the learning rate, r t+1 is the immediate reward, s t is the state of the mowing boat at time t, a t is the action selected by the mowing boat in state s t , s t+1 is the state after executing action a t at time t + 1, max Q(s t+1 , a t+1 ) is the maximum Q value of all possible actions corresponding to state s t+1 .

[0028] Step S45: Enter the initial state s1 of the mowing boat.

[0029] Step S46: Obtain the adjustment factor in the current episode according to formula (3) in step S43, and select an action a t under the current state s t according to the ε-greedy strategy and execute it;

[0030] After executing action a t , observe the new state s t+1 entered and the immediate reward r t+1 ;

[0031] Update the Q-value table using formula (4), and update the current state s t to the new state s t+1 .

[0032] Step S47: Determine whether the current episode reaches the target point. If so, enter step S48 to continue training; otherwise, return to step S46.

[0033] Step S48: Determine whether the number of iterations n exceeds the maximum number of iterations. If so, generate the final Q-value table and execute step S5; if not, return to step S45 to continue training for the next episode.

[0034] Furthermore, the specific steps of the said step S5 are:

[0035] Step S51: Import the rasterized actual environment map.

[0036] Step S52: Load the optimal Q-value table generated through pre-training in Step S4. This table contains the expected future rewards from each state to each possible action.

[0037] Step S53: Obtain the positions of the current return point and the grass unloading dock. Starting from the return point s, select the action a with the highest Q-value.

[0038] Step S54: Execute the action a, move to the next state s', and select the action a' with the highest Q-value. Repeat this step until reaching the position of the grass unloading dock. If a dynamic obstacle is encountered during the return journey, return to Step S46 to reselect the action and update the Q-value table to avoid the dynamic obstacle, and then continue to execute this step.

[0039] Step S55: Record the state and action at each step to generate a collision-free optimal path from the return point to the grass unloading dock.

[0040] Step S56: The mowing boat executes the return task according to the optimal path planned in Step S55. Since it is difficult to mow all the grass in the entire pond in one operation, after the mowing boat completes this grass unloading operation, it needs to return to the position where the previous operation was interrupted along the original route and continue with the mowing operation.

[0041] Step S57: Repeat Steps S53 to S56 until the mowing operation of the entire pond is completed.

[0042] The beneficial effects of the present invention are as follows:

[0043] The present invention proposes an efficient operation path planning method for an unmanned mowing boat in aquaculture. First, the actual map environment is converted into a grid map with obstacles, and the position of the optimal grass unloading dock is selected. To overcome the problems of slow convergence speed and the exploration-exploitation balance in the traditional Q-Learning algorithm, the improved algorithm introduces an eight-direction exploration mechanism, redesigns the reward function, and dynamically adjusts the exploration factor. After training until convergence, an optimal Q-value table is generated to plan an optimal path with the shortest path, fewer turns, and no collisions in real time. In addition, for the problem of dynamic obstacles that are difficult to solve by the traditional point-to-point algorithm during the return journey, the mowing boat reselects the action and updates the Q-value table with the help of the improved Q-Learning algorithm to ensure a smooth return. The efficient operation path planning method for an unmanned mowing boat in aquaculture provided by the present invention designs a multiple return strategy by improving the Q-Learning algorithm, and in real time plans a return path with a short path and few turns covering the entire pond mowing task, which can significantly reduce energy consumption and improve operation efficiency. Description of the Drawings

[0044] To more clearly illustrate the technical solution of the present invention, the following will briefly introduce the attached drawings required in the description. Obviously, the attached drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other attached drawings can also be obtained based on these attached drawings.

[0045] Figure 1 It is a flowchart of an efficient operation path planning method for an unmanned grass cutting boat in aquaculture.

[0046] Figure 2 It is a schematic diagram of the grass cutting boat hitting an obstacle during oblique movement due to the limitation of the working width.

[0047] Figure 3 It is a flowchart of the Q-Learning algorithm for performing the entire pond grass cutting task.

[0048] Figure 4 It is a comparison of the paths planned by the original Q-Learning algorithm and the improved Q-Learning algorithm under different obstacle ratios. Specific embodiments

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the attached drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0050] The present invention provides an efficient operation path planning method for an unmanned grass cutting boat in aquaculture, as Figure 1 shown, and the specific steps are as follows:

[0051] Step S1: The GPS / IMU integrated navigation system obtains the coordinate information of the pond boundary and obstacles, and completes the conversion of longitude and latitude coordinates to plane coordinates by means of the Gauss-Kruger projection formula.

[0052] Step S2: The grid method is used to model the pond environment, and the grid size is determined by the size of the grass cutting boat and the size of the operation area. For some boundary obstacles, after rasterization, they only occupy a small part of the grid area, and it is impossible to directly set them as obstacle areas. Whether the grid on the boundary line is an obstacle area is determined by checking whether the center point of the grid is located in the obstacle area.

[0053] Step S3: Select the location of the grass unloading dock. Considering that the mowing boat operation needs to cover the entire area of the pond, in order to improve the operation efficiency and energy consumption of the mowing boat, the middle position of the long side of the rectangular pond is selected as the grass unloading dock because these positions are more convenient for management operations compared to the vertices of the pond, facilitating the docking and operation of the mowing boat.

[0054] Step S4: Improve the direction, reward function, and exploration rate of the traditional Q-Learning algorithm, and train it to generate the optimal Q-value table. The steps are as follows:

[0055] Step S41: Change the four-direction exploration environment map of the traditional Q-Learning algorithm to an eight-direction exploration, increasing the number of available step samples. Since the mowing boat is relatively large, in order to solve problems such as easy missed mowing during right-angle turns and excessive energy consumption due to too long path lengths during operation, in this embodiment, the original four directions are extended to eight directions for exploration, adding four exploration directions: upper left, lower left, upper right, and lower right. In this embodiment, the path length of translating one unit is set to 1. If translating one unit in the diagonal direction, the path unit length is

[0056] Step S42: Redesign the reward function R. The reward mechanism of the traditional Q-learning algorithm only includes rewards for normal walking, encountering obstacles, and reaching the target point. In this embodiment, the reward function is improved as shown in formula (1). A corresponding reward value is set for each step during the movement of the unmanned mowing boat, reducing the influence of misleading rewards, providing more detailed and specific feedback information to guide the unmanned mowing boat to learn, and enabling it to converge to the optimal strategy faster. The improved rewards are as follows:

[0057] (1) Give a penalty of -4 for turning. In order to reduce the energy consumption of the mowing boat operation, a turning penalty is given to reduce the number of turns in the path.

[0058] (2) Give a penalty of -1 for staying in place when hitting the map boundary. The mowing boat will continuously hit the map boundary during the training process. Therefore, in order to accelerate the convergence speed, a penalty of -1 is given after hitting the obstacle boundary.

[0059] (3) Give a penalty of -3.5 for hitting an obstacle diagonally. As Figure 2 shown, due to the limitation of the mowing width of the mowing boat, diagonal movement is prone to hitting obstacles. Therefore, a penalty is given to avoid this situation.

[0060]

[0061] Step S43: To balance the relationship between exploration and exploitation, this embodiment designs a dynamic adjustment function ξ(n) with the number of iterations n as the independent variable to dynamically adjust the greedy factor. The action selection of the traditional Q-learning algorithm uses the ε-greedy strategy, and the exploration rate ε ∈ (0, 1) is generally a fixed value, as shown in formula (2). A represents the action selected according to the ε-greedy strategy, a * represents the action that selects the maximum Q value, a r represents randomly selecting an action, and P represents the probability of selecting action a * or a r .

[0062]

[0063] For the problem in the traditional Q-learning algorithm that when the value of ε is too large, the mowing boat has been exploring the environment all the time, resulting in a slow convergence speed; when the value of ε is too small, the mowing boat will use the existing knowledge too much to select the action with the largest Q value during the learning process, and may fall into a local optimum due to insufficient learning of the unknown environment, the present invention proposes a dynamic adjustment function ξ(n). The value of the dynamic adjustment factor ξ gradually decreases as the number of iterations n increases. After the algorithm iterates 6000 times, the value of ε remains unchanged at 0.1. The optimized dynamic adjustment function enables the algorithm to fully explore the environment in the early stage. As the number of iterations increases, the exploration rate gradually decreases, and the mowing boat is more inclined to use the known environmental information, which can accelerate the convergence speed of the algorithm and avoid the algorithm falling into a local optimum at the same time.

[0064] Among them, the dynamic adjustment function ξ(n) is defined as follows:

[0065]

[0066] Step S44: Based on the above improvements, first create a Q-table to store each state-action pair, and initialize the Q-table to 0. Initialize the learning rate (α), discount factor (γ), and exploration rate (ε). Among them, the Q-value function update rule is as follows:

[0067] Q(s t , a t ) ← Q(s t , a t ) + α[r t+1 + γmax Q(s t+1 , a t+1 ) - Q(s t , a t )] (4)

[0068] In the formula, γ ∈ (0, 1) is the discount rate, α ∈ (0, 1) is the learning rate, r t+1 is the immediate reward, and st is the state of the mowing boat at time t, a t is the action selected by the mowing boat in the s t state, s t+1 is the state after executing action a at time t + 1, maxQ(s t , a t+1 ) is the maximum Q value of all possible actions corresponding to state s t+1 . t+1 The maximum Q value of all possible actions corresponding to state s

[0069] Step S45: Enter the initial state s1 of the mowing boat.

[0070] Step S46: Obtain the adjustment factor in the current episode according to formula (3) in Step S43, and select an action a according to the ε-greedy strategy in the current state s t and execute it. t After executing action a

[0071] observe the new state s t entered and the immediate reward r t+1 and t+1 .

[0072] Update the Q-value table using formula (4), and update the current state s t to the new state s t+1 .

[0073] Step S47: Determine whether the current episode reaches the target point. If so, enter Step S48 to continue training; otherwise, return to Step S46.

[0074] Step S48: Determine whether the number of iterations n exceeds the maximum number of iterations (8000 episodes). If so, generate the final Q-value table and execute Step S5; if not, return to Step S45 to continue training for the next episode.

[0075] Step S5: Use the optimal Q-value table generated by the improved Q-Learning algorithm to plan the collision-free optimal operation path from the return point to the unloading dock in real time. If a dynamic obstacle is encountered during the return journey, reselect the action and update the Q-value table to avoid this obstacle. After the mowing boat finishes unloading the grass, drive back along the original path to the interruption point of the previous operation and continue to execute the mowing operation until the mowing task of the entire pond is completed;

[0076] Step S51: Import the rasterized actual environment map.

[0077] Step S52: Load the optimal Q-value table pre-trained in Step S4. This table contains the expected future rewards from each state to each possible action.

[0078] Step S53: Obtain the positions of the current return point and the grass unloading dock. Starting from the return point s, select the action a with the highest Q value.

[0079] Step S54: Execute the action a, move to the next state s', and select the action a' with the highest Q value. Repeat this step until reaching the position of the grass unloading dock. If a dynamic obstacle is encountered during the return journey, execute Step S46 to reselect the action and update the Q value table to avoid the dynamic obstacle, and then continue to execute this step.

[0080] Step S55: Record the state and action of each step to generate a collision-free optimal path from the return point to the grass unloading dock.

[0081] Step S56: The mowing boat executes the return task according to the optimal path planned in Step S55. Since it is difficult to mow all the grass in the pond in one operation, after the mowing boat finishes this grass unloading operation, it needs to return to the position where the previous operation was interrupted along the original route and continue with the mowing operation.

[0082] Step S57: Repeat Steps S53 to S56 until the mowing operation of the entire pond is completed.

[0083] Figure 3 The flowchart of the Q-Learning algorithm for executing the mowing task of the entire pond is given.

[0084] Figure 4 The difference in the path planning effect between the traditional Q-learning algorithm and the Q-learning algorithm in the present invention is shown.

[0085] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for planning an efficient operation path for an unmanned mowing boat for aquaculture, characterized in that: The following steps are involved: Step S1: Scan the overall environment of the pond, collect the longitude and latitude coordinates of the pond vertices and obstacles, and convert them into plane coordinates; Step S2: Model the environment map using the grid method and establish an obstacle matrix; Step S3: Select the location of the hay unloading dock; Step S4: Improve the direction, reward function and exploration rate of the traditional Q-Learning algorithm, train it and generate an optimal Q value table; Step S5: Using the optimal Q-value table generated by the improved Q-Learning algorithm, the optimal operation path from the return point to the unloading dock is planned in real time; if a dynamic obstacle is encountered during the return process, the action is reselected and the Q-value table is updated to avoid the obstacle; After the unmanned mowing boat has finished unloading the grass, it will follow the original route to the place where the last operation was interrupted and continue the mowing operation until the mowing task of the entire pond is completed.

2. The method for planning an efficient operation path of an unmanned aquaculture mowing boat according to claim 1, characterized in that: The step S2 adopts the grid method to model the pond environment. The grid size is determined by the size of the mowing boat and the size of the working area. Some boundary obstacles only occupy a small part of the grid area after rasterization and cannot be directly set as an obstacle area. Whether the grid on the boundary line is an obstacle area is determined by checking whether the center point of the grid is in the obstacle area.

3. The method for efficient operation path planning of an unmanned aquaculture mowing boat according to claim 1 is characterized in that: The step S4 specifically comprises the following steps: Step S41: Change the four-direction exploration of the environment map of the traditional Q-Learning algorithm to eight-direction exploration, so that the number of step length samples available for selection is increased. Due to the large size of the mowing boat, in order to solve the problems of missing mowing when turning at right angles during operation and more energy consumption due to the long path length, the present invention expands the original four directions to eight directions for exploration, and adds four new exploration directions: upper left, lower left, upper right, and lower right. In the present invention, the unit path length of one grid translation is set to d. If one grid is translated in the diagonal direction, the unit path length is Step S42: Redesign the reward function R. The traditional Q-learning algorithm reward mechanism only includes rewards for normal walking, encountering obstacles, and reaching the target point. The present invention improves the reward function. As shown in formula (1), a corresponding reward value is set for each step in the movement of the unmanned lawn mowing boat, reducing the impact of misleading rewards and providing more detailed and specific feedback information to guide the unmanned lawn mowing boat to learn, so that it converges to the optimal strategy faster. The improved rewards include the following aspects: (1) Giving a turning penalty r4, in order to reduce the energy consumption of the mowing boat, giving a turning penalty reduces the number of turns in the path; (2) Give a penalty of r5 for standing still after hitting the map boundary. The mowing boat will constantly hit the map boundary during training. Therefore, in order to speed up the convergence speed, a penalty of r5 is given after hitting the obstacle boundary. (3) Give a penalty of r6 for hitting obstacles in an oblique direction. Due to the limited width of the mowing boat, it is easy to hit obstacles in an oblique direction, so a penalty is given to avoid such a situation. Step S43: In order to balance the relationship between exploration and utilization, the present invention designs a dynamic adjustment function ξ(n) with the number of iterations n as the independent variable, dynamically adjusts the greed factor, and the action of the traditional Q-learning algorithm uses the ε-greedy strategy. The exploration rate ε∈(0,1) is generally a constant, as shown in formula (2), A represents the action selected according to the ε-greedy strategy, a * represents the action of selecting the maximum Q value, a r represents random selection of actions, and P represents the selection of action a * or a r probability; In the traditional Q-learning algorithm, when the value of ε is too large, the mowing boat is always in the process of exploring the environment, which will lead to slow convergence; when the value of ε is too small, the mowing boat will make excessive use of existing knowledge to select the action with the largest Q value during the learning process, and may fall into the problem of local optimality due to insufficient learning of the unknown environment. The present invention proposes a dynamic adjustment function ξ(n), and the value of the dynamic adjustment factor ξ gradually decreases with the increase of the number of iterations n. When the algorithm iterates to k times, the value of ε remains unchanged; the optimized dynamic adjustment function enables the algorithm to fully explore the environment in the early stage. As the number of iterations increases, the exploration rate gradually decreases, and the mowing boat is more inclined to use known environmental information, which can speed up the convergence of the algorithm and avoid the algorithm from falling into the local optimality; The dynamic adjustment function ξ(n) is defined as follows, i is an adjustable parameter, and N is the maximum number of iterations; Step S44: Based on the above improvements, first create a Q table to store each state-action pair, and initialize the Q table to 0, and initialize the learning rate (α), discount factor (γ), and exploration rate (ε). The Q value function update rule is as follows: Q(s t ,a t )←Q(s t ,a t )+α[r t+1 +γmaxQ(s t+1 ,a t+1 )-Q(s t ,a t )](4) In the formula, γ∈(0,1) is the discount rate, α∈(0,1) is the learning rate, r t+1 For immediate rewards, t is the state of the mowing boat at time t, a t For mowing boats in s t The action selected in the state, s t+1 To perform action a at time t+1 t The state after maxQ(s t+1 ,a t+1 ) is state s t+1 The corresponding maximum Q value of all possible actions; Step S45: entering the initial state s1 of the mowing boat; Step S46: Obtain the adjustment factor for the current episode according to formula (3) in step S43. t Next, we select an action a according to the ε-greedy strategy t implement; Execute action a t After that, observe the new state s t+1 and instant rewards t+1 ; Use formula (4) to update the Q value table and set the current state s t Update to new status t+1 ; Step S47: Determine whether the current episode (round) has reached the target point. If so, proceed to step S48 to continue training; otherwise, return to step S46; Step S48: Determine whether the number of iterations n exceeds the maximum number of iterations. If so, generate a final Q value table and execute step S5; if not, return to step S45 and continue training for the next episode.

4. The method for planning an efficient operation path of an unmanned aquaculture mowing boat according to claim 1, characterized in that: The step S5 specifically comprises the following steps: Step S51: importing the rasterized actual environment map; Step S52: Load the optimal Q value table generated by pre-training in step S4, which contains the expected future rewards from each state to each possible action; Step S53: Obtain the current return point and the location of the unloading dock, start from the return point s, and select the action a with the highest Q value; Step S54: Execute action a, move to the next state s', and select action a' with the highest Q value. Repeat this step until the unloading dock is reached; if a dynamic obstacle is encountered during the return journey, return to step S46 to reselect actions and update the Q value table to avoid dynamic obstacles, and then continue to execute this step; Step S55: Record the status and action of each step, and generate a collision-free optimal path from the return point to the hay unloading dock; Step S56: the mowing boat performs the return task according to the optimal path planned in step S55; since it is difficult to mow the entire pond in one operation, after the mowing boat completes the unloading operation, it needs to return to the location where the last operation was interrupted according to the original route and continue the mowing operation; Step S57: Repeat steps S53 to S56 until the mowing operation of the entire pond is completed.