An AUV cluster underwater communication detection data collection system and method

By using a centralized UCDN network architecture and reinforcement learning algorithms, combined with integrated communication and detection signals and reflected echo signals, the problems of unsafe path planning and high energy consumption in underwater data collection by AUV clusters have been solved, achieving efficient, low-power, and safe data collection.

CN117459156BActive Publication Date: 2026-08-04SHANGHAI INST OF MICROSYSTEM & INFORMATION TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI INST OF MICROSYSTEM & INFORMATION TECH CHINESE ACAD OF SCI
Filing Date
2023-09-20
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing underwater data collection technologies for AUV swarms, path planning is unsafe and energy consumption is high. Current research has not effectively solved the problems of AUV collisions with obstacles and energy consumption.

Method used

The system adopts a centralized UCDN network architecture, combining integrated communication and detection signals with reinforcement learning algorithms. It plans the data collection path of the AUV cluster through Q-learning and SAC algorithms, and uses reflected echo signals to detect obstacles and optimize the path, thus realizing the integration of data communication and environmental detection.

Benefits of technology

It enables efficient, low-power, and safe underwater data collection by AUV clusters, improves the network's data collection efficiency and security, avoids obstacle collisions, and optimizes energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117459156B_ABST
    Figure CN117459156B_ABST
Patent Text Reader

Abstract

The application relates to an AUV cluster underwater communication detection data collection system, which comprises a data center, at least one AUV cluster and a plurality of sensor nodes fixed on the water bottom, the data center allocates a data collection area for the AUV cluster, the AUV cluster comprises a control node and a plurality of AUVs, the AUVs sail according to a data collection path obtained from the control node, communicate with the sensor nodes to obtain sensing data through communication detection integrated signals along the way, receive reflection echo signals of the communication detection integrated signals in the surrounding environment, analyze the environment information, and transmit the sensing data, the environment information and the AUV state information of the AUVs to the control node, and the control node plans the data collection path according to the related information of the sensor nodes in the data collection area, the environment information and the AUV state information. The application can efficiently, lowly and safely collect underwater data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underwater communication and detection technology, and in particular to an AUV cluster underwater communication and detection data collection system and method. Background Technology

[0002] Underwater acoustic communication networks have become the main method for underwater data collection, and using AUV clusters to collect data collaboratively is currently a relatively reliable, efficient and widely used UACN data collection solution.

[0003] Current research on collaborative data collection among multiple autonomous underwater vehicles (AUVs) mainly focuses on two aspects: AUV path planning and data transmission between nodes. Regarding path planning, most existing studies only consider fixed path planning based on sensor node positions, which is unreliable and unsafe. To avoid collisions with obstacles, AUVs must explore the unknown environment and make online decisions, which also consumes significant energy. Existing research on data transmission and communication largely focuses on reducing the energy consumption of sensor nodes, neglecting the energy consumption of the AUVs themselves. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an AUV cluster underwater communication detection data collection system and method that can collect underwater data efficiently, with low power consumption and in a safe manner.

[0005] The technical solution adopted by this invention to solve its technical problem is as follows: An AUV cluster underwater communication and detection data collection system is provided, comprising a data center, at least one AUV cluster, and multiple sensor nodes fixed on the seabed. The data center allocates a data collection area for the AUV cluster. The AUV cluster includes a control node and multiple AUVs communicating with the control node. Each AUV obtains a data collection path from the control node, navigates along the data collection path, sends integrated communication and detection signals along the way, receives sensing data uploaded by the sensor nodes and reflected echo signals from obstacles in the surrounding environment, analyzes the reflected echo signals to obtain environmental information, and transmits the sensing data, environmental information, and its own AUV status information to the control node. The control node plans the data collection path based on the relevant information of the sensor nodes within the data collection area, the environmental information, and the AUV status information.

[0006] Furthermore, the control node calculates the obstacle location information using a bistatic or multistatic cooperative detection method based on the environmental information contained in the reflected echoes obtained by different AUVs.

[0007] Furthermore, based on the relevant information of the sensor nodes in the data collection area, the control node uses the Q-learning algorithm to obtain the traversal order of the AUV to the sensor nodes. Based on the environmental information and the AUV state information, the control node uses the SAC algorithm to obtain the optimal path between two adjacent nodes in the traversal order. The optimal path is the path that minimizes the AUV's travel distance and avoids obstacles along the way.

[0008] Furthermore, the data center communicates with the control node via underwater acoustic signals, and the sensor node uploads the sensed data via underwater acoustic signals.

[0009] Furthermore, the control node periodically broadcasts data packets, determines the responsible AUV cluster network based on the AUV's response, and uploads the relevant information to the data center. The data center allocates data collection areas based on the received information from the AUV cluster network and sends the relevant information of the sensor nodes within the data collection areas to the control node.

[0010] Furthermore, a time-division multiple access (TDMA) packet transmission strategy is employed to coordinate the communication between the control node, sensor node, and AUV, as well as the AUV's reception of the reflected echo signal.

[0011] The technical solution adopted by this invention to solve its technical problem is: to provide a method for collecting underwater communication and detection data of an AUV cluster, characterized by including the following steps:

[0012] S1. The data center allocates a data collection area for the AUV cluster;

[0013] S2. The control node uses the Q-learning algorithm to obtain the traversal order of each AUV to the sensor nodes based on the relevant information of the sensor nodes in the data collection area, and plans the data collection path.

[0014] S3. The AUV travels along the data collection path, sends communication and detection integrated signals along the way, receives the sensing data uploaded by the sensor node and the reflected echo signals of the communication and detection integrated signals from obstacles in the surrounding environment, analyzes the reflected echo signals to obtain environmental information, and transmits the sensing data, environmental information and its own AUV status information to the control node.

[0015] S4. The control node calculates the obstacle location information based on the environmental information, uses the SAC algorithm to obtain the optimal path for the AUV from its current location to the target node, and updates the data collection path.

[0016] S5. Repeat steps S3-S4 until all sensor data within the data collection area has been collected.

[0017] Furthermore, the step of obtaining environmental information based on the reflected echo signal analysis includes:

[0018] The reflected echo signal is demodulated and the packet header information is extracted;

[0019] Based on the packet header information, the signal transmission distance is calculated, which includes the transmission distance of the integrated communication and detection signal transmitted by the local unit and the transmission distance of the integrated communication and detection signal transmitted by the friendly unit.

[0020] The environmental information includes the signal transmission distance.

[0021] Furthermore, the control node calculates the obstacle location information based on the environmental information, including:

[0022] Analyze the environmental information uploaded by the first AUV to obtain the straight-line distance d from the first AUV to the first obstacle. AS ;

[0023] Analyzing the environmental information uploaded by the second AUV, the propagation distance d of the integrated communication and detection signal emitted by the first AUV to the second AUV is obtained. ASB Thus, the straight-line distance d from the second AUV to the first obstacle is obtained. SB =d ASB -d AS ;

[0024] Based on the straight-line distance d from the first AUV to the first obstacle AS The straight-line distance d from the second AUV to the first obstacle SB and the straight-line distance d between the first AUV and the second AUV AB The location information of the obstacle is calculated.

[0025] Furthermore, obtaining the traversal order of each AUV to the sensor nodes using the Q-learning algorithm includes:

[0026] With the goal of minimizing the sum of straight-line distances between nodes when the AUV sequentially collects node data, the current node is the state, the target node is the action, and the negative value of the distance between the current node and the target node is the action reward, a Q-value table is constructed.

[0027] According to a preset probability, a non-collected sensor node is randomly selected as the target node; otherwise, the non-collected sensor node that gives the highest Q value is selected as the target node, the action is executed, and the Q value table is updated.

[0028] Repeat the previous step until there are zero uncollected sensor nodes, thus obtaining the traversal order.

[0029] Furthermore, obtaining the optimal path between two adjacent nodes in the traversal order using the SAC algorithm includes:

[0030] A reinforcement learning framework is constructed, which includes a state space, an action space, and a reward function. The state space includes the AUV state information, the target node position information, and obstacle-related state information. The action space includes the AUV's velocity in the x-axis and y-axis directions. The reward function includes rewards for avoiding obstacles and approaching the target node, as well as penalties for collision, disconnection, and timeout.

[0031] Based on the aforementioned reinforcement learning framework, the optimal path is obtained using the SAC algorithm. The optimal path is the path that minimizes the AUV's travel distance and avoids obstacles along the way.

[0032] Beneficial effects

[0033] By adopting the above-mentioned technical solution, the present invention has the following advantages and positive effects compared with the prior art:

[0034] (1) This invention proposes an underwater communication and detection network (UCDN) architecture based on centralized control. The controller can obtain detection information and collection status from multiple AUVs, comprehensively analyze the obtained information, and make decisions on AUV data collection and path planning.

[0035] (2) This invention proposes a data transmission strategy and a dual-base / multi-base collaborative detection scheme based on integrated communication and detection. While communicating with sensor nodes to acquire data, the integrated communication and detection signal determines the relevant information of obstacles by receiving reflected echo signals from obstacles in the surrounding environment. This allows the network to use the same acoustic signal to perform data communication and environmental detection simultaneously.

[0036] (3) This invention proposes a multi-AUV collaborative data collection algorithm based on reinforcement learning. The collaborative data collection problem of AUV clusters is modeled as a combined optimization problem including node traversal problem and obstacle avoidance problem, so that AUVs can avoid obstacles while collecting data. The path planning problem is solved by combining Q-learning algorithm and SAC algorithm to minimize the path length of AUV cluster data collection. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the network structure according to the first embodiment of the present invention;

[0038] Figure 2This is a schematic diagram illustrating the use of UCD signal echoes to calculate the location of obstacles in the first and second embodiments of the present invention. Detailed Implementation

[0039] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.

[0040] The first embodiment of the present invention relates to an AUV cluster underwater communication detection data collection system, which adopts a centralized controlled UCDN network architecture, such as... Figure 1 As shown, the network consists of four parts: a surface data center, underwater control nodes, AUV clusters, and sensors deployed on the seabed. The data center stores the location information of all fixed sensor nodes in the entire network, assigns a responsible area to each controller and its AUV cluster, and collects relevant information from all sensor nodes and control nodes to determine the network topology and overall status. The underwater control nodes, acting as the control center for each AUV cluster, are responsible for acquiring data collection and environmental detection information from each AUV within the cluster. The control nodes also store the location information of the assigned sensor nodes within their cluster, arranging the order in which AUVs collect data from each sensor node. Furthermore, all AUV decisions are made by the control nodes, which plan the AUVs' paths to collect data from sensor nodes and avoid obstacles. The AUVs are equipped with an integrated communication and detection module, using integrated underwater communication and detection (UCD) signals to communicate between nodes and detect obstacles in the surrounding environment. Numerous AUVs are divided into different clusters based on different control nodes, collecting data from sensor nodes in different areas. The sensor nodes are randomly distributed at different locations on the seabed. These sensor nodes are equipped with various sensing devices, such as salinity sensors and temperature monitoring sensors. The system uses a time-division multiple access-based packet transmission strategy to coordinate data packet transmission and echo signal reception between sensor nodes, AUVs, control nodes, and obstacles. Furthermore, an active bistatic / multistatic detection scheme is designed to detect the distance to obstacles by analyzing the echo signal of the UCD signal, thereby determining the location of the obstacles.

[0041] The communication scheme consists of three steps: First, the control node periodically broadcasts data packets. Each AUV responds to the broadcast packets, and the controller determines the AUV cluster it is responsible for based on the responses, obtains a view of the entire network, and uploads relevant information to the data center. The data center then distributes sensor node information for its assigned area. Based on this information, the controller plans a path for each AUV, directing it towards sensor nodes to collect data and avoid collisions. During navigation, each AUV periodically sends UCD signals to search for sensor nodes and detect the environment. Upon receiving the UCD signals, the sensor nodes upload data packets to the AUVs. Within one work cycle, the AUVs and sensor nodes repeat the above steps until the next work cycle.

[0042] The AUV sends search data packets via UCD signals, and sensor nodes upload data packets to the AUV via underwater acoustic communication signals. In addition to searching for nearby sensor nodes, the search data packets sent by the AUV also have environmental detection capabilities. The controller communicates with the AUV via underwater acoustic communication signals, including downlink data packets sent by the controller to the AUV and uplink data packets sent by the AUV to the controller.

[0043] like Figure 2 As shown, the UCD search packets emitted by an AUV may be reflected by obstacles. Other AUVs will receive these echo packets and demodulate the packet header to obtain the source node ID and transmission time. Based on this information, the AUV can determine the distance the packet has traveled and send it to the controller. The controller can calculate the distance and angle of obstacles associated with each AUV, thereby determining the location of the obstacles.

[0044] The calculation formula is as follows:

[0045]

[0046] In the above formula, α, β, θ and γ are as follows: Figure 2 As shown, d AB This is the distance between AUVA and AUVB, monitored by the controller. AS and d SB These are the distances between AUV A, AUV B, and S, respectively, which can be calculated from AUV A and AUV B based on the echo signals. The calculation formula is as follows:

[0047]

[0048] In the above formula, and These are the times when AUVA and AUVB receive echo data packets, respectively. It is the transmission time of the UCD data packet, v s It is the speed of sound underwater.

[0049] In the path planning problem, the combinatorial optimization problem of minimizing the path length for AUV cluster data collection is decomposed into two sub-problems. The first sub-problem is the node traversal problem, which determines the traversal order of sensor nodes, transforming it into a Multiple Traveling Salesman Problem (MTSP), and solving it using the Q-learning algorithm. The second sub-problem is the online path planning and obstacle avoidance problem, which uses a softactor-critic (SAC) algorithm to make continuous path planning decisions to reach the target node and avoid obstacles.

[0050] The second embodiment of the present invention relates to a method for collecting underwater communication detection data of an AUV cluster, comprising the following steps:

[0051] The control node periodically broadcasts data packets. Each AUV responds to the broadcast packets, and the controller determines the AUV cluster it is responsible for based on the responses, obtains a view of the entire network, and uploads relevant information to the data center. The data center then distributes sensor node information for its assigned area. Based on this information, the controller plans a path for each AUV, directing it towards sensor nodes to collect data and avoid collisions. During navigation, each AUV periodically sends a UCD signal to search for sensor nodes and probe the environment. Upon receiving the UCD signal, the sensor node uploads the data packet to the AUV. Within one work cycle, the AUVs and sensor nodes repeat the above steps until the next work cycle.

[0052] The AUV sends search data packets via UCD signals, and sensor nodes upload data packets to the AUV via underwater acoustic communication signals. In addition to searching for nearby sensor nodes, the search data packets sent by the AUV also have environmental detection capabilities. The controller communicates with the AUV via underwater acoustic communication signals, including downlink data packets sent by the controller to the AUV and uplink data packets sent by the AUV to the controller.

[0053] like Figure 2 As shown, the UCD search packets emitted by an AUV may be reflected by obstacles. Other AUVs will receive these echo packets and demodulate the packet header to obtain the source node ID and transmission time. Based on this information, the AUV can determine the distance the packet has traveled and send it to the controller. The controller can calculate the distance and angle of obstacles associated with each AUV, thereby determining the location of the obstacles.

[0054] The calculation formula is as follows:

[0055]

[0056] In the above formula, α, β, θ and γ are as follows: Figure 2 As shown. d AB This is the distance between AUVA and AUVB, monitored by the controller. AS and d SB These are the distances between AUVA, AUV B, and S, respectively, which can be calculated from AUVA and AUV B based on the echo signals. The calculation formula is as follows:

[0057]

[0058] In the above formula, and These are the times when AUVA and AUVB receive echo data packets, respectively. It is the transmission time of the UCD data packet, v s It is the speed of sound underwater.

[0059] In the path planning problem, the combinatorial optimization problem of minimizing the path length for AUV cluster data collection is broken down into two sub-problems. The first sub-problem is the node traversal problem, which determines the traversal order of sensor nodes, transforming it into a multiple traveling salesman problem, and solving it using the Q-learning algorithm. The second sub-problem is the online path planning and obstacle avoidance problem, which uses the SAC algorithm to make continuous path planning decisions to reach the target node and avoid obstacles.

[0060] The first sub-problem is determining the traversal order of sensor nodes, which can be modeled as a multiple traveling salesman problem. The objective of this sub-problem is to minimize the sum of the straight-line distances between all nodes traversed by the AUV when collecting data from all nodes sequentially. To solve this optimization problem, we propose a node traversal order algorithm based on Q-learning. Assuming there are K sensor nodes, we first construct a K×K Q-value table. The (s, a)-th element in the table represents the Q-value for selecting sensor node a as the next traversal node when starting from node s. The state refers to the node the AUV is currently in, and the action corresponds to selecting the next node. After executing the action, the AUV's state changes to the selected node. The reward for the action is the negative of the distance between the starting node and the target node. The larger the total reward, the shorter the distance the AUV traverses. The Q-value table is updated according to the Bellman equation. The decision to select the next node is based on a random number. If the random number is less than a threshold, the AUV will randomly select a previously uncollected target node. Otherwise, the selected next target node will be the maximum Q-value from the current node to the uncollected node; the threshold decreases as the training epochs increase.

[0061] The second sub-problem is planning the shortest path for an AUV from the starting node to the target node and detecting and avoiding obstacles during the journey. To solve this optimization problem, we propose a path planning algorithm based on SAC, whose state space, action space, reward function, and training process are as follows:

[0062] State Space: In addition to collecting information from each sensor node, the AUV also needs to automatically avoid obstacles. The system's state space mainly consists of three parts: the AUV's own state, the target node's position, and obstacle-related states. The AUV's own state includes its current position, whether it is within the controller's communication range, and its operating time. Obstacle-related states include whether obstacles exist around the AUV, whether the AUV has collided with an obstacle, the distance between the AUV and the obstacle, and the relative motion angle.

[0063] The motion space is the velocity of the AUV in the x and y directions.

[0064] Reward function: Rewards consist of those for avoiding obstacles and those for approaching the target node. In addition, penalties include collision penalties, disconnection penalties, and timeout penalties.

[0065] Training Process: The system employs the SAC algorithm, where each AUV corresponds to four neural networks: an action network, a target action network, a comment network, and a target comment network. First, the parameters of the four neural networks are initialized. In each step of each round, the action network determines the action to be taken in the current state. After the AUV executes the action, it updates its own state and receives a reward from the environment. The state, action, reward, and next state are stored as a set of data in an experience buffer. After each training round, several sets of data are retrieved from the experience buffer, and backpropagation is performed according to the loss function of each neural network to update the parameters of each neural network.

[0066] In summary, this invention combines underwater communication and detection technology with an underwater acoustic communication network, proposing the concept of an underwater communication and detection network. It utilizes reinforcement learning technology to plan the movement path of an AUV swarm online, enabling the AUV swarm to detect the surrounding marine environment and avoid obstacles while collecting data. This data collection scheme based on an AUV swarm's integrated underwater acoustic communication and detection network overcomes the low energy efficiency of traditional AUV-based underwater wireless sensor networks, improving both the data collection efficiency of the sensor network and the safety of the AUV swarm during data collection.

Claims

1. An AUV cluster underwater communication and detection data collection system, characterized in that, It includes a data center, at least one AUV cluster, and multiple sensor nodes fixed on the seabed; The data center allocates a data collection area for the AUV cluster, which includes a control node and multiple AUVs that communicate with the control node. The control node obtains the traversal order of each AUV to the sensor nodes based on the relevant information of the sensor nodes in the data collection area, and plans the data collection path. The AUV obtains the data collection path from the control node, navigates along the data collection path, sends communication and detection integrated signals along the way, receives the sensing data uploaded by the sensor node and the reflected echo signals of the communication and detection integrated signals from obstacles in the surrounding environment, analyzes the reflected echo signals to obtain environmental information, and transmits the sensing data, environmental information and its own AUV status information to the control node. The control node calculates the obstacle location information based on the environmental information, obtains the optimal path between two adjacent nodes in the traversal order, the optimal path is the path that minimizes the AUV's travel distance and avoids obstacles along the way, and updates the data collection path.

2. The system according to claim 1, characterized in that, The control node calculates the obstacle location information using a bistatic or multistatic cooperative detection method based on the environmental information contained in the reflected echoes obtained by different AUVs.

3. The system according to claim 1, characterized in that, The traversal order of each AUV to the sensor nodes is obtained based on the Q-learning algorithm, and the optimal path between two adjacent nodes in the traversal order is obtained based on the SAC algorithm.

4. The system according to claim 1, characterized in that, The data center communicates with the control node via underwater acoustic signals, and the sensor node uploads the sensed data via underwater acoustic signals.

5. The system according to claim 1, characterized in that, A time-division multiple access (TDMA) packet transmission strategy is used to coordinate the communication between the control node, sensor node, and AUV, as well as the AUV's reception of the reflected echo signal.

6. A method for collecting underwater communication and detection data in an AUV cluster, characterized in that, Applied to the system as described in any one of claims 1-5, the method includes the following steps: S1. The data center allocates a data collection area for the AUV cluster; S2. The control node uses the Q-learning algorithm to obtain the traversal order of each AUV to the sensor nodes based on the relevant information of the sensor nodes in the data collection area, and plans the data collection path. S3. The AUV travels along the data collection path, sends communication and detection integrated signals along the way, receives the sensing data uploaded by the sensor node and the reflected echo signals of the communication and detection integrated signals from obstacles in the surrounding environment, analyzes the reflected echo signals to obtain environmental information, and transmits the sensing data, environmental information and its own AUV status information to the control node. S4. The control node calculates the obstacle location information based on the environmental information, uses the SAC algorithm to obtain the optimal path between two adjacent nodes in the traversal order, the optimal path is the path that minimizes the AUV's travel distance and avoids obstacles along the way, and updates the data collection path. S5. Repeat steps S3-S4 until all sensor data within the data collection area has been collected.

7. The method according to claim 6, characterized in that, The process of obtaining environmental information based on the reflected echo signal analysis includes: The reflected echo signal is demodulated and the packet header information is extracted; Based on the packet header information, the signal transmission time is calculated, which includes the transmission time of the integrated communication and detection signal transmitted by the local machine and the transmission time of the integrated communication and detection signal transmitted by the friendly machine. The environmental information includes the signal transmission time.

8. The method according to claim 6, characterized in that, The environmental information obtained based on the analysis of the reflected echo signal includes: The reflected echo signal is demodulated and the packet header information is extracted; Based on the packet header information, the signal transmission distance is calculated, which includes the transmission distance of the integrated communication and detection signal transmitted by the local unit and the transmission distance of the integrated communication and detection signal transmitted by the friendly unit. The environmental information includes the signal transmission distance.

9. The method according to claim 7 or 8, characterized in that, The control node calculates the obstacle location information based on the environmental information, including: Analyze the environmental information uploaded by the first AUV to obtain the straight-line distance from the first AUV to the first obstacle. ; By analyzing the environmental information uploaded by the second AUV, the propagation distance of the integrated communication and detection signal emitted by the first AUV to the second AUV can be obtained. This allows us to obtain the straight-line distance from the second AUV to the first obstacle. ; Based on the straight-line distance from the first AUV to the first obstacle The straight-line distance from the second AUV to the first obstacle and the straight-line distance between the first AUV and the second AUV. The location information of the obstacle is calculated.

10. The method according to claim 6, characterized in that, The step of obtaining the traversal order of each AUV to the sensor nodes using the Q-learning algorithm includes: With the goal of minimizing the sum of straight-line distances between nodes when the AUV sequentially collects node data, the current node is the state, the target node is the action, and the negative value of the distance between the current node and the target node is the action reward, a Q-value table is constructed. According to a preset probability, a non-collected sensor node is randomly selected as the target node; otherwise, the non-collected sensor node that gives the highest Q value is selected as the target node, the action is executed, and the Q value table is updated. Repeat the previous step until there are zero uncollected sensor nodes, thus obtaining the traversal order.

11. The method according to claim 6, characterized in that, The step of using the SAC algorithm to obtain the optimal path between two adjacent nodes in the traversal order includes: A reinforcement learning framework is constructed, which includes a state space, an action space, and a reward function. The state space includes the AUV state information, the target node position information, and obstacle-related state information. The action space includes the AUV's velocity in the x-axis and y-axis directions. The reward function includes rewards for avoiding obstacles and approaching the target node, as well as penalties for collision, disconnection, and timeout. Based on the aforementioned reinforcement learning framework, the optimal path is obtained using the SAC algorithm. The optimal path is the path that minimizes the AUV's travel distance and avoids obstacles along the way.