Method for detecting ship hull surface coating of tethered underwater robot based on deep reinforcement learning
By combining deep reinforcement learning and sensor systems, a control framework for a cabled underwater robot was designed, which solved the problems of automation and full traversal in the detection of ship surface coatings in existing technologies, and achieved efficient and stable ship hull detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2026-03-27
AI Technical Summary
Existing methods for detecting surface coatings on hulls using tethered underwater robots rely on manual operation, resulting in inconsistent detection quality. Furthermore, existing reinforcement learning methods are not applicable to surface coating detection and cannot achieve full-coverage detection of hulls with different shapes.
The control framework of a cable-stayed underwater robot is designed using the flexible actor-critic algorithm in deep reinforcement learning. Combined with a sensor system, a full-traversal detection path is planned, and path tracking is achieved through a reinforcement learning model to resist ocean current interference and complete the detection of the hull surface coating.
It enables comprehensive inspection of ship surface coatings by tethered underwater robots, improving inspection efficiency and performance, adapting to different target ship hulls, reducing the risk of human intervention, and possessing the advantages of high efficiency, stability, and low risk in inspection.
Smart Images

Figure CN120631027B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of underwater robots, in particular to a tethered underwater robot hull surface coating detection method based on deep reinforcement learning. BACKGROUND
[0002] With the deepening development of world multipolarity and economic globalization, the shipping industry has gradually become an important foundation for economic and social development, and also an important bridge and link for international trade. Shipping activities are widespread, mainly covering cargo transportation, resource exploitation, and entertainment and tourism industries. Ships are affected by complex marine environments, and phenomena such as corrosion and cracking may occur on the hull surface. At the same time, marine biofouling will increase fuel consumption and generate frictional resistance, which will seriously affect the sailing efficiency and safety of the ship. The hull surface is usually coated with a coating to protect the ship from seawater corrosion and biofouling, thereby improving the service life and sailing safety of the ship. Therefore, domestic and foreign regulations require regular hull surface coating detection to inspect the condition and integrity of the hull, which is an important task to ensure the sustainability of the ship. Traditional hull surface coating detection is carried out in the dock or by divers inspecting the outside of the hull underwater, while the use of a tethered underwater robot can improve detection efficiency and safety factor, reduce operation and maintenance costs, and establish a database by collecting real-time high-definition video images to realize full life cycle management of the ship.
[0003] Currently, most tethered underwater robots are manually operated and controlled by operators through handles based on real-time images and data transmitted back, requiring operators to have a certain level of proficiency, and the quality of the detected data is often uneven.
[0004] Reinforcement learning is a machine learning algorithm that trains an agent to maximize cumulative rewards in a complex and uncertain environment. In this process, the agent learns the optimal strategy through continuous interaction with the environment. When the agent obtains a certain state in the environment, it will select and execute an action based on the current state, and the environment will output the next state and the reward or punishment brought by the action. Through this trial-and-error process, the agent gradually learns to adopt behavior patterns that maximize rewards. CN114995468A discloses an underwater robot intelligent control method based on Bayesian deep reinforcement learning, including the following steps: sensing underwater environment information according to the sensor system carried by the underwater robot; constructing a Bayesian deep reinforcement learning intelligent control model for the underwater robot; completing underwater robot intelligent control model learning according to interactive training; and deploying and applying the underwater robot intelligent control method. However, this method is mainly aimed at underwater structure maintenance of offshore wind power piles, focusing on autonomous navigation and obstacle avoidance of underwater robots near specific structures (such as wind power piles), and is not applicable to hull surface coating detection scenarios, and cannot achieve full traversal coating detection of different shaped hulls. SUMMARY
[0005] The purpose of the present application is to provide a tethered underwater robot hull surface coating detection method based on deep reinforcement learning, which adopts the Soft Actor-Critic (SAC) algorithm in the deep reinforcement learning method to design the control framework of the tethered underwater robot, establishes the robot simulation and training environment, designs the corresponding state, action space and reward function, etc., and gives the tethered underwater robot learning ability, combined with the sensor system carried, so that it can track the path according to the target path, and can resist the current disturbance in the environment, so as to efficiently and stably realize the hull full traversal detection of the tethered underwater robot.
[0006] The purpose of the present application can be realized by the following technical solutions:
[0007] A tethered underwater robot hull surface coating detection method based on deep reinforcement learning, comprising the following steps:
[0008] The sensor carried on the tethered underwater robot is used to perceive underwater environmental information and robot state information;
[0009] The hull full traversal detection path of the tethered underwater robot is planned;
[0010] The hull full traversal detection path is taken as the target path, and is input into the reinforcement learning model constructed based on the deep reinforcement learning method, combined with the real-time perceived underwater environmental information and robot state information, to obtain the control signal of the tethered underwater robot; wherein the simulation and training environment constructed based on the underwater environmental disturbance model and the motion control model is used to obtain the underwater environmental information and robot state information in the training process, so as to train the reinforcement learning model;
[0011] The tethered underwater robot is controlled based on the control signal to track the target path, and the hull surface coating detection is completed.
[0012] The sensor carried on the tethered underwater robot is used to perceive underwater environmental information and robot state information, specifically:
[0013] A turbidity sensor, ranging sonar, and UGPS underwater positioning system are installed on an existing tethered underwater robot. The turbidity sensor collects the turbidity of the underwater environment, and the distance between the tethered underwater robot and the target vessel is set based on the turbidity to ensure the clarity of the video image acquisition. The ranging sonar collects the distance between the tethered underwater robot and the obstacle to prevent collision with the target vessel during the detection process. The UGPS underwater positioning system obtains the real-time position information of the tethered underwater robot. Combined with the sensors on the tethered underwater robot body, the robot's attitude and velocity information s1=[x,y,z,ψ,u,v,w,r] is obtained, where (x,y,z) represents the three-dimensional position information of the tethered underwater robot, ψ represents the angle information of the tethered underwater robot, (u,v,w) represents the linear velocity information of the tethered underwater robot, and r represents the angular velocity information of the tethered underwater robot.
[0014] The specific plan for the full-coverage detection path of the tethered underwater robot is as follows: based on the information of the target hull, an improved Dijkstra algorithm is used to determine the full-coverage detection path of the target hull according to underwater stability and real-time shooting range. The detection path is a grating-type path. After the tethered underwater robot completes one lateral detection, it dives or moves according to the working mode to perform the next lateral detection. The working mode is distinguished according to whether the detection area is the hull or the bottom of the hull.
[0015] The underwater environmental disturbance model is represented as: τ d =[τ dx ,τ dy ,τ dz ] T , τ d τ represents the vector of disturbance force or torque caused by ocean currents in the underwater environment. dx ,τ dy ,τ dz Let τ represent the disturbance forces or moments caused by ocean currents in the underwater environment in the x, y, and z directions, respectively. By using random functions to enhance the model's resistance to external disturbances, we denot them as τ. dx =3rand(-1,1)sint,τ dy =2rand(-1,1)sin2t,τ dz = rand(-1,1)sin3t, where rand represents a random function.
[0016] The motion control model includes kinematic equations and dynamic equations, wherein the six-degree-of-freedom kinematic equations of the underwater robot are as follows: The dynamic equations used to describe the robot's position, velocity, and acceleration are as follows: for analyzing the force and motion state of a robot, wherein J(η) is a transformation matrix between an inertial coordinate system and a carrier coordinate system, η represents a displacement vector, v represents a velocity vector, τ represents a force and torque vector, τ d represents an interference force or torque vector caused by ocean current in underwater environment, v = [u, v, w, p, q, r] T , τ = [X, Y, Z, K, M, N] T , ξ, η, ζ are horizontal, vertical and vertical axes positions of the underwater robot in the inertial coordinate system, with the unit of m; θ, ψ are attitude angles of the underwater robot, representing roll angle, pitch angle and yaw angle respectively, with the unit of rad; u, v, w are displacement velocities in the carrier coordinate system, representing longitudinal, lateral and vertical velocities respectively, with the unit of m / s; p, q, r are angular velocities around x, y, z axes in the carrier coordinate system, with the unit of rad / s; X, Y, Z and K, M, N are thrust and torque along x, y, z axes in the carrier coordinate system, with the units of N and N·m respectively; M is an inertia matrix, C(v) is a Coriolis centripetal force matrix, D(v) is a fluid damping matrix, and g(η) is a restoring force vector.
[0017] The state space s of the reinforcement learning model t is defined as s t = {x, y, z, ψ, u, v, w, r, e d , e ψ , e z}, wherein (x, y, z) represents three-dimensional position information of the tethered underwater robot, ψ represents angle information of the tethered underwater robot, (u, v, w) represents linear velocity information of the tethered underwater robot, r represents angular velocity information of the tethered underwater robot, e d , e ψ , e z represents tracking error determined by the line-of-sight navigation method, e d represents lateral error of the underwater robot from the target path, e ψ represents directional error to be minimized, e z represents depth error; wherein the range of ψ is [-π, π], and when the value of ψ exceeds the range, the calculation of ψ = ((ψ + π) % (2π)) - π is used to map the value of ψ to the range of [-π, π]; the current state s t In training, {x, y, z, ψ, u, v, w, r} in the state space is obtained from the motion control model of the tethered underwater robot; in actual use, {x, y, z, ψ, u, v, w, r} in the state space is perceived by the sensors carried on the tethered underwater robot.
[0018] The action space a of the reinforcement learning modelt The channel control signal commands of the 8 thrusters of the tethered underwater robot 8 are defined as a continuous action space: a t = {u1, u2, u3, u4, u5, u6, u7, u8}, where u1, u2, u3, u4 are the control signal commands of the thrusters in the horizontal direction, providing the power of horizontal movement and rotation; u5, u6, u7, u8 are the control signal commands of the thrusters in the vertical direction, providing the power of vertical movement; the values of all control signal commands are normalized in the range of [-1, 1].
[0019] The reward function r t of the reinforcement learning model is defined as r t = r d + r z + r ψ + r target , where r d is a lateral distance error penalty to reduce the relative lateral distance of the tethered underwater robot from the target path, r d = k1·exp(-e d )-c1, where k1 is the weight of the penalty of the lateral distance component, c1 is a constant penalty to avoid the tethered underwater robot from taking no action to avoid the penalty, e d is the lateral error of the underwater robot from the target path; r z is a depth error penalty to reduce the relative vertical distance of the tethered underwater robot from the target path, r z = k2·exp(-e z )-c2, where k2 is the weight of the penalty of the vertical distance component, e z is the depth error, and c2 is a constant penalty; r ψ is a heading error penalty to reduce the heading error of the tethered underwater robot from the target path, r ψ = k3·exp(-k4·|e ψ |)-c3, where k3 and k4 are the weights of the penalty of the heading error, e ψ represents the heading error to be minimized, and c3 is a constant penalty; r target is a target waypoint reward, where k5 is the target waypoint reward weight, k is the serial number of the current target waypoint, and n is the total number of target path waypoints.
[0020] The reinforcement learning model comprises five neural networks, namely one policy network and four Q value evaluation networks, the policy network gives an action according to the input current state, for a continuous action space, the mean and standard deviation of the output action are used to generate a Gaussian distribution of the action, the Q value evaluation network gives a value estimation according to the input current state and action, and the advantages and disadvantages of the current policy are evaluated, wherein the input of the policy network is the state space of the agent, passes through four linear layers of full connection network, two output layers output the action mean and action logarithmic standard deviation respectively, finally the action taken under the current state is obtained from the probability distribution and is limited in the range of [-1, 1], and the input of the Q value evaluation network is the state space and action space of the agent, passes through three linear layers of full connection network, and the output dimension is 1, that is, the value estimation of the action-state pair.
[0021] In the training process of the reinforcement learning model, a set of experience data (s t ,a t ,r t ,s t+1 ) generated after the agent interacts with the environment is saved to an experience replay buffer, s t ,a t ,r t ,s t+1 respectively represent the state, action and reward of time step t, when the amount of experience data in the buffer reaches a threshold, the experience data entering the buffer earliest is automatically deleted, and when the reinforcement learning model is trained, a batch of experience data is sampled from the buffer to update the Q function in the Q value evaluation network, wherein the sampling probability of a set of experience data i is Wherein p i >0 represents the priority of the experience data i, and the index alpha is used to balance the degree of uniform and greedy, alpha=0 represents uniform sampling, and alpha=1 represents greedy sampling, that is, the maximum value is selected.
[0022] Compared with the prior art, the present application has the following beneficial effects:
[0023] 1、The present application adopts a grating type path as the ship body full traversal detection path and the target path of the reinforcement learning model according to the target ship body information, combines the sensor system carried on the tethered underwater robot, automatically controls the tethered underwater robot to track the path to complete the detection, can submerge or translate according to the working mode, and this can ensure that the ship body surface coating is comprehensively detected, avoids missed areas, and improves the performance and efficiency of the underwater operation of the tethered underwater robot.
[0024] 2、The reinforcement learning model of the present application can be trained according to different target paths, adapt to diversified target ship body full traversal detection, and realize flexible adjustment to different task requirements.
[0025] 3、The ship hull surface coating detection method of the tethered underwater robot provided by the application, the reward function comprehensively considers multiple factors such as lateral distance error, depth error and direction error, can more effectively guide the robot to track the target path, so that the underwater robot can complete the ship hull detection task and video image acquisition instead of manual work, facilitate the digital management and subsequent maintenance of the ship, and has the advantages of high efficiency, stability, low risk and the like.
[0026] 4、By constructing a simulation and training environment including an underwater environment interference model, the application can simulate interference factors such as ocean currents, so that the reinforcement learning model can learn how to maintain stability in a complex environment during the training process, and enhance its adaptability and anti-interference ability.
[0027] 5、The application combines the tracking error determined by the line-of-sight navigation method, optimizes the tracking control of the detection path, and the design of the reward function also promotes the robot to efficiently track the target path, thereby improving the detection efficiency.
[0028] 6、The application is equipped with a turbidity sensor, a ranging sonar and a UGPS underwater positioning system on the tethered underwater robot, and the combination of these sensors can more comprehensively perceive underwater environmental information and robot state information, and provide more accurate data support for path planning and motion control. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 The method flowchart of the application;
[0030] Figure 2 The basic model diagram of the tethered underwater robot;
[0031] Figure 3 The connection diagram of the tethered underwater robot and the sensor system;
[0032] Figure 4 The ship hull full traversal detection path schematic diagram;
[0033] Figure 5 The inertial coordinate system and carrier coordinate system schematic diagram;
[0034] Figure 6 The model diagram designed based on the deep reinforcement learning method;
[0035] Figure 7 The horizontal plane path tracking schematic diagram. DETAILED DESCRIPTION
[0036] The application will be described in detail below in combination with the drawings and specific embodiments. The embodiments are implemented on the premise of the technical solution of the application, and detailed implementation modes and specific operation processes are given, but the protection scope of the application is not limited to the following embodiments.
[0037] The application provides a tethered underwater robot hull surface coating detection method based on deep reinforcement learning, including the following contents: a sensor is carried on the tethered underwater robot to perceive underwater environment information, realize shooting distance self-adaptation, safe distance keeping and obtain the attitude information of the tethered underwater robot; a motion control model and an underwater environment model of the tethered underwater robot are established, such as kinematics and dynamics models, underwater environment interference models and the like, to provide a training environment for reinforcement learning; a reinforcement learning model is constructed based on a deep reinforcement learning method as a controller of the tethered underwater robot, including the design of state, action space, reward function and SAC algorithm model; a hull full traversal detection path is designed, and a grating type path is used to perform full traversal detection on the target hull; the model is trained to track the target path. The method uses the automatic control technology of the agent and the sensor system, combines the deep reinforcement learning method, enables the agent to make optimal action decisions based on the environment, effectively improves the performance and efficiency of the tethered underwater robot in hull surface coating detection, reduces the safety risk of manual operation in a complex underwater environment, and also enables the full life cycle management of the ship detection data through digital technology.
[0038] Specifically, as shown in the method includes the following steps: Figure 1
[0039] S1, the sensor carried on the tethered underwater robot perceives underwater environment information and robot state information.
[0040] As shown in the tethered underwater robot, the turbidity sensor, the ranging sonar and the positioner in the UGPS underwater positioning system are assembled, and the connection diagram is composed as shown in Figure 2 The tethered underwater robot based on multi-sensor fusion is composed according to the connection diagram as shown in Figure 3 First, the tethered underwater robot and the underwater receiving array in the UGPS underwater positioning system are released from both sides of the target hull, and it is ensured that they are completely immersed in the water. Secondly, the upper computer communicates with the tethered underwater robot and the sensor, tests the communication state and stability.
[0041] According to the turbidity information returned by the turbidity sensor, the shooting distance range is judged to ensure the definition of video image acquisition. Specifically, when the turbidity is less than or equal to 1 NTU, the distance between the tethered underwater robot and the target hull is set to 5 meters; when the turbidity is less than or equal to 10 NTU, the distance between the tethered underwater robot and the target hull is set to 3 meters; when the turbidity is less than or equal to 30 NTU, the distance between the tethered underwater robot and the target hull is set to 2 meters; when the turbidity exceeds 30 NTU, it is not recommended to perform detection. The turbidity sensor, the ranging sonar and the UGPS underwater positioning system are assembled on the existing tethered underwater robot, the upper computer communicates with the sensor through the serial port, transmits the measured data and performs data processing.
[0042] The ranging sonar is used to collect the distance between the tethered underwater robot and the obstacle, so as to keep the safe distance between the tethered underwater robot and the target ship body, prevent the collision with the target ship body, and ensure the definition of the video image collection.
[0043] The UGPS underwater positioning system is used to obtain the real-time position information of the tethered underwater robot, and the sensors of the tethered underwater robot body are combined to obtain the attitude and speed information s1=[x, y, z, ψ, u, v, w, r] of the robot, wherein (x, y, z) represents the three-dimensional position information of the tethered underwater robot, ψ represents the angle information of the tethered underwater robot, (u, v, w) represents the linear speed information of the tethered underwater robot, and r represents the angular speed information of the tethered underwater robot.
[0044] S2, the ship body full traversal detection path of the tethered underwater robot is planned.
[0045] According to the information of the target ship body, the improved Dijkstra algorithm is used to determine the full coverage detection path of the target ship body according to the underwater stability and the real-time shooting range. The Dijkstra algorithm is to start from the starting point, adopt the strategy of the greedy algorithm, traverse to the adjacent node of the vertex closest to the starting point and not visited each time, and expand to the terminal point. The detection path is a raster type path, after the tethered underwater robot completes one horizontal detection, it will dive or translate according to the working mode to perform the next horizontal detection. The position after diving / translation will be used as the starting point of the Dijkstra algorithm path exploration. This method can efficiently and comprehensively and stably complete the ship body surface coating detection task. The working mode is distinguished according to the detection area of the ship body or the ship bottom.
[0046] According to the information of the target ship body, including the length, width and draft of the ship body, a raster type path as shown in Figure 4 is designed. After the tethered underwater robot completes one horizontal detection, it will dive or translate according to the working mode to perform the next horizontal detection. Considering the factors such as the wide-angle of the gimbal of the tethered underwater robot and the shape of the ship body, the detection range in the path planning should be reduced by a part of offset to improve the working efficiency of the tethered underwater robot and be closer to the actual application scene.
[0047] In the ship body detection, the tethered underwater robot performs horizontal detection according to the path shown by the arrow in the route in Figure 4 When the detection distance reaches the length of the ship body, it will dive to a certain depth and perform reverse horizontal detection according to the raster path planning scheme. The diving depth is determined by the distance of the wide-angle of the gimbal and the ranging sonar. In order to make the image content shot by the gimbal have no omission, the diving depth is controlled to make the upper and lower edges of the image have a part of overlap. The image collection area is as shown in Figure 4The shadow area in FIG. 6B shows the detection range of the first detection. Similarly, the detection is performed until the detection depth reaches the draft of the ship, and the full detection of the ship body is completed.
[0048] In the ship bottom detection, the tethered underwater robot first adjusts the gimbal angle to look up to shoot a video image of the ship bottom, and is submerged to a suitable depth range for detection. When the detection distance reaches the length of the ship body, horizontal movement is performed, and reverse lateral detection is performed according to the raster path planning scheme. Similarly, the full detection of the ship bottom is completed when the detection width reaches the width of the ship body.
[0049] S3, taking the full detection path of the ship body as a target path, inputting a reinforcement learning model constructed based on a depth reinforcement learning method, combining real-time perception of underwater environment information and robot state information, to obtain a control signal of the tethered underwater robot; wherein the underwater environment information and the robot state information in the training process are obtained by using a simulation and training environment constructed based on an underwater environment disturbance model and a motion control model, to train the reinforcement learning model.
[0050] First, a simulation and training environment is constructed by using an underwater robot motion control model (including its kinematics and dynamics model) and an underwater environment disturbance model, to train the reinforcement learning agent.
[0051] First, two coordinate systems are established to describe the motion of the underwater robot, namely the inertial coordinate system E-ξηζ and the carrier coordinate system O-xyz, as shown in FIG. 6A. Figure 5 The inertial coordinate system establishes the coordinate system origin E at a point on the sea surface or in the sea, the ξ axis and the η axis are perpendicular to each other and located in the same horizontal plane, and the ζ axis points to the center of the earth and is perpendicular to the Eξη coordinate plane; the carrier coordinate system is established on the underwater robot, taking the center of gravity as the origin O, the x axis represents the forward direction of the underwater robot, pointing to the bow, the y axis is positive through the center of gravity pointing to the right side, and the z axis is vertically downward. The symbols are defined as follows: ξ, η, ζ are the horizontal, vertical and vertical positions of the underwater robot in the inertial coordinate system, with a unit of m; θ, ψ are the attitude angles of the underwater robot, respectively roll angle, pitch angle and yaw angle, with a unit of rad; u, v, w are the displacement velocities in the carrier coordinate system, respectively longitudinal, lateral and vertical velocities, with a unit of m / s; p, q, r are the angular velocities around the x, y, z axes in the carrier coordinate system, with a unit of rad / s; X, Y, Z and K, M, N are the thrust and moment along the x, y, z axes in the carrier coordinate system, with a unit of N and N·m respectively. Thus, the displacement vector of the underwater robot in the inertial coordinate system is defined as η1=[ξ,η,ζ] T , the velocity vector in the carrier coordinate system is v1=[u,v,w] T , and v2=[p,q,r] T, the force and moment vector is τ1 = [X, Y, Z] T , τ2 = [K, M, N] T .
[0052] Secondly, in the process of underwater robot motion, its motion speed and angular velocity state parameters are real-time changes, so it is necessary to describe its kinematics equation to convert between inertial coordinate system E-ξηζ and body coordinate system O-xyz, so as to obtain its position and attitude parameter information in the body coordinate system. The dynamic modeling mainly reflects the mutual conversion of various degrees of freedom input force and output moment, speed and other parameters. When external force and external moment act on underwater robot, its position and attitude will change. By analyzing the parameter conversion relationship between the two coordinate systems, the dynamic model is constructed. The six-degree-of-freedom kinematics and dynamics equations of underwater robot are:
[0053]
[0054] Where J(η) is the conversion matrix between the inertial coordinate system and the carrier coordinate system, η represents the displacement vector, v represents the velocity vector, τ represents the force and moment vector, τ d represents the interference force or moment vector brought by ocean current in underwater environment, v = [u, v, w, p, q, r] T , τ = [X, Y, Z, K, M, N] T , M is the inertia matrix, C(v) is the Coriolis centripetal force matrix, D(v) is the fluid damping matrix, and g(η) is the restoring force vector.
[0055] The force and moment of the eight thrusters in six degrees of freedom are:
[0056] τ = T(α) Ku,
[0057] Where T = [t1, t2, t3, t4, t5, t6, t7, t8] is the thrust distribution matrix, α is the thrust rotation angle vector, K = diag[K1, K2, K3, K4, K5, K6, K7, K8] is the thrust coefficient matrix, and u = [u1, u2, u3, u4, u5, u6, u7, u8] T is the control input vector.
[0058] Since the position of the buoyancy center of the tethered underwater robot is accurately measured, the stability of the roll and pitch two degrees of freedom can be maintained, that is Therefore, it is simplified to use four-degree-of-freedom kinematics and dynamics model.
[0059] The underwater environment disturbance model is represented as: τ d = [τ dx , τ dy , τdz ] T ,τ d ,τ dx ,τ dy ,τ dz ,τ dx ,τ dy ,τ dz ,τ
[0060] During the simulation and training process, the underwater environment disturbance model works as follows: at each step (simulation step) of state transition calculation, the time-varying disturbance force τ d is dynamically superimposed on the dynamics model of the underwater robot as an external force term, thereby affecting the dynamic response of the robot (i.e. the state at the next time). In addition, when selecting the fluid dynamics parameters, an uncertainty range of ±20% is added to simulate the water dynamic parameter disturbance of the underwater robot, ensuring that the strategy learned by the agent has universal applicability and good generalization ability in the actual working environment with modeling errors and environmental disturbances.
[0061] The reinforcement learning training environment is constructed, including the design of state space, action space and reward function, etc., so that the tethered underwater robot can complete the tracking control task of the target path. The overall module designed based on deep reinforcement learning method is shown in Figure 6 In the training environment, the tethered underwater robot as an agent interacts with the environment to obtain the control strategy. At each time step t, the agent perceives the environmental information to obtain the current state s t , and selects an action a t , which will then cause the environment to transition to the next state s t+1 , and get the reward r t from the environment, thereby quantifying the performance of the agent.
[0062] For the state space s t , the goal of the designed deep reinforcement learning-based controller is to reduce the error between the tethered underwater robot and the target path, and to track the desired heading angle. To achieve this goal, the line-of-sight (LOS) method is used to calculate the direction that the tethered underwater robot needs to track in the horizontal plane, as shown in Figure 7 For the total target path, a quadratic polynomial interpolation method is used to decompose it into a series of waypoints {P1,…,P k-1P k P k+1 P n P k-1 P k P k-1 P k P k P k-1 P k P k P k+1 P k P k+1 P n e d e ψ e z L ψ α k P k+1 P
[0063] Based on this, the current state s t is defined as s t ={x,y,z,ψ,u,v,w,r,e d ,e ψ ,e z}, where (x,y,z) represents the three-dimensional position information of the tethered underwater robot, ψ represents the angle information of the tethered underwater robot, (u,v,w) represents the linear velocity information of the tethered underwater robot, and r represents the angular velocity information of the tethered underwater robot. Among them, the range of ψ is [-π, π], when the value of ψ exceeds this range, the calculation method of ψ = ((ψ+π)%(2π))-π is used to map the value of ψ to the range of [-π, π], to prevent the value of ψ from jumping at the critical point, causing the algorithm to make an error judgment that the underwater robot has made a sharp turn. The current state s t During training, {x,y,z,ψ,u,v,w,r} in the state space is obtained from the motion control model of the tethered underwater robot; during actual use, {x,y,z,ψ,u,v,w,r} in the state space is perceived by the sensors mounted on the tethered underwater robot.
[0064] For the action space a tThe tethered underwater robot is powered by 8 thrusters, so the passage control signal command of the thrusters is defined as a continuous action space: a t = {u1, u2, u3, u4, u5, u6, u7, u8}, where u1, u2, u3, u4 are the control signal commands of the thrusters in the horizontal direction, providing the power of horizontal movement and rotation; u5, u6, u7, u8 are the control signal commands of the thrusters in the vertical direction, providing the power of vertical movement; the values of all control signal commands are normalized in the range of [-1, 1].
[0065] For the reward function r t , the smaller the error, the greater the value of the reward function, so as to guide the agent to complete the given task. For the path tracking problem of the tethered underwater robot, the reward function r t is designed as follows: r t = r d + r z + r ψ + r target , where r d is the lateral distance error penalty to reduce the relative lateral distance between the tethered underwater robot and the target path, r d = k1·exp(-e d )-c1, where k1 is the weight of the lateral distance component penalty, c1 is a constant penalty to avoid the tethered underwater robot from taking no action to avoid the penalty; r z is the depth error penalty to reduce the relative vertical distance between the tethered underwater robot and the target path, r z = k2·exp(-e z )-c2, where k2 is the weight of the vertical distance component penalty, c2 is a constant penalty; r ψ is the direction error penalty to reduce the heading error between the tethered underwater robot and the target path, r ψ = k3·exp(-k4·|e ψ |)-c3, where k3 and k4 are the weights of the direction error penalty, c3 is a constant penalty; r target is the target waypoint reward, where k5 is the target waypoint reward weight, k is the serial number of the current target waypoint, and n is the total number of target path waypoints.
[0066] The reinforcement learning model includes five neural networks, one policy network and four Q-value evaluation networks. The policy network gives the action according to the input current state, and for the continuous action space, the output action mean and standard deviation are used to generate a Gaussian distribution of actions; the Q-value evaluation network gives the value estimate according to the input current state and action, to evaluate the advantages and disadvantages of the current policy.
[0067] The input of the policy network is the state space of the agent, i.e., s t ={x, y, z, ψ, u, v, w, r, e d , e ψ , e z}, which is passed through a fully connected network with four linear layers, where the hidden layers contain 256 neural units with ReLU activation function, and two output layers output the action mean and action log standard deviation, respectively, and the action taken in the current state is obtained from the probability distribution and is limited in the range of [-1, 1]. The input of the Q-value evaluation network is the state space s t and the action space a t , which is 19-dimensional, passed through a fully connected network with three linear layers, where the hidden layers contain 256 neural units with ReLU activation function, and the output dimension is 1, which is the action-state pair value estimate.
[0068] During the training process of the reinforcement learning model, a set of experience data (s t , a t , r t , s t+1 ) generated after the agent interacts with the environment is saved in the experience replay buffer, where s t , a t , r t , s t+1 represent the state, action, reward at time step t, respectively; when the amount of experience data in the buffer reaches a threshold, the oldest experience data in the buffer is automatically deleted; when the reinforcement learning model is trained, a batch of experience data is sampled from the buffer to update the Q function in the Q-value evaluation network. In the sampling process, the method of prioritized experience replay is used to select experience data that is more helpful to the update of the network model parameters. The sampling probability of a set of experience data i is where p i >0 represents the priority of the experience data i, and the exponent α is used to balance the degree of uniformity and greediness, where α=0 represents uniform sampling, and α=1 represents greedy sampling, i.e., selecting the maximum value.
[0069] The specific update steps of the model are as follows:
[0070] Step S31, initialize network parameters θ1, θ2, and the experience replay buffer D;
[0071] Step S32, copy the parameters to the target network
[0072] Step S33, in each round, the following large loop is executed:
[0073] Step S331, reset the environment and obtain the initial state st ;
[0074] Step S332: Execute the following small loop at each time step:
[0075] according to Sampling action a t ;
[0076] Environmental feedback rewards and the next state t+1 ~p(s t+1 |s t ,a t );
[0077] Store the experience data in the experience replay cache, D←D∪{(s t+1 ,a t ,r(s t ,a t ),s t+1 )};
[0078] Update environment status s t+1 ←s t ;
[0079] Update strategy;
[0080] Step S333: When the time step reaches the threshold, the small loop ends;
[0081] Step S34: When the number of rounds reaches the maximum threshold M, the large loop ends.
[0082] Furthermore, the specific steps for updating the strategy are as follows:
[0083] 1) Update the Q function: for i∈{1,2},
[0084] in,
[0085] 2) Update policy weights: in,
[0086]
[0087] 3) Update the temperature coefficient:
[0088] in,
[0089] 4) Update the target network weights: for i∈{1,2}.
[0090] In this embodiment, the hyperparameters of the reinforcement learning model are set as follows: a discount factor of 0.99, a learning rate of the policy network of 0.0003, a learning rate of the Q value network of 0.0003, a soft update parameter of the target network of 0.01, a size of the experience replay buffer of 1,000,000, a sampling batch size of 128, a maximum time step of each round of 1,000, and an exploration step of 1,000. In the reward function, k1=k3=2, k2=k4=3, k5=10, c1=c3=1, and c2=2.
[0091] The end condition of each round in the training environment is set as follows: the lateral distance error between the tethered underwater robot and the target path exceeds 10 meters or the maximum time step length is 1,000, indicating that the robot cannot complete the tracking task or has reached the maximum time step length, the current round will be terminated, and a new round will start at the starting position.
[0092] S4, based on the control signal, controlling the tethered underwater robot to track the target path and complete the ship surface coating detection.
[0093] The preferred embodiments of the present application are described in detail above. It should be understood that those skilled in the art can make many modifications and changes without creative labor based on the concept of the present application. Therefore, any technical solutions obtained by logical analysis, reasoning, or limited experiments based on the prior art according to the concept of the present application shall be within the scope of protection defined by the claims.
Claims
1. A tethered underwater robot hull surface coating detection method based on deep reinforcement learning, characterized in that, The method comprises the following steps: sensing underwater environment information and robot state information by using sensors carried on the tethered underwater robot; planning a hull full-traversal detection path of the tethered underwater robot; inputting the hull full-traversal detection path as a target path into a reinforcement learning model constructed based on a deep reinforcement learning method, combining the real-time sensed underwater environment information and robot state information, and obtaining a control signal of the tethered underwater robot; wherein the underwater environment information and robot state information in the training process are obtained by using a simulation and training environment constructed based on an underwater environment interference model and a motion control model, so as to train the reinforcement learning model; controlling the tethered underwater robot to track the target path based on the control signal, and completing hull surface coating detection. The planning of the hull full-traversal detection path of the tethered underwater robot specifically comprises: according to the information of the target hull, using an improved Dijkstra algorithm, determining the full-coverage detection path of the target hull according to the underwater stability and the real-time shooting range, wherein the detection path is a grating type path; after the tethered underwater robot completes one horizontal detection, performing submersion or translation according to the working mode to perform the next horizontal detection, and the working mode is distinguished according to whether the detection area is a ship body or a ship bottom.
2. The method of claim 1, wherein the method is based on deep reinforcement learning. The sensing of the underwater environment information and the robot state information by using the sensors carried on the tethered underwater robot specifically comprises: assembling a turbidity sensor, a ranging sonar and a UGPS underwater positioning system on an existing tethered underwater robot, collecting the turbidity of the underwater environment by using the turbidity sensor, setting the distance between the tethered underwater robot and the target hull based on the turbidity to ensure the clarity of the video image collection; collecting the distance between the tethered underwater robot and the obstacle by using the ranging sonar to prevent collision with the target hull during the detection process; obtaining the real-time position information of the tethered underwater robot by using the UGPS underwater positioning system, combining the sensors of the tethered underwater robot body, obtaining the attitude and speed information s1=[x, y, z, ψ, u, v, w, r] of the robot, wherein (x, y, z) represents the three-dimensional position information of the tethered underwater robot, ψ represents the yaw angle of the tethered underwater robot, (u, v, w) represents the linear velocity information of the tethered underwater robot, and r represents the angular velocity of the tethered underwater robot around the z axis.
3. The method of claim 1, wherein the method is based on deep reinforcement learning. The underwater environment disturbance model is expressed as: , represents the disturbance force or torque vector brought by ocean current in the underwater environment, respectively represent the disturbance force or torque in x, y, z directions brought by ocean current in the underwater environment, by using a random function to enhance the ability of the model to resist external disturbance, respectively set as: , , , wherein rand represents a random function.
4. The method of claim 1, wherein, The motion control model includes kinematic equations and dynamic equations, wherein the six-degree-of-freedom kinematic equations of the underwater robot are as follows: This is used to describe the robot's position, velocity, and acceleration; the dynamic equations are as follows: It is used to analyze the forces acting on and motion states of a robot, among which, This is the transformation matrix between the inertial coordinate system and the vehicle coordinate system. Represents the displacement vector. Represents the velocity vector. Represents force and torque vectors. This represents the vector of disturbance force or torque caused by ocean currents in the underwater environment. =[ξ,η,ζ,φ,θ,ψ] T , =[u,v,w,p,q,r] T , =[X,Y,Z,K,M,N] T ξ, η, ζ represent the positions of the underwater robot along the horizontal, vertical, and inertial axes in the inertial coordinate system, in meters (m); φ, θ, ψ represent the attitude angles of the underwater robot, namely roll angle, pitch angle, and yaw angle, respectively, in rad; u, v, w represent the linear velocities of the underwater robot in the carrier coordinate system, namely longitudinal, lateral, and vertical velocities, respectively, in m / s; p, q, r represent the angular velocities about the x, y, and z axes in the carrier coordinate system, in rad / s; X, Y, Z and K, M, N represent the thrust and torque along the x, y, and z axes in the carrier coordinate system, respectively, in N and N·m. The inertia matrix, The Coriolis centripetal force matrix, Here is the fluid damping matrix. This is the restoring force vector.
5. The method of claim 1, wherein, A state space s of the reinforcement learning model t is defined as s t ={x,y,z,ψ,u,v,w,r,e d ,e ψ ,e z}, wherein (x,y,z) represents three-dimensional position information of the tethered underwater robot, ψ represents a yaw angle of the tethered underwater robot, (u,v,w) represents linear velocity information of the tethered underwater robot, r represents angular velocity information of the tethered underwater robot around the z-axis, e d ,e ψ ,e z represents a tracking error determined by using a line-of-sight navigation method, e d represents a lateral error of the underwater robot from a target path, e ψ represents a directional error that needs to be minimized, and e z represents a depth error; wherein the range of ψ is [-π, π], and when the value of ψ exceeds the range, the value of ψ is mapped to the range of [-π, π] by using a calculation method of ψ=((ψ+Π)%(2Π))-Π; the current state s t In training, {x,y,z,ψ,u,v,w,r} in the state space is obtained by a motion control model of the tethered underwater robot; in actual use, {x,y,z,ψ,u,v,w,r} in the state space is perceived by sensors carried on the tethered underwater robot.
6. The method of claim 1, wherein, The action space a of the reinforcement learning model t The action space a of the reinforcement learning model t = {u1, u2, u3, u4, u5, u6, u7, u8}, where u1, u2, u3, u4 are the control signal commands of the thrusters in the horizontal direction, providing the power of horizontal movement and rotation; u5, u6, u7, u8 are the control signal commands of the thrusters in the vertical direction, providing the power of vertical movement; the values of all control signal commands are normalized in the range [-1, 1].
7. The method of claim 1, wherein, a reward function r of the reinforcement learning model t is defined as wherein, is a lateral distance error penalty to reduce the relative lateral distance of the tethered underwater vehicle from the target path, wherein, k1 is a weight to penalize the lateral distance component, c1 is a constant penalty to avoid the tethered underwater vehicle from taking no action to avoid the penalty, e d is the lateral error of the underwater vehicle from the target path; is a depth error penalty to reduce the relative vertical distance of the tethered underwater vehicle from the target path, wherein, k2 is a weight to penalize the vertical distance component, e z is the depth error, is a constant penalty; is a heading error penalty to reduce the heading error of the tethered underwater vehicle from the target path, wherein, k3 and k4 are weights to penalize the heading error, e ψ denotes the heading error to be minimized, is a constant penalty; is a target waypoint reward, wherein, k5 is a target waypoint reward weight, k is the index of the current target waypoint, n is the total number of waypoints of the target path.
8. The method of claim 1, wherein, The reinforcement learning model comprises five neural networks, namely one policy network and four Q value evaluation networks, the policy network gives an action according to the input current state, for a continuous action space, the mean and standard deviation of the output action are used to generate a Gaussian distribution of the action, the Q value evaluation network gives a value estimation according to the input current state and action, and the advantages and disadvantages of the current policy are evaluated, wherein the input of the policy network is the state space of the agent, passes through four linear layers of the full connection network, two output layers output the action mean and action logarithmic standard deviation respectively, finally the action taken under the current state is obtained from the probability distribution and is limited in the range of [-1, 1], the input of the Q value evaluation network is the state space and action space of the agent, passes through three linear layers of the full connection network, and the output dimension is 1, that is, the value estimation of the action-state pair.
9. The method of claim 1, wherein, A set of experience data generated after the agent interacts with the environment in the training process of the reinforcement learning model saved into an experience replay buffer, respectively represent the state, action and reward of time step t; when the amount of experience data in the buffer reaches a threshold, the experience data that enters the buffer earliest is automatically deleted; During training of the reinforcement learning model, a batch of experience data is sampled from the buffer to update the Q function in the Q-value evaluation network, where the sampling probability of a set of experience data i is where p i >0 represents the priority of the experience data i , and the exponent a is used to balance the degree of uniform and greedy, a=0 represents uniform sampling, and a=1 represents greedy sampling, i.e., selecting the maximum value.
Citation Information
Patent Citations
Underwater robot intelligent control method based on Bayesian deep reinforcement learning
CN114995468A
Underwater cleaning robot path planning method and system based on hull model
CN107918396A
Novel mobile aquaculture monitoring system
CN109471445A