An Underwater Robot Formation Control Method Based on Digital Twin and Reinforcement Learning

By combining digital twins and reinforcement learning methods in the control of underwater robot formations, a high-precision digital twin model and virtual underwater environment is built, which solves the control and monitoring problems of underwater robot formations in complex environments, and achieves real-time and accurate formation control and monitoring effects.

CN116679733BActive Publication Date: 2025-05-30YANSHAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310796400.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-05-30
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problems of precise control and real-time monitoring of underwater robot formations in complex underwater environments, especially due to the influence of nonlinearity and external factors of the dynamic model, which makes it difficult to monitor the motion posture information of underwater robots in real time.

Method used

The underwater robot formation control method based on digital twins and reinforcement learning is adopted. By constructing the underwater robot digital twin model and virtual underwater environment, the digital twin model is trained using integral reinforcement learning, controlling strategies and perturbation inputs are generated, and the system design and optimization is carried out through two-way real-time data interaction.

Benefits of technology

It reduces the number of underwater robot tests and equipment wear, ensures the real-time and effectiveness of data, and realizes accurate control and real-time monitoring of underwater robot formations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116679733B_ABST
    Figure CN116679733B_ABST
Patent Text Reader

Abstract

The present invention discloses an underwater robot formation control method based on digital twin and reinforcement learning, belonging to the field of underwater robot control, which includes collecting the data required for underwater robots and water area environment and preparing them into a data set; constructing digital twins of underwater robots and virtual underwater environments to complete the preliminary establishment of the digital twin model; training the digital twin model using integral reinforcement learning, and constructing a value function, a control strategy and a disturbance input at the initial moment; using the method of neural network structure and iterative weights to update the weights corresponding to the value function, control strategy and disturbance input in the neural network, and transmitting the newly generated control strategy and disturbance input to the underwater robot; correcting the digital twin and selecting the optimal solution; monitoring the underwater robot in real time through the digital twin model, and judging whether the underwater robots form an expected formation. If not, repeat steps 3-5; if so, the formation is completed. The present invention can ensure the real-time and effectiveness of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of underwater robot control, and in particular to a formation control method for underwater robots based on digital twin and reinforcement learning. Background Art

[0002] With the advancement of ocean resource development, underwater robot control technology has witnessed unprecedented development. In underwater cooperative target search and rescue, and non-cooperative target early warning monitoring systems, multiple underwater robots are required to form a specific formation to achieve continuous monitoring of intrusion targets. Compared with static underwater sensor network monitoring, underwater robot formation monitoring improves the flexibility and adaptability of the network, and through multi-hop transmission of underwater robots, it enhances the rapidity and persistence of monitoring underwater moving targets. However, compared with land robots, the dynamic models of underwater robots exhibit characteristics such as strong nonlinearity and high coupling, making it very difficult to establish an accurate dynamic model of underwater vehicles. Moreover, affected by external factors such as water flow, complex underwater terrain, and suspended matter in water, it is difficult to monitor the motion attitude information of underwater robots in real time.

[0003] After searching the existing literature, it is found that the Chinese patent application number is 201910274101.0, and the name is "A Formation Control Method for Multiple Underwater Robots Based on Reinforcement Learning". In this invention, after each underwater robot obtains its own position, the control center provides the trajectory information of the virtual leader and sends it to the neighbor nodes of the virtual leader; the underwater robot formation uses the current control strategy to track the trajectory, and each node calculates a one-step cost function by interacting with the environment and adjacent nodes, and uses the optimal control strategy to find the expected tracking target. However, the reinforcement learning method requires a large amount of training for underwater robots in a complex underwater environment, increasing the entire control cycle, and frequent training will cause secondary damage to underwater equipment.

[0004] After further searching, it is found that the Chinese patent application number is 202211002374.8, and the name is "Underwater Vehicle Modeling Method and System Based on Digital Twin". This invention discloses an underwater vehicle modeling method and system based on digital twin, and its twin model application system can ensure the accuracy of various working conditions and numerical simulations of underwater vehicles, and improve the credibility of underwater vehicle design optimization. However, this system only designs the corresponding digital twin model of the underwater vehicle, does not combine with the control algorithm, and does not consider the control effectiveness and real-time performance of underwater robots.

[0005] Therefore, considering the above deficiencies, it is particularly important to design a formation control method for underwater robots based on digital twin and reinforcement learning. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an underwater robot formation control method based on digital twin and reinforcement learning. By combining digital twin and reinforcement learning, training the twin model in the twin system reduces the number of underwater robot underwater tests and the wear of corresponding underwater equipment. And two-way real-time data interaction runs through the entire digital twin process, which can ensure the timeliness and effectiveness of data.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0008] An underwater robot formation control method based on digital twin and reinforcement learning, comprising the following steps:

[0009] Step 1, collect the required data of underwater robots and the water area environment, and prepare the collected data into a data set;

[0010] Step 2, based on the data of underwater robots and the water area environment in the data set of Step 1, construct underwater robot digital twins and virtual underwater environments in the same proportion to complete the preliminary establishment of the digital twin model;

[0011] Step 3, use the position, attitude, linear velocity, and angular velocity in the data set of Step 1 as state data quantities for training the digital twin model. Adopt integral reinforcement learning to train the digital twin model, and construct a value function and the control strategy and disturbance input at the initial moment;

[0012] Step 4, adopt the method of neural network structure and iterative weights to update the weights corresponding to the value function, control strategy, and disturbance input in the neural network, generate a new control strategy and disturbance input, and then transmit the newly generated control strategy and disturbance input to the underwater robot through the communication module;

[0013] Step 5, perform actions based on the received control strategy and disturbance input in Step 4. The underwater robot updates its own state, and at the same time feeds back the state data information of position, attitude, linear velocity, and angular velocity to the digital twin model. Further, based on the feedback state data, make real-time comparisons to correct the digital twin and select the optimal system design scheme;

[0014] Step 6, monitor the underwater robot in real time through the digital twin model, and judge whether the underwater robot forms an expected formation. If not, repeat Steps 3-5; if so, the formation is completed.

[0015] A further improvement of the technical solution of the present invention lies in: in Step 1, the data set includes the size characteristics, structural characteristics, system dynamics, hydrodynamics, buoyancy, propeller propulsion performance, structural stress analysis, input-output characteristics of the underwater robot, water flow velocity, fluid density, water temperature, and water area volume.

[0016] A further improvement of the technical solution of the present invention lies in: in step 2, characteristic parameters are used to represent the design scheme, and the specific construction process of the digital twin model is as follows:

[0017] 3) Design of the digital twin model:

[0018] D = D k (M, B, F h , V w , T, V s , ρ w , p, a, v, ω) (1)

[0019] In the formula, D represents the system design scheme, and D k represents the design scheme of the kth subsystem, and k ∈ N * , M, B, F h , V w , T, V s , ρ w , p, a, v, ω are the characteristic parameters of the underwater robot's size characteristics, buoyancy, resultant external force, water flow velocity, water temperature, water area volume, liquid density, underwater robot position, attitude, linear velocity, and angular velocity respectively;

[0020] 4) Expression for the net buoyancy of the underwater robot:

[0021] ΔB k = B k - G k (2)

[0022] In the formula, B and G represent the buoyancy and gravity of the underwater robot respectively, B = ρ w gv,

[0023] 3) V X , V Y , V Z are the water flow velocities in the x, y, and z axis directions respectively;

[0024] 4) The expression for the resultant force on the underwater robot is:

[0025] F h = G + F r + F c , (3)

[0026] In the formula, F r and F c are the repulsive force and the ocean current flow field force on the underwater robot respectively. At the same time, through computational fluid dynamics, F c is roughly considered as F c = αF f , α is the proportionality coefficient, Ff is the ocean current force;

[0027] 5) In the input-output characteristics of the underwater robot, the relationship between the input x and the output speed y is approximately:

[0028] y = ax 3 + bx 2 + cx + d (4)

[0029] In the formula, the output speed y = y(ω, v ac , v up ), and the above cubic function relationship is obtained by linearly fitting and organizing the measured operation data of the underwater robot. Moreover, ω, v ac , v up respectively represent the angular velocity, the front-back linear velocity, and the up-down linear velocity of the underwater robot, and all can approximately represent the input-output relationship by the cubic function relationship.

[0030] A further improvement of the technical solution of the present invention lies in: in step 3, an integral reinforcement learning is used to train the digital twin model, and the cost function is defined as:

[0031]

[0032] In the formula, e pij = x pi - x pj - p ij , x pi , x pj respectively represent the real position states of the i-th and j-th underwater robots, and p ij is the difference between the expected position states of the i-th and j-th underwater robots, that is e pi represents the error between the real position state and the expected position state, that is

[0033] A further improvement of the technical solution of the present invention lies in: in step 4, a neural network structure is used to estimate the unknown value function, control strategy, and disturbance input, which is expressed as:

[0034]

[0035] In the formula, are respectively the basis functions corresponding to the value function control strategy disturbance input ;

[0036] The least squares method is used to update the weights. Based on Y pi = W pi X pi the weight update is obtained:

[0037]

[0038] A further improvement of the technical solution of the present invention lies in that: in step 5, based on the position p, attitude a, linear velocity v, and angular velocity ω state data feedback by the underwater robot, real-time comparison is performed to correct the digital twin, and at the same time, the performance index Y of the current design solution c and the design index Y o are compared. Based on ||Y c -Y o ||≤ε, this formula is used for judgment. If the difference between Y c and Y o is greater than the maximum index threshold ε, then steps 4-5 need to be repeated. When the difference between the two is less than or equal to the maximum index threshold ε, the performance index Y of the current design solution c is the optimal system design solution D * .

[0039] Due to the adoption of the above technical solution, the technical progress achieved by the present invention is as follows:

[0040] 1. The present invention adopts the digital twin method to construct a high-precision digital twin model system, which can simulate the state of the underwater robot in the real underwater environment and provides a visual online monitoring platform.

[0041] 2. The present invention adopts the reinforcement learning method, which does not require the system model parameters to be completely known and further estimation of the dynamic model parameters, can get rid of the dependence on the system parameters, and belongs to the model-free data-driven type.

[0042] 3. The present invention combines the digital twin and reinforcement learning methods. Training the twin model in the twin system reduces the number of underwater trials of the underwater robot and the wear of the corresponding underwater equipment, and the two-way real-time data interaction runs through the entire digital twin process, which can ensure the real-time and effectiveness of the data. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings;

[0044] Figure 1 is the flowchart of the formation control of the underwater robot based on digital twin and reinforcement learning in the embodiment of the present invention;

[0045] Figure 2 is the overall block diagram of the development platform in the embodiment of the present invention;

[0046] Figure 3 It is a schematic diagram of the formation of underwater robots in the embodiments of the present invention. Specific implementation manners

[0047] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0049] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:

[0050] As Figure 1 , Figure 2 and Figure 3 shown, the present invention describes a formation control method for underwater robots based on digital twin and reinforcement learning, including the following steps:

[0051] Step 1: Collect the required data of the underwater robot and the water area environment, and prepare the collected data into a data set. The above data set includes the size characteristics, structural characteristics, system dynamics, hydrodynamics, buoyancy, propeller propulsion performance, structural force analysis, input-output characteristics of the underwater robot, water flow velocity, fluid density, water temperature, and water area volume. The devices used for data collection include velocity sensors, vision sensors, attitude sensors, pressure sensors, temperature sensors, and locators.

[0052] Step 2: Based on the data of the underwater robot and the water area environment in the data set in Step 1, construct a digital twin of the underwater robot and a virtual underwater environment in the same proportion to complete the preliminary establishment of the digital twin system. Use characteristic parameters to represent the design scheme. The specific construction process of the digital twin model is as follows:

[0053] 1) Design of the digital twin model:

[0054] D = D k (M, B, F h , V w , T, V s , ρ w , p, a, v, ω) (1)

[0055] In the formula, D represents the system design scheme, and D k represents the design scheme of the kth subsystem, and k ∈ N * , M, B, F h , V w , T, V s , ρ w , p, a, v, ω are the characteristic parameters of the underwater robot's size characteristics, buoyancy, resultant external force, water flow velocity, water temperature, water area volume, liquid density, underwater robot position, attitude, linear velocity, and angular velocity respectively.

[0056] 2) Net buoyancy expression of the underwater robot:

[0057] ΔB k = B k - G k (2)

[0058] In the formula, B and G represent the buoyancy and gravity of the underwater robot respectively, B = ρ w gv, In this embodiment, assuming that the underwater robot is a cuboid and completely submerged in water, is the depth at which the underwater robot is submerged in water, x, y, z are the length, height, and width of the underwater robot respectively, and that is, if the object is completely submerged in water, then is equal to the object's height; if only part of the object is submerged in water, then is equal to the depth at which the object is submerged in water.

[0059] y d = y s - y p + y b , where y s is the horizontal plane, y d is the depth of the underwater robot in water, y p is the position of the center point of the underwater robot, y b is half of the height of the underwater robot's boundary, that is, y b = y / 2.

[0060] 3) V X , VY , V Z are the water flow velocities in the x, y, and z axis directions respectively.

[0061] 4) The expression for the resultant force on the underwater robot is:

[0062] F h = G + F r + F c , (3)

[0063] F r and F c are the repulsive force and the ocean current flow field force on the underwater robot respectively. At the same time, through computational fluid dynamics, F c can be roughly considered as F c = αF f , where α is a proportionality coefficient, and F f is the ocean current force.

[0064] 5) The relationship between the input x (PWM) and the output speed y in the input-output characteristics of the underwater robot can be approximated as:

[0065] y = ax 3 + bx 2 + cx + d (4)

[0066] In the formula, the output speed y = y(ω, v ac , v up ), and the above cubic function relationship is obtained by linear fitting and organizing the measured operation data of the underwater robot. Moreover, ω, v ac , v up represent the angular velocity, the front-back linear velocity, and the up-down linear velocity of the underwater robot respectively, and all can be approximately expressed by the cubic function relationship between the input and the output; PWM refers to the pulse value, which is equivalent to the input, such as the size of stepping on the accelerator when driving a car.

[0067] Step 3: Use the position, attitude, linear velocity, and angular velocity in the dataset of Step 1 as state data quantities for training the digital twin model. Adopt integral reinforcement learning to train the digital twin model, and construct the value function and the control strategy u pi and the perturbation input δ pi , and define the cost function as:

[0068]

[0069] In the formula, e pij = x pi - x pj - p ij , x pi , x pj represent the true position states of the i-th and j-th underwater robots respectively, and pij is the difference between the expected position states of the i-th and j-th underwater robots, that is e pi represents the error between the actual position state and the expected position state, that is

[0070] Step 4: Adopt the method of neural network structure and iterative weights to update the weights corresponding to the value function, control strategy, and disturbance input in the neural network, and generate a new control strategy u pi and disturbance input δ pi , the neural network estimates the unknown value function, control strategy, and disturbance input, which can be expressed as:

[0071]

[0072] In the formula, are the value function control strategy disturbance input corresponding basis functions respectively.

[0073] Update the weights using the least squares method. Based on Y pi = W pi X pi The weight update formula is obtained as:

[0074]

[0075] Furthermore, transmit the newly generated control strategy and disturbance input to the underwater robot through the communication module and the buoy, so as to realize the next generated control strategy and disturbance input to drive the underwater robot.

[0076] Step 5: Based on the received control strategy and disturbance input in Step 4, perform actions. The underwater robot updates its own state, and at the same time feeds back the state data information of position p, attitude a, linear velocity v, and angular velocity ω to the digital twin model through the buoy. Further, based on the feedback state data, make a real-time comparison to correct the digital twin, and at the same time compare the performance index Y c of the current design scheme with the design index Y o Make a comparison based on ||Y c - Y o || ≤ ε. If the difference between Y c and Y o is greater than the maximum index threshold ε, then Steps 4 - 5 need to be repeated. When the difference between the two is less than or equal to the maximum index threshold ε, the performance index Y c of the current design scheme is the optimal system design scheme D * .

[0077] Step 6, monitor the underwater robot in real time through the digital twin model, and determine whether the underwater robot forms the expected formation. If not, repeat Steps 3-5 until the system weight approaches a certain fixed value, the weight change range is less than the threshold 0.001, and the control strategy perturbation input At this time, V pi is the optimal value function, u pi and δ pi are the optimal control strategy and the perturbation input, and the underwater robot forms the expected formation.

[0078] The above-described implementation only describes the preferred implementation manner of the present invention, and does not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solution of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. An underwater robot formation control method based on digital twin and reinforcement learning, characterized in that, it includes the following steps: Step 1, collect the required data of the underwater robot and the water area environment, and prepare the collected data into a data set; Step 2, based on the data of the underwater robot and the water area environment in the data set in Step 1, construct digital twins of underwater robots and virtual underwater environments in the same proportion to complete the preliminary establishment of the digital twin model; Step 3, use the position, attitude, linear velocity, and angular velocity in the data set in Step 1 as state data quantities for training the digital twin model. Use integral reinforcement learning to train the digital twin model, and construct a value function, a control strategy at the initial moment, and a disturbance input; Step 4, use the method of neural network structure and iterative weights to update the weights corresponding to the value function, control strategy, and disturbance input in the neural network, generate a new control strategy and disturbance input, and then transmit the newly generated control strategy and disturbance input to the underwater robot through the communication module; Step 5, perform actions based on the received control strategy and disturbance input in Step 4. The underwater robot updates its own state, and at the same time feeds back the position, attitude, linear velocity, and angular velocity state data information to the digital twin model, and further makes real-time comparisons based on the feedback state data to correct the digital twin, and selects the optimal system design scheme; Step 6, monitor the underwater robot in real time through the digital twin model, and judge whether the underwater robot forms an expected formation. If not, repeat Steps 3-5; if so, the formation is completed.

2. An underwater robot formation control method based on digital twin and reinforcement learning according to claim 1, characterized in that: In Step 1, the data set includes the size characteristics, structural characteristics, system dynamics, hydrodynamics, buoyancy, propeller propulsion performance, structural force analysis, input-output characteristics of the underwater robot, water flow velocity, fluid density, water temperature, and water area volume.

3. An underwater robot formation control method based on digital twin and reinforcement learning according to claim 1, characterized in that: In Step 2, use characteristic parameters to represent the design scheme. The specific construction process of the digital twin model is as follows: 1) Design of the digital twin model: D = D k (M, B, F h , V w , T, V s , ρ w , p, a, v, ω) (1) Wherein, D represents the system design solution, D k represents the design solution of the k-th subsystem, and k ∈ N * , M, B, F h , V w , T, V s , ρ w , p, a, v, ω are the characteristic parameters of the underwater robot's size characteristics, buoyancy, resultant external force, water flow velocity, water temperature, water area volume, liquid density, underwater robot position, attitude, linear velocity, and angular velocity, respectively; 2) Net buoyancy expression of the underwater robot: ΔB k = B k - G k (2) where B and G respectively represent the buoyancy and gravity of the underwater robot, and B = ρ w gv, 3) V X 、V Y 、V Z are the water flow velocities in the x, y, and z axis directions respectively; 4) Expression of the resultant force on the underwater robot: F h = G + F r + F c , (3) Where, F r and F c are the repulsive force and the ocean current force received by the underwater robot respectively. Meanwhile, through computational fluid dynamics, F c is roughly considered as F c = αF f , where α is the proportionality coefficient and F f is the ocean current force; 5) The relationship between the input x and the output velocity y in the input-output characteristics of the underwater robot is approximately: y = ax 3 + bx 2 + cx + d (4) where the output speed y = y(ω, v ac , v up ), and the above cubic function relationship is obtained by linearly fitting and organizing the measured operation data of the underwater robot. Moreover, ω, v ac , v up respectively represent the angular velocity, the front-back linear velocity, and the up-down linear velocity of the underwater robot, and all of them can approximately represent the input-output relationship with a cubic function relationship.

4. An underwater robot formation control method based on digital twin and reinforcement learning according to claim 1, characterized in that: In Step 3, use integral reinforcement learning to train the digital twin model, and define the cost function as: where e pij = x pi - x pj - p ij , x pi , x pj represent the true position states of the i-th and j-th underwater robots respectively, and p ij is the difference between the expected position states of the i-th and j-th underwater robots, that is e pi represents the error between the true position state and the expected position state, that is 5. An underwater robot formation control method based on digital twin and reinforcement learning according to claim 1, characterized in that: In Step 4, use the neural network structure to estimate the unknown value function, control strategy, and disturbance input, expressed as: In the formula, are the value functions control strategy perturbation input corresponding basis functions; Update the weights using the least squares method, based on Y pi = W pi X pi The weight update formula is obtained as follows:

6. An underwater robot formation control method based on digital twin and reinforcement learning according to claim 1, characterized in that: In step 5, based on the position p, attitude a, linear velocity v, and angular velocity ω state data feedback from the underwater robot, real-time comparison is carried out to correct the digital twin, and at the same time, the performance index Y of the current design scheme c is compared with the design index Y o . Based on ||Y c -Y o ||≤ε, this formula is used for judgment. If the difference between Y c and Y o is greater than the maximum index threshold ε, then steps 4-5 need to be repeated. When the difference between the two is less than or equal to the maximum index threshold ε, the performance index Y of the current design scheme c is the optimal system design scheme D * .

Citation Information

Patent Citations

  • Underwater vehicle modeling method and system based on digital twinning

    CN115329459A

  • Multi-underwater-robot formation control method based on reinforcement learning

    CN109947131A

  • Robot motion planning method based on digital twinning and reinforcement learning

    CN115903825A