Deep-sea sediment collection system and method based on adaptive robust control
Through a deep-sea sediment acquisition system combining DDPG, AFC and PID algorithms, high-precision and low-perturbation sediment acquisition in complex deep-sea environments is achieved, solving the problems of low automation and insufficient sampling accuracy in the prior art, ensuring sample stability and data reliability.
Patent Information
- Application Number
- CN202510635422.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-29
AI Technical Summary
The prior art does not have high degree of automation when collecting sediment in deep-sea environments, has low sampling accuracy and efficiency, and is prone to disturb the sediment layer, affecting the original state of the sample and the accuracy of the research results.
A deep-sea sediment acquisition system based on adaptive robust control is adopted, combining the depth determination strategy gradient (DDPG) algorithm, an adaptive fuzzy control (AFC) algorithm and a PID control algorithm to achieve high-precision and low-perturbation sediment acquisition through intelligent control and adaptive adjustment.
It improves the stability and accuracy of the sampling process, reduces disturbances to sediments, ensures the original state of the sample, improves sampling efficiency and data reliability, and enhances the adaptability to complex deep-sea environments.
Smart Images

Figure CN120560076A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep-sea seabed sediment collection, and more particularly to a deep-sea sediment collection system and method based on adaptive robust control. Background Art
[0002] As marine science research continues to advance, the demand for rapid and accurate collection of deep-sea sediments is growing. Deep-sea sediments are a vital component of the marine environment and ecosystem, containing a wealth of geological, chemical, and biological information. This information is crucial for understanding marine geological processes, climate change, biodiversity, and the migration and transformation of pollutants. Deep-sea sediments exhibit a complex hierarchical structure from the surface to the depths, including an organic-rich surface layer, a middle layer with active redox reactions, and a deep layer with anaerobic conditions. Key components in each layer, such as organic carbon, nitrogen, phosphorus compounds, heavy metals, and microbial communities, have a significant impact on the composition and function of marine ecosystems. Analysis of these sediments can reveal the historical changes in the marine environment, the dynamics of the ecosystem, and the migration patterns of pollutants.
[0003] In deep-sea environments, collecting sediments faces many challenging technical difficulties, especially under high pressure and complex seabed topography. Existing technologies such as gravity samplers, box samplers, and column samplers are significantly limited in their automation. Furthermore, these traditional devices can easily disturb the sediment layer during sampling, compromising the original state and integrity of the samples and affecting the accuracy of research results. Furthermore, these devices struggle to maintain optimal posture for sample collection in complex and changing marine environments, further reducing sampling accuracy and efficiency. Summary of the Invention
[0004] In order to overcome the problems of low automation, low sampling accuracy and efficiency in the above-mentioned prior art when collecting deep-sea sediments, the present invention provides a deep-sea sediment collection system and method based on adaptive robust control. The system can not only achieve high-precision, low-disturbance sediment collection in complex deep-sea environments, but also greatly improve the robustness and automation of the system through intelligent control algorithms and adaptive adjustment mechanisms, respond to environmental changes and complex seabed topography in real time, ensure that the sampling device maintains the optimal posture in the deep-sea environment, better adapt to the complex and changeable marine environment, and provide important support for marine scientific research.
[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0006] A deep-sea sediment collection system based on adaptive robust control, comprising: an above-water control module, a lifting and deployment module, and a deep-sea sediment collection module;
[0007] The above-water control module is respectively connected to the lifting and placing module and the deep-sea sediment collection module in communication; the lifting and placing module is also connected to the deep-sea sediment collection module;
[0008] The above-water control module has built-in DDPG, AFC and PID algorithms. The above-water control module is used to calculate the sampling strategy based on the environmental information and posture information transmitted by the deep-sea sediment acquisition module, and issue instructions to the lifting and deployment module;
[0009] The lifting and deploying module is used to drive the deep-sea sediment collection module to move vertically in the seawater according to the received instructions;
[0010] The deep-sea sediment collection module is used to collect environmental information and posture information in real time and send them to the above-water control module, and to collect samples of deep-sea sediments according to a sampling strategy.
[0011] Preferably, the deep-sea sediment collection module includes a sealed cabin, and an underwater control submodule, a sediment collection device, a sensor submodule and a propeller arranged in the sealed cabin;
[0012] The sealed cabin is connected to the lifting and deploying module;
[0013] The underwater control submodule is communicatively connected to the surface control module; the underwater control submodule is also electrically connected to the sediment collection device, the sensor submodule and the propeller respectively;
[0014] The sensor submodule is used to collect environmental information and posture information and send them to the underwater control submodule;
[0015] The underwater control submodule is used to control the propeller and sediment collection device to sample deep-sea sediments according to the sampling strategy, and to send environmental information and posture information to the surface control module;
[0016] The thruster is used to adjust the magnitude and direction of its thrust, thereby adjusting the sampling posture of the sediment collection device, and the sediment collection device is used to perform the sampling action of deep-sea sediments.
[0017] Preferably, the underwater control submodule includes: a power supply device, a storage device, a processor, a data acquisition device and a data transmission device;
[0018] The power supply device is used to power the deep-sea sediment collection module; the data collection device is used to summarize and collect the environmental information and posture information collected by the sensor submodule; the data transmission device is used to transmit the environmental information and posture information to the storage, processor and water control module; the processor is also used to receive the sampling strategy sent by the water control module and control the thruster and sediment collection device to sample deep-sea sediments.
[0019] Preferably, the sensor submodule includes at least: a positioning sensor, an inertial sensor (IMU), a three-axis magnetometer, an acoustic Doppler current profiler (ADCP), a multibeam echosounder (MBES), a side scan sonar, an altitude sensor, an image sensor, a particle size analyzer, a pressure sensor, a temperature sensor, a salinity sensor, a conductivity sensor, a methane concentration sensor, a carbon dioxide sensor, a dissolved oxygen sensor, a chlorophyll sensor and a pH sensor.
[0020] The present invention also provides a deep-sea sediment collection method based on adaptive robust control, which is based on the above-mentioned deep-sea sediment collection system based on adaptive robust control and includes the following steps:
[0021] S1: The above-water control module controls the lifting and deploying module to drive the deep-sea sediment collection module to dive to the seabed;
[0022] S2: The deep-sea sediment collection module collects environmental information and posture information in real time and sends it to the surface control module. The surface control module compares the real-time collected posture information with a preset target posture to obtain a posture error.
[0023] S3: Based on the environmental information and attitude error, the surface control module generates a sampling strategy using DDPG, AFC and PID algorithms and converts it into a control command and sends it to the deep-sea sediment collection module;
[0024] S4: The deep-sea sediment collection module adjusts the thrust and direction of the propeller according to the control instruction to achieve precise attitude control, and collects deep-sea sediment samples according to the sampling strategy;
[0025] S5: Repeat steps S2 to S4. After the deep-sea sediment sample collection is completed, the above-water control module controls the lifting and placing module to drive the deep-sea sediment collection module to rise to the water surface.
[0026] Preferably, in step S2, the environmental information includes at least latitude and longitude position, depth, height from the bottom, seawater flow speed and direction, temperature, and salinity information;
[0027] The attitude information at least includes real-time underwater pitch angle, roll angle, yaw angle, angular velocity, linear velocity and acceleration information of the deep-sea sediment acquisition module.
[0028] Preferably, in step S3, the DDPG algorithm is provided with an Actor network and a Critic network, and generates a preliminary sampling strategy by interacting with environmental information;
[0029] The parameters of the DDPG algorithm include: state s t 、Action a t , reward function r t , Q value and sampling strategy π(s t );
[0030] The state s t Contains the environment information and posture information at time t, s t Expressed as:
[0031] s t =[θ t ,ω t , p t , v t , x t , z t ]
[0032] Among them, θ t is the attitude angle vector, ω t is the angular velocity vector, p t is the position vector, v t is the water velocity vector, x t and z t first and second data collected by the sensor;
[0033] The action a t Contains the control measures taken by the deep-sea sediment collection module at time t, a t Expressed as:
[0034] a t =[F x , F y , F z , Δθ, Δφ, Δψ]
[0035] Among them, F x , F y , F z are the first, second and third components of the thrust of the propeller, Δθ, Δφ and Δψ are the adjustment amounts of the roll, pitch and yaw angles respectively;
[0036] The reward function r t Used to evaluate the t The state after s t+1 The good or bad, r t Expressed as:
[0037] r t =-(α*attitude error+β*energy consumption+γ*instability)
[0038] Among them, α, β, γ are the first, second and third weight coefficients;
[0039] The attitude error is expressed as:
[0040] Attitude error = |e φ |+|e θ |+|e ψ |
[0041] e θ =θ target -θ current
[0042] e φ =φ target -φ current
[0043] e ψ =ψ target -ψ current
[0044] Among them, e θ 、e φ and e ψ are the error values of roll angle, pitch angle and yaw angle respectively; θ target and θ current are the target value and current value of the roll angle respectively; φ target and φ current are the target value and current value of the pitch angle respectively; ψ target and ψ current are the target value and current value of the yaw angle respectively;
[0045] The energy consumption is expressed as:
[0046] Energy consumption = F x 2 +F y 2 +F z 2
[0047] The instability is expressed as:
[0048] Instability = |Δe φ |+|Δe θ |+|Δe ψ |
[0049]
[0050] The Actor network is used to map states into deterministic actions a t =π(s t |θ π ), the Critic network is used to evaluate a given state-action pair (st ,a t )’s Q value Q(s) t ,a t |θ Q ), where θ π and θ Q These are the parameters of the Actor network and the Critic network respectively;
[0051] The calculation formula of the Q value is:
[0052] Q(s t ,a t )=E[r t +γQ(s t+1 ,π(s t+1 |θ π ))]
[0053] Where γ is the discount factor, satisfying 0<γ<1;
[0054] The sampling strategy is expressed as:
[0055] π(s t )=arg max a Q(s t ,a)
[0056] By continuously iteratively training and updating the parameters of the Actor network and the Critic network, the sampling strategy under the optimal parameters is obtained and output after the training is completed.
[0057] Preferably, in step S3, the AFC algorithm includes:
[0058] Get the parameter a in the sampling strategy output by the DDPG algorithm t , as well as attitude errors and instabilities;
[0059] According to the preset fuzzy set, the trapezoidal membership function is used to fuzzify the attitude error and instability;
[0060] According to the preset fuzzy rule base, fuzzy reasoning is used to calculate the control quantity, generate the fuzzy control quantity, and transform the fuzzy control quantity into the actual continuous control quantity through defuzzification.
[0061] Preferably, in step S3, the PID algorithm includes:
[0062]
[0063] Among them, u PID (t) is the control quantity output by the PID algorithm at time t; e(t) is the control deviation; K p , K i and Kd are proportional, integral and differential gains respectively;
[0064] In step S3, the actual continuous control quantity U(t) generated by the AFC algorithm and the control quantity u output by the PID algorithm are combined. PID (t), generate the final control instruction u final (t), expressed as:
[0065] u final (t)=U(t)+u PID (t)
[0066] In step S4, the deep-sea sediment collection module collects the sediment according to the final control instruction u final (t) Adjust the thrust and direction of the propeller to achieve precise attitude control.
[0067] Preferably, in step S5, the sampling process also includes real-time feedback and optimization:
[0068] The gain of the PID algorithm is dynamically adjusted through the synergy of the DDPG algorithm and the AFC algorithm, which is expressed as:
[0069] K′ p ←K p +ΔK p
[0070] K′ i ←K i +ΔK i
[0071] K′ d ←K d +ΔK d
[0072] Among them, K′ p , K′ i and K′ d They are proportional, integral and differential gains after dynamic adjustment of PID algorithm; ΔK p , ΔK i , ΔK d are the proportional, integral, and differential gain adjustments determined by the DDPG algorithm and the AFC algorithm, respectively.
[0073] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0074] The present invention provides a deep-sea sediment collection system and method based on adaptive robust control. The present invention combines a Deep Deterministic Policy Gradient (DDPG) algorithm, an Adaptive Fuzzy Controller (AFC) algorithm, and a PID (Proportional Integral Derivative) control algorithm to achieve intelligent comprehensive control of the sediment collection process. The present invention utilizes the DDPG algorithm for efficient learning and adaptive adjustment, uses the AFC algorithm to process nonlinearities and uncertainties in the system, and uses the PID control algorithm for real-time and precise control to ensure the stability and accuracy of the system under different environmental conditions. The sampling strategy is dynamically adjusted based on real-time feedback data, and the position and posture of the thruster and sampling device are optimized to ensure the stability and accuracy of the sampling process. The present invention utilizes advanced mechanical structures and control algorithms to minimize disturbances to the sediment, maintain the original state of the sample, and improve data reliability.
[0075] Compared to existing sediment collection technologies that rely on manual control and a single sampling method, the present invention uses the DDPG algorithm for efficient learning and adaptive adjustment, combined with the AFC algorithm to handle nonlinearity and uncertainty, and then uses the PID algorithm for dynamic control to achieve comprehensive and precise control of the sediment collection process. The system can monitor the sampling process in real time and dynamically adjust the sampling strategy and location based on real-time feedback data, significantly improving sampling efficiency and data reliability. Compared with existing technologies, the present invention significantly enhances its adaptability to complex deep-sea environments and ensures the stability and accuracy of the sampled data.
[0076] Compared with existing sampling technologies that require manual control and monitoring, this invention proposes an intelligent and automated control system for the sampling process, which eliminates the need for human intervention and greatly reduces operational difficulty. The system has a real-time monitoring function, allowing operators to monitor the sampling progress and status in real time through remote monitoring, and intervene and adjust when necessary, improving the user experience. By combining DDPG, AFC, and PID algorithms, the system can adaptively adjust to different environmental conditions, ensuring the efficiency and stability of the sampling process and improving the convenience and reliability of the overall operation.
[0077] The present invention specifically addresses the problem that existing technologies are difficult to meet the needs of sediment sampling in special sea areas. Through the fusion of multiple algorithms, precise control and dynamic adjustment of the sampling process are achieved, ensuring efficient and accurate collection in complex environments. The system can accurately and efficiently collect deep-sea sediment samples, provide reliable data support, and assist in deep-sea environmental monitoring and resource exploration. This will help deepen the understanding of deep-sea ecosystems, promote in-depth related scientific research, and lay the foundation for future resource development and utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 This is a structural diagram of a deep-sea sediment collection system based on adaptive robust control provided in Example 1.
[0079] Figure 2 This is a schematic diagram of the structure of the deep-sea sediment collection module provided in Example 1.
[0080] Figure 3 This is a structural diagram of the water control module provided in Example 1.
[0081] Figure 4 This is a flow chart of a deep-sea sediment collection method based on adaptive robust control provided in Example 2.
[0082] Figure 5 This is a schematic diagram of the DDPG algorithm flow provided in Example 2.
[0083] Figure 6 This is a flow chart of the AFC algorithm provided in Example 2.
[0084] Figure 7 This is a schematic diagram of the PID algorithm flow provided in Example 2. DETAILED DESCRIPTION
[0085] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;
[0086] In order to better illustrate this embodiment, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product size;
[0087] It is understandable to those skilled in the art that some well-known structures and descriptions thereof may be omitted in the drawings.
[0088] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0089] Example 1
[0090] like Figure 1As shown, this embodiment provides a deep-sea sediment collection system based on adaptive robust control, comprising: an above-water control module, a lifting and deployment module, and a deep-sea sediment collection module;
[0091] The above-water control module is respectively connected to the lifting and placing module and the deep-sea sediment collection module in communication; the lifting and placing module is also connected to the deep-sea sediment collection module;
[0092] The above-water control module has built-in DDPG, AFC and PID algorithms. The above-water control module is used to calculate the sampling strategy based on the environmental information and posture information transmitted by the deep-sea sediment acquisition module, and issue instructions to the lifting and deployment module;
[0093] The lifting and deploying module is used to drive the deep-sea sediment collection module to move vertically in the seawater according to the received instructions;
[0094] The deep-sea sediment collection module is used to collect environmental information and posture information in real time and send it to the water control module, and collect deep-sea sediment samples according to the sampling strategy;
[0095] like Figure 2 As shown, the deep-sea sediment collection module includes a sealed cabin, and an underwater control submodule, a sediment collection device, a sensor submodule and a propeller arranged in the sealed cabin;
[0096] The sealed cabin is connected to the lifting and deploying module;
[0097] The underwater control submodule is communicatively connected to the surface control module; the underwater control submodule is also electrically connected to the sediment collection device, the sensor submodule and the propeller respectively;
[0098] The sensor submodule is used to collect environmental information and posture information and send them to the underwater control submodule;
[0099] The underwater control submodule is used to control the propeller and sediment collection device to sample deep-sea sediments according to the sampling strategy, and to send environmental information and posture information to the surface control module;
[0100] The propeller is used to adjust the magnitude and direction of its thrust, thereby adjusting the sampling posture of the sediment collection device, and the sediment collection device is used to perform the sampling action of deep-sea sediments;
[0101] The underwater control submodule includes: a power supply device, a storage device, a processor, a data acquisition device and a data transmission device;
[0102] The power supply device is used to supply power to the deep-sea sediment collection module; the data collection device is used to collect and summarize the environmental information and posture information collected by the sensor submodule; the data transmission device is used to transmit the environmental information and posture information to the storage, processor and water control module; the processor is also used to receive the sampling strategy sent by the water control module and control the propeller and sediment collection device to sample deep-sea sediments;
[0103] The sensor submodule includes at least: a positioning sensor, an inertial sensor, a three-axis magnetometer, a Doppler acoustic current profiler, a multi-beam detector, a side-scan sonar, an altitude sensor, an image sensor, a particle size analyzer, a pressure sensor, a temperature sensor, a salinity sensor, a conductivity sensor, a methane concentration sensor, a carbon dioxide sensor, a dissolved oxygen sensor, a chlorophyll sensor and a pH sensor.
[0104] In a specific implementation process, this embodiment proposes an adaptive robust control deep-sea sediment collection system, which includes an above-water control module, a lifting and deployment module, and a deep-sea sediment collection module; the above-water control module is respectively connected to the lifting and deployment module and the deep-sea sediment collection module via optical cables; the lifting and deployment module is also connected to the deep-sea sediment collection module via a cable;
[0105] In this embodiment, the surface control module has built-in DDPG, AFC and PID algorithms. The surface control module is used to calculate the sampling strategy based on the environmental information and posture information transmitted by the deep-sea sediment acquisition module, and issue instructions to the lifting and deployment module;
[0106] The lifting and deploying module is used to drive the deep-sea sediment collection module to move vertically in the seawater using a cable according to the received instructions;
[0107] The deep-sea sediment collection module is used to collect environmental information and posture information in real time and send it to the surface control module, as well as collect deep-sea sediment samples according to the sampling strategy;
[0108] In this embodiment, the deep-sea sediment collection module includes a pressure-resistant titanium alloy sealed cabin, and an underwater control submodule, a sediment collection device, a sensor submodule and a thruster arranged in the sealed cabin;
[0109] The sealed cabin is connected to the lifting and deploying module;
[0110] The underwater control submodule is communicatively connected to the surface control module; the underwater control submodule is also electrically connected to the sediment collection device, the sensor submodule and the propeller respectively;
[0111] The sensor submodule is used to collect environmental information and attitude information and send it to the underwater control submodule;
[0112] The underwater control submodule is used to control the thruster and sediment collection device to sample deep-sea sediments according to the sampling strategy, and to send environmental information and attitude information to the surface control module;
[0113] The thruster is used to adjust the magnitude and direction of its thrust, thereby adjusting the sampling posture of the sediment collection device. In this embodiment, the sediment collection device is specifically a sediment collection robot, and the control end of the robot is electrically connected to the underwater control submodule to perform the deep-sea sediment sampling action.
[0114] More specifically, the underwater control submodule includes: a power supply device, a storage device, a processor, a data acquisition device, and a data transmission device;
[0115] The power supply device is used to power the deep-sea sediment collection module; the data collection device is used to summarize and collect the environmental information and posture information collected by the sensor submodule; the data transmission device is used to transmit the environmental information and posture information to the storage, processor and water control module; the processor is also used to receive the sampling strategy sent by the water control module and control the propeller and sediment collection device to sample deep-sea sediments;
[0116] In this embodiment, the sensor submodule includes at least: a positioning sensor, an inertial sensor, a three-axis magnetometer, a Doppler acoustic current profiler, a multi-beam sounder, a side-scan sonar, an altitude sensor, an image sensor, a particle size analyzer, a pressure sensor, a temperature sensor, a salinity sensor, a conductivity sensor, a methane concentration sensor, a carbon dioxide sensor, a dissolved oxygen sensor, a chlorophyll sensor, and a pH sensor;
[0117] like Figure 3 As shown, the surface control module in this embodiment further includes: a central processing unit, a central memory, a display, a power supply, and a surface data acquisition and transmission device; the central memory stores the DDPG, AFC, and PID algorithms, and the central processing unit is used to execute the three algorithms based on the environmental information and posture information transmitted by the deep-sea sediment acquisition module, thereby calculating and generating a sampling strategy, and issuing instructions to the lifting and deployment module;
[0118] This system achieves high-precision control of the sediment collection process through the integration of multiple algorithms; by integrating multiple sensors, the system can collect multi-dimensional data in real time and perform efficient data processing to provide accurate posture and environmental information; further combining DDPG, AFC and PID algorithms to generate and execute optimal control strategies to ensure the efficient and stable operation of the sampling system in complex environments; this system can automatically adjust the control strategy according to real-time environmental data to achieve efficient automated operation, improve sampling efficiency and data quality, and at the same time, this system can adjust the posture of the sediment collection device in real time to adapt to the complex and changeable deep-sea environment, realize adaptive robust control of sampling, and ensure the stability of the system during the sample collection process.
[0119] Example 2
[0120] like Figure 4 As shown, this embodiment provides a deep-sea sediment collection method based on adaptive robust control, based on the deep-sea sediment collection system based on adaptive robust control described in Example 1, comprising the following steps:
[0121] S1: The above-water control module controls the lifting and deploying module to drive the deep-sea sediment collection module to dive to the seabed;
[0122] S2: The deep-sea sediment collection module collects environmental information and posture information in real time and sends it to the surface control module. The surface control module compares the real-time collected posture information with a preset target posture to obtain a posture error.
[0123] S3: Based on the environmental information and attitude error, the surface control module generates a sampling strategy using DDPG, AFC and PID algorithms and converts it into a control command and sends it to the deep-sea sediment collection module;
[0124] S4: The deep-sea sediment collection module adjusts the thrust and direction of the propeller according to the control instruction to achieve precise attitude control, and collects deep-sea sediment samples according to the sampling strategy;
[0125] S5: Repeat steps S2 to S4. After the deep-sea sediment sample collection is completed, the above-water control module controls the lifting and deployment module to drive the deep-sea sediment collection module to rise to the water surface;
[0126] In step S2, the environmental information includes at least latitude and longitude, depth, height from the bottom, seawater flow speed and direction, temperature, and salinity information;
[0127] The attitude information includes at least the real-time underwater pitch angle, roll angle, yaw angle, angular velocity, linear velocity and acceleration information of the deep-sea sediment acquisition module;
[0128] In step S3, the DDPG algorithm is configured with an Actor network and a Critic network, and generates a preliminary sampling strategy by interacting with environmental information;
[0129] The parameters of the DDPG algorithm include: state s t 、Action a t , reward function r t , Q value and sampling strategy π(s t );
[0130] The state s t Contains the environment information and posture information at time t, s t Expressed as:
[0131] s t =[θ t ,ω t , p t , v t , x t , z t ]
[0132] Among them, θ t is the attitude angle vector, ω t is the angular velocity vector, p t is the position vector, v t is the water velocity vector, x t and z t first and second data collected by the sensor;
[0133] The action a t Contains the control measures taken by the deep-sea sediment collection module at time t, a t Expressed as:
[0134] a t =[F x , F y , F z , Δθ, Δφ, Δψ]
[0135] Among them, F x , F y , F z are the first, second and third components of the thrust of the propeller, Δθ, Δφ and Δψ are the adjustment amounts of the roll, pitch and yaw angles respectively;
[0136] The reward function r t Used to evaluate the t The state after s t+1 The good or bad, r t Expressed as:
[0137] r t=-(α*attitude error+β*energy consumption+γ*instability)
[0138] Among them, α, β, γ are the first, second and third weight coefficients;
[0139] The attitude error is expressed as:
[0140] Attitude error = |e φ |+|e θ |+|e ψ |
[0141] e θ =θ target -θ current
[0142] eφ=φ target -φ current
[0143] eψ=ψ target -ψ current
[0144] Among them, e θ 、e φ and e ψ are the error values of roll angle, pitch angle and yaw angle respectively; θ target and θ current are the target value and current value of the roll angle respectively; φ target and φ current are the target value and current value of the pitch angle respectively; φ target and ψ current are the target value and current value of the yaw angle respectively;
[0145] The energy consumption is expressed as:
[0146] Energy consumption = F x 2 +F y 2 +F z 2
[0147] The instability is expressed as:
[0148] Instability = |Δe φ |+|Δe θ |+|Δe ψ |
[0149]
[0150] The Actor network is used to map states into deterministic actions a t =π(s t |θ π), the Critic network is used to evaluate a given state-action pair (s t ,a t )’s Q value Q(s) t ,a t |θ Q ), where θ π and θ Q These are the parameters of the Actor network and the Critic network respectively;
[0151] The calculation formula of the Q value is:
[0152] Q(s t ,a t )=E[r t +γQ(s t+1 ,π(s t+1 |θ π ))]
[0153] Where γ is the discount factor, satisfying 0<γ<1;
[0154] The sampling strategy is expressed as:
[0155] π(s t )=argmax a Q(s t ,a)
[0156] By continuously iteratively training and updating the parameters of the Actor network and the Critic network, the sampling strategy under the optimal parameters is obtained and output after the training is completed;
[0157] In step S3, the AFC algorithm includes:
[0158] Get the parameter a in the sampling strategy output by the DDPG algorithm t , as well as attitude errors and instabilities;
[0159] According to the preset fuzzy set, the trapezoidal membership function is used to fuzzify the attitude error and instability;
[0160] According to the preset fuzzy rule base, fuzzy reasoning is used to calculate the control quantity, generate the fuzzy control quantity, and transform the fuzzy control quantity into the actual continuous control quantity through defuzzification;
[0161] In step S3, the PID algorithm includes:
[0162]
[0163] Among them, u PID (t) is the control quantity output by the PID algorithm at time t; e(t) is the control deviation; K p , Ki and K d are proportional, integral and differential gains respectively;
[0164] In step S3, the actual continuous control quantity U(t) generated by the AFC algorithm and the control quantity u output by the PID algorithm are combined. PID (t), generate the final control instruction u final (t), expressed as:
[0165] u final (t)=U(t)+u PID (t)
[0166] In step S4, the deep-sea sediment collection module collects the sediment according to the final control instruction u final (t) Adjust the thrust and direction of the propeller to achieve precise attitude control;
[0167] In step S5, the sampling process also includes real-time feedback and optimization:
[0168] The gain of the PID algorithm is dynamically adjusted through the synergy of the DDPG algorithm and the AFC algorithm, which is expressed as:
[0169] K′ p ←K p +ΔK p
[0170] K′ i ←K i +ΔK i
[0171] K′ d ←K d +ΔK d
[0172] Among them, K′ p , K′ i and K′ d They are proportional, integral and differential gains after dynamic adjustment of PID algorithm; ΔK p , ΔK i , ΔK d are the proportional, integral, and differential gain adjustments determined by the DDPG algorithm and the AFC algorithm, respectively.
[0173] In the specific implementation process, this embodiment provides a method for deep-sea sediment collection based on adaptive robust control, which is implemented based on the deep-sea sediment collection system in Example 1, and the specific steps are as follows:
[0174] First, the environmental information and attitude information of the sediment collection system are collected in real time. The environmental information includes latitude and longitude position, depth, height from the bottom, seawater flow speed and direction, temperature, salinity and other information. The attitude information includes the underwater real-time pitch angle, roll angle, yaw angle, angular velocity, linear velocity, acceleration and other information of the collection system.
[0175] Then, the environmental information collected in real time and the posture information of the sediment collection system are compared with the set posture of the sediment collection system to obtain the error between the set posture and the set posture;
[0176] After that, the control strategy is generated. This method uses DDPG, AFC and PID control algorithms to generate the attitude adjustment strategy, and loads all parameter data into the obtained model for calculation;
[0177] After that, the control instructions are executed and the thrust of the thruster is adjusted according to the control strategy to achieve precise control of the system attitude;
[0178] Finally, real-time feedback and optimization are performed. By real-time monitoring of system status, feedback data is generated, and DDPG, AFC, and PID control algorithms are continuously optimized to ensure efficient and stable operation of the system.
[0179] like Figure 5 The figure shows the flow chart of DDPG. In this embodiment, the DDPG algorithm is used to optimize the sampling strategy and posture control through interaction with the environment. The parameters of the algorithm include state s t 、Action a t , reward function r t , Q-value update and strategy π;
[0180] The DDPG algorithm includes an actor network and a critic network. The hidden layers of both networks use ReLU as the activation function. The output layer of the actor network uses Tanh as the activation function, and the output layer of the critic network uses a linear activation function. After the action is output, random exploration noise Euler-Markov noise (Ornstein-Uhlenbeck Noise, OU Noise) is added to increase the exploration probability in the early stages of training. The maximum number of training rounds is 2000, and the maximum number of steps per round is 1000. The target network uses a soft update mode to gradually update the parameters of the training network to the target network. The soft update coefficient tau is set to 0.001 to ensure smooth update of the target network parameters.
[0181] In this embodiment, the state s tContains environmental information and attitude information at time t, including current attitude angles (pitch angle, roll angle, heading angle), current attitude angular velocity (three components of angular velocity), current system position (depth, height from the bottom, etc.), current environmental parameters (water flow velocity, water flow direction, etc.), sensor data (temperature, salinity, etc.), s t Expressed as:
[0182] s t =[θ t ,ω t , p t , v t , x t , z t ]
[0183] Among them, θ t is the attitude angle vector, ω t is the angular velocity vector, p t is the position vector, v t is the water velocity vector, x t and z t first and second data collected by the sensor;
[0184] Action a t Contains the control measures taken by the deep-sea sediment collection module at time t, including the magnitude and direction adjustment of the thrust of the propeller, and the attitude adjustment instructions, a t Expressed as:
[0185] a t =[F x , F y , F z , Δθ, Δφ, Δψ]
[0186] Among them, F x , F y , F z are the first, second and third components of the thrust of the propeller, Δθ, Δφ and Δψ are the adjustment amounts of the roll, pitch and yaw angles respectively;
[0187] Reward function r t Used to evaluate the t The state after s t+1 The goal is to maximize the cumulative reward. The reward function is designed based on the deviation degree of the target posture, energy consumption, stability and safety. t Expressed as:
[0188] r t =-(α*attitude error+β*energy consumption+γ*instability)
[0189] Among them, α, β, γ are the first, second and third weight coefficients;
[0190] Attitude error (using e n represents the error between the target posture and the current posture, which is expressed as:
[0191] Attitude error = |e φ |+|e θ |+|e ψ |
[0192] e θ =θ target -θ current
[0193] eφ=φ target -φ current
[0194] eψ=ψ target -ψ current
[0195] Among them, e θ 、e φ and e ψ are the error values of roll angle, pitch angle and yaw angle respectively; θ target and θ current are the target value and current value of the roll angle respectively; φ target and φ current are the target value and current value of the pitch angle respectively; ψ target and ψ current are the target value and current value of the yaw angle respectively;
[0196] Energy consumption is expressed as:
[0197] Energy consumption = F x 2 +F y 2 +F z 2
[0198] Instability (using Δe n Expressed as, that is, the error change rate) is expressed as:
[0199] Instability = |Δe φ |+|Δe θ |+|Δe ψ |
[0200]
[0201] In this embodiment, the Actor network is used to map the state into a deterministic action a t =π(s t |θπ ), the hidden layer uses ReLU activation function, and the output layer uses Tanh activation function to ensure that the action output range is bounded; the Critic network is used to evaluate the given state-action pair (s t ,a t )’s Q value Q(s) t ,a t |θ Q ), the hidden layer uses the ReLU activation function, and the output layer uses the linear activation function to obtain an estimate of any Q value range; where θ π and θ Q These are the parameters of the Actor network and the Critic network respectively;
[0202] The parameter update process is as follows:
[0203] Use the replay buffer (R) to store the past state, action, reward, and next state data; after each time step, a batch of data (s i ,a i ,r i ,s i+1 ,done i ) is used to update network parameters;
[0204] Critic Network Update:
[0205] First, use the target network to calculate the target Q value:
[0206] y i =r i +γQ′(s i+1 ,π′(s i+1 |θ π′ )|θ Q′ )
[0207] Where Q′ and π′ are the Critic network and the target Actor network respectively, γ is the discount factor, 0<γ<1;
[0208] The loss function of the Critic network is:
[0209]
[0210] Use gradient descent to minimize the loss function:
[0211]
[0212] Among them, α C is the learning rate of the Critic network;
[0213] Update the Actor network using deterministic policy gradients:
[0214]
[0215] Update Actor parameters using gradient ascent:
[0216]
[0217] Among them, α A is the learning rate of the Actor network;
[0218] The target network parameters adopt a soft update strategy, slowly moving closer to the current training network parameters to ensure the stability of training:
[0219] θ Q′ ←τθ Q +(1-τ)θ Q′
[0220] θ π′ ←τθ π +(1-τ)θ π′
[0221] Among them, τ is the soft update coefficient;
[0222] The relationship between behavior strategy and value function uses the Q-value function Bellman equation:
[0223] Q(s t ,a t )=E[r t +γQ(s t+1 ,π(s t+1 |θ π ))]
[0224] Among them, the strategy π is deterministic and directly maps the state to the action. The optimal strategy is:
[0225] π(s t )=ar gmax a Q(s t ,a)
[0226] like Figure 6 The following is the fuzzy logic control process of AFC: the control strategy parameter a generated by DDPG is t Perform fuzzy processing, use fuzzy reasoning to calculate the control quantity according to the fuzzy rule base, and generate the fuzzy control quantity u; convert the fuzzy control quantity u into the actual control quantity U through defuzzification;
[0227] Specifically, for each input variable, a set of fuzzy sets is defined, and each fuzzy set is described using a trapezoidal membership function. For each input variable (error e n and error change rate Δen ), define the following 7 fuzzy sets:
[0228] NB (Negative Big), NM (Negative Medium), NS (Negative Small), Z (Zero), PS (Positive Small), PM (Positive Medium), PB (Positive Big); for fuzzy set A i , its membership function μ Ai (x) is:
[0229]
[0230] Among them, a, b, c, and d are the four key points of the trapezoidal membership function;
[0231] Construct a fuzzy rule base, specify fuzzy control rules based on expert experience and system characteristics, and convert the error e n and error change rate Δe n The fuzzy sets are arranged into rows and columns to form a two-dimensional table, where each cell corresponds to an output fuzzy set U; the constructed fuzzy rule table is shown in Table 1:
[0232] Table 1 Fuzzy rules table
[0233]
[0234] In Table 1, row (left): represents the error e n Fuzzy sets, from NB to PB;
[0235] Column (top): represents the error change rate Δe n Fuzzy sets, from NB to PB;
[0236] Cell content: fuzzy set representing the output control quantity U;
[0237] Then, fuzzy reasoning is performed: for each rule, its activation strength is calculated, the corresponding fuzzy output is generated, and the fuzzy outputs of all rules are aggregated; among them, for the kth rule, the activation strength is:
[0238]
[0239] in, is the input variable a t In the fuzzy set A i The degree of membership under
[0240] Apply the activation strength to the corresponding output set fuzzy set B j :
[0241]
[0242] in, Is the output quantity in fuzzy set B j The membership function under
[0243] Aggregate the fuzzy outputs of all rules using the "maximum method":
[0244]
[0245] Use the centroid method to defuzzify, take the membership function of the fuzzy set as the mass distribution function, and determine the final output by calculating the centroid of the mass distribution:
[0246]
[0247] Among them, u i is the discrete control output value; N is the number of discrete points; μ output (u i ) is the corresponding membership value;
[0248] like Figure 7 The control flow of the PID algorithm is shown. The PID controller output u is calculated based on the control error. PID (t):
[0249]
[0250] Where, e(t) = SP-PV(t) is the control deviation (the difference between the set point and the current process variable); K p , K i , K d are proportional, integral, and derivative gains respectively;
[0251] For attitude control, the PID control outputs are calculated for the pitch angle (φ), roll angle (θ), and yaw angle (ψ):
[0252]
[0253] Calculate the control gain parameters according to AFC adjustment PID parameters:
[0254]
[0255] Among them, FuzzyController represents fuzzy control;
[0256] Thruster control command calculation:
[0257] Thrust φ=CalculateThrust(u φ )
[0258] Thrust θ =CalculateThrust(u θ )
[0259] Thrust ψ =CalculateThrust(u ψ )
[0260] Among them, CalculateThrust represents thrust calculation; Thrust represents thrust;
[0261] In order to adapt to the discrete thrust levels of the thruster, the thrust command of the thruster is quantized into a quantized :
[0262]
[0263] Among them, a i is the possible thrust level; u(a i ) is the membership degree of the corresponding thrust level;
[0264] Final control command generation: Combine the control quantity U generated by AFC and PID output u PID (t), generate the final control instruction u final (t):
[0265] u final (t)=U(t)+u PID (t)
[0266] Execute control instructions, according to the final control instruction u final (t), adjust the thrust of the thrusters to achieve precise control of the system attitude;
[0267] For real-time feedback and optimization, this embodiment dynamically adjusts the gain and fuzzy rule parameters of the PID controller through the synergistic effect of DDPG and AFC. Under the feedback loop:
[0268] K′ p ←K p +ΔK p
[0269] K′ i ←K i +ΔK i
[0270] K′ d ←K d +ΔK d
[0271] Among them, K′ p , K′ i and K′ d They are proportional, integral and differential gains after dynamic adjustment of PID algorithm; ΔK p , ΔK i , ΔK d are the proportional, integral, and differential gain adjustments determined by the DDPG algorithm and the AFC algorithm, respectively;
[0272] This method innovatively integrates multiple algorithms to achieve intelligent control and optimization of the sediment collection process, improving sampling quality and efficiency.
[0273] Example 3
[0274] This embodiment provides a deep-sea sediment collection method based on adaptive robust control, based on the deep-sea sediment collection system based on adaptive robust control described in Example 1, comprising the following steps:
[0275] S1: The above-water control module controls the lifting and deploying module to drive the deep-sea sediment collection module to dive to the seabed;
[0276] S2: The deep-sea sediment collection module collects environmental information and posture information in real time and sends it to the surface control module. The surface control module compares the real-time collected posture information with a preset target posture to obtain a posture error.
[0277] S3: Based on the environmental information and attitude error, the surface control module generates a sampling strategy using DDPG, AFC and PID algorithms and converts it into a control command and sends it to the deep-sea sediment collection module;
[0278] S4: The deep-sea sediment collection module adjusts the thrust and direction of the propeller according to the control instruction to achieve precise attitude control, and collects deep-sea sediment samples according to the sampling strategy;
[0279] S5: Repeat steps S2 to S4. After the deep-sea sediment sample collection is completed, the above-water control module controls the lifting and placing module to drive the deep-sea sediment collection module to rise to the water surface.
[0280] In the specific implementation process, the water control module is first used to control the lifting and deployment module to drive the deep-sea sediment collection module to dive to the seabed (this example sets the mission scene in the deep-sea cold spring area, which can efficiently and accurately collect deep-sea sediment samples in the area), and collect the environmental information and posture information of the deep-sea sediment collection module: including state s t 、Action is a t :
[0281] s t =[θ t ,ω t , p t , v t , x t , z t ]
[0282] a t =[F x , F y , F z , Δθ, Δφ, Δψ]
[0283] Compare the real-time collected environmental information and posture information with the set posture information to obtain the error e n and Δe n ;
[0284] Generate attitude adjustment strategies using DDPG, AFC, and PID control algorithms;
[0285] The strategy is trained and updated using the DDPG algorithm. In this embodiment, the neural network structure is as follows: the Actor and Critic networks each have two hidden layers, each with 256 neurons, the hidden layer activation function is ReLU, the Actor output layer is Tanh, and the Critic output layer is linear activation; reinforcement learning parameters: γ = 0.99, α A =10 -4 , α C =10 -3 , τ = 0.001; OU noise parameters: μ = 0, θ = 0.15, σ = 0.2, to increase the initial exploration of the action; Among them, the reward function r t =-(α*attitude error + β*energy consumption + γ*instability), set α=1.0, β=0.01, γ=0.5; where attitude error = |e φ |+|e θ |+|e ψ |, if the current total attitude error is 1°, the energy consumption is F x 2 +F y 2 +F z 2 =400 (e.g. 20N thrust is 20 2 =400), and the instability (total error rate of change) is about 1, then:
[0286] r t ≈-(1.0×1+0.01×400+0.5×1)=-5.5
[0287] As training progresses, DDPG tends to reduce error and energy consumption, thereby increasing cumulative rewards. After training, the DDPG strategy enables the sediment collection robot module to maintain a posture deviation of less than 1° under most water flow disturbances in simulation.
[0288] Introducing the AFC control algorithm, the input quantity e n is the attitude error (±5°), Δe n is the error change rate (±2° / s); define 7 fuzzy sets: NB, NM, NS, Z, PS, PM, PB, and use trapezoidal membership function to describe them; construct a fuzzy rule table based on expert experience, such as when e n =NS (slightly positive small error), Δe n =NM (moderate decrease in error), then the output U = PS (slightly positive), indicating that the control strength should be slightly increased;
[0289] All rule outputs are aggregated through fuzzy reasoning and the maximum method, and then defuzzified using the center of gravity method. Assuming the final defuzzification result is U = 2.0, it corresponds to an increase of approximately 10N of thrust. AFC flexibly adjusts the DDPG output action to reduce overshoot and oscillation.
[0290] Then, the PID gain is adjusted adaptively, and the initial PID parameters are set to: K p =1.0,K i =0.1, K d =0.5; Gain adaptive strategy: When U>0, it indicates that a faster response is required, then ΔK p =0.01U,ΔK i =0.001U,ΔK d =0.005U; if U=2.0, then K p ←1.0+0.02=1.02K i ←0.1+0.002=0.102K d ←0.5+0.01=0.51;
[0291] When the system error suddenly increases, the PID gain increases to respond faster. After the error stabilizes, U decreases and the gain decreases to avoid over-regulation. The final control instruction is generated and executed: the DDPG action is combined with the AFC output U, and the updated PID output u PID (t):
[0292] u final (t)=U(t)+u PID (t)
[0293] Assuming the DDPG output is a recommended thrust of 15N in the x-direction, the PID output is an additional 5N, and the AFC gives U = 2.0 (≈10N), the total is 30N. Based on the discrete gear, 30N is rounded up to the 20N gear to meet the actual discrete characteristics of the thruster.
[0294] The control strategy then adjusts the thrust and direction of the robot to achieve precise posture control and begin sediment collection. The sediment collection robot executes the command, achieving precise posture and position control. When water disturbances increase, the system automatically increases the gain and U value to quickly reduce errors. When the environment stabilizes, the gain and U value are reduced to maintain low energy consumption and low oscillation.
[0295] Implementing feedback and dynamic optimization: The deep-sea sediment collection module continuously monitors its status throughout the sampling process. When a micro-perturbation occurs while the deep-sea sediment collection module contacts the seabed sediment, such as an error increasing from 0.5° to 2°, the AFC and PID adaptive control immediately respond, increasing control strength and reducing the error to within ±1° within seconds, keeping the sediment collection robot module in a stable hover and improving sampling success rate and sample quality.
[0296] The above process is repeated, and when the sediment sample collection is completed, the sediment collection robot module returns to the water surface;
[0297] This method achieves high-precision control of the sediment collection process through the integration of multiple algorithms; this method can automatically adjust the control strategy according to real-time environmental data, realize efficient automated operation, improve sampling efficiency and data quality, and adapt to the complex and changeable deep-sea environment; in addition, this method demonstrates higher robustness, accuracy and automation level in the deep-sea sediment collection process, and can provide reliable technical support for marine scientific research.
[0298] The same or similar reference numerals correspond to the same or similar components;
[0299] The terms used in the drawings to describe positional relationships are for illustrative purposes only and should not be construed as limiting this patent;
[0300] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.
Claims
1. A deep-sea sediment collection system based on adaptive robust control, characterized in that: include: Surface control module, lifting and deployment module, and deep-sea sediment collection module; The above-water control module is respectively connected to the lifting and placing module and the deep-sea sediment collection module in communication; the lifting and placing module is also connected to the deep-sea sediment collection module; The above-water control module has built-in DDPG, AFC and PID algorithms. The above-water control module is used to calculate the sampling strategy based on the environmental information and posture information transmitted by the deep-sea sediment acquisition module, and issue instructions to the lifting and deployment module; The lifting and deploying module is used to drive the deep-sea sediment collection module to move vertically in the seawater according to the received instructions; The deep-sea sediment collection module is used to collect environmental information and posture information in real time and send them to the above-water control module, and to collect samples of deep-sea sediments according to a sampling strategy.
2. A deep-sea sediment collection system based on adaptive robust control according to claim 1, characterized in that: The deep-sea sediment collection module includes a sealed cabin, and an underwater control submodule, a sediment collection device, a sensor submodule and a propeller arranged in the sealed cabin; The sealed cabin is connected to the lifting and deploying module; The underwater control submodule is communicatively connected to the surface control module; the underwater control submodule is also electrically connected to the sediment collection device, the sensor submodule and the propeller respectively; The sensor submodule is used to collect environmental information and posture information and send them to the underwater control submodule; The underwater control submodule is used to control the propeller and sediment collection device to sample deep-sea sediments according to the sampling strategy, and to send environmental information and posture information to the surface control module; The thruster is used to adjust the magnitude and direction of its thrust, thereby adjusting the sampling posture of the sediment collection device, and the sediment collection device is used to perform the sampling action of deep-sea sediments.
3. A deep-sea sediment collection system based on adaptive robust control according to claim 2, characterized in that: The underwater control submodule includes: a power supply device, a storage device, a processor, a data acquisition device and a data transmission device; The power supply device is used to power the deep-sea sediment collection module; the data collection device is used to summarize and collect the environmental information and posture information collected by the sensor submodule; the data transmission device is used to transmit the environmental information and posture information to the storage, processor and water control module; the processor is also used to receive the sampling strategy sent by the water control module and control the thruster and sediment collection device to sample deep-sea sediments.
4. The deep-sea sediment collection system based on adaptive robust control according to claim 2, characterized in that: The sensor submodule includes at least: a positioning sensor, an inertial sensor, a three-axis magnetometer, a Doppler acoustic current profiler, a multi-beam detector, a side-scan sonar, an altitude sensor, an image sensor, a particle size analyzer, a pressure sensor, a temperature sensor, a salinity sensor, a conductivity sensor, a methane concentration sensor, a carbon dioxide sensor, a dissolved oxygen sensor, a chlorophyll sensor and a pH sensor.
5. A method for collecting deep-sea sediments based on adaptive robust control, based on the deep-sea sediment collection system based on adaptive robust control according to any one of claims 1 to 4, characterized in that: The following steps are involved: S1: The above-water control module controls the lifting and deploying module to drive the deep-sea sediment collection module to dive to the seabed; S2: The deep-sea sediment collection module collects environmental information and posture information in real time and sends it to the surface control module. The surface control module compares the real-time collected posture information with a preset target posture to obtain a posture error. S3: Based on the environmental information and attitude error, the surface control module generates a sampling strategy using DDPG, AFC and PID algorithms and converts it into a control command and sends it to the deep-sea sediment collection module; S4: The deep-sea sediment collection module adjusts the thrust and direction of the propeller according to the control instruction to achieve precise attitude control, and collects deep-sea sediment samples according to the sampling strategy; S5: Repeat steps S2 to S4. After the deep-sea sediment sample collection is completed, the above-water control module controls the lifting and placing module to drive the deep-sea sediment collection module to rise to the water surface.
6. The method for collecting deep-sea sediments based on adaptive robust control according to claim 5, characterized in that: In step S2, the environmental information includes at least latitude and longitude, depth, height from the bottom, seawater flow speed and direction, temperature, and salinity information; The attitude information at least includes real-time underwater pitch angle, roll angle, yaw angle, angular velocity, linear velocity and acceleration information of the deep-sea sediment acquisition module.
7. The method for collecting deep-sea sediments based on adaptive robust control according to claim 5, characterized in that: In step S3, the DDPG algorithm is configured with an Actor network and a Critic network, and generates a preliminary sampling strategy by interacting with environmental information; The parameters of the DDPG algorithm include: state s t 、Action a t , reward function r t , Q value and sampling strategy π(s t ); The state s t Contains the environment information and posture information at time t, s t Expressed as: s t =[θ t ,ω t ,p t ,v t ,x t ,z t ] Among them, θ t is the attitude angle vector, ω t is the angular velocity vector, p t is the position vector, v t is the water velocity vector, x t and z t first and second data collected by the sensor; The action a t Contains the control measures taken by the deep-sea sediment collection module at time t, a t Expressed as: a t =[F x ,F y ,F z , Δθ, Δφ, Δψ] Among them, F x , F y , F z are the first, second and third components of the thrust of the propeller, Δθ, Δφ and Δψ are the adjustment amounts of the roll, pitch and yaw angles respectively; The reward function r t Used to evaluate the t The state after s t+1 The good or bad, r t Expressed as: r t =-(α*attitude error+β*energy consumption+γ*instability) Among them, α, β, γ are the first, second and third weight coefficients; The attitude error is expressed as: Attitude error = |e φ |+|e θ |+|e ψ | e θ =θ target -θ current e φ =φ target -f current e ψ =ψ target -ψ current Among them, e θ 、e φ and e ψ are the error values of roll angle, pitch angle and yaw angle respectively; θ target and θ current are the target value and current value of the roll angle respectively; φ target and φ current are the target value and current value of the pitch angle respectively; ψ target and ψ current are the target value and current value of the yaw angle respectively; The energy consumption is expressed as: Energy consumption = F x 2 +F y 2 +F z 2 The instability is expressed as: Instability = |Δe φ |+|Δe θ |+|Δe ψ | The Actor network is used to map states into deterministic actions t =π(a t |θ π ), the Critic network is used to evaluate a given state-action pair (s t ,a t )’s Q value Q(s) t ,a t |θ Q ), where θ π and θ Q These are the parameters of the Actor network and the Critic network respectively; The calculation formula of the Q value is: Q(s t ,a t )=E[r t +γQ(s t+1 ,π(s t+1 |θ π ))] Where γ is the discount factor, satisfying 0<γ<1; The sampling strategy is expressed as: π(s t )=argmax a Q(s t ,a) By continuously iteratively training and updating the parameters of the Actor network and the Critic network, the sampling strategy under the optimal parameters is obtained and output after the training is completed.
8. The method for collecting deep-sea sediments based on adaptive robust control according to claim 7, characterized in that: In step S3, the AFC algorithm includes: Get the parameter a in the sampling strategy output by the DDPG algorithm t , as well as attitude errors and instabilities; According to the preset fuzzy set, the trapezoidal membership function is used to fuzzify the attitude error and instability; According to the preset fuzzy rule base, fuzzy reasoning is used to calculate the control quantity, generate the fuzzy control quantity, and transform the fuzzy control quantity into the actual continuous control quantity through defuzzification.
9. The method for collecting deep-sea sediments based on adaptive robust control according to claim 8, characterized in that: In step S3, the PID algorithm includes: Among them, u PID (t) is the control quantity output by the PID algorithm at time t; e(t) is the control deviation; K p , K i and K d are proportional, integral and differential gains respectively; In step S3, the actual continuous control quantity U(t) generated by the AFC algorithm and the control quantity u output by the PID algorithm are combined. PID (t), generate the final control instruction u final (t), expressed as: u final (t)=U(t)+u PID (t) In step S4, the deep-sea sediment collection module collects the sediment according to the final control instruction u final (t) Adjust the thrust and direction of the propeller to achieve precise attitude control.
10. The method for collecting deep-sea sediments based on adaptive robust control according to claim 9, characterized in that: In step S5, the sampling process also includes real-time feedback and optimization: The gain of the PID algorithm is dynamically adjusted through the synergy of the DDPG algorithm and the AFC algorithm, which is expressed as: K′ p ←K p +ΔK p K′ i ←K i +ΔK i K′ d ←K d +ΔK d Among them, K′ p , K′ i and K′ d They are proportional, integral and differential gains after dynamic adjustment of PID algorithm; ΔK p , ΔK i , ΔK d are the proportional, integral, and differential gain adjustments determined by the DDPG algorithm and the AFC algorithm, respectively.