DPPO reinforcement learning-based dredging mechanical arm intelligent control method and system

Through the intelligent control method of dredging robot arm based on DPPO reinforcement learning, data is collected in real time, life expectancy is predicted and maintenance strategies is dynamically adjusted, which solves the problems of traditional maintenance efficiency and high cost, and achieves efficient and economical dredging operations and equipment maintenance.

CN120170720AActive Publication Date: 2025-06-20HUANJIAN ECOLOGICAL RESTORATION (BEIJING) CO LTD

Patent Information

Application Number
CN202510592340.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-20
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The maintenance of traditional dredging robotic arms has problems of low efficiency and high cost, lacks real-time monitoring methods, cannot accurately predict the remaining life of the equipment, and the maintenance strategy is not optimized enough.

Method used

The intelligent control method of dredging robotic arm based on DPPO reinforcement learning is adopted. Through hardware-algorithm collaborative architecture, dynamic constraint processing and industrial-level deployment, robotic arm data is collected in real time, the remaining service life is predicted, maintenance strategies are dynamically adjusted, and control strategies are optimized through reward functions.

Benefits of technology

It has achieved improvements in dredging operation efficiency, reduced maintenance costs, extended the service life of the equipment, and solved the problems of maintenance lag, energy efficiency imbalance and infeasible instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120170720A_ABST
    Figure CN120170720A_ABST
Patent Text Reader

Abstract

The invention discloses a dredging mechanical arm intelligent control method and system based on DPPO reinforcement learning, and relates to the technical field of dredging engineering equipment intellectualization. The method comprises the steps that hydraulic pressure, vibration spectrum and temperature data of a dredging and desilting mechanical arm of a dredger are collected in real time and preprocessed; the preprocessed data are input into an LSTM-Attention model, feature extraction is carried out based on a long short-term memory network and an attention mechanism, and a predicted value of the remaining service life of the mechanical arm is output; a degradation-efficiency coupling model is established, and the health index and the operation efficiency of the mechanical arm are calculated based on the predicted value of the remaining service life of the mechanical arm; and a DPPO reinforcement learning controller is established, the mechanical arm health index, the operation efficiency, the energy consumption and the mechanical arm joint angle serve as a state space, an action space of a dynamic load adjusting instruction is output, a control strategy is optimized through a reward function, and the load adjusting instruction is subjected to dynamic constraint processing and then drives a mechanical arm hydraulic executing mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent dredging engineering equipment, and more specifically, to an intelligent control method and system for a dredging robotic arm based on DPPO reinforcement learning. Background Art

[0002] Dredging robotic arms are used for dredging operations in ports, river channels, etc. They belong to heavy machinery, operate in harsh environments, and are difficult to maintain. Traditional maintenance methods may be regular maintenance or repair after a failure. Both of these methods have problems of low efficiency and high cost. Regular maintenance may lead to over-maintenance or under-maintenance, while repair after a failure will cause downtime losses and safety risks.

[0003] The existing technology lacks effective real-time monitoring means, cannot accurately predict the remaining life of the equipment, and cannot dynamically adjust the maintenance strategy. In addition, traditional prediction models may not consider the mutual influence between equipment degradation and operation efficiency, resulting in inaccurate predictions or sub-optimal maintenance strategies. Therefore, the traditional maintenance of dredging vessel robotic arms has the following limitations: 1. The lag of the traditional maintenance mode: Dredging robotic arms usually adopt regular maintenance or repair after a failure (passive maintenance), which cannot predict the equipment degradation trend, easily leading to sudden failure shutdowns, operation interruptions, or high repair costs. 2. The separate analysis of equipment degradation and operation efficiency: Existing methods only focus on a single dimension of equipment degradation (such as component wear) or operation efficiency (such as task completion rate), without quantifying the dynamic coupling relationship between the two, making it difficult to formulate a maintenance strategy that takes into account both efficiency and reliability. 3. Insufficient prediction accuracy under complex working conditions: In dredging operations, multi-parameter non-linear coupling such as hydraulic pressure, vibration, and temperature makes it difficult for traditional statistical models (such as linear regression) or single LSTM networks to capture key degradation characteristics, resulting in large prediction errors for the remaining useful life (RUL).

[0004] Therefore, how to provide an intelligent control method for a dredging robotic arm based on DPPO reinforcement learning to solve the above drawbacks, improve the efficiency of dredging operations, reduce costs, and extend the service life of the equipment is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides an intelligent control method and system for a dredging robotic arm based on DPPO reinforcement learning. The present invention provides an efficient intelligent control solution for heavy dredging equipment through a hardware-algorithm collaborative architecture (perception-decision-execution closed loop), dynamic constraint processing (action safety mapping), and industrial-level deployment (TensorRT acceleration, OPC UA protocol).

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The intelligent control method of dredging robot arm based on DPPO reinforcement learning includes:

[0008] Collect the hydraulic pressure, vibration spectrum and temperature data of the dredging robot arm in real time and perform pre-processing;

[0009] The preprocessed data is input into the LSTM-Attention model, and features are extracted based on the long short-term memory network and attention mechanism to output the predicted value of the remaining service life of the robot arm;

[0010] Establish a degradation-performance coupling model, and calculate the robot arm health index and operating performance based on the predicted value of the remaining service life of the robot arm;

[0011] A DPPO reinforcement learning controller is established, which takes the robot health index, working efficiency, energy consumption and robot joint angle as the state space, outputs the action space of dynamic load adjustment instructions, and optimizes the control strategy through the reward function. The load adjustment instructions are processed by dynamic constraints to drive the hydraulic actuator of the robot.

[0012] Preferably, the pretreatment comprises:

[0013] Missing value processing: Lagrange interpolation method is used to fill missing data;

[0014] Outlier processing: Tukey's Test method was used to test outliers and eliminate error values;

[0015] Normalization: Min-Max normalization is used to normalize the collected multi-dimensional parameters;

[0016] Substitute the vibration spectrum after the above preprocessing into the formula:

[0017]

[0018] Among them, V rms is the root mean square value of the vibration velocity at time t, and N is the sample selected during calculation.

[0019] Number of points, v i is the instantaneous vibration velocity measurement value of the i-th sampling point;

[0020]

[0021] Among them, F peak (t) is the maximum amplitude of the vibration signal in the 1-5kHz frequency band at time t, is the spectrum obtained by Fourier transforming the vibration velocity signal v(t), max(...) 1-5kHz It is the maximum amplitude value in the frequency range of 1-5kHz.

[0022] Preferably, the preprocessed data is input into the LSTM-Attention model, and feature extraction is performed based on the long short-term memory network and the attention mechanism to output the predicted remaining service life of the robotic arm, including:

[0023] Receiving the preprocessed multi-source time series data through the input layer, including hydraulic pressure, the root mean square value V of vibration velocity rms , the maximum vibration amplitude F peak and temperature parameters;

[0024] And sequentially extracting local time series features through the convolutional layer, performing max pooling through the pooling layer, and flattening the three-dimensional tensor output by the pooling layer into one-dimensional vector data;

[0025] The flattened one-dimensional vector data is input into the LSTM layer, and the long-term time series dependence relationship is captured through the forget gate, input gate, output gate and cell state, and the hidden state sequence is output; combined with the multi-head parallel attention mechanism, the weights are dynamically allocated to focus on the key degradation stages, and finally the predicted remaining service life value of the robotic arm is generated.

[0026] Preferably, the degradation-efficiency coupling model is:

[0027]

[0028] E = E0·e -0.25HI

[0029] where HI represents the health index of the robotic arm, E represents the operation efficiency, S vib is the vibration entropy, and the vibration entropy is obtained by extracting high-frequency features through F peak (t) and calculating through wavelet packet decomposition or spectral entropy, ΔP represents the change in hydraulic pressure, E0 represents the initial value of efficiency, and RUL represents the predicted remaining service life value of the robotic arm.

[0030] Preferably, the reward function is:

[0031] R = -0.3×energy consumption + 0.5×E + 0.2×HI

[0032] where R is the reward function.

[0033] Preferably, the training process of the DPPO reinforcement learning controller:

[0034] The policy network and the value network adopt a double-delay update mechanism;

[0035] Prioritized sampling of the critical state data with more than 10% decrease in HI in the experience replay pool;

[0036] Calculating the policy gradient update amount through generalized advantage estimation.

[0037] Preferably, driving the hydraulic actuator of the robotic arm after processing the load adjustment instruction through dynamic constraints, including:

[0038] Eliminating the mutation signal through a first-order low-pass filter, restricting the single adjustment amplitude not to exceed 5%, to prevent the robotic arm from being impacted due to the jump of the load adjustment instruction; at the same time, mapping the load adjustment instruction to the joint torque safety range based on the Jacobian matrix to ensure that the action conforms to the mechanical dynamics constraints;

[0039] The processed load adjustment instruction realizes millisecond-level inference through the TensorRT acceleration engine, and is converted into a standard industrial signal, which is transmitted to the PLC controller through the OPC UA protocol to drive the hydraulic proportional valve of the robotic arm to dynamically adjust the oil pressure and flow rate, so as to achieve precise load control.

[0040] Preferably, it also includes health dynamic monitoring and early warning:

[0041] When HI≥0.8, trigger a first-level early warning, automatically record the current working condition snapshot and prompt manual intervention for inspection;

[0042] When HI≥0.9, forcibly switch to the low-load protection mode, reduce the operation efficiency to delay the equipment degradation, and at the same time push the maintenance work order through the cloud, and associate the fault type with the spare parts inventory information.

[0043] The intelligent control system of the dredging robotic arm based on DPPO reinforcement learning includes:

[0044] Data acquisition and preprocessing module: Real-time collect the hydraulic pressure, vibration spectrum and temperature data of the dredging robotic arm, and perform preprocessing;

[0045] Model construction module: Input the preprocessed data into the LSTM-Attention model, extract features based on the long short-term memory network and attention mechanism, and output the predicted value of the remaining service life of the robotic arm;

[0046] Degradation-efficiency coupling model module: Establish a degradation-efficiency coupling model, and calculate the health index and efficiency attenuation coefficient of the robotic arm based on the predicted value of the remaining service life of the robotic arm;

[0047] Prediction target module: Establish a DPPO reinforcement learning controller, take the health index, operation efficiency, energy consumption and joint angle of the robotic arm as the state space, output the action space of the dynamic load adjustment instruction, optimize the control strategy through the reward function, and drive the hydraulic actuator of the robotic arm after processing the load adjustment instruction through dynamic constraints.

[0048] Preferably, it also includes:

[0049] Maintenance strategy module: Used for health dynamic monitoring and early warning, specifically:

[0050] When HI≥0.8, a first-level early warning is triggered, the current working condition snapshot is automatically recorded, and manual intervention for inspection is prompted.

[0051] When HI≥0.9, it is forced to switch to the low-load protection mode, the operation efficiency is reduced to delay equipment degradation, and at the same time, a maintenance work order is pushed through the cloud, associating the fault type with the spare parts inventory information.

[0052] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses an intelligent control method and system for a dredging robotic arm based on DPPO reinforcement learning. Through the full-link design of "state perception - policy optimization - safe execution - closed-loop early warning", the present invention solves the three major pain points of lagging maintenance, energy efficiency imbalance, and infeasible instructions in traditional methods, and realizes the engineering implementation of intelligent control of heavy equipment. Brief Description of the Drawings

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0054] Figure 1 It is a flowchart of the intelligent control method for a dredging robotic arm based on DPPO reinforcement learning provided by the present invention;

[0055] Figure 2 It is an internal structure diagram of the LSTM-Attention model provided by the present invention;

[0056] Figure 3 It is a schematic diagram of the internal structure of the LSTM network provided by the present invention;

[0057] Figure 4 It is a flowchart of multi-head parallel attention;

[0058] Figure 5 It is an internal structure diagram of the application of the DPPO reinforcement learning algorithm provided by the present invention;

[0059] Figure 6 It is a structural block diagram of the intelligent control system for a dredging robotic arm based on DPPO reinforcement learning provided by the present invention. Detailed Embodiments

[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0061] An intelligent control method for a dredging manipulator based on DPPO reinforcement learning is disclosed in an embodiment of the present invention. As Figure 1 shown, it includes:

[0062] Real-time collect the hydraulic pressure, vibration spectrum and temperature data of the dredging and silt removal manipulator of the dredger, and perform preprocessing;

[0063] Input the preprocessed data into the LSTM-Attention model, extract features based on the long short-term memory network and the attention mechanism, and output the predicted remaining service life;

[0064] Establish a degradation-effectiveness coupling model, and calculate the manipulator health index and operation effectiveness based on the predicted remaining service life;

[0065] Establish a DPPO reinforcement learning controller, use the manipulator health index, operation effectiveness, energy consumption and manipulator joint angles as the state space, output the action space of the dynamic load regulation instruction, optimize the control strategy through the reward function, and drive the manipulator hydraulic actuator after processing the load regulation instruction by dynamic constraints.

[0066] Next, the above steps will be further described.

[0067] S1. Real-time collect multi-modal time-series data of the manipulator operation state through the hydraulic pressure sensor (range 0-40 MPa, accuracy ±0.5% FS), three-axis vibration accelerometer (bandwidth 5-10 kHz), and infrared temperature measurement module (accuracy ±1 °C) carried by the dredger, including the hydraulic pressure, vibration spectrum and temperature data of the dredging and silt removal manipulator, and perform preprocessing. Specifically:

[0068] S101. Data format definition

[0069] The format of each training sample is <t, X_m, Y>, where:

[0070] The time window identifier t: represents the time series window to which the data belongs (for example, 60 seconds is a window, and it slides with a step of 10 seconds); the input feature X_m: contains m multi-source sensor and operation parameter features; the label Y: the classification label for the manipulator health state, and the one-hot encoding is used to define the label of each category, that is, the data 0 and 1 are used to distinguish the three states of the manipulator.

[0071] Normal state → [1, 0, 0]

[0072] Slight degradation → [0, 1, 0]

[0073] Severe degradation → [0, 0, 1]

[0074] S102. Data acquisition

[0075] In this example, the normal operation data of the robotic arm during 30 consecutive days of dredging operations is collected, with a total of 518,400 pieces (60 seconds / piece × 24 hours × 30 days × 60 minutes / hour ÷ 60 seconds). Through a countermeasure network (CGAN), simulated robotic arm degradation scenarios are generated, and 1,036,800 pieces of synthetic data containing faults such as hydraulic leakage, bearing wear, and high-temperature overload are generated, making the ratio of normal to abnormal data 1:2 to alleviate the problem of class imbalance. The total sample size of this example is 1,555,200 pieces.

[0076] S103. Data augmentation

[0077] All data is processed and normalized. Specifically:

[0078] Missing value processing: Lagrange interpolation method is used to fill in the missing sensor data;

[0079] Outlier processing: Tukey's Test method is used to detect outliers and eliminate error values;

[0080] Normalization processing: Min-Max normalization is used to normalize the collected multi-dimensional parameters;. Subsequently, the preprocessed dataset is randomly divided into a training set, a validation set, and a test set in a ratio of 7:2:1 for model training, hyperparameter tuning, and generalization ability evaluation respectively. The format is as follows:

[0081] Training samples:

[0082]

[0083] Validation samples:

[0084]

[0085] Test samples:

[0086]

[0087] Among them, n represents the nth piece of data.

[0088] Substitute the vibration frequency after the above preprocessing into the formula:

[0089]

[0090] Among them, V rms represents the root mean square value of the vibration velocity at time t, indicating the effective energy level of the vibration (unit such as m / s); N is the number of sample points selected during calculation (corresponding to the amount of data within a certain time window); v i is the measured value of the instantaneous vibration velocity at the i-th sampling point;

[0091]

[0092] Among them, F peak (t) is the maximum amplitude of the vibration signal at time t within the frequency band of 1 - 5 kHz (which may correspond to the amplitude of the fault characteristic frequency); is the spectrum obtained by performing Fourier transform on the vibration velocity signal v(t); max(...) 1-5kHz is to take the maximum value of the amplitude within the frequency range of 1 - 5 kHz.

[0093] Substituting the vibration spectrum into the formulas of V rms (t) and F peak (t) has the following core functions:

[0094] (1) Feature extraction: V rms The root mean square value of the vibration velocity is used to quantify the overall energy level of the vibration signal in the time domain, reflecting the vibration intensity of the robotic arm. F peak (t) (the maximum amplitude in the 1 - 5 kHz frequency band) is used to capture high-frequency vibration characteristics and identify specific fault modes (such as bearing wear, abnormal gear meshing, etc.).

[0095] (2) Data dimensionality reduction and standardization: The original vibration spectrum data has a high dimension and is complex. By calculating V rms and F peak (t), the high-dimensional spectrum signal is compressed into key feature values, which is convenient for subsequent model processing.

[0096] S2. As Figures 2 - 4 shown, an LSTM-Attention prediction model is constructed. Its network architecture design includes an LSTM layer that captures the temporal dependence relationship of parameters such as hydraulic pressure and vibration through forget gates, input gates, output gates, and cell states, and an Attention layer that dynamically allocates weights to focus on key degradation stages such as sudden increases in high-frequency vibrations. The input of the model is the preprocessed multi-source sensor temporal data (hydraulic pressure, vibration spectrum characteristics, temperature at time t), and the output is the predicted value of the remaining useful life (RUL) of the robotic arm.

[0097] Introduce a multi-modal degradation feature adaptive attention mechanism into the LSTM-Attention model. Aiming at the coupling characteristics of multi-source sensing data (hydraulic pressure, vibration, temperature) of the dredging manipulator, optimize the traditional single-head attention structure, and finally use it to predict the remaining useful life (RUL) of the manipulator. The specific improvements are as follows:

[0098] S201. Multi-head parallel attention:

[0099] Aiming at the physical property differences of multi-modal sensing data such as hydraulic pressure, vibration, and temperature of the dredging manipulator, design an independent attention head sub-modal analysis mechanism: The three-branch attention network designed in the present invention includes three specialized feature extraction modules: The hydraulic pressure fluctuation head constructs a pressure time series analysis model to focus on capturing abnormal fluctuation characteristics such as sudden increases and drops in the hydraulic system pressure (for example, a sudden increase in pressure can accurately identify valve blockage faults), and realizes the quantitative analysis of the pressure fluctuation frequency and amplitude; The vibration frequency domain energy head adopts wavelet packet energy entropy analysis technology to monitor the energy mutation mode in a specific frequency band (such as 10 - 15 kHz) (typically, a sudden increase in the frequency band energy caused by bearing wear), and quantitatively evaluates the modal abnormality degree of the vibration signal; The temperature trend head captures the slow change trend characteristics of temperature based on the sliding window algorithm (such as the change in the temperature rise slope caused by motor overload), and effectively identifies the long-term heat load accumulation effect.

[0100] S202. Degradation feature adaptive weighting:

[0101] Based on the manipulator historical fault database (including 6 types of typical faults such as hydraulic failure, bearing wear, and motor overheating), construct a dynamic allocation strategy for the weights of the attention heads. This strategy realizes the intelligent fusion of multi-source signals through two-stage optimization: First, in the offline training stage, use the fault type label supervised learning method to determine the initial weight allocation of each attention head (such as 40% for the hydraulic pressure fluctuation head, 35% for the vibration frequency energy head, and 25% for the temperature trend head); Second, in the online inference stage, by real-time monitoring of the sensor data characteristics (such as vibration energy entropy value, pressure fluctuation amplitude, etc.), when a specific fault mode symptom is detected (such as a sudden increase in the vibration entropy value indicating bearing wear), automatically trigger the weight dynamic adjustment mechanism.

[0102] S203. Spatiotemporal feature fusion:

[0103] The spatio-temporal feature fusion architecture realizes the complementary enhancement of multi-dimensional degradation information. Its core fusion mechanism includes three key links: First, use the LSTM network to capture the long-term evolution trends of parameters such as hydraulic pressure and temperature (such as the slow pressure drop curve caused by hydraulic leakage); Second, based on the attention mechanism, extract the wavelet packet energy entropy features of vibration signals (such as the sudden increase in the energy ratio in the 2000-4000Hz frequency band can accurately identify the pitting fault of the gearbox); Finally, perform feature-level splicing on the time series feature vector output by the LSTM and the frequency domain feature vector generated by the attention module, and realize the deep fusion of multi-modal features through a fully connected network to generate the final RUL prediction value.

[0104] The trained model can be used to predict the RUL of the manipulator of the dredger.

[0105] Among them, the generation process of predicting the RUL of the manipulator includes:

[0106] (1) Input layer: Receive preprocessed multi-source time series data, including hydraulic pressure, vibration spectrum features (V rms 、F peak ) temperature, etc.

[0107] (2) Convolution layer (CNN) processing: Use a one-dimensional convolution kernel (kernel_size = 3, filters = 32) to slide along the time axis to extract local time series features to generate a feature map; Activation function: ReLU, enhance the non-linear expression ability, and the output feature map dimension is [batch_size, sequence_length = 20, filters = 32].

[0108] Among them, (kernel_size = 3, filters = 32) are the parameters of the convolution layer. kernel_size = 3 means that the time span (time window size) of the one-dimensional convolution kernel is 3 time steps, and filters = 32 means that 32 different convolution kernels are used, and each convolution kernel is responsible for extracting a specific local feature. [batch_size, sequence_length = 20, filters = 32] is the output dimension of the convolution layer. batch_size represents the number of samples input to the model at one time (i.e., the batch size), sequence_length = 20 represents the time series length (number of time steps) of each sample. After the convolution operation, the time steps of the original time series data may be compressed or remain 20 due to padding or stride adjustment. filters = 32 represents the number of feature channels corresponding to each time step (i.e., the number of convolution kernels). Each channel corresponds to a feature map extracted by a convolution kernel.

[0109] (3) Pooling layer: Perform max pooling along the time axis (pool_size = 2, stride = 2), compress the feature dimension to [batch_size, sequence_length = 10, filters = 32], select the maximum value within each window, retain significant features, and reduce the computational complexity.

[0110] (pool_size = 2, stride = 2) are the parameters of the pooling layer, where pool_size = 2 indicates that the time span of the pooling window is 2 time steps. The pooling operation slides on the time axis, covering 2 consecutive time steps each time. stride = 2 indicates that the sliding step of the pooling window is 2. It moves 2 time steps each time, and the windows do not overlap. [batch_size, sequence_length = 10, filters = 32] is the output dimension of the pooling layer. Among them, sequence_length = 10 means that the pooling operation compresses the time steps from 20 to 10, batch_size and filters = 32 remain unchanged, and only the time axis is compressed.

[0111] (4) Flatten the three-dimensional tensor ([batch_size, 10, 32]) output by the pooling layer into a one-dimensional vector with a shape of [batch_size, 320] to adapt to the input of the subsequent time series modeling module.

[0112] Unfold the three-dimensional tensor [batch_size, 10, 32] along the last dimension to obtain a one-dimensional vector. The calculation formula is: 10 (time steps) × 32 (feature channels) = 320.

[0113] (5) Input to the LSTM-Attention model:

[0114] The flattened one-dimensional vector data is input to the LSTM layer. The long-term time series dependencies are captured through the forget gate, input gate, output gate, and cell state, and a hidden state sequence is output. Combined with the multi-head parallel attention mechanism (processing of sub-modalities of hydraulic pressure, vibration, and temperature), the weights are dynamically allocated to focus on the key degradation stages (such as the period of sudden increase in high-frequency vibration). Finally, the predicted value of the remaining useful life (RUL) of the robotic arm is generated. The specific steps are as follows:

[0115] Sa. Temporal feature extraction:

[0116] The vibration signal enters to calculate 15-dimensional features including the root mean square value (RMS), peak factor, kurtosis, impulse factor, waveform factor, margin factor, band energy (low frequency), band energy (medium frequency), band energy (high frequency), main frequency component, spectral centroid, spectral entropy, zero crossing rate, time domain variance, and envelope spectrum peak of the vibration signal in real time based on the Fourier transform. The output shape is [batch_size, sequence_length = 20, feature_dim = 15], where batch_size: the batch size (the number of samples processed at one time); sequence_length = 20: each sample contains data of 20 consecutive time points; feature_dim = 15: 15 features are extracted at each time point.

[0117] Sb, multi-head parallel attention mechanism:

[0118] Based on multi-modal sensing data such as hydraulic pressure, vibration, and temperature, three independent attention heads for hydraulic pressure fluctuation, vibration frequency energy, and temperature trend are designed to process the data in a modal-separated manner, thereby dynamically focusing on key degradation features. Among them, the hydraulic pressure fluctuation head is responsible for extracting the time series dimension related to hydraulic pressure in the LSTM output, and then capturing fluctuation features such as sudden pressure increase and decrease through time series analysis, and then generating a pressure fluctuation feature vector to identify abnormal situations in the hydraulic system such as valve blockage; the vibration frequency energy head extracts the time series dimension related to vibration in the LSTM output to detect bearing wear or gearbox pitting faults; after the temperature trend head extracts the time series dimension related to temperature in the LSTM output, it uses a sliding window smoothing algorithm to extract slow-changing trends such as the change of temperature rise slope, generates temperature trend parameters, and finally achieves the purpose of capturing motor overload or long-term heat load cumulative effect.

[0119] Sc, dynamic allocation of attention weights:

[0120] The initial weights are allocated based on the statistics of historical fault data, and the hydraulic pressure fluctuation head, vibration frequency energy head, and temperature trend head account for 40%, 35%, and 25% respectively; during dynamic adjustment, the system monitors the features of each modality in real time (such as a sudden increase in vibration entropy value), and automatically increases the weight of the corresponding attention head (for example, the weight of the vibration head can be increased to 50%); the weight calculation generates attention scores through a SoftMax-like normalization, and the formula is:

[0121]

[0122] Among them, F(Q i , K i ) is the similarity function between the query (Query) and the key (Key).

[0123] Sd, multi-modal feature fusion:

[0124] Feature fusion and prediction are completed through the following steps: First, using the feature concatenation method, the outputs of three parallel attention heads (with a shape of [batch_size, 64]) are concatenated along the feature dimension to form a fused feature vector (with a shape of [batch_size, 192]); subsequently, through a mapping network composed of two fully connected layers (with the ELU activation function), the fused features are non-linearly mapped to the target space, and finally the RUL prediction value (with a shape of [batch_size, 1]) is output. GPU is used to accelerate the calculation of the attention weight matrix (with a complexity of O(n 2 )) Output processing: The output of the fully connected layer is activated by the exponential linear unit (ELU) to obtain the RUL prediction value.

[0125] It may also include:

[0126] Se, establish a three-dimensional evaluation framework including statistical indicators, trend matching degree, and fault warning rate:

[0127] Statistical indicators:

[0128]

[0129] Trend matching degree: The dynamic time warping (DTW) algorithm is used to calculate the similarity between the predicted curve and the actual degradation curve, and a trend deviation threshold (such as a ±15% prediction interval) is set.

[0130] Fault warning rate: Early warning time: When the predicted RUL is less than T0, an early warning is triggered (T0 = 24 hours)

[0131]

[0132] In this embodiment, the model is iteratively optimized to construct a closed-loop optimization mechanism, including:

[0133] Online learning: 1000 newly collected samples (including 50 fault samples) are added to the training set every week

[0134] Hyperparameter tuning: Use the Bayesian optimization algorithm to search for the optimal parameter combination:

[0135] Set the number of LSTM layers: [1, 3]; the number of attention heads: [2, 8]; the learning rate: [1e-4, 1e-3].

[0136] The present invention predicts the remaining life through the LSTM-Attention neural network, and predicts the remaining service life of the robotic arm (RUL, 40%), dynamically monitors the hydraulic pressure change rate (ΔP, 30%), and the vibration entropy (S vib, 30%) to construct a composite health index (HI). Among them, the decrease in RUL reversely represents the intensification of equipment aging, the abnormal fluctuation of ΔP reflects hydraulic system failures, and S vib The increase in entropy indicates mechanical structure abnormalities, and a dual-threshold early warning mechanism is combined to achieve hierarchical control of equipment risks.

[0137] S3. Establish a degradation-performance coupling model, and calculate the health index and operation performance of the robotic arm based on the predicted remaining service life of the robotic arm;

[0138] The operation performance of the robotic arm (such as dredging efficiency) is not fixed, but dynamically decays with the health status of the equipment. The degradation-performance coupling model maps the HI value to the operation performance through an exponential function:

[0139]

[0140] E = E0·e -0.25HI

[0141] Among them, HI represents the health index of the robotic arm, E represents the operation performance, and S vib is the vibration entropy. Substitute the vibration spectrum into the formula Extract high-frequency features, and further calculate S through wavelet packet decomposition or spectral entropy vib , ΔP represents the change in hydraulic pressure, which is obtained by calculating the pressure difference between adjacent time windows (or the pressure change rate within the sliding window) during preprocessing.

[0142] When HI = 0 (optimal health), the performance maintains the initial value E0; when HI = 1 (worst health), the performance drops to about 78% of E0, forcing the system to balance between efficiency and lifespan.

[0143] S4. The DPPO-based reinforcement learning controller achieves dynamic optimization by integrating the health index (HI) of the robotic arm and the degradation-performance coupling model: real-time perception of 10-dimensional state space data such as the health index HI, operation performance, energy consumption, and joint angles of the robotic arm, dynamically adjusting the hydraulic load and vibration gain to match the health status; using operation performance (50%), energy consumption (30%), and the health index of the robotic arm (20%) as the objective function to balance multiple objectives, combining kinematic constraints and low-pass filtering to limit the single adjustment amplitude ≤ 5%; adopting a double-delay strategy for training, combining prioritized replay and TensorRT acceleration, and achieving real-time control through OPCUA. Among them,

[0144] a. The health index (HI) of the robotic arm: The core index of the comprehensive health status;

[0145] b. The operation performance (E): The percentage of the current operation performance to the initial value;

[0146] c. Energy consumption index: Energy consumption per unit time (such as the power of the hydraulic system);

[0147] d. Robotic arm joint angles: Real-time angle values of 6 joints, reflecting the motion state.

[0148] Update the state space data every 5 seconds, input the data into the DPPO controller, and generate an adaptive control strategy.

[0149] As Figure 5 shown, specifically:

[0150] a. State perception:

[0151] The controller receives a 10-dimensional state vector in real time, including the robotic arm health index (HI), the current job efficiency (E(t)), the energy consumption index, and the 6 joint angles of the robotic arm, comprehensively perceiving the equipment health, efficiency, and motion state;

[0152] Input of the state space: The joint angles, as one of the parameters of the state space, are used to characterize the real-time motion state of the robotic arm (such as load distribution, motion trajectory).

[0153] b. Action decision-making:

[0154] Output continuous load adjustment instructions, dynamically adjusting the load of the hydraulic system (±20% of the rated pressure) and the vibration suppression gain to ensure dynamic matching of the load and the health state;

[0155] c. Policy optimization:

[0156] Drive the controller to balance three goals through the reward function (energy consumption ratio -30%, job efficiency ratio +50%, robotic arm health index ratio +20%) - prioritize ensuring efficiency, while suppressing the excessive growth of energy consumption and delaying equipment degradation;

[0157] The reward function is:

[0158] R = -0.3 × energy consumption + 0.5 × E + 0.2 × HI

[0159] The robotic arm joint angles do not directly appear in the reward function, but they indirectly affect the control strategy in the following ways:

[0160] Implicit influence of policy optimization: The joint angles are transmitted to the DPPO controller through the state space, affecting the generation of the action space (such as load adjustment instructions).

[0161] d. Safety constraints:

[0162] The safety constraint mechanism, through the collaborative action of the robotic arm joint angles and the Jacobian matrix, maps the load adjustment instruction to the safe range of joint torques, thus effectively preventing mechanical structures from being damaged due to overloaded forces. Specifically, the robotic arm joint angles represent in real time the geometric configuration of the robotic arm - physical states such as the bending degree of each joint, etc. This parameter directly determines the torque requirements of each joint. Since different joint angle configurations change the mechanical characteristics of the robotic arm (for example, the change in the length of the moment arm), the same load adjustment instruction may trigger significantly different joint torque responses under different configurations. As a mathematical tool for establishing the mapping relationship between the motion speed of the end effector of the robotic arm and the joint angular velocity, the transpose matrix of the Jacobian matrix can accurately convert the end-effector load force vector into the torque vectors of each joint. By combining the real-time collected joint angle data with the Jacobian matrix model, the system can dynamically calculate the torque safety threshold under the current configuration, ensuring that the execution of the load adjustment instruction always remains within the safe working range of the mechanical structure.

[0163] Prevent mechanical over-limit damage, and smooth the instruction through a low-pass filter (cutoff frequency 5Hz), restricting the single-step adjustment amplitude ≤ 5%;

[0164] e. Training and deployment:

[0165] The policy network and the value network adopt a dual-delay update mechanism (50 steps for the policy network / 20 steps for the value network) and prioritized experience replay (focusing on data in the HI sudden drop condition), calculate the policy gradient update amount through generalized advantage estimation, and achieve millisecond-level inference by combining with the TensorRT acceleration engine. Finally, the instruction is sent to the industrial PLC through the OPC UA protocol to complete the control closed-loop. This design breaks through the static limitations of traditional rule-based control. Among them, the policy network is responsible for generating dynamic load adjustment instructions (action space), that is, selecting the optimal control strategy according to the current state (robotic arm health index, operation efficiency, energy consumption, and joint angles); the value network is responsible for evaluating the value of the current state (i.e., the expected cumulative reward) to optimize the decision-making of the policy network.

[0166] The S5, DPPO controller generates maintenance strategies and efficiency calculations through multi-dimensional state perception (health index HI, real-time efficiency E(t), energy consumption, and joint angles). Among them, there is no time sequence for the S501 and S502 maintenance measurement and efficiency calculation, and they are carried out simultaneously:

[0167] S501, Maintenance strategy:

[0168] a. First-level warning (HI ≥ 0.8):

[0169] When the robotic arm health index (HI) reaches or exceeds 0.8, the system automatically triggers the following operations:

[0170] Data recording: Save a snapshot of the current working condition, including vibration spectrum, hydraulic pressure curve, and temperature trend.

[0171] Manual intervention prompt: Send an alarm through the Human-Machine Interface (HMI) to prompt the operator for on-site inspection or remote diagnosis.

[0172] b. Secondary protection mode (HI ≥ 0.9):

[0173] When HI further rises to 0.9, the system executes forced protection measures:

[0174] Load limit: Switch to the low-load mode, reduce the operation efficiency to 60% of the initial value E0 to delay equipment degradation. At the same time, automatically generate a maintenance work order, associate the fault type (such as bearing wear, hydraulic leakage) and spare part inventory information, and push it to the management platform through the cloud.

[0175] c. Intelligent management of maintenance work orders;

[0176] Fault location and spare part matching: Based on the historical fault database, combined with real-time sensor data, accurately locate the faulty component (such as pitting corrosion of the gearbox), and recommend the appropriate spare part model.

[0177] Dynamic adjustment of priority: Dynamically adjust the execution priority of the maintenance work order according to the change rate of the HI value and the urgency of the operation task.

[0178] d. Safety constraints and instruction execution guarantee

[0179] Map the instruction to the safe range of joint torque through the Jacobian matrix, and smooth the instruction using a first-order low-pass filter (cutoff frequency 5Hz), limiting the single adjustment amplitude ≤ 5% to avoid mechanical shock;

[0180] S502, Efficiency calculation:

[0181] a. State perception and multi-objective optimization

[0182] Input parameters:

[0183] Health Index (HI): Comprehensively reflect the equipment degradation state.

[0184] Operation efficiency (E): The ratio of the real-time dredging efficiency to the initial efficiency.

[0185] Energy consumption index: Hydraulic system power per unit time (kWh / m 3 )

[0186] Joint angle: Real-time configuration parameters of 6 joints.

[0187] After the input parameters are determined, the reward function is designed. Based on the principle of multi-objective optimization, the reward function R quantitatively couples energy consumption, operation efficiency, and the manipulator health index, and is specifically defined as:

[0188] This function reflects the priority strategy through weight allocation: with the highest weight of 50%, it gives priority to ensuring operation efficiency to ensure that the system can complete the dredging task efficiently; at the same time, it takes into account energy consumption suppression and manipulator health index maintenance with weights of 30% and 20% respectively, controlling energy consumption and delaying equipment degradation while improving operation efficiency, and realizing the balanced optimization of performance, energy efficiency, and reliability.

[0189] b. Dynamic load regulation and energy consumption optimization

[0190] In the link of generating hydraulic load commands, the system relies on the output results of the DPPO reinforcement learning controller to dynamically generate load regulation commands, achieving precise adjustment of the hydraulic system load within the range of ±20% of the rated pressure, and synchronously optimizing the vibration suppression gain parameters to balance operation efficiency and system stability.

[0191] During the industrial-level deployment process, the following technical solutions are adopted to ensure the efficient transmission and precise execution of control commands:

[0192] S503. Comparison of technical effects: The end-to-end control delay is <100ms, supporting high-frequency regulation of 20Hz, and the response speed is doubled compared with traditional PID control (delay >200ms); the sudden failure rate is reduced by 67% (the average annual failure is 12 times under traditional rule control → it is reduced to 4 times in the present invention), and the maintenance response time is shortened from 2 hours to 15 minutes; through dynamic load regulation, the manipulator life is extended by 41.6% (MTBF 1200h → 1700h), and the unit operation energy consumption is reduced by 23% (2.8 → 2.15kWh / m 3 )

[0193] Compared with the existing technology, the present invention solves the three major pain points of maintenance lag, energy efficiency imbalance, and infeasible commands in traditional methods through the full-link design of "state perception - policy optimization - safe execution - closed-loop warning", and realizes the engineering implementation of intelligent control of heavy equipment.

[0194] An intelligent control system for a dredging manipulator based on DPPO reinforcement learning is disclosed in an embodiment of the present invention. As Figure 6 shown, it includes:

[0195] Data acquisition and preprocessing module: It collects the hydraulic pressure, vibration spectrum, and temperature data of the dredging and silt removal manipulator of the dredger in real time and performs preprocessing.

[0196] Model construction module: Input the preprocessed data into the LSTM-Attention model, extract features based on the long short-term memory network and the attention mechanism, and output the predicted remaining useful life.

[0197] Degradation-performance coupling model module: Establish a degradation-performance coupling model, and calculate the manipulator health index and the performance decay coefficient based on the predicted remaining useful life.

[0198] Prediction target module: Establish a DPPO reinforcement learning controller. With the manipulator health index, operation performance, energy consumption, and manipulator joint angles as the state space, output the action space of the dynamic load adjustment instruction. Optimize the control strategy through the reward function, and drive the manipulator hydraulic actuator after processing the load adjustment instruction through dynamic constraints.

[0199] Preferably, it further includes:

[0200] Maintenance strategy module: Used for health dynamic monitoring and warning. Specifically:

[0201] When HI≥0.8, trigger a first-level warning, automatically record the current working condition snapshot, and prompt manual intervention for inspection.

[0202] When HI≥0.9, forcibly switch to the low-load protection mode, reduce the operation performance to delay equipment degradation, and at the same time push a maintenance work order through the cloud, associating the fault type with the spare parts inventory information.

[0203] The implementation processes and methods of each part of the system of the present invention are the same, and will not be elaborated here.

[0204] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0205] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. The intelligent control method of dredging robot arm based on DPPO reinforcement learning is characterized by: include: Collect the hydraulic pressure, vibration spectrum and temperature data of the dredging robot arm in real time and perform pre-processing; The preprocessed data is input into the LSTM-Attention model, and features are extracted based on the long short-term memory network and attention mechanism to output the predicted value of the remaining service life of the robot arm; Establish a degradation-performance coupling model, and calculate the robot arm health index and operating performance based on the predicted value of the remaining service life of the robot arm; A DPPO reinforcement learning controller is established, which takes the robot health index, working efficiency, energy consumption and robot joint angle as the state space, outputs the action space of dynamic load adjustment instructions, and optimizes the control strategy through the reward function. The load adjustment instructions are processed by dynamic constraints to drive the hydraulic actuator of the robot.

2. The dredging mechanical arm intelligent control method based on DPPO reinforcement learning according to claim 1 is characterized in that: Preprocessing includes: Missing value processing: Lagrange interpolation method is used to fill missing data; Outlier processing: Tukey's Test method was used to test outliers and eliminate error values; Normalization: Min-Max normalization is used to normalize the collected multi-dimensional parameters; Substitute the vibration spectrum after the above preprocessing into the formula: Among them, V rms is the root mean square value of the vibration velocity at time t, N is the number of sample points selected during calculation, v i is the instantaneous vibration velocity measurement value of the i-th sampling point; Among them, F peak (t) is the maximum amplitude of the vibration signal in the 1-5kHz frequency band at time t, is the spectrum obtained by Fourier transforming the vibration velocity signal v(t), max(...) 1-5kHz It is the maximum amplitude value in the frequency range of 1-5kHz.

3. The intelligent control method of dredging mechanical arm based on DPPO reinforcement learning according to claim 1 is characterized in that: The preprocessed data is input into the LSTM-Attention model, and feature extraction is performed based on the long short-term memory network and attention mechanism to output the predicted value of the remaining service life of the robot arm, including: The input layer receives preprocessed multi-source time series data, including hydraulic pressure, vibration velocity root mean square value V rms , maximum vibration amplitude F peak and temperature parameters; The convolution layer is used to extract local temporal features, the pooling layer is used to perform maximum pooling, and the flattening operation is used to flatten the three-dimensional tensor output by the pooling layer into one-dimensional vector data. The flattened one-dimensional vector data is input into the LSTM layer, which captures long-term temporal dependencies through the forget gate, input gate, output gate and cell state, and outputs a hidden state sequence. Combined with the multi-head parallel attention mechanism, weights are dynamically allocated to focus on the key degradation stages, and finally a predicted value of the remaining service life of the robot arm is generated.

4. The intelligent control method of dredging mechanical arm based on DPPO reinforcement learning according to claim 1 is characterized in that: The degradation-efficiency coupling model is: E=E0·e -0.25HI Among them, HI represents the health index of the robot arm, E represents the operation efficiency, S vib is the vibration entropy, which is expressed by F peak (t) After extracting high-frequency features and calculating through wavelet packet decomposition or spectral entropy, ΔP represents the change in hydraulic pressure, E0 represents the initial value of effectiveness, and RUL represents the predicted value of the remaining service life of the robot arm.

5. The intelligent control method of dredging mechanical arm based on DPPO reinforcement learning according to claim 1 is characterized in that: The reward function is: R=-0.3×energy consumption+0.5×E+0.2×HI Among them, R is the reward function, HI represents the health index of the robot arm, and E represents the operation efficiency.

6. The intelligent control method of dredging mechanical arm based on DPPO reinforcement learning according to claim 1 is characterized in that: The training process of DPPO reinforcement learning controller: The policy network and value network adopt a double-delay update mechanism; In the experience replay pool, priority is given to sampling critical state data where HI drops by more than 10%; Compute the policy gradient update via the generalized advantage estimate.

7. The intelligent control method of dredging mechanical arm based on DPPO reinforcement learning according to claim 1 is characterized in that: The load adjustment command is processed by dynamic constraints to drive the hydraulic actuator of the robot arm, including: The first-order low-pass filter is used to eliminate sudden changes in signals and limit the single adjustment range to no more than 5%, preventing the robot arm from being impacted by load adjustment command jumps. At the same time, the load adjustment command is mapped to the joint torque safety range based on the Jacobian matrix to ensure that the action complies with the mechanical dynamics constraints. The processed load adjustment instructions are accelerated by the TensorRT engine to achieve millisecond-level inference, converted into standard industrial signals, and transmitted to the PLC controller via the OPC UA protocol, driving the hydraulic proportional valve of the robotic arm to dynamically adjust the oil pressure and flow to achieve precise load control.

8. The intelligent control method of dredging mechanical arm based on DPPO reinforcement learning according to claim 4 is characterized in that: It also includes dynamic health monitoring and early warning: When HI ≥ 0.8, a first-level warning is triggered, automatically recording a snapshot of the current working condition and prompting manual intervention for inspection; When HI ≥ 0.9, it is forced to switch to low-load protection mode, reducing operating efficiency to delay equipment degradation. At the same time, maintenance work orders are pushed through the cloud, linking the fault type with spare parts inventory information.

9. The intelligent control system of dredging robot arm based on DPPO reinforcement learning is characterized by: include: Data acquisition and preprocessing module: real-time acquisition of hydraulic pressure, vibration spectrum and temperature data of the dredging robot arm, and preprocessing; Model building module: input the preprocessed data into the LSTM-Attention model, extract features based on the long short-term memory network and attention mechanism, and output the predicted value of the remaining service life of the robot arm; Degradation-efficiency coupling model module: establishes a degradation-efficiency coupling model and calculates the health index and efficiency attenuation coefficient of the robot arm based on the predicted value of the remaining service life of the robot arm; Prediction target module: Establish a DPPO reinforcement learning controller, use the robot health index, work efficiency, energy consumption and robot joint angle as the state space, output the action space of the dynamic load adjustment command, optimize the control strategy through the reward function, and drive the robot hydraulic actuator after the load adjustment command is processed by dynamic constraints.

10. The dredging mechanical arm intelligent control system based on DPPO reinforcement learning according to claim 9 is characterized in that: Also includes: Maintenance strategy module: used for dynamic health monitoring and early warning, specifically: When HI ≥ 0.8, a first-level warning is triggered, automatically recording a snapshot of the current working condition and prompting manual intervention for inspection; When HI ≥ 0.9, it is forced to switch to low-load protection mode, reducing operating efficiency to delay equipment degradation. At the same time, maintenance work orders are pushed through the cloud, linking the fault type with spare parts inventory information.

Citation Information

Patent Citations

  • System residual life prediction method and device, equipment and storage medium

    CN116151088A

  • Control method of cluster unmanned aerial vehicle system based on DPPO deep reinforcement learning

    CN119002518A

  • Machine-learned models for electric vehicle component health monitoring

    US12073668B1

  • Battery State of Health Assessment System

    US20100090650A1

Cited By

  • Artificial intelligence monitoring model method based on thermal power plant edge side single node deployment

    CN120543870A

  • Detection method and system for electrical automation equipment

    CN120606419A

  • Intelligent grabbing method and system of mechanical arm based on Internet of Things

    CN120620229A

  • Industrial equipment data recovery method, system, equipment, medium and product

    CN120653641A

  • Wharf portal crane operation efficiency optimization method and system based on digital twinning and wharf portal crane

    CN120851309A