Interferential offline reinforcement learning data set recovery method for mechanical arm grabbing task in industrial automation

By building a damaged data model and using the environmental diffusion model for denoising training, dynamically adjusting the denoising strategy and optimizing the diffusion parameters, the existing offline reinforcement learning methods performed poorly in data pollution situations, significantly improving the robustness and data quality of the data set, and improving the strategy learning ability of the robotic arm grab task.

CN120144934APending Publication Date: 2025-06-13TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510278350.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing offline reinforcement learning methods perform poorly in the face of data pollution, especially when the data volume is limited. Direct filtering out noise samples may lead to the loss of key information required for decision-making, and the existing diffusion model cannot effectively handle some of the pollution problems of the training data, which is easy to lead to overfitting.

Method used

A disturbed offline reinforcement learning data set recovery method for robotic arm grabbing tasks in industrial automation is adopted. By building a damaged data model, the environmental diffusion model is trained to obtain a noise predictor, distinguish damaged samples from clean samples, and train the denoiser with clean samples, dynamically adjust the denoising strategy, optimize the parameters of the diffusion process, and merge samples to form a recovery data set.

Benefits of technology

It significantly improves the robustness of offline reinforcement learning algorithms on contaminated data sets, avoids overfitting, enhances data quality, improves the strategy learning ability of robotic arm crawling tasks, and improves the crawling success rate and reliability of task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144934A_ABST
    Figure CN120144934A_ABST
Patent Text Reader

Abstract

An interfered off-line reinforcement learning data set recovery method for a mechanical arm grabbing task in industrial automation comprises the steps that for an off-line reinforcement learning data set of the mechanical arm grabbing task, a damaged data model is constructed through a training environment diffusion model (Ambient DDPM), and samples are divided into damaged samples and clean samples through a noise predictor; naive DDPM is adopted to carry out denoising training by using a clean sample, and meanwhile, a denoising strategy is dynamically adjusted to adapt to different noise levels. And the diffusion time step length is optimized through a self-adaptive method, so that the denoising effect is improved and overfitting is prevented. And combining the de-noised samples with the clean samples to form a high-quality recovery data set which is used for offline reinforcement learning training. The method does not need manual annotation, can adapt to a complex noise mode in a mechanical arm grabbing task, improves the data quality, has wide applicability, and effectively enhances the robustness of off-line reinforcement learning on a pollution data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to machine learning technologies, and particularly to a method for recovering a dataset of disturbed offline reinforcement learning for robotic arm grasping tasks in industrial automation. Background Art

[0002] Offline reinforcement learning (Offline RL) is an important method for learning decision-making strategies through offline datasets. However, due to its high dependence on data, Offline RL faces significant challenges when dealing with randomly noisy or adversarially contaminated data. These data contaminations may lead to a serious decline in policy performance and even cause a significant deviation between the learned policy and the expected goal. Therefore, ensuring the robustness of the policy is crucial for the effectiveness of Offline RL in practical applications.

[0003] Currently, existing research has explored the theoretical properties and robustness guarantees of Offline RL under data contamination conditions, and proposed methods such as the Q-ensemble method, Huber loss, and quantized Q-estimation to enhance robustness. In addition, sequence modeling methods have also been used to iteratively correct noisy actions and rewards in the dataset. However, recovering high-dimensional observation data remains a challenge.

[0004] In recent years, diffusion models have received attention in the field of Offline RL due to their powerful data modeling capabilities. Some research has attempted to utilize the denoising characteristics of diffusion models to reduce data noise. For example, existing methods reduce observation noise during the test phase, but these methods assume that the training data is clean and cannot handle partial contamination in the training dataset. In addition, existing diffusion models usually assume that the dataset is either completely clean or has a consistent noise distribution, so overfitting may occur when dealing with partially contaminated data.

[0005] Therefore, the main deficiencies of the existing technologies are as follows:

[0006] 1. Existing Offline RL methods perform poorly under data contamination conditions. Especially when the amount of data is limited, directly filtering out noisy samples may result in the loss of key information required for decision-making.

[0007] 2. Existing sequence modeling methods have limited effectiveness when facing incomplete trajectories. Simply filtering out noisy samples is not sufficient to address the data contamination problem.

[0008] 3. Existing denoising methods based on diffusion models mainly target the test phase and cannot effectively handle partial contamination problems in training data, which easily leads to overfitting.

[0009] For the robotic arm grasping tasks in industrial automation, the challenges faced by offline reinforcement learning (Offline RL) are particularly prominent because such tasks have extremely high requirements for the accuracy and integrity of data. When performing fine operations such as clamping, handling, and assembling, sensor data is extremely vulnerable to interference from noise, errors, and environmental factors, especially when dealing with difficult and complex tasks involving friction, contact forces, and small objects. These interferences not only lead to inaccurate data but also directly affect the training effect and inference performance of the offline reinforcement learning model, causing the learned policy to deviate significantly from the expected goal, thereby affecting the grasping success rate of the robotic arm and the reliability of task execution.

[0010] It should be noted that the information disclosed in the above background art section is only used for understanding the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0011] The main purpose of the present invention is to overcome the defects existing in the above background art, and provide a method for restoring a disturbed offline reinforcement learning dataset for robotic arm grasping tasks in industrial automation.

[0012] To achieve the above purpose, the present invention adopts the following technical solutions:

[0013] A method for restoring a disturbed offline reinforcement learning dataset for robotic arm grasping tasks in industrial automation, comprising the following steps:

[0014] S1. Construct a damaged data model: For the offline reinforcement learning dataset of the robotic arm grasping task, obtain a noise predictor by training an ambient diffusion model (Ambient DDPM). Based on the prediction results of the noise predictor and in combination with a set threshold, divide the samples in the dataset into damaged samples and clean samples, thereby distinguishing data with different noise levels.

[0015] S2. Denoising training based on the diffusion model: Utilize the idea of the ambient diffusion model to estimate the noise level of the robotic arm sensor data through an approximate distribution. Train a denoiser (naive DDPM) using the clean samples in the robotic arm grasping task, and filter the training data using a set threshold. During the training process, dynamically adjust the denoising strategy of the diffusion model to adapt to different degrees of noise pollution.

[0016] S3. Optimize the parameters of the diffusion process: Determine the optimal diffusion time step through an adaptive method, enabling the damaged data to gradually approach the true data distribution during the denoising process while avoiding overfitting.

[0017] S4. Merge the samples to form a restored dataset: Merge the denoised samples with the clean samples to form a restored dataset for offline reinforcement learning training of the robotic arm grasping task.

[0018] Further, in step S1, obtaining the noise predictor by training the environmental diffusion model specifically means training a partially perturbed offline RL dataset. After the training converges, a noise predictor is obtained, which is used to evaluate whether a sample contains noise. Here, the partially perturbed offline RL dataset refers to a dataset in which some samples are contaminated while other samples remain clean.

[0019] Further, in step S1, dividing the samples based on the prediction result of the noise predictor combined with a set threshold means mapping the evaluation result of the noise predictor for the samples to the interval [0, 1], comparing it with the set threshold. Samples greater than the threshold are determined to be damaged samples, and samples less than or equal to the threshold are determined to be clean samples.

[0020] Further, in step S2, estimating the noise level using the idea of the environmental diffusion model means simulating the noise change of the data during the diffusion process based on an approximate distribution, so as to estimate the noise level of the data. Here, the approximate distribution is used to estimate the noise distribution through the environmental diffusion model.

[0021] Further, in step S2, when training the denoiser using clean samples, a binary indicator variable is defined to mark whether the sample is perturbed. When calculating the training loss, the binary indicator variable is used to perform a Hadamard product with the corresponding calculation terms to filter the training data and avoid overfitting.

[0022] Further, in step S2, dynamically adjusting the denoising strategy of the diffusion model means, through an adaptive method, adjusting the denoising intensity and method at each stage of the denoising model according to the noise level of different data points during the diffusion process. Specifically, when the signal-to-noise ratio of the data point is low and the difference from the target distribution is large, the denoising intensity is increased and a more complex denoising method is adopted; when the signal-to-noise ratio is high and it is close to the target distribution, the denoising intensity is reduced and the denoising method is simplified.

[0023] Further, in step S3, determining the optimal diffusion time step through an adaptive method means calculating the noise level of different data points and automatically adjusting the diffusion time step until a time step that best denoises the damaged data and avoids overfitting is found.

[0024] Further, in step S4, merging the denoised samples with the clean samples means performing the reverse DDPM process on all damaged samples to complete denoising, and then integrating them with the clean samples to form a restored dataset.

[0025] Further, when the restored dataset is used for offline reinforcement learning training in step S4, a method of repeating the experiment multiple times and reporting the standard deviation is adopted to ensure the stability of the training results.

[0026] A computer program product includes a computer program which, when executed by a processor, implements the described method.

[0027] The present invention has the following beneficial effects:

[0028] The present invention proposes a method for recovering a disturbed offline reinforcement learning dataset for robotic arm grasping tasks in the field of industrial automation. It innovatively applies a diffusion model to recover the dataset, thereby significantly improving the robustness of the offline reinforcement learning algorithm on contaminated datasets. This method first constructs a damaged model for robotic arm sensor data, trains an environmental diffusion model to obtain a noise predictor, and then effectively distinguishes damaged samples and clean samples in the dataset. On this basis, it uses the idea and approximate distribution of the environmental diffusion model to accurately estimate the noise level of robotic arm sensor data, trains a denoiser with clean samples, and filters the training data using a set threshold to avoid overfitting. In addition, the present invention also optimizes the parameters of the diffusion process through an adaptive method and dynamically adjusts the denoising strategy to adapt to different degrees of noise pollution encountered by the robotic arm in different operating environments. Finally, the denoised samples are merged with the clean samples to form a high-quality recovered dataset for offline reinforcement learning training of robotic arm grasping tasks. This method not only does not require manual annotation of clean data, avoiding the costs and errors brought by manual annotation, but also can adapt to complex noise patterns and handle different degrees and types of noise pollution. By denoising, it improves the training effect of the reinforcement learning model, making it more robust, and at the same time has wide applicability and can be applied to various tasks of cleaning damaged data. Experimental results show that the method of the present invention can significantly improve the performance of the offline reinforcement learning algorithm in handling common random data pollution and adversarial data pollution environments in robotic arm grasping tasks, and has strong generalization ability and practical application value.

[0029] Other beneficial effects in the embodiments of the present invention will be further described below. Description of the Drawings

[0030] Figure 1 It is the overall flowchart of the method for recovering a disturbed offline reinforcement learning dataset for robotic arm grasping tasks in industrial automation according to the present invention.

[0031] Figure 2 It is the algorithm architecture diagram of the method for recovering a disturbed offline reinforcement learning dataset for robotic arm grasping tasks in industrial automation according to the embodiments of the present invention. Detailed Embodiments

[0032] The following provides a detailed description of the embodiments of the present invention. It should be emphasized that the following description is merely exemplary and not intended to limit the scope of the present invention and its applications.

[0033] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0034] Referring to Figure 1 , an embodiment of the present invention provides a method for recovering a disturbed offline reinforcement learning dataset for a robotic arm grasping task in industrial automation (hereinafter referred to as the ADG method), including the following steps:

[0035] Step S1. Construct a damaged data model: For the offline reinforcement learning dataset of the robotic arm grasping task, a noise predictor is obtained by training an ambient diffusion model (Ambient DDPM). Based on the prediction results of the noise predictor and a set threshold, the samples in the dataset are divided into damaged samples and clean samples to distinguish data with different noise levels.

[0036] For example, for the offline reinforcement learning dataset in the robotic arm grasping task of industrial automation, the dataset contains data collected by visual sensors (such as RGB / D cameras, lidars, depth cameras), force / tactile sensors, and joint position sensors (such as encoders and inertial measurement units). A noise predictor is obtained by training an ambient diffusion model (Ambient DDPM). Based on the prediction results of the noise predictor and a set threshold, the samples in the dataset are divided into damaged samples and clean samples to distinguish data with different noise levels. These noises are caused by factors such as light changes, reflections, occlusions, sensor drifts, external impacts, mechanical vibrations, external force disturbances, and high-speed movements on sensor data.

[0037] In some embodiments, obtaining the noise predictor by training the ambient diffusion model specifically means training a partially disturbed offline RL dataset, and obtaining the noise predictor after the training converges, which is used to evaluate whether a sample contains noise. The partially disturbed offline RL dataset refers to some samples in the dataset being contaminated while other samples remain clean. Dividing the samples based on the prediction results of the noise predictor and the set threshold means mapping the evaluation results of the noise predictor for the samples to the interval [0, 1], comparing with the set threshold, and samples greater than the threshold are determined as damaged samples, while samples less than or equal to the threshold are determined as clean samples.

[0038] Step S2. Denoising training based on the diffusion model: Using the idea of the environmental diffusion model, estimate the noise level of the robotic arm sensor data through approximate distribution; use the clean samples in the robotic arm grasping task to train the denoiser (naive DDPM), and filter the training data using a set threshold; during the training process, dynamically adjust the denoising strategy of the diffusion model to adapt to different levels of noise pollution in the robotic arm grasping task.

[0039] For example, this noise pollution may affect the robotic arm's judgment of the object's position, shape, and posture, as well as the perception of contact force, friction force, and the measurement of the motion state. In some embodiments, the estimation of the noise level using the idea of the environmental diffusion model is based on simulating the noise change of the data during the diffusion process through an approximate distribution, so as to estimate the noise level of the data, where the approximate distribution is used to estimate the noise distribution through the environmental diffusion model. When training the denoiser using clean samples, define a binary indicator variable to mark whether the sample is disturbed. When calculating the training loss, perform a Hadamard product with the corresponding calculation terms through this binary indicator variable to filter the training data and avoid overfitting. The dynamic adjustment of the denoising strategy of the diffusion model is through an adaptive method, adjusting the denoising intensity and method at each stage of the denoising model according to the noise level of different data points during the diffusion process. Specifically, when the signal-to-noise ratio of the data point is low and the difference from the target distribution is large, the denoising intensity can be increased and a more complex denoising method can be adopted; when the signal-to-noise ratio is high and it is close to the target distribution, the denoising intensity can be reduced and the denoising method can be simplified.

[0040] Step S3. Optimize the parameters of the diffusion process: Determine the optimal diffusion time step through an adaptive method, so that the damaged robotic arm grasping task data gradually approaches the true data distribution during the denoising process, while avoiding overfitting; thereby improving the prediction accuracy of the interaction force between the object and the robotic arm in the robotic arm grasping task.

[0041] In some embodiments, the determination of the optimal diffusion time step through an adaptive method is to calculate the noise level of different data points and automatically adjust the diffusion time step until the time step that makes the denoising effect of the damaged data the best and avoids overfitting is found.

[0042] Step S4. Combine samples to form a restored dataset: Combine the denoised samples with the clean samples to form a restored dataset for offline reinforcement learning training of the industrial automation robotic arm grasping task. Thus, the stability of the robotic arm grasping strategy can be improved, its generalization ability in different grasping objects and complex environments can be enhanced, and the grasping success rate and task execution reliability can be improved.

[0043] In some embodiments, the merging of the denoised samples with the clean samples is to perform the reverse DDPM process on all the damaged samples to complete denoising, and then integrate them with the clean samples to form the restored dataset. Further, when the restored dataset is used for offline reinforcement learning training, the method of repeating the experiment multiple times and reporting the standard deviation is adopted to ensure the stability of the training results.

[0044] The method for recovering the interference-affected offline reinforcement learning dataset of the present invention can enhance the robustness of offline reinforcement learning in the face of data contamination. This method uses the diffusion probability model, combines the environmental diffusion denoising strategy, and effectively improves the training effect of the model on damaged data by constructing a complete process of noise estimation, diffusion training, and adaptive denoising. Compared with traditional denoising methods, the present invention does not require manual annotation of clean data, reduces the cost and error rate, and can adapt to complex noise patterns and handle different degrees of noise contamination. In addition, this method improves the data quality through denoising, enhances the robustness of the reinforcement learning model, and has wide applicability, and can be applied to various damaged data cleaning tasks.

[0045] The specific embodiments, algorithm examples, and experimental verifications of the present invention are further described below.

[0046] The core method of the present invention is based on the denoising diffusion probability model (DDPM) and combines the idea of the ambient diffusion model to perform denoising training under different noise conditions.

[0047] The method of the present invention includes the following main steps:

[0048] Construct a damaged data model: Since the present invention cannot directly obtain the noise-free real data, the present invention first establishes a hypothesis model of data corruption, that is, some samples in the dataset are contaminated by an unknown noise. The present invention separates these damaged samples from the clean samples so as to perform different processing during the training process.

[0049] Denoising training based on the diffusion model: Traditional diffusion model training requires a consistent noise distribution, while the method of the present invention allows the noise level to vary with the samples. The present invention uses the method of approximate distribution to estimate the noise degree of the data and adjusts the training strategy of the diffusion model accordingly.

[0050] Adjusting Denoising Ability through Approximate Distribution: During the training process of the standard diffusion model, each data point undergoes a series of random noise addition and denoising processes. The method of the present invention further optimizes this process to ensure that the model can adapt to different degrees of noise pollution. The present invention calculates the noise level of different data points during the diffusion process and dynamically adjusts the denoising strategy of the diffusion model to find the optimal balance point between damaged samples and clean samples.

[0051] Optimizing Parameter Selection for the Diffusion Process: In the model of the present invention, the time step of the diffusion process (i.e., the level of data denoising) plays a crucial role. The present invention designs an adaptive method to select the optimal time step, enabling the damaged data to gradually approach the true data distribution during the denoising process while avoiding the model overfitting to the incorrect information of the damaged samples.

[0052] To implement the above method, the present invention also designs the following system architecture:

[0053] Diffusion Model Training Module: Adopting the idea of the environmental diffusion model, it conducts denoising training without directly using noiseless data. By means of approximate distribution, it estimates the noise level of data points during the diffusion process and adjusts the denoising strategy.

[0054] Adaptive Denoising Module: According to the noise conditions of different samples, it dynamically selects suitable diffusion step sizes to ensure the denoising effect.

[0055] Training Optimization and Inference Module: After training is completed, it uses the diffusion model to denoise new data and conducts reinforcement learning training on the denoised data.

[0056] In the robotic arm grasping tasks in the field of industrial automation, sensor data is often disturbed by noise, errors, and environmental factors when performing fine operations such as clamping, handling, and assembly. Especially when dealing with high - difficulty complex tasks involving friction, contact force, and grasping of small objects, this impact is particularly significant, resulting in data inaccuracy, which in turn affects the training effect and inference performance of the offline reinforcement learning model.

[0057] In these grasping tasks, sensor data mainly relies on visual sensors, force / tactile sensors, and joint position sensors. Visual sensors (such as RGB / D cameras, lidars, depth cameras) are responsible for detecting the position, shape, and posture of objects, but are prone to measurement errors under conditions such as lighting changes, reflections, and occlusions. For example, when the robotic arm attempts to grasp a transparent or reflective object, the depth camera may not be able to accurately measure its actual position, resulting in deviation of control instructions. In addition, in low - light environments, the signal quality of RGB visual sensors may decline, possibly leading to object recognition errors or target detection failures.

[0058] Force / tactile sensors are used to sense the contact force and frictional force between the robotic arm and the object. However, due to factors such as sensor drift, external shocks, and mechanical vibrations, the data often contains noise. When the robotic arm moves at high speed or finely adjusts the grasping force, the feedback signal of the force sensor may become unstable, making it difficult for the system to accurately determine whether the object has been successfully grasped. In addition, frictional force plays a crucial role in grasping stability, but in reality, the friction coefficient is affected by factors such as temperature, material, and surface roughness, making it difficult to accurately model. Since the force sensor can only measure the resultant force and cannot directly obtain the true value of the frictional force, the reinforcement learning algorithm may be severely disturbed when learning friction modeling, resulting in policy instability.

[0059] Joint position and velocity sensors (such as encoders and inertial measurement units) are used to measure the motion state of the robotic arm, but their measurement errors may increase when subjected to external force disturbances or high-speed motion. For example, when the robotic arm grasps a dynamically moving object, the encoder data may lag behind the actual motion, resulting in inaccurate prediction of the object's trajectory and affecting the execution effect of the grasping strategy.

[0060] In the robotic arm grasping task based on offline reinforcement learning, these damaged data may significantly affect the training effect of the policy. First, due to inaccurate friction estimation, the reinforcement learning model may mistake some failed grasping operations as successful or misinterpret a low-quality policy as efficient, thus affecting the learning direction of the model. Second, sensor noise may lead to unstable data distribution, making it difficult for the policy to generalize to different grasping objects and environmental conditions. In addition, damaged data may affect the convergence speed of the reinforcement learning model, making the training process more difficult and increasing the uncertainty of policy learning.

[0061] The disturbed data recovery method based on the diffusion model proposed in the present invention can effectively solve the above problems, improve the quality of the offline reinforcement learning dataset, and thus enhance the stability of the robotic arm grasping strategy. First, the method models the distribution of damaged data by training an ambient diffusion model (Ambient DDPM), learning the noise patterns, friction errors, and contact force fluctuations in the data, thereby constructing a damaged data model. Then, a noise predictor is used to detect abnormal data and divide the data into clean samples and damaged samples to avoid damaged data affecting policy learning. During the data recovery process, the method uses a denoising diffusion model (DDPM) for data repair, training the denoising model only with clean samples to ensure that the denoised data is more realistic.

[0062] In addition, this method improves the denoising effect and prevents overfitting by adaptively optimizing the diffusion time step. Since the intensity of sensor noise may vary under different environmental conditions, this method can dynamically adjust the denoising strategy according to the data noise level, making the data recovery more accurate. Finally, the denoised data is merged with the original clean data to form a high-quality reinforcement learning dataset, so as to improve the policy learning ability of the robotic arm grasping task.

[0063] The advantages of this method are that it can improve the accuracy of friction modeling, enabling the offline reinforcement learning policy to more accurately predict the interaction force between the object and the robotic arm, thereby improving the grasping success rate. At the same time, this method can reduce the impact of sensor noise, automatically remove the noise of the force sensor and the vision sensor, and improve the data quality. In addition, this method does not require manual intervention and does not need to manually label damaged data. Instead, it automatically screens and recovers high-quality data through the diffusion model. Finally, this method significantly enhances the stability of the reinforcement learning policy, enabling it to better generalize to different grasping objects and complex environments, and improving the grasping success rate and task execution reliability of the robotic arm in industrial automation scenarios.

[0064] Algorithm Example

[0065] Specifically, as Figure 2 shown, the sliding window technique can be used to process a partially disturbed offline reinforcement learning (RL) dataset. A noise predictor is obtained by training an ambient diffusion model (Ambient DDPM), and the noise predictor is used to evaluate whether a sample contains noise. A restorer is trained using the vanilla DDPM. Among them, a binary indicator variable (Mask) is used to mark whether a sample is disturbed, and a Hadamard product is performed between the binary indicator variable and the corresponding calculation term when calculating the training loss to filter the training data and avoid overfitting. A detector and a restorer are used to repair a partially disturbed dataset, including: processing the disturbed part of the trajectory, identifying the disturbed samples through the detector; using the restorer to denoise the detected disturbed samples to generate a repaired trajectory; and deciding whether to replace the repaired trajectory with the original data copy by judging the quality of the repaired trajectory. Among them, the sliding window technique is used to slide and select a continuous sample subset in the dataset for processing, and the processing flow of the detector and the restorer includes noise evaluation and denoising repair of the samples in each sliding window, so as to gradually recover the entire dataset.

[0066] In the context of the score-based continuous diffusion model, the ambient diffusion model can train a diffusion model on a dataset contaminated by a consistent noise scale. The inventors observed that this conclusion can be seamlessly extended to the discrete DDPM framework. The present invention derives the following lemma to extend this result to discrete DDPM.

[0067] (Environmental DDPM) Assume a given sample Let be the sample with further added noise, where 1 ≤ k a < k. Then, the unique minimization objective of the target is:

[0068]

[0069] where for all k ≥ k a .

[0070] This equation cannot be directly applied because data with a consistent noise distribution is not accessible. Some parts of the dataset are contaminated while other parts remain clean. An ideal alternative is to approximate these distributions and use the approximate distributions to train the diffusion model. Let q(x k |x 0 ) represent the true distribution of the samples generated by the DDPM forward process at diffusion time step k, and ρ(x k |x 0 ) represent the approximation of this distribution. To effectively learn under this approximation, the following assumptions can be made:

[0071] There exists a positive constant c such that for any k ≥ k a , if the KL divergence satisfies then the environmental DDPM can be effectively learned from the samples ρ(x k |x 0 ), and k ≥ k a .

[0072] In a typical robust offline RL setting, the samples provided by the present invention may or may not contain the scaled Gaussian noise a·∈. The distribution of these samples at diffusion time step k is denoted as q(x k |x 0 ).

[0073] For any bounded noise scale a, one can always find a diffusion time step ka such that the environmental DDPM learned from q(x k |x 0 ), for k ≥ k a , can be effectively learned from q(x k |x 0 ). That is, for any c > 0, one can always find ka such that for any k ≥ k a , the following inequality holds:

[0074] Given a sample that may contain Gaussian noise a·∈ The present invention proposes to use the squared Frobenius norm of a noise predictor to determine the presence of noise.

[0075] Assume that the noise prediction error ∈ θ (·, k) follows a Gaussian distribution Define the difference between the noisy

[0076] and noise-free cases as:

[0077]

[0078] where and represent the noisy and noise-free samples respectively, which are different under the original information consistency operation. The signal-to-noise ratio (SNR) of noise prediction can be expressed as:

[0079]

[0080] Assume that the noise prediction error is the same at all diffusion times, i.e., for any k, σ k = σ, then when k = k a , SNR(k) reaches its maximum value.

[0081] The proposed method for repairing a dataset of perturbed offline reinforcement learning based on a diffusion model (ADG) in the present invention is used for interference-robust offline RL. Since two time instances are involved, the superscript k is used to represent the diffusion model time step, and the subscript t is used to represent the reinforcement learning time step for clear expression.

[0082] The true trajectory matrix is defined as where represents the set of perturbed elements (which can correspond to an observation s or a state-action-reward triple (s, a, r)), M represents the dimension of x, and H specifies the time slice size. Let represent the RL elements that may contain noise a·∈. Only the observed (partially perturbed) trajectory

[0083] Given a partially perturbed offline RL dataset, k is predefined a and the ambient diffusion model (Ambient DDPM) is trained as follows:

[0084]

[0085] where After training converges, the noise predictor is obtained, which reaches the maximum SNR at k = k a . Note that for each sample Only focus on whether the RL elements at the central position are interfered, rather than evaluating the entire trajectory segment. To this end, further define where (·)H+1 represents the (H + 1)-th column of the matrix. Subsequently, use to evaluate each sample in the partial interference dataset and rescale the prediction range to between [0, 1]. By manually defining a threshold ζ, when the sample is classified as a noisy sample The remaining samples are regarded as clean samples

[0086] Once the interference samples are identified, the remaining non-interfered samples can be used to train the denoiser (through the naive DDPM), and then can be applied to repair the interference samples. To avoid the DDPM overfitting to misclassified noisy data, ζ is reused to filter the training data of the naive DDPM. Define as a binary indicator variable whether it is interfered (when then otherwise ). Given the mask defined as the training loss of the naive DDPM is

[0087]

[0088] where ⊙ is the Hadamard product. After the training converges, the reverse DDPM process will be performed on all samples in Finally, the denoised samples Dn and Dc are merged to form the final dataset.

[0089] Experimental section

[0090] This section verifies the effectiveness of the ADG method of the present invention through a series of experiments and focuses on the following three core issues: (1) How the ADG method improves the performance of non-robust and robust offline reinforcement learning methods under different data contamination scenarios; (2) What are the respective contributions of the two key components adopted by the ADG method, namely the ambient loss and the selective training, to the final effect; (3) What are the advantages of the structure of the ADG method adopting two independent diffusion models compared with other methods.

[0091] Experimental settings

[0092] Evaluate the performance of ADG on multiple common offline reinforcement learning benchmark datasets, including MuJoCo, Kitchen, and Adroit, etc. Existing studies have shown that when the amount of data is small, the impact of data contamination will be more serious. Therefore, the present invention further explores the performance of the ADG method on datasets of different scales. For this purpose, 10% and 1% of the data are selected for experiments in MuJoCo and Adroit tasks, and additional experimental results with 2% and 5% data volumes are provided in the appendix. In addition, for the MuJoCo task, the "medium-replay" dataset is mainly used, and the experimental results of the "medium" and "expert" datasets are shown in the appendix.

[0093] Regarding data contamination, consider two main types of contamination: random contamination and adversarial contamination. These contaminations can act on information containing only states, or on complete data points of state-action-reward. The specific implementation refers to existing studies, setting the data contamination rate to 30% and the contamination amplitude to 1.0. More implementation details about data contamination are shown in the appendix. In addition, the performance of ADG under Gaussian noise contamination and the effects of different noise ratios and intensities are further verified.

[0094] This experiment evaluates a series of representative offline reinforcement learning methods, including the CQL method based on pessimistic value estimation, the IQL based on policy constraints and its robust variant RIQL, and the DT and RDT methods based on sequence modeling. To ensure the stability of the experimental results, the experiment is repeated on four different random seeds and the standard deviation is reported.

[0095] Evaluation under different data contamination conditions

[0096] Results under random data contamination

[0097] Test the improvement effect of ADG on a variety of offline reinforcement learning algorithms in an environment of random data contamination. It can be seen from the experimental results that ADG can effectively improve the performance of non-robust and robust algorithms in all scenarios. Especially for algorithms based on Markov decision processes (MDPs) (such as CQL, IQL, and RIQL), ADG improves by about 69.1% on average, indicating its significant advantage in complementing missing information in sequences. In addition, ADG also significantly improves methods based on sequence modeling (such as DT and RDT), with an average improvement amplitude of 17.4%.

[0098] It is worth noting that with the help of ADG, IQL and DT can outperform their corresponding robust variants RIQL and RDT in almost all cases. This phenomenon indicates that the core advantage of ADG lies in enabling standard algorithms to outperform specially designed robust variants in terms of noise resistance by simply modifying the dataset. In addition, ADG can also enhance the performance of existing robust algorithms, improving the effect of RIQL by nearly 60%. On the other hand, after being equipped with ADG, most algorithms perform similarly in both the case of only state contamination and complete data contamination, indicating that ADG has strong generalization ability. To further explain the effectiveness of ADG, a visual analysis of the data detection and noise reduction process by ADG is provided, and the effect of data filtering using only ADG is explored.

[0099] Table 1 Performance under Random Data Contamination

[0100]

[0101] Results under Adversarial Data Contamination

[0102] Further analyze the robustness of ADG in the environment of adversarial data contamination. It can be seen from the experimental results that the average improvement of ADG over all baseline methods reaches 24.41%. Notably, IQL and DT equipped with ADG outperform their robust variants RIQL and RDT in all scenarios, and this trend is consistent with the results under random data contamination. This indicates that ADG can effectively adapt to and mitigate the impact of adversarial data contamination.

[0103] Table 2 Performance under Adversarial Data Contamination

[0104]

[0105] To sum up, based on the diffusion probability model and combined with the method of environmental diffusion denoising, the present invention proposes a denoising strategy suitable for corrosion-resistant offline reinforcement learning for the sensor data of the robotic arm grasping task in the field of industrial automation. By constructing a complete process of noise estimation, diffusion training, and adaptive denoising, this method can effectively improve the training effect of the reinforcement learning model on damaged sensor data, while avoiding the limitations of traditional supervised learning methods when dealing with such complex data.

[0106] Compared with traditional denoising methods, the method of the present invention has the following advantages:

[0107] No need for manual annotation of clean data: It avoids the costs and errors brought by manual annotation, especially when dealing with the force / tactile sensor data and visual sensor data of the robotic arm, reducing the dependence on manual annotation.

[0108] Adapt to complex noise patterns: capable of handling different degrees and types of noise pollution, such as the effects of light changes, reflections, occlusions, sensor drift, external shocks, mechanical vibrations, etc. on sensor data.

[0109] Improve data quality: enhance the training effect of the reinforcement learning model through denoising, making it more robust when dealing with key factors such as friction estimation and contact force fluctuations in robotic arm grasping tasks.

[0110] Wide applicability: can be applied to various tasks of cleaning damaged data.

[0111] The invention embodiment also provides a storage medium for storing a computer program, which when executed performs at least the method described above.

[0112] The invention embodiment also provides a control device, including a processor and a storage medium for storing a computer program; wherein, the processor is used to execute the computer program to perform at least the method described above.

[0113] The invention embodiment also provides a processor, which executes a computer program to perform at least the method described above.

[0114] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, Ferromagnetic Random Access Memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the invention embodiment is intended to include but not limited to these and any other suitable types of memories.

[0115] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings, direct couplings, or communication connections between the components shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.

[0116] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0117] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0118] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage media include: removable storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical disks and other various media that can store program codes.

[0119] Alternatively, if the above-mentioned integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present invention. The foregoing storage media include: removable storage devices, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.

[0120] In the method disclosed in several method embodiments provided by the present invention, they can be arbitrarily combined without conflict to obtain new method embodiments.

[0121] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0122] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0123] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the technical field to which the present invention pertains, without departing from the concept of the present invention, several equivalent substitutions or obvious variations can be made, and as long as the performance or use is the same, they should all be regarded as falling within the protection scope of the present invention.

Claims

1. A method for recovering disturbed offline reinforcement learning datasets for robotic arm grasping tasks in industrial automation, characterized in that: The following steps are involved: S1. Constructing a damaged data model: For the offline reinforcement learning dataset of the robot grasping task, a noise predictor is obtained by training the ambient diffusion model (Ambient DDPM). Based on the prediction results of the noise predictor and the set threshold, the samples in the dataset are divided into damaged samples and clean samples to distinguish data with different noise levels. S2. Denoising training based on diffusion model: Using the idea of ​​environmental diffusion model, the noise level of robot arm sensor data is estimated by approximate distribution; clean samples in the robot arm grasping task are used to train the denoiser (naive DDPM), and the training data is filtered by setting a threshold; during the training process, the denoising strategy of the diffusion model is dynamically adjusted to adapt to different levels of noise pollution; S3. Optimize the parameters of the diffusion process: determine the optimal diffusion time step through an adaptive method, so that the damaged data gradually approaches the real data distribution during the denoising process, while avoiding overfitting; S4. Merge samples to form a restored dataset: Merge the denoised samples with the clean samples to form a restored dataset for offline reinforcement learning training of the robotic arm grasping task.

2. The method for recovering a disturbed offline reinforcement learning dataset according to claim 1, characterized in that: In step S1, the noise predictor is obtained by training the environment diffusion model, specifically, training is performed on a partially disturbed offline RL dataset. After the training converges, a noise predictor is obtained to evaluate whether the sample contains noise, wherein the partially disturbed offline RL dataset refers to a dataset in which some samples are contaminated while other samples remain clean.

3. The method for recovering a disturbed offline reinforcement learning dataset according to claim 1, characterized in that: In step S1, the sample is divided based on the prediction result of the noise predictor combined with the set threshold, that is, the evaluation result of the noise predictor on the sample is mapped to the interval [0,1], and compared with the set threshold. Samples greater than the threshold are determined as damaged samples, and samples less than or equal to the threshold are determined as clean samples.

4. The method for recovering a disturbed offline reinforcement learning dataset according to claim 1, characterized in that: In step S2, the noise level is estimated by using the environmental diffusion model concept, which is based on the noise change of the approximate distribution simulation data during the diffusion process to estimate the noise level of the data, wherein the approximate distribution is to estimate the noise distribution through the environmental diffusion model.

5. The method for recovering a disturbed offline reinforcement learning dataset according to claim 1, characterized in that: In step S2, when the denoiser is trained using clean samples, a binary indicator variable is defined to mark whether the sample is disturbed. When calculating the training loss, the binary indicator variable is multiplied by the corresponding calculation item to filter the training data and avoid overfitting.

6. The method for recovering a disturbed offline reinforcement learning dataset according to claim 1, characterized in that: In step S2, the denoising strategy of the dynamic adjustment diffusion model is to adjust the denoising strength and method of the denoising model at each stage according to the noise level of different data points in the diffusion process through an adaptive method.

7. The method for recovering a disturbed offline reinforcement learning dataset according to claim 1, characterized in that: In step S3, the optimal diffusion time step is determined by an adaptive method, which is to automatically adjust the diffusion time step by calculating the noise level of different data points until a time step is found that achieves the best denoising effect on the damaged data and avoids overfitting.

8. The method for recovering a disturbed offline reinforcement learning dataset according to claim 1, characterized in that: In step S4, the denoised samples are merged with the clean samples, which is to perform the reverse DDPM process on all damaged samples to complete denoising, and then integrate them with the clean samples to form a restored data set.

9. The method for recovering a disturbed offline reinforcement learning dataset according to claim 1, characterized in that: In step S4, when the restored data set is used for offline reinforcement learning training, multiple repeated experiments are performed and the standard deviation is reported to ensure the stability of the training results.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.