A robot control method and system

CN122807855APending Publication Date: 2026-09-25XI AN JIAOTONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610817912.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]为提升在上述变化下的稳定性,现有方案多采用基于相机输入的端到端强化学习,并在训练阶段引入随机裁剪、亮度与对比度扰动、模糊与遮挡等数据增强以增强泛化能力;部分方法通过引入对比学习分支或外部网络以稳固特征,导致算力与参数开销增加,或依赖复杂标注或外部先验知识,维护成本高、跨产线迁移缓慢

Benefits of technology

本发明通过计算数据增强后的双视图特征的特征一致性损失作为感知稳健性正则项,以约束编码器对观测扰动的不变性,使同一状态在不同增强条件下的特征更为集中;通过计算学生价值估计与教师价值估计之间的自蒸馏损失作为决策稳健性正则项,以平滑价值函数的学习过程,为价值评估提供时间维度的稳定参照;通过联合两项约束与主强化学习目标经合并优化,使训练曲线更平滑、无效震荡显著减少,达到可用策略所需的交互步数下降,样本效率提升;在面向强反光、光照波动、轻度遮挡与相机微振等真实工况,强化训练流程的鲁棒性,在不改变网络结构、几乎不增加参数的前提下,同时稳定感知表征与价值评估,提高部署模型的确定性与可预期性,进而使得在光照变化、金属反光与轻度遮挡场景中,抓取、对准与旋拧动作更为连贯:末端试探与微抖动减少,接触过程更顺畅,越界保护与力矩超限触发次数下降,稳定性提高,产线流畅度更好。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807855A_ABST
    Figure CN122807855A_ABST
Patent Text Reader

Abstract

The application discloses a mechanical arm control method and system, and relates to the technical field of artificial intelligence and intelligent control.The method comprises the following steps: generating two paths of enhanced views consistent with a baseline for observation at the same time; performing consistency constraint on the features obtained through a shared encoder by distance measurement to weaken appearance disturbance irrelevant to a task; constructing a fixed reference or an exponential moving average teacher value network to provide a stable reference in the time dimension for the current value output and inhibit numerical fluctuations caused by data enhancement; jointly optimizing the two constraints and the original reinforcement learning goal in one back propagation, which only takes effect in the training stage to complete the training of the deployment model, and outputting control instructions to drive the mechanical arm to complete grabbing, alignment or pressing operation; the method can significantly reduce gradient and value fluctuations, improve sample efficiency and convergence speed, and improve operation success rate and running stability under complex lighting and shielding conditions in tasks such as mechanical arm alignment grabbing, alignment and pressing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence and intelligent control technology, specifically to a robotic arm control method and system. Background Technology

[0002] When robotic arms in industrial settings perform operations such as grasping, aligning, twisting, and pressing under complex visual conditions, they often face changes in appearance caused by factors such as lighting switching, mirror reflection of metal parts, camera vibration, partial obstruction, and batch differences in workpieces.

[0003] To improve stability under the above changes, existing solutions mostly adopt end-to-end reinforcement learning based on camera input and introduce data augmentation such as random cropping, brightness and contrast perturbation, blurring and occlusion during the training phase to enhance generalization ability; some methods introduce contrastive learning branches or external networks to stabilize features, resulting in increased computing power and parameter overhead, or rely on complex annotations or external prior knowledge, resulting in high maintenance costs and slow cross-production line migration. However, such methods generally have the following shortcomings: (1) After different augmentations, the feature encoding results of the same physical state are quite different, resulting in the same scene being regarded as different samples, requiring more interaction steps for alignment during training, and the sample efficiency is low; (2) Even with slight augmentation, value assessment is still prone to significant fluctuations, the training process is oscillating, and the strategy frequently changes the path when approaching the workpiece stage, affecting the rhythm and yield; (3) During the training period, a large amount of data augmentation is used while the input during the deployment stage is close to the original image, resulting in inconsistent distribution, which easily leads to a gap between stable laboratory and unstable online.

[0004] In summary, existing methods are prone to drastic fluctuations in value assessment during the end-to-end reinforcement learning training phase based on camera input. This causes control commands to switch frequently between forward, backward, and minor adjustments, resulting in end-effector jitter, overcorrection, and alignment failure, which leads to low stability in grasping, aligning, or pressing operations. Summary of the Invention

[0005] To address the shortcomings of existing technologies in the end-to-end reinforcement learning training phase based on camera input, where value assessment is prone to drastic fluctuations, causing frequent switching of control commands between forward, backward, and minor adjustments, resulting in end-effector jitter, overcorrection, and alignment failure, this invention proposes a robotic arm control method and system. Without changing the network structure or adding almost any parameters, it simultaneously stabilizes the perception representation and value assessment through a dual-view steady-state perception module and a robotic arm operation smoothing control module at the end of the robotic arm, thereby solving the problems existing in the prior art.

[0006] A robotic arm control method includes the following steps: Real-time acquisition of observation data from the robotic arm's camera; The observation data is input into the pre-trained deployment model, and the feature representation of the observation data is extracted through the shared visual encoder. This feature representation is then input into the policy network to generate control commands to drive the robotic arm to complete the corresponding operation tasks. The training process of the deployment model includes the following steps: Two independent data augmentations are performed on the observation data at the same time to generate dual views; feature representations of the dual views are extracted using a shared visual encoder; and feature consistency loss of the dual view features is calculated based on the feature representations of the dual views. Construct a student value network and a teacher value network with the same network structure as the student value network; perform forward computation on the feature representations of the two views through the student value network and the teacher value network respectively to obtain the corresponding student value estimate and teacher value estimate; and calculate the self-distillation loss between the student value estimate and the teacher value estimate. The reinforcement learning main loss, calculated by weighting the feature consistency loss, self-distillation loss, and the value estimate based on the action parameters output by the policy network and the value estimate output by the student value network, is generated by summing the results to form a joint optimization objective. During backpropagation, the parameters of the shared visual encoder, policy network, and student value network are updated based on the joint optimization objective to complete the training of the deployment model.

[0007] Furthermore, the step of calculating the feature consistency loss of the two-view features based on the feature representation of the two views specifically includes the following steps: Views were generated from robotic arm camera observations at the same time using two independent random data augmentation methods. , ; View , All inputs are to a visual encoder with shared parameters. In this process, the corresponding feature representation is obtained. , ; Based on the feature representation of the two views, the feature consistency loss of the two view features is calculated using the L2 alignment method. , is represented as: .

[0008] Furthermore, the teacher value network generates teacher values ​​using either a fixed benchmark or an exponential moving average method; wherein, the fixed benchmark method for the teacher value network involves freezing the weights of the student value network according to a predetermined number of rounds or a stability threshold; the exponential moving average method is... ,in The attenuation coefficient is... For teacher parameters, .

[0009] Furthermore, the calculation process for the self-distillation loss between the student value estimate and the teacher value estimate is expressed as follows: ; in, For the teacher value network, Value network for students.

[0010] Furthermore, the joint optimization objective is expressed as: ; in, This is a feature consistency coefficient used to adjust the proportion of the feature consistency loss. The self-value distillation coefficient is used to adjust the proportion of self-distillation loss. The main loss for reinforcement learning is calculated from the action parameters output by the policy network and the value estimates output by the student value network.

[0011] Furthermore, the reinforcement learning main loss calculated based on the action parameters output by the policy network and the value estimation output by the student value network is compatible with any of the PPO, SAC, and DrQv2 algorithms.

[0012] Furthermore, it also includes dynamically adjusting the self-distillation loss based on the oscillation amplitude during the training process of the deployed model by monitoring cross-view value differences, return curve variance and gradient magnitude, and directional stability indicators. Update strategies for weighting or switching teacher value networks.

[0013] Furthermore, the data enhancement includes at least one or more of the following: random cropping, brightness and contrast perturbation, mild blurring, translation, rotation, and partial occlusion.

[0014] The present invention also includes a robotic arm control system, comprising: The data acquisition module is used to acquire real-time observation data from the robotic arm's camera. The deployment module is used to input observation data into a pre-trained deployment model, extract feature representations of the observation data through a shared visual encoder, input these feature representations into a policy network, and generate control commands to drive the robotic arm to complete corresponding operational tasks; wherein, the training process of the deployment model includes: The feature consistency loss calculation unit is used to perform two independent data augmentations on the observation data at the same time to generate dual views; extract the feature representations of the dual views using a shared visual encoder; and calculate the feature consistency loss of the dual view features based on the feature representations of the dual views. The self-distillation loss calculation unit is used to construct a student value network and a teacher value network with the same network structure as the student value network; to perform forward calculation on the feature representations of the two views through the student value network and the teacher value network respectively to obtain the corresponding student value estimate and teacher value estimate; and to calculate the self-distillation loss between the student value estimate and the teacher value estimate. The joint optimization objective generation unit is used to generate a joint optimization objective by weighted summing of the feature consistency loss, self-distillation loss, and the reinforcement learning main loss calculated based on the action parameters output by the policy network and the value estimate output by the student value network. The update unit is used to update the parameters of the shared visual encoder, policy network, and student value network based on the joint optimization objective during backpropagation, in order to complete the training of the deployment model.

[0015] This invention provides a robotic arm control method, which has the following beneficial effects: This invention uses the feature consistency loss of the augmented dual-view features as a perceptual robustness regularization term to constrain the encoder's invariance to observational perturbations, making the features of the same state more concentrated under different augmentation conditions. It also uses the self-distillation loss between student and teacher value estimates as a decision robustness regularization term to smooth the learning process of the value function, providing a stable time-dimensional reference for value assessment. By combining and optimizing the two constraints with the main reinforcement learning objective, the training curve is smoother, ineffective oscillations are significantly reduced, the number of interaction steps required to achieve a usable strategy decreases, and sample efficiency is improved. In real-world scenarios such as strong reflections, lighting fluctuations, slight occlusion, and camera vibrations, the robustness of the training process is enhanced. Without changing the network structure or increasing parameters, it simultaneously stabilizes perceptual representation and value assessment, improving the determinism and predictability of the deployed model. This results in more coherent grasping, alignment, and twisting actions in scenarios with changing lighting, metallic reflections, and slight occlusion: reduced end-effector probing and micro-jittering, smoother contact processes, fewer out-of-bounds protection and torque over-limit triggers, improved stability, and better production line smoothness. Attached Figure Description

[0016] Figure 1 This is a block diagram of a robust reinforcement learning training system for robotic arm control in an embodiment of the present invention; Figure 2 This is a visual perturbation robustness evaluation diagram in the robotic arm operation scenario of an embodiment of the present invention; Figure 3 This is a distribution diagram of the visual feature encoding of the FC module in the robotic arm operation sub-task in this embodiment of the invention, arranged from strong to weak in each direction. Figure 4This is a comparison chart of the DrQv2 baseline for robotic arm operation in this embodiment of the invention and the multi-task learning curve of this invention; wherein, each sub-graph represents a task, the task name is indicated by the sub-graph name, and several curves for each task represent the performance of different methods; Figure 5 This is a comparison chart of the SAC baseline of the robotic arm operation in this embodiment of the invention and the multi-task learning curve of the present invention; wherein, each sub-graph represents a task, the task name is indicated by the sub-graph name, and several curves for each task represent the performance of different methods; Figure 6 This is a comparison diagram of gradient intensity and directional stability of the robotic arm during training based on a fixed benchmark teacher in an embodiment of the present invention; Figure 7 This is a comparison diagram of gradient intensity and directional stability during the training period of a robotic arm based on an EMA teacher, as shown in this embodiment of the invention. Detailed Implementation

[0017] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0018] This invention proposes a robotic arm control method for industrial robotic arm operations such as grasping, alignment, and pressing. Addressing complex conditions including lighting changes, metal mirror reflections, equipment vibration, and occlusion, the method employs a dual-view steady-state perception module to constrain and solidify perception features, a smoothing control module to smooth the value function learning, and lightweight joint optimization with the original reinforcement learning loss. During training, a cross-view data augmentation consistency index is introduced as a process monitoring and selection criterion. This method does not alter the existing network structure and adds almost no learnable parameters: it improves sample utilization and convergence stability during training; on the deployment side, it maintains the original inference path and latency without increasing computational burden. It is suitable for common robotic arm tasks such as grasping, alignment, and pressing, enabling rapid deployment and stable operation. Figure 1 As shown, the method includes the following stages:

[0019] (1) Design a dual-view steady-state sensing module for the end effector of a robotic arm.

[0020] In real production line robotic arm operations, situations such as high light reflection from metal parts, large variations in light intensity, equipment micro-vibration, and local occlusion are common. Training paradigms that rely solely on data augmentation tend to treat observed disturbances in the same physical state as independent samples, disrupting state equivalence and representation consistency. This leads to perceptual instability, slow policy convergence, and manifests as trajectory instability and motion jitter after deployment.

[0021] Based on this, during the training phase, two enhanced views are generated from the original image at the same time, including cropping, brightness and contrast perturbations, slight blurring, and partial occlusion. Both views are then input into the same shared encoder. Without altering the main algorithm flow and network structure, a feature consistency constraint is introduced to make the encoded features of the two views as similar as possible. This suppresses irrelevant differences caused by lighting variations, reflections, and minor jitter, guiding the model to stably focus on and capture geometric and texture information related to grasping, alignment, and twisting. This consistency constraint is only enabled on the training side; no new parameters or inference overhead are added on the deployment side.

[0022] Compared with existing technologies, this design has the following differentiated features: (1) No new branch networks are added, the main structure is not changed, and the engineering modification is small; (2) Existing data augmentation in training is directly used to transform it from random disturbances into optimized feature signals, thereby improving training stability and sample efficiency; (3) A dual-view design is carried out around the actual working conditions of the robotic arm, such as reflection, occlusion, and micro-vibration, emphasizing deployability and avoiding the limitation of being effective only in the simulation environment.

[0023] The workflow of this module specifically includes the following steps: S1. Acquire camera image data from the robotic arm, covering actual operation scenarios such as grasping, alignment, and twisting, with the acquisition rate consistent with the production line.

[0024] S2. Generate two enhanced views, A and B, for each frame simultaneously, with the enhancement intensity consistent with the existing baseline scheme to ensure fair comparison.

[0025] S3. Input View A and View B into the same set of shared encoders and connect them to the same policy network and value network to form two forward computing paths shared by the front end.

[0026] S4. Calculate the feature consistency index between the feature vectors of view A and view B, and perform joint optimization with the main algorithm training objective for parameter updates.

[0027] (2) Design a smooth control module for robotic arm operation.

[0028] In robotic arm assembly and handling operations in industrial settings, the last few centimeters of the end effector near the workpiece are a high-risk area. Factors such as localized glare and minute changes in viewing angle can cause drastic fluctuations in value assessment, leading to frequent switching of control commands between forward, backward, and minor adjustments. This results in problems such as end effector jitter, overcorrection, and alignment failure. Conventional data augmentation alone is insufficient to eliminate these decision-making jumps, and the convergence curve during training also exhibits significant oscillations.

[0029] Based on this, a teacher value network is introduced during the training phase to provide a stable reference for current value assessment. The engineered implementation of the teacher network includes: first, a fixed benchmark teacher, where historical weights that have shown stable performance during training are frozen as teachers at predetermined intervals; second, an exponential moving average (EMA) teacher, which obtains smooth output by performing a time-weighted average of the value network parameters. During training, the student network, while completing regular reinforcement learning updates, further aligns with the teacher output, thereby imposing temporal consistency constraints on each decision step and significantly reducing value fluctuations caused by factors such as changes in illumination, reflections, and slight blurring. Teacher branches are only enabled during the training phase and are not loaded during the deployment phase, ensuring that inference paths and latency are not affected.

[0030] Compared to existing technologies, this design does not add any new learning branches or modify the original network structure; the teacher network is only used to stabilize the reference and does not participate in deployment inference. Unlike approaches that rely on contrastive learning branches or additional networks, this invention achieves smooth value assessment with low modification costs, can be directly integrated with mainstream algorithms, and complements the dual-view steady-state perception module, with the former used to stabilize representations and the latter used to smooth values.

[0031] The workflow of this module specifically includes the following steps: S1. Initialize the teacher value network, determine whether to use a fixed benchmark teacher or an exponential moving average (EMA) teacher, and set the corresponding update frequency and decay coefficient.

[0032] S2. Read a batch of robotic arm camera data, generate two augmented views synchronously for each frame, use the same shared encoder and value and policy network for forward computation, and obtain the value assessment results of the student network.

[0033] S3. Construct two types of training objectives: one is a conventional reinforcement learning objective, such as a temporal difference objective; the other is a teacher alignment objective, used to measure the difference between student value and teacher output.

[0034] S4. Perform joint optimization of the two types of objectives according to the preset weights, execute backpropagation and parameter update, and only update the student network parameters.

[0035] S5. Continuously monitor the value difference across views and the fluctuation range of the training curve; when the fluctuation exceeds the threshold, increase the teacher alignment weight or switch to a smoother teacher source; when the training stabilizes, gradually reduce the weight.

[0036] S6. Maintain the teacher network as planned, with fixed baseline teachers updated in rounds; EMA teachers updated parameters smoothly based on the set attenuation coefficient.

[0037] (3) Joint constraint training for robotic arm operation.

[0038] In robotic arm operations within industrial production environments, control and decision-making algorithms typically employ PPO, SAC, and DrQv2. However, existing technologies often rely on modifying network structures, adding branches, or introducing auxiliary tasks to suppress instabilities caused by factors such as lighting variations, specular reflection, and occlusion. This results in high modification costs, long deployment cycles, and insufficient transferability between different algorithms. Therefore, there is an urgent need for a unified method that requires no changes to the network structure, has a low parameter tuning burden, and can be directly integrated into existing training scripts.

[0039] Based on this, without modifying the encoder, policy network, and value network, the dual-view steady-state perception module and the smooth control module for robotic arm operation are used as regularization modules and jointly optimized with the original reinforcement learning objective. Only two auxiliary constraints are added on the training side, and no additional branches are loaded on the deployment side, maintaining the existing inference path and ensuring unchanged latency. This solution provides a standardized interface in the form of an adapter, compatible with on-policy algorithms such as PPO and off-policy algorithms such as SAC and DrQv2, and compatible with existing engineering solutions.

[0040] Compared to existing technologies, this method does not require the introduction of new branch networks and does not alter the existing network structure or inference path. By simultaneously solidifying representations and value assessments through regularization during training, it achieves zero additional overhead during deployment. This method employs an interface-based approach, making it compatible with multiple algorithms; its weighting and scheduling strategies are concise and clear, facilitating cross-production line migration.

[0041] The workflow for this joint constraint training includes the following steps: S1. Read a batch of robotic arm camera data and sample it according to a predetermined strategy; among them, PPO adopts round sampling, and SAC and DrQv2 sample through the playback buffer.

[0042] S2. Two enhanced views, A and B, are generated synchronously for each frame and forward computation is completed through the same encoder and the existing policy value network.

[0043] Specifically, two data augmentation views, A and B, consistent with the baseline, are constructed on the robotic arm's camera observations at the same time, and A and B are input into a visual encoder with shared parameters. In order to obtain the corresponding feature vector. , .

[0044] Data augmentation includes at least one or more of the following: random cropping, brightness and contrast perturbation, mild blurring, translation, rotation, and partial occlusion. The augmentation intensity must be consistent with the existing baseline training to ensure comparability.

[0045] S3. Calculate the main training objective. PPO constructs the strategy and value objective based on advantages and rewards. SAC and DrQv2 construct the objective based on temporal differences.

[0046] S4, Computational View Figure 1 Consistency constraints are applied to the encoding features of views from the same source, and a consistency index is calculated to form a feature loss.

[0047] Specifically based on , Calculate the feature consistency loss between two views This is used to constrain the characterization stability of co-source observations under illumination fluctuations, specular reflection, slight occlusion, and camera micro-shakes. It can be either the L2 norm distance or the cosine distance, preferably using... .

[0048] S5. Calculate the self-value distillation constraint, using the value output of the fixed benchmark teacher or the exponential moving average (EMA) teacher as a reference to form the value loss.

[0049] Specifically, teacher value is generated using either a fixed benchmark or an exponential moving average method. Output of student value network Apply alignment constraints to form self-value distillation loss The fixed benchmark method for the teacher value network is to freeze the student value network weights according to a predetermined round or stability threshold and use them as teachers.

[0050] Teacher selection strategy for scenarios with different visual complexity: EMA teachers are preferred for environments with complex visual perturbations or large value variance; fixed benchmark teachers are preferred for environments with dense rewards and simple visuals.

[0051] EMA method is based on Smooth updates, in which The attenuation coefficient is... .

[0052] Self-value distillation loss to measure and The differences between them.

[0053] S6. The weight manager merges the three objectives into one update, performs backpropagation, and only updates the student network parameters; the weights are automatically scheduled according to the rule of larger weights in the early stage and annealing in the later stage.

[0054] Specifically, this will strengthen the main learning objectives. and , In one backpropagation, the parameters of the student network are jointly minimized to update the student network parameters, and only the student network participates in gradient backpropagation.

[0055] The joint optimization objective is ,in, To adjust feature consistency loss The characteristic consistency coefficient of the proportion To adjust self-distillation loss The self-value distillation coefficient of the proportion; set before the strategy converges. and The value of makes the auxiliary regularization term and The numerical magnitude is no less than that of the main loss in reinforcement learning. The magnitude of this is to ensure that the gradient generated by the auxiliary loss dominates in backpropagation, thereby quickly establishing representation invariance and time stability; in the later stages of training, , The main task is to optimize the center of gravity regression to maximize the task reward by using a linear or stepwise decay strategy.

[0056] Strengthen learning objectives It is compatible with any of the algorithms such as PPO, SAC, and DrQv2. The sample sampling strategy is consistent with the selected algorithm; the PPO algorithm is based on round-based sampling, while the SAC and DrQv2 algorithms are based on replay buffer sampling. After sampling, views A and B are generated synchronously in batches and calculated. and .

[0057] S7, Record Return Curve, View Figure 1 Consistency indicators and value fluctuation indicators are used for training monitoring and selection, and teacher versions are maintained according to plan.

[0058] The training process records and monitors cross-view value differences, return curve variance and gradient magnitude, and directional stability metrics. When the oscillation amplitude exceeds a threshold, improvements are made. The weights may be shifted to a smoother source of teachers, and the weights will be reduced after the training stabilizes.

[0059] During the deployment phase, the regularization branches corresponding to steps S4 and S5 are not loaded, and inference is performed along the original main path of camera observation → encoder → policy and value network.

[0060] This invention addresses the problems of unstable characterization and value assessment, low sample efficiency, and slow convergence caused by strong metal reflection, lighting fluctuations, slight occlusion, and camera micro-vibrations in actual production lines. Without altering the deployment and inference structure, it introduces two types of regularization constraints during the training phase, forming a unified joint training process: First, a dual-view steady-state perception module (FC) at the robotic arm end effector generates two enhanced views consistent with the baseline for observations at the same time. Features obtained through a shared encoder are constrained for consistency using distance metrics, reducing appearance perturbations irrelevant to the task. Second, a smoothing control module (SD) for robotic arm operation constructs a fixed benchmark or exponential moving average teacher value network, providing a stable temporal reference for the current value output and suppressing numerical fluctuations caused by data augmentation. These two constraints are jointly optimized with the original reinforcement learning objective in a single backpropagation, taking effect only during the training phase. This method is adapted to mainstream algorithms such as PPO, SAC, and DrQv2 via an adapter approach. In tasks such as robotic arm positioning, grasping, alignment, and pressing, it can significantly reduce gradient and value fluctuations, improve sample efficiency and convergence speed, and enhance the success rate and operational stability under complex lighting and occlusion conditions.

[0061] The present invention has the following advantages: (1) It enhances the robustness of the training process in real working conditions such as strong reflection, light fluctuation, slight occlusion and camera micro-vibration. It stabilizes the perception representation and value assessment at the same time without changing the network structure or increasing the parameters. (2) Through the dual-view steady-state perception module at the end of the robotic arm, the features of the same state under different enhancement conditions are more concentrated; through the smooth control module of the robotic arm operation, a stable reference for the time dimension of value assessment is provided. The above two constraints and the main reinforcement learning objective are merged and optimized to make the training curve smoother, reduce invalid oscillations significantly, reduce the number of interaction steps required to reach the usable strategy, and improve sample efficiency. (3) Only the original encoder and the reasoning path of the strategy and value network are retained, without introducing new branches. The reasoning delay and computing power consumption are consistent with the existing scheme. After going online, in the scenarios of light change, metal reflection and slight occlusion, the grasping, alignment and twisting actions are more coherent: the end probe and micro-shaking are reduced, the contact process is smoother, the number of out-of-bounds protection and torque over-limit triggering is reduced, the stability is improved, and the production line is smoother. (4) The plug-and-play training paradigm is compatible with mainstream algorithms such as PPO, SAC, and DrQv2. It requires minimal modification, facilitates cross-production line migration, reduces the cost of tuning and maintenance, and improves the determinism and predictability of the model.

[0062] Based on the above, this invention proposes an embodiment of a robust reinforcement learning training system for robotic arm control, specifically including: (1) Dual-view steady-state sensing module at the end of the robotic arm This invention proposes a visual representation stabilization module. This module is used during pixel-level reinforcement learning training with data augmentation in robotic arm operations to ensure consistent encoded features for observations at the same time under different data augmentation conditions, thus suppressing learning instability introduced by data augmentation at its source. The module is only activated during the training phase and does not add any computational overhead during deployment.

[0063] Views are generated by augmenting the same observation using two independent random data sets. , The enhancement type is consistent with the baseline, such as random cropping, translation, and rotation, to ensure comparability; both views are input to a visual encoder with shared parameters. , thus obtaining the feature representation , The module output is the consistency loss. This branch is used for joint backpropagation with the main training objective; when deploying online, this branch is not exported, and only the main path of camera → encoder → policy and value target network is retained.

[0064] The module uses L2 alignment as a consistency target: ; Its engineering significance lies in transforming the appearance differences generated by data augmentation into a regularized signal that can be optimized during training, constraining the encoder to maintain a stable, task-related representation under perturbations such as brightness fluctuations, specular reflections, slight occlusion, and slight blurring. Unlike contrastive learning methods, this consistency constraint does not rely on negative samples or additional projection branches, requires no changes to the network structure, and does not require the addition of learnable parameters, making it easy to integrate with various reinforcement learning algorithms.

[0065] like Figure 2 As shown, under perturbations such as viewpoint micro-rotation, relative displacement, and field-of-view clipping, common data augmentation significantly increases the difficulty of representation learning; in grasping and alignment scenarios, this problem is often caused by high metallic reflectivity, shadows, and occlusion. By directly zooming in... , The distance allows for the smoothing of task-independent appearance variations within the latent space, enabling subsequent strategy and value learning to be built upon more stable inputs. Correspondingly, Figure 3 Show introduction Afterwards, the feature orientation distribution becomes more compact and sparse, high-variance but task-irrelevant orientations are suppressed, the representation becomes more concentrated, which is conducive to improving sample utilization efficiency.

[0066] Each iteration is for batch size of Two enhancements are generated simultaneously from the sample, forming Two images; after forward propagation of the two views through a shared encoder, the pairwise consistency loss is computed locally at a specified feature layer. It is then jointly minimized with the main reinforcement learning objective in a single backpropagation, achieving synchronous optimization of the encoder, policy, and value network.

[0067] (2) Smooth control module for robotic arm operation This invention addresses the problem of drastic fluctuations in value assessment and training instability caused by data augmentation in pixel-level robotic arm training. It provides a temporal teacher constraint that is configurable during the training phase and requires no modification during the deployment phase. The core idea is to use a smoother historical value estimate as a benchmark to align the current value output, thereby suppressing numerical jitter caused by data augmentation and stabilizing the optimization process.

[0068] In tasks such as grasping, aligning, and rotating, the same semantic state often manifests in multiple appearances due to light reflection, slight occlusion, or slight blurring. During training, this is generated by the student's value network. Simultaneously, a teacher value network is being constructed. The value is obtained by time-weighted averaging of student parameters and used as a reference. The SD branch applies a mean-square constraint to both values ​​to obtain the distillation loss. Participate in joint reverse propagation.

[0069] During training, from the current moment t Iterative Path Student Value Network produce Simultaneously, a teacher value network is being constructed. produce This network represents a historical iteration moment. t - i The stable state, whose parameters are obtained based on historical parameters of the student value network through exponential moving average or periodic freezing, is used as a value reference benchmark; the distillation loss between the two outputs is calculated. This can effectively suppress numerical fluctuations caused by data augmentation.

[0070] The teacher uses an exponential moving average (EMA) to update the parameters: ; Teacher parameters do not participate in backpropagation, and the regularization term is defined as: ; It is used to update the current value assessment towards a historically stable reference, reducing transient biases caused by data augmentation.

[0071] In robotic arm tasks, while data augmentation increases sample diversity, it amplifies value uncertainty: the same latent state is assigned significantly different values ​​under different data augmentations, and the error accumulates through Bellman recursion, leading to unstable training and decreased sample efficiency. SD (Simplified Data Augmentation) uses historical networks as anchors to effectively reduce the variance introduced by data augmentation and improve learning efficiency. This method does not rely on external experts and can be formed solely using the model's own historical trajectories, making it particularly suitable for scenarios with disturbances such as reflections and shadows.

[0072] Based on different robotic arm operation tasks, two update strategies are proposed. The first is the EMA teacher (SD-p), characterized by continuous time smoothing, suitable for more complex and dynamically unstable environments such as those with strong reflections, significant viewpoint jitter, etc. The second is the fixed reference teacher (SD-r), characterized by periodic replacement, suitable for tasks with simpler vision and denser rewards, and can avoid the slight oscillations of the EMA target itself. Static teachers are preferred for visually simple scenarios such as DMControl, while EMA teachers are preferred for complex vision scenarios such as Atari. Depending on the complexity of the situation, the more stable teacher type can be selected as a reference.

[0073] Experiments show that introducing SD significantly reduces the variance of value assessment, smooths the training curve, and improves convergence speed and sample efficiency. Specifically, for example... Figure 4 As shown, in multiple tasks including inverted pendulum swing, bipedal walking, and quadrupedal walking, the FCSD-r and FCSD-p curves proposed in this invention exhibit significantly higher slopes in the early 0.2M-0.4M steps compared to baseline algorithms such as DrQv2. This demonstrates the effectiveness of feature consistency loss. and self-distillation loss With the joint constraints, the model can learn stable representations from the original pixels more quickly, significantly reducing the number of interaction steps required to reach a usable policy. After integrating DrQv2, the overall performance on 20 DMControl tasks is improved across the board, achieving higher performance at the same number of times. Figure 5This paper presents a performance comparison between the FCSD-r method of this invention and the native SAC algorithm in eight typical visual dynamics tasks. In all eight tasks, the slope of the reward increase of the red curve (the method of this invention) is significantly higher than that of the blue curve (the baseline algorithm), demonstrating the superior performance of this invention in improving sample utilization efficiency. Especially in the "Walker Walk" and "Reacher Hard" tasks, the native SAC algorithm, limited by complex pixel-level inputs, shows almost no increase in reward within 1.0M steps, exhibiting obvious training stagnation. In contrast, the method of this invention achieves a significant improvement in reward in the above tasks, proving that the dual-view steady-state perception module (FC), by constraining feature consistency, enables the model to extract key geometric and texture information from the original observations more quickly. Integrating SD into the native SAC also significantly improves sample efficiency and the final score, indicating that SD not only enhances pixel RL but also stabilizes the training of the basic algorithm. At the same sampling frequency, the number of successful grabbing and alignment rounds increases, the number of failed rounds decreases, and the training iteration fluctuation is reduced.

[0074] (3) Joint constraint training process for robotic arm operation This invention proposes a training scheme for robotic arm control tasks. Without altering the deployment inference network structure or adding new online branches, the main reinforcement learning objective, the robotic arm end-effector dual-view steady-state perception module, and the robotic arm operation smoothing control module are jointly optimized in the same round of backpropagation. Specifically, two data-augmented views are fed into a shared encoder, and the encoder output is used to compute the dual-view... Figure 1 Consistency constraints Policy networks and value networks are used to compute the main task objective. and self-value regularization constraint The three losses have been unified, merged, optimized, and updated.

[0075] In the process of unifying the objective function and weights, the joint objective is written as: ; in, The primary objective is the selected reinforcement learning algorithm, such as TD error or PPO shearing target. For binocular vision Figure 1 Coherence term; This is the self-valued distillation term. To reduce the complexity of parameter tuning, , The same value can be used; moderately increase the value in the early stages of training. , To accelerate the establishment of the invariance of the representation and the time stability of the value, the latter part is gradually reduced according to the linear or stepwise strategy, so that the optimization focus returns to the main task.

[0076] This paradigm, as a training-phase plugin, can be directly integrated with commonly used pixel-based reinforcement learning (RL). Represented by off-policy algorithms such as DrQv2 and SAC, it reads value and forms a policy network after pixel enhancement and the encoder. And calculate at the encoder output. And perform parallel computation on the value network side. The three loss terms are combined and backpropagated once. The above constraints only participate in gradient updates during the training period; Figure 4 , Figure 5 The study demonstrates the significant benefits of DrQv2 and SAC in DMControl multitasking, achieving higher scores and faster convergence at the same number of synchronizations, and achieving higher success rates and shorter completion times in robotic arm tasks.

[0077] The specific steps involved in implementation are as follows: S. Generate two data augmentation views consistent with the baseline for the same observation for each sample. , .

[0078] S2, will , Input shared encoder , to obtain features , It also completes the regular forward computation of the policy network and the value network.

[0079] S3. Calculate the three losses. , ,as well as .

[0080] S4, with Combine backpropagation with optimizer update, and only update student network parameters; , Initially, use the same values ​​to simplify parameter tuning, and later anneal according to the preset schedule.

[0081] S5. Maintain teacher parameters after each update. Updates are performed using either exponential moving average (SD-p) or periodic versioning (SD-r). The teacher network is used only to generate supervisory signals and does not participate in backpropagation or deployment inference.

[0082] This invention aims to address the stable operation of robotic arms under complex visual conditions by proposing a training paradigm that does not alter the existing network structure and adds almost no parameters. It solves the problem of maintaining feature consistency under perturbation conditions, improving sample utilization efficiency; it makes value assessment smoother and more predictable across views and time dimensions, reducing decision jumps during training and inference; and it provides a highly portable training process that can be directly integrated into existing algorithms such as PPO, SAC, and DrQv2, reducing modification and migration costs. To adapt to stable training and rapid convergence for visual operation tasks such as robotic arm alignment and trajectory following, this invention introduces an FCSD stabilization mechanism during training. Training instability is typically manifested by drastic fluctuations in gradient magnitude and frequent changes in gradient direction, leading to inefficient learning and slow convergence. Figure 6 This paper demonstrates a comparison of the gradient magnitude (Gradient Norm) and gradient direction stability (Gradient Direction) of the proposed method (FCSD-r, fixed benchmark teacher strategy) and the baseline algorithm (DrQv2) during training. The experiments selected two typical tasks: high-precision robotic arm target tracking and multi-joint dynamic stabilization control. Figure 6 In all subplots, the gradient magnitude (GradientNorm) of the red curve (FCSD-r) is significantly lower than that of the blue baseline, and the magnitude variance (shaded area) is greatly suppressed. This indicates that the present invention achieves this through self-distillation loss. It provides a stable anchor point for training, avoiding gradient explosion or violent fluctuations caused by visual perturbations. In the gradient direction stability comparison, the red curve shows a smaller fluctuation amplitude, which means that the model chooses a more consistent and clear direction when updating parameters, effectively reducing invalid attempts during training. This smooth gradient behavior directly corresponds to a more robust policy update process. Figure 7This paper demonstrates a comparison of the gradient behavior during training between the proposed method (FCSD-p, employing an exponential moving average (EMA) teacher strategy) and the baseline algorithm (DrQv2). The experiment was conducted on high-precision robotic arm target tracking and multi-joint dynamic stabilization control tasks to verify the effectiveness of the smoothing control module under different teacher update strategies. In both tasks, the gradient magnitude (GradientNorm) of the red curve (FCSD-p) was lower than that of the blue baseline, and the shaded portion (magnitude variance) significantly contracted. This indicates that even in environments with complex visual perturbations, the EMA teacher, through time-weighted averaging of the value network parameters, can still provide a smooth supervisory signal to the student network, preventing drastic gradient oscillations. In the gradient direction index, the red curve exhibits lower fluctuation frequency and amplitude than the baseline, meaning a clearer parameter update path. This reduces decision-making iterations caused by instantaneous biases introduced by data augmentation, thereby improving the reliability of convergence. Figure 6 , Figure 7 The results show that, compared with the baseline, the overall gradient magnitude is significantly reduced, the magnitude variance is suppressed, and the directional variation is significantly reduced after adopting this invention; among them, the fixed benchmark teacher (FCSD-r) has a stronger stabilization effect than the EMA teacher (FCSD-p). This phenomenon is consistent with the performance comparison results, both methods improve sample efficiency, but FCSD-r has a more significant improvement.

[0083] In summary, this invention achieves smoother policy updates and a more predictable learning process by stabilizing gradient behavior, thereby obtaining faster and more reliable convergence in robotic arm scenarios. In engineering implementation, gradient magnitude and direction stability can be used as the basis for monitoring during training and model export.

[0084] Based on the same inventive concept, this invention also proposes a robotic arm control system, comprising: The acquisition module is used to acquire real-time observation data from the robotic arm's camera.

[0085] The deployment module is used to input observation data into a pre-trained deployment model, extract feature representations of the observation data through a shared visual encoder, input these feature representations into a policy network, and generate control commands to drive the robotic arm to complete corresponding operational tasks. The training process of the deployment model includes: The feature consistency loss calculation unit is used to perform two independent data augmentations on the observation data at the same time to generate dual views; extract the feature representations of the dual views using a shared visual encoder; and calculate the feature consistency loss of the dual view features based on the feature representations of the dual views.

[0086] The self-distillation loss calculation unit is used to construct a student value network and a teacher value network with the same network structure as the student value network; the feature representations of the two views are forward-computed through the student value network and the teacher value network respectively to obtain the corresponding student value estimate and teacher value estimate; and the self-distillation loss between the student value estimate and the teacher value estimate is calculated.

[0087] The joint optimization objective generation unit is used to generate a joint optimization objective by weighted summing of the feature consistency loss, self-distillation loss, and the reinforcement learning main loss calculated from the action parameters output by the policy network and the value estimate output by the student value network.

[0088] The update unit is used to update the parameters of the shared visual encoder, policy network, and student value network based on the joint optimization objective during backpropagation, in order to complete the training of the deployment model.

[0089] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A robotic arm control method, characterized in that, Includes the following steps: Real-time acquisition of observation data from the robotic arm's camera; The observation data is input into the pre-trained deployment model, and the feature representation of the observation data is extracted through the shared visual encoder. This feature representation is then input into the policy network to generate control commands to drive the robotic arm to complete the corresponding operation tasks. The training process of the deployment model includes the following steps: Two independent data augmentations are performed on the observation data at the same time to generate dual views; feature representations of the dual views are extracted using a shared visual encoder; and feature consistency loss of the dual view features is calculated based on the feature representations of the dual views. Construct a student value network and a teacher value network with the same network structure as the student value network; perform forward computation on the feature representations of the two views through the student value network and the teacher value network respectively to obtain the corresponding student value estimate and teacher value estimate; and calculate the self-distillation loss between the student value estimate and the teacher value estimate. The reinforcement learning main loss, calculated by weighting the feature consistency loss, self-distillation loss, and the value estimate based on the action parameters output by the policy network and the value estimate output by the student value network, is generated by summing the results to form a joint optimization objective. During backpropagation, the parameters of the shared visual encoder, policy network, and student value network are updated based on the joint optimization objective to complete the training of the deployment model.

2. The robotic arm control method according to claim 1, characterized in that, The step of calculating the feature consistency loss of the two-view features based on the feature representation of the two views specifically includes the following steps: Views were generated from robotic arm camera observations at the same time using two independent random data augmentation methods. , ; View , All inputs are to a visual encoder with shared parameters. In this process, the corresponding feature representation is obtained. , ; Based on the feature representation of the two views, the feature consistency loss of the two view features is calculated using the L2 alignment method. , represented as: 。 3. The robotic arm control method according to claim 2, characterized in that, The teacher value network generates teacher values ​​using either a fixed benchmark or an exponential moving average method. The fixed benchmark method involves freezing the student value network weights according to a predetermined number of rounds or a stability threshold. The exponential moving average method involves... ,in The attenuation coefficient is... For teacher parameters, .

4. The robotic arm control method according to claim 3, characterized in that, The calculation process for the self-distillation loss between student valuation and teacher valuation is as follows: ; in, For the teacher value network, Value network for students.

5. The robotic arm control method according to claim 4, characterized in that, The joint optimization objective is expressed as: ; in, This is a feature consistency coefficient used to adjust the proportion of the feature consistency loss. The self-value distillation coefficient is used to adjust the proportion of self-distillation loss. The main loss for reinforcement learning is calculated from the action parameters output by the policy network and the value estimates output by the student value network.

6. The robotic arm control method according to claim 1, characterized in that, The reinforcement learning main loss calculated based on the action parameters output by the policy network and the value estimates output by the student value network is compatible with any of the PPO, SAC, and DrQv2 algorithms.

7. The robotic arm control method according to claim 1, characterized in that, This also includes dynamically adjusting the self-distillation loss based on the oscillation amplitude during the training process of the deployed model by monitoring cross-view value differences, return curve variance and gradient magnitude, and directional stability indicators. Update strategies for weighting or switching teacher value networks.

8. The robotic arm control method according to claim 1, characterized in that, The data augmentation includes at least one or more of the following: random cropping, brightness and contrast perturbation, mild blurring, translation, rotation, and partial occlusion.

9. A robotic arm control system, characterized in that, include: The data acquisition module is used to acquire real-time observation data from the robotic arm's camera. The deployment module is used to input observation data into a pre-trained deployment model, extract feature representations of the observation data through a shared visual encoder, input these feature representations into a policy network, and generate control commands to drive the robotic arm to complete corresponding operational tasks; wherein, the training process of the deployment model includes: The feature consistency loss calculation unit is used to perform two independent data augmentations on the observation data at the same time to generate dual views; extract the feature representations of the dual views using a shared visual encoder; and calculate the feature consistency loss of the dual view features based on the feature representations of the dual views. The self-distillation loss calculation unit is used to construct a student value network and a teacher value network with the same network structure as the student value network; to perform forward calculation on the feature representations of the two views through the student value network and the teacher value network respectively to obtain the corresponding student value estimate and teacher value estimate; and to calculate the self-distillation loss between the student value estimate and the teacher value estimate. The joint optimization objective generation unit is used to generate a joint optimization objective by weighted summing of the feature consistency loss, self-distillation loss, and the reinforcement learning main loss calculated based on the action parameters output by the policy network and the value estimate output by the student value network. The update unit is used to update the parameters of the shared visual encoder, policy network, and student value network based on the joint optimization objective during backpropagation, in order to complete the training of the deployment model.