Self-learning joint control method and system based on multi-modal feedback and reinforcement learning
By employing a self-learning splice control method based on multimodal feedback and reinforcement learning, the adaptability and perception issues of the splice system in the spinning workshop under environmental changes were solved. This resulted in high-precision and stable splice control, reduced manual intervention, and improved the system's intelligence level.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU WEI RUIXIN ROAD TECHNOLOGY CO LTD
- Filing Date
- 2026-03-17
- Publication Date
- 2026-04-14
AI Technical Summary
Existing twin-arm splicing robots or automatic splicing systems in spinning workshops have poor environmental adaptability, limited perception dimensions, and lack of autonomous learning capabilities when faced with changes in yarn raw materials, yarn count, and workshop environment. This leads to problems such as missing the target, breaking the yarn, and substandard splicing quality.
A self-learning connector control method based on multimodal feedback and reinforcement learning is adopted. Real-time sensing data is acquired through a multimodal synchronous sensing array, and multi-level feature extraction and deep fusion are performed to establish an adaptive decision engine based on deep reinforcement learning. Combined with a multi-objective composite reward function and an online training optimization mechanism, adaptive connector control is achieved.
It improves the adaptability of joint control, enhances the accuracy and stability of joint control, reduces manual debugging costs, and strengthens the intelligence level and adaptability of the joint process.
Smart Images

Figure CN121859945A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automation control technology, and in particular to a self-learning connector control method and system based on multimodal feedback and reinforcement learning. Background Technology
[0002] Currently, the core technologies of dual-arm splicing robots or automatic splicing systems for ring spinning workshops are mostly based on preset fixed programs. The workflow is typically as follows: rough positioning is performed through simple vision, and actions such as grasping, yarn feeding, and knotting are executed according to pre-programmed fixed trajectories and parameters (such as clamping force, yarn feeding speed, and knotting force). Existing technologies specifically include mechanical cam mechanisms, track-based fixed-program robots, and teach-and-playback robots.
[0003] Mechanical automatic splicing devices primarily utilize cam mechanisms, programmable controllers, and fixed actuators (such as yarn guides and splicers). Their principle is based on a preset rigid mechanical program and motion trajectory, sequentially executing standardized actions such as yarn finding, yarn feeding, and splicing. Track-based robots are used in multi-spindle scenarios. Their structure includes a track-moving platform, a multi-axis robotic arm, and a machine vision system. Their principle is to visually identify the breakage location, and then the motion control system drives the robotic arm to execute a pre-programmed fixed splicing path. Teach-and-playback robots allow an operator to manually guide the robotic arm to complete a successful splice and record the parameters of the entire motion trajectory. Their principle is to accurately reproduce these recorded standard actions to mimic human operation.
[0004] The main problems and structural reasons for the existing technology are as follows: Extremely poor environmental adaptability: Due to its fixed mechanical structure (such as rigid cams) and control program, the system cannot sense or adapt to changes in yarn raw materials (cotton, linen, chemical fibers), yarn count (fineness), workshop temperature and humidity (affecting yarn strength and coefficient of friction), and airflow disturbances. This directly leads to frequent problems such as missing yarn, yarn breakage, knotting failure, or substandard joint quality (tightness, size) when conditions change, resulting in large fluctuations in success rate. The mechanical rigid cam structure and track-based fixed program prevent it from sensing environmental changes such as yarn tension and spindle position deviation. Therefore, when the yarn type or workshop environment changes, the original preset parameters are no longer applicable, exhibiting extremely poor versatility.
[0005] Limited Perception and Blind Control: Most systems rely solely on two-dimensional vision for initial positioning, severely lacking real-time perception of key physical quantities during execution. For example, the absence of force feedback makes it impossible to determine whether the clamping force is appropriate (excessive force damages the yarn, insufficient force causes slippage); the lack of tension perception leads to blind yarn feeding; and the lack of proximity perception makes it impossible to precisely control the relative position of the actuator and the tiny yarn guide hook. This "open-loop" control makes robot operation clumsy and highly dependent on the positioning accuracy and consistency of precision machinery. Because existing structures generally lack real-time force feedback mechanisms or only use fixed impedance parameters, their position control systems cannot respond to the complex contact dynamics changes of the yarn, a flexible body, during splicing, resulting in inaccurate control of yarn feeding speed and twisting force, frequently leading to yarn breakage or unstable splice quality.
[0006] Completely lacking autonomous learning and evolutionary capabilities: System performance depends entirely on initialization parameters. Any optimization must be done manually by engineers analyzing logs, adjusting code, and redeploying—a tedious process highly dependent on expert experience. The robot itself cannot learn autonomously from massive amounts of successful or failed experiences, and cannot achieve "experience accumulation" or self-performance improvement. Whether it's a track-based programming path or a teach-and-playback recorded trajectory, its system structure is inherently closed, lacking the ability to learn from interactive data and self-optimize. Therefore, it cannot autonomously adjust and optimize its connector strategy to adapt to new situations or achieve multi-objective performance improvements through accumulated experience, relying entirely on manual intervention and parameter retuning. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention provides a self-learning connector control method and system based on multimodal feedback and reinforcement learning.
[0008] To achieve the above objectives, in a first aspect, this invention provides a self-learning joint control method based on multimodal feedback and reinforcement learning. The method includes the following steps: constructing a multimodal synchronous sensing array to acquire real-time multimodal sensing data during the joint process; performing multi-level feature extraction and deep fusion on the real-time multimodal sensing data to obtain a high-dimensional state vector; constructing an adaptive decision engine based on deep reinforcement learning, inputting the high-dimensional state vector into the adaptive decision engine to output a joint control command; establishing a multi-objective composite reward function, obtaining a composite reward quantification index for the joint control command based on the multi-objective composite reward function; introducing an online training optimization mechanism, training and optimizing the adaptive decision engine in conjunction with the composite reward quantification index to obtain an optimized joint control command; executing the optimized joint control command and acquiring full-state actual joint data, iterating the adaptive decision engine in a closed loop to complete the self-learning joint control. This invention effectively overcomes the dependence of traditional control on fixed scenarios, improves the adaptive capability of joint control under complex working conditions, balances control accuracy and stability, reduces manual debugging costs, and significantly enhances the intelligence level and adaptability of the joint process.
[0009] Optionally, the construction of the multimodal synchronous sensing array, and the acquisition of multimodal real-time sensing data during the jointing process through the multimodal synchronous sensing array, includes: constructing the multimodal synchronous sensing array through a visual sensing array, a force sensing array, and a pose sensing array; during the jointing process, acquiring visual image data, force signal data, and pose state data through the multimodal synchronous sensing array as the multimodal real-time sensing data. This invention ensures the comprehensiveness and accuracy of data acquisition during the jointing process, providing high-quality data support for subsequent feature processing and decision-making, reducing control deviations caused by data gaps, and improving the controllability of the jointing process.
[0010] Optionally, the step of performing multi-level feature extraction and deep fusion on the multimodal real-time sensing data to obtain a high-dimensional state vector includes: extracting visual features, force features, and pose features from the multimodal real-time sensing data to obtain heterogeneous feature vectors, which include visual state vectors, force state vectors, and pose state vectors; establishing a multimodal feature fusion network to concatenate the heterogeneous feature vectors in a preset order to form an original feature vector; and inputting the original feature vectors into a multilayer perceptron for fusion mapping to generate the high-dimensional state vector. This invention accurately represents the joint condition through a high-dimensional state vector, providing comprehensive state information for the decision engine, improving the scientific nature of decision commands, and helping to achieve more precise joint control.
[0011] Optionally, the step of constructing an adaptive decision engine based on deep reinforcement learning, and inputting the high-dimensional state vector into the adaptive decision engine to output joint control commands, includes: constructing the adaptive decision engine based on the actuator-evaluator algorithm framework in deep reinforcement learning; inputting the high-dimensional state vector into the adaptive decision engine to output incremental adjustment commands of the underlying actuator control parameters as the joint control commands. This invention rapidly responds to changes in the high-dimensional state vector, dynamically optimizes the underlying execution parameters, avoids sudden changes in control commands impacting the joint, and improves the smoothness and reliability of joint actions.
[0012] Optionally, establishing a multi-objective composite reward function and obtaining a composite reward quantification index for the joint control command based on the multi-objective composite reward function includes: acquiring a multi-dimensional reward quantification index; constructing the multi-objective composite reward function based on the multi-dimensional reward quantification index; calculating the total reward value of the joint control command based on the multi-objective composite reward function; and using the total reward value as the composite reward quantification index. This invention scientifically evaluates the comprehensive performance of control commands, provides reasonable guidance for decision engine optimization, balances multi-dimensional requirements such as joint quality and efficiency, and ensures that the optimized control commands are more suitable for actual application scenarios.
[0013] Optionally, the step of obtaining the multi-dimensional reward quantification index and constructing the multi-objective composite reward function based on the multi-dimensional reward quantification index includes: performing a joint action according to the joint control command and determining the joint result to output a success flag; when the success flag indicates joint success, allocating a basic success reward; based on the joint success, obtaining the strength ratio and diameter error rate at the joint to calculate a comprehensive joint quality score, and allocating a joint quality reward; calculating the total operation time of the joint action and obtaining an efficiency penalty reward by combining it with a time penalty coefficient; quantifying the smoothness index of the joint action and obtaining a smoothing penalty reward based on a smoothing penalty coefficient; setting an energy-saving reward based on the action energy consumption of the joint action, and using the energy-saving reward as other auxiliary rewards; and using the success flag, the basic success reward, the joint quality reward, the efficiency penalty reward, the smoothing penalty reward, and the other auxiliary rewards as the multi-dimensional reward quantification index to construct the multi-objective composite reward function. This invention guides the decision engine towards multi-objective optimization through a multi-objective composite reward function, improving joint success rate and quality stability, helping to reduce energy consumption and mechanical losses, and effectively improving the practicality of the control method.
[0014] Optionally, the introduction of an online training optimization mechanism, combined with the composite reward quantification index, to train and optimize the adaptive decision engine to obtain joint optimization control instructions includes: constructing a digital twin simulation environment; pre-training the adaptive decision engine in the digital twin simulation environment to obtain an adaptive decision training engine; deploying the adaptive decision training engine to a physical robot controller and acquiring experience data for each control cycle and storing it in an experience playback buffer; updating the parameters of the adaptive decision training engine online based on the experience data to obtain an adaptive decision update engine; and continuously monitoring and evolving the adaptive decision update engine to output the joint optimization control instructions. This invention reduces the training cost of physical scenarios through pre-training, and ensures that the engine adapts to dynamic changes in working conditions through online updates and continuous evolution, enhancing system robustness and enabling continuous optimization of joint control to adapt to complex and changing working environments.
[0015] Optionally, the step of updating the parameters of the adaptive decision training engine online based on the empirical data to obtain an adaptive decision update engine includes: sampling micro-batch data from the experience replay buffer according to the priority distribution of the empirical data; calculating the loss function value of the adaptive decision training engine based on the micro-batch data; updating the parameters of the adaptive decision training engine using gradient descent based on the loss function value to obtain network update parameters; performing a soft update on the target network using the network update parameters to synchronize the target network parameters; calculating the update error value corresponding to the empirical data using the target network parameters; updating the sampling priority of the empirical data in the experience replay buffer according to the update error value, and synchronously adjusting the priority distribution of the empirical data in the summation tree data structure to obtain the adaptive decision update engine. This invention improves the targeting and efficiency of gradient descent updates by prioritizing high-value empirical data through priority sampling, avoids network oscillations through a soft update strategy, improves engine stability, accelerates convergence speed, and enables rapid iterative optimization of the decision engine.
[0016] Optionally, the step of executing the joint optimization control command and acquiring full-state actual joint data to iterate the adaptive decision engine in a closed loop and complete self-learning joint control includes: parsing the joint optimization control command to obtain a target joint angle value; smoothly fitting the target joint angle value to obtain an instantaneous target value; combining the instantaneous target value and acquiring the actual joint angle value to calculate the torque command required to reach the target joint angle value; converting the torque command into servo drive signals for each joint servo driver; and during the execution of the servo drive signals, collecting the full-state actual joint data and transmitting it back to the experience playback buffer in real time to complete closed-loop feedback and parameter updates. This invention ensures effective and reliable iteration of the adaptive decision engine, continuously optimizes control accuracy and consistency, reduces the probability of joint failure, achieves autonomous evolution of joint control, and is beneficial for improving the uniformity of batch joints.
[0017] Secondly, this invention provides a self-learning connector control system based on multimodal feedback and reinforcement learning. The system executes the self-learning connector control method based on multimodal feedback and reinforcement learning provided by this invention. The system includes an input device, an output device, a processor, and a memory, which are interconnected. The memory stores a computer program, which includes program instructions, and the processor is configured to call the program instructions. This invention achieves automated operation of the self-learning connector control through high-performance hardware collaboration, improving system practicality and scalability, and facilitating engineering deployment. Attached Figure Description
[0018] Figure 1 This is a flowchart of the self-learning connector control method based on multimodal feedback and reinforcement learning according to an embodiment of the present invention; Figure 2 This is a framework diagram of a self-learning connector control system based on multimodal feedback and reinforcement learning, according to an embodiment of the present invention. Detailed Implementation
[0019] Specific embodiments of the present invention will now be described in detail. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the invention. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, well-known circuits, software, or methods have not been specifically described to avoid obscuring the invention.
[0020] Throughout this specification, references to "an embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with that embodiment or example is included in at least one embodiment of the invention. Therefore, the phrases "in an embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily refer to the same embodiment or example. Furthermore, specific features, structures, or characteristics can be combined in one or more embodiments or examples in any suitable combination and / or sub-combination. Moreover, those skilled in the art will understand that the illustrations provided herein are for illustrative purposes and are not necessarily drawn to scale.
[0021] Please see Figure 1 One embodiment of the present invention provides a self-learning connector control method based on multimodal feedback and reinforcement learning, the method comprising the following steps: S1. Construct a multimodal synchronous sensing array and acquire real-time multimodal sensing data during the joint process through the multimodal synchronous sensing array.
[0022] In this embodiment, the robot's end effector and body undergo a comprehensive upgrade of their sensing systems, integrating multiple types of high-precision, miniaturized sensors and achieving microsecond-level synchronization through a unified clock, providing high-quality raw data for intelligent decision-making.
[0023] All sensors are connected via a high-speed synchronous bus, and the main controller sends a unified clock pulse to synchronize data acquisition, ensuring that visual, force, and pose information are strictly aligned on the time axis, providing a spatiotemporally consistent fused data stream for digital twin simulation environments and decision-making.
[0024] The multimodal synchronous sensing array includes a visual sensing array, a force sensing array, and a pose sensing array to acquire real-time multimodal sensing data, as detailed below: 1) Acquiring visual image data through a visual perception array: Layout: A binocular stereo vision solution of "global monitoring + end-effector guidance" is adopted; a high-resolution global shutter industrial camera (keyence CV-X200 series or Hikvision MV-CH series recommended, with high frame rate, vibration resistance and gigabit Ethernet interface) is fixedly installed on a crossbeam about 1.2 meters above the robot's working area, looking vertically down to cover multiple spindles for rapid initial positioning of broken yarn; another high-precision miniature camera of the same level is mounted laterally on the end effector of the robotic arm through a rigid bracket, with its optical axis at an angle of about 25 degrees to the tool axis. This design is intended to avoid self-obstruction and to capture the millimeter or even sub-millimeter relative position between the yarn end, shape (bending angle, twisting angle) and fuzz state and the tip of the actuator in real time and at close range (10mm-50mm).
[0025] Key parameters and configuration: Camera resolution no less than 1280x960 pixels, frame rate no less than 60fps, and global shutter used to avoid motion blur. A ring-shaped light-emitting diode (LED) must be equipped as an active fill light to overcome reflections from the yarn surface and interference from ambient light changes, meeting the requirements for dynamic capture and real-time transmission. Visual image data is transmitted to the control module in real time via a gigabit Ethernet interface.
[0026] 2) Acquire force signal data through a force sensing array: Core force feedback: A six-dimensional force / torque sensor is embedded in the gripper base of the robot, with a measurement range covering ±10N (force) and ±0.5Nm (torque), and a sampling frequency of no less than 1kHz. This sensor is used to measure the three-dimensional contact force of the gripper on the yarn and the dynamic tension and torque change trend generated during yarn feeding and movement in real time.
[0027] Auxiliary Tension Sensing: To more directly monitor the yarn's own tension, a miniature tension sensor is integrated inside or near the inlet of the pneumatic splicer's guide ceramic nozzle (ME-Meßsysteme's KD43s series or Triton's Dillon Compact series are recommended, as they are compact, have high resolution, and strong overload protection). This sensor measures the micro-Newton bending stress generated when the yarn passes through using a miniature strain gauge and converts it into a linear tension signal. To ensure signal accuracy and eliminate interference from the robotic arm's inertial forces, the contact path between the sensor and the yarn is designed to be as short as possible, and it is decoupled from the robotic arm's vibrating components through mechanical isolation.
[0028] 3) Acquire pose state data through a pose sensing array: Joint feedback: To achieve fully closed-loop precision control, an absolute multi-turn encoder (Tamagawa TS5700 series or Heidenhain EQN 400 series are recommended, as they possess high vibration resistance and electromagnetic interference resistance) is installed at the rear end of the servo motors of each of the three rotary joints of the robotic arm to measure the absolute angle of the joints. Simultaneously, a high-resolution single-turn encoder is installed at the output end of the servo motor or after the reducer, forming a fully closed-loop position feedback system that directly detects the actual position of the connecting rods to compensate for transmission errors such as gear backlash and elastic deformation. This arrangement is called the full joint feedback scheme.
[0029] End-effector motion state: An additional miniature inertial measurement unit is installed at the end flange of the robotic arm to directly measure the linear acceleration and angular velocity of the end effector in three-dimensional space, providing more accurate raw data for state estimation, so as to more accurately calculate its actual motion state.
[0030] Microscopic proximity sensing: A laser triangulation sensor or capacitive proximity sensor is installed at the tip of the end effector (such as a gripper or yarn guide). Its measurement range covers 0mm-10mm, with a resolution better than 0.01mm and a repeatability of ±0.02mm. It is used to accurately sense the instantaneous absolute distance between the tip of the actuator and key components such as the yarn guide hook and the ring rail, guiding fine alignment operations.
[0031] In this embodiment, a multi-sensor synchronization mechanism is constructed: all the aforementioned sensors are connected to the main controller via a high-speed synchronization bus (such as EtherCAT). The main controller sends a unified distributed clock synchronization pulse, and all data acquisition is performed under the drive of this clock and is stamped with a unified microsecond-level precision timestamp, ensuring that visual image data, force signal data, and pose state data are strictly aligned on the time axis, providing a spatiotemporally consistent data stream for subsequent feature fusion.
[0032] S2. Perform multi-level feature extraction and deep fusion on the multimodal real-time sensing data to obtain a high-dimensional state vector.
[0033] In this embodiment, the synchronously acquired multimodal real-time sensing data is preprocessed and feature extracted, and a joint state vector that comprehensively encodes the current junction scene information is generated as a high-dimensional state vector through a fusion algorithm. A hybrid data fusion framework combining an Extended Kalman Filter (EKF) and a convolutional neural network is adopted.
[0034] Multi-level feature extraction includes visual feature extraction, force feature extraction, and pose feature extraction to obtain heterogeneous feature vectors, as detailed below: 1) Obtain visual feature vectors through visual feature extraction: 1.1) Image preprocessing and segmentation: For the image acquired by the end-guided camera, Gaussian filtering is first performed to remove noise. Then, an adaptive threshold segmentation algorithm is used to separate the yarn foreground from the complex background and extract the clear yarn outline.
[0035] 1.2) Subpixel-level positioning: The yarn contour is analyzed using a subpixel edge detection algorithm (such as the Zernike moment-based method) to calculate the subpixel precision coordinates of the yarn endpoints or center lines in the image coordinate system.
[0036] 1.3) Deep Feature Extraction: The preprocessed image region is input into a lightweight Convolutional Neural Network (CNN) for feature extraction. Architectures such as MobileNetV3 can be used, with the input image uniformly scaled to 224x224 pixels. A 3x3 convolutional kernel visual attention mechanism can be executed in the end tool region to focus on key areas. The network ends with a global pooling layer, outputting a fixed-length 128-dimensional deep visual feature vector, which contains abstract semantic information such as yarn texture and morphology.
[0037] 1.4) 3D Geometric Feature Extraction: A binocular stereo vision system consisting of a fixed monitoring camera and an end-effector guidance camera is used. For image pairs after epipolar correction, a disparity map is calculated using a stereo matching algorithm. Then, combined with camera calibration parameters, the 3D coordinate point cloud of the yarn target point (e.g., a broken end) in the robot's base coordinate system is calculated using triangulation principles. After filtering this point cloud using a random sampling consensus algorithm to remove background noise, it is spatially registered with the end-effector pose estimation output by the EKF. Derivative features can then be calculated, such as the curvature of the yarn in the current local space (3D vector) and the dominant frequency of the drift caused by airflow (2D scalar).
[0038] 2) Obtain the force feature vector through force feature extraction: 2.1) Signal filtering: The raw signals output by the six-dimensional force sensor and the miniature tension sensor are first subjected to low-pass filtering (the cutoff frequency is usually set to 50Hz) to remove high-frequency electrical noise and mechanical resonance interference.
[0039] 2.2) Time-domain feature calculation: Within a short sliding time window (e.g., 100ms), calculate the mean, peak, and standard deviation of the clamping force to quantify the stability and strength of the gripping. Simultaneously, extract the tension curve of the yarn feeding process from the tension sensor signal, calculate the mean tension and the real-time first derivative (tension change rate) to determine tension abrupt changes, and extract the dominant frequency components in its spectrum through fast Fourier transform to determine whether periodic jitter exists.
[0040] 3) Obtain the pose feature vector through pose feature extraction: 3.1) State Estimation: The angles and angular velocities fed back by the encoders of each joint, as well as the accelerations and angular velocities measured by the end effector (IMU), are used as observations and input into an EKF (Execution Frame). The state vector of the EKF is defined as 12-dimensional, containing the three-dimensional position, three-dimensional velocity, yarn tension value, and their first derivatives of the end effector. The process noise covariance matrix is set as a diagonal matrix diag(1e-5,1e-5,1e-5,1e-3,1e-3,1e-3,1e-2,1e-1) based on the measured joint servo error values, and the observation noise covariance matrix is set as diag(5e-4,5e-4,5e-4,0.1,0.1,0.1,0.05) based on the sensor accuracy calibration values. Through the prediction-update loop of the EKF, a smooth, optimally estimated spatiotemporal state vector is output every 2 milliseconds.
[0041] 3.2) Distance characteristics: The instantaneous distance between the proximity sensor and the target is read directly, and the rate of change of this distance over time is calculated as a feedback characteristic for fine approximation action.
[0042] In this embodiment, a multimodal feature fusion network is established to deeply fuse heterogeneous feature vectors to generate a high-dimensional state vector.
[0043] First, feature concatenation: All the extracted heterogeneous feature vectors are concatenated to form a high-dimensional original feature vector, satisfying the following relationship: in, The original feature vector, For visual feature vectors, As a derived feature, For force-feeling feature vectors, It is a spatiotemporal state vector. This is a feedback feature.
[0044] It should be noted that, (128 dimensions) (5-dimensional: 3-dimensional curvature + 2-dimensional main frequency of oscillation) (n-dimensional, depending on the number of force features extracted) (12-dimensional) (2-dimensional); meanwhile, assume that the total dimension is m-dimensional, where m is determined by the sum of the dimensions of each feature vector.
[0045] Second, fusion mapping: Design a multimodal feature fusion network, the core of which is a multilayer perceptron (MLP). The number of neurons in the input layer of this MLP is equal to m. The network structure can contain multiple fully connected layers, for example: the first layer has 512 neurons, the second layer has 256 neurons, and the third layer has 128 neurons. Each layer is followed by a normalization operation and a ReLU activation function. Finally, the output layer outputs a unified high-dimensional state vector through a linear transformation.
[0046] Specifically, the high-dimensional state vector is obtained by sequentially processing the original feature vector through multiple fully connected layers, performing linear transformations, layer normalization, and activation functions.
[0047] It should be noted that the high-dimensional state vector fully encodes the coupled state of the robot's motion, the yarn's dynamic shape, and environmental disturbances, and serves as the real-time driving input for the digital twin model and the decision-making basis for the reinforcement learning policy network.
[0048] S3. Construct an adaptive decision engine based on deep reinforcement learning, and input the high-dimensional state vector into the adaptive decision engine to output connector control commands.
[0049] In this embodiment, an adaptive decision engine is established based on deep reinforcement learning. As a deep reinforcement learning agent, its policy network acts as the "brain" of the system, responsible for mapping high-dimensional state vectors into optimal action instructions. The deep reinforcement learning policy network is based on a flexible executor-critic algorithm framework, specifically a hybrid architecture combining a task-adapted deep deterministic policy gradient algorithm with maximum entropy reinforcement learning, and employs an executor-critic framework (Actor-Critic).
[0050] Actor Network (Policy Network): The input is a fused high-dimensional state vector. The main network structure can adopt a deep fully connected network with residual connections, containing three hidden layers with 512, 256, and 128 neurons respectively. Each layer is followed by a normalization layer and a ReLU activation function. For continuous motion control, the output layer is designed with two parallel heads: one outputs the mean vector of the motion, and the other outputs the logarithmic standard deviation vector of the motion, used to define the Gaussian distribution of the motion. The output layer uses the Tanh activation function to constrain the final sampled motion within the normalization interval [-1, 1], corresponding to the normalized values of the incremental motion commands, speed scaling factors, and splicer pressure settings in the three axes of the robotic arm end effector.
[0051] Critic Network (Value Network): The input is also a high-dimensional state vector, and its hidden layer structure can be similar to or the same as that of the Actor Network. The output layer is a single neuron that outputs a scalar value representing an estimate of the expected cumulative reward obtained by performing a certain action under the high-dimensional state vector. Key hyperparameters include: a learning rate of 3e-4, using the Adam optimizer, a discount factor of 0.99, a soft update coefficient of 0.005, an experience replay buffer size of 1e6, and a batch size of 256.
[0052] In this embodiment, the action space is defined and interpreted as follows: the action output by the policy network is a normalized vector, representing an incremental adjustment instruction for the control parameters of the underlying actuator, and the incremental adjustment instruction is used as the joint control instruction.
[0053] Specifically, the mapping relationship needs to be predefined and satisfy the following conditions: in, The action instructions output by the policy network. For the amount of clamping force adjustment, The amount of adjustment for the yarn feeding speed. This is the fine-tuning amount for the trajectory position of the endpoint in the XY horizontal plane. Adjust the motor speed by a percentage for knotting.
[0054] It should be noted that the clamping force adjustment is mapped to the actual force value change (e.g., ±2N). The yarn feed speed adjustment is mapped to the actual speed change (e.g., ±10mm / s). The fine-tuning of the end's trajectory position in the XY horizontal plane is (e.g., ±0.5mm). The knotting motor speed adjustment percentage is (e.g., ±5%).
[0055] The above adjustments will be added to the original nominal trajectory or parameter settings generated by the upper-level task planner in each control cycle (e.g., 10ms), thereby achieving adaptive fine control based on real-time perception.
[0056] In an optional embodiment, to encourage exploration, adaptive Ornstein-Uhlenbeck process noise is added to the policy network output actions, with its variance decreasing linearly from 0.3 to 0.05 over the training phase.
[0057] S4. Establish a multi-objective composite reward function, and obtain the composite reward quantification index of the joint control command based on the multi-objective composite reward function.
[0058] In this embodiment, a multi-objective composite reward function is designed to drive policy learning. This function is used to quantitatively evaluate the merits of each action during the training phase, guiding the adaptive decision engine to learn an optimization policy that meets multiple requirements. This function combines sparse and dense rewards, employing a multi-objective optimization framework that combines a hierarchical reward structure with conditional policy gradients to handle the collaborative optimization of objectives such as joint success rate, joint strength, operation time, and energy consumption.
[0059] The multi-objective composite reward function satisfies the following relationship: in, As a quantitative indicator for composite rewards, This is a success flag. Basic success reward As a weight for quality rewards, For the overall quality score of the joint, This is the time penalty coefficient. The total operation time for the joint action. To smooth out the penalty coefficient, As an indicator of the smoothness of joint operation, For other auxiliary rewards.
[0060] In this embodiment, when the joint fails, the total reward value is directly taken as the preset failure penalty value; when the joint succeeds, the total reward value is determined by the basic success reward, the joint quality reward, the efficiency penalty reward, the smoothing penalty reward, and other auxiliary rewards.
[0061] It should be noted that the joint quality comprehensive score, total operation time, smoothness index and motion energy consumption are normalized before entering the multi-objective composite reward function to form dimensionless evaluation quantities.
[0062] 1) Success Flag: This is a binary variable, taking the value 1 (success) or 0 (failure). Success is determined by a combination of a photoelectric sensor and a micro-tension sensor installed on the yarn path. The specific logic is as follows: After the splicing action is completed, a 2-second observation period is initiated. If, during this period, the yarn continues to pass smoothly through the detection point (the photoelectric sensor signal is stable), and the yarn tension remains stable within the preset normal process range (e.g., 15-25 centinewtons), then the splicing is considered successful. Otherwise, it is considered a failure. Successful connection is set as a mandatory constraint and the highest priority objective.
[0063] 2) Basic Success Reward: If and only if At this point, the adaptive decision engine receives a fixed, high positive reward (e.g., +10.0). Its core objective is to quickly establish a strong association between "completing the rendezvous task" and "obtaining a high reward" in the early stages of training, guiding the adaptive decision engine to first learn to achieve the basic objective. In the reward function, this is reflected in a large positive reward upon success and a strong negative reward (e.g., -5) upon failure, terminating the current training round.
[0064] 3) The joint quality bonus is a bonus given based on the normalized estimated or measured joint strength.
[0065] 3.1) Joint Quality Overall Score: This is a value between 0.0 and 1.0 used to quantify the quality of the joint. It is calculated by weighting the strength ratio and diameter error rate, and satisfies the following relationship: in, For the overall quality score of the joint, As the weight of the intensity ratio, For strength ratio, As the weight of the diameter error rate, This represents the diameter error rate.
[0066] Strength ratio: Obtained through an integrated miniature online tensile testing device. After splicing, the system applies a brief standard test force (e.g., 60% of the original yarn breaking strength) to the yarn splice to detect whether the splice breaks or slips. The strength ratio = (test force at the splice / original yarn breaking strength) * 100%, and is normalized to 0-1.
[0067] Diameter error rate: The diameter at the joint is measured using a laser diameter gauge and compared with the standard diameter. The percentage error is calculated and normalized to 0-1. A smaller value indicates a smoother joint appearance and more uniform thickness.
[0068] It should be noted that the weights of the strength ratio and the diameter error rate are usually set to add up to 1, for example, 0.7 and 0.3, to emphasize the importance of the joint's mechanical properties.
[0069] 3.2) Quality Reward Weights: An adjustable hyperparameter (e.g., 2.0) used to balance the relative importance of basic success rewards and quality rewards. Increasing this value allows the adaptive decision engine to focus more on improving joint quality while ensuring success. The weights are not fixed but are dynamically adjusted during training based on the multi-objective preference vector or through a Pareto front search algorithm, allowing for the exploration of Pareto optimal solutions for tradeoffs between different objectives in later stages of training.
[0070] 4) Efficiency penalty and reward encourage quick completion of contact actions.
[0071] 4.1) Total operation time: The total time (in seconds) from the start to the end of this splicing operation. Timing starts from when the robotic arm receives the "start splicing" command and initiates the first yarn-finding action, and ends when the splicer completes its action and fully resets to the standby position.
[0072] 4.2) Time penalty coefficient: A positive coefficient (e.g., 0.05) converts operation time into a negative reward (penalty), incentivizing the adaptive decision engine to improve job efficiency.
[0073] 5) Smoothness penalty reward is used to constrain the smoothness of the joint action.
[0074] The smoothness index can be constructed based on the change in acceleration within adjacent control cycles. The smaller the smoothness index, the smoother the motion.
[0075] 6) Other auxiliary rewards: Energy-saving rewards are preferred. The energy consumption of the motion is estimated based on the servo motor current and running time, and the auxiliary reward value is determined based on the energy consumption of the motion, satisfying the following relationship: in, For other auxiliary rewards, This refers to the energy consumption of the joint operation.
[0076] It should be noted that the motion energy consumption is estimated based on the servo motor current and running time, and is a very small negative value. This is intended to subtly cultivate energy-saving habits during long-term training. The energy consumption estimate is derived from the motor current and motion amplitude.
[0077] In this embodiment, the multi-objective composite reward function explicitly guides the learning strategy of the adaptive decision engine: the primary objective is to ensure joint success (to obtain a basic success reward); based on success, efforts are made to improve the strength and appearance quality of the joint (maximizing the overall joint quality score); at the same time, the operation speed (minimizing the total operation time) and the smoothness of the action (minimizing the smoothness index) are also taken into account, thereby learning to make the best trade-offs among multiple competing objectives.
[0078] In terms of policy network architecture, a conditional policy network under a multi-task learning framework can be adopted. In addition to the high-dimensional state vector, its input also includes an additional target preference vector. This vector explicitly specifies the relative importance of each optimization objective in the current scenario (e.g., [intensity weight, time weight, energy consumption weight]). The network learns through the conditional layer to output the corresponding optimal action policy according to different preferences, so that in the deployment phase, the preference vector can be adjusted to flexibly adapt to different production requirements.
[0079] At the optimization algorithm level, a near-end policy optimization algorithm is adopted and a constraint optimization technique based on Lagrange multipliers is integrated. The objectives other than success rate are formalized into inequality constraints. The penalties for violating these constraints are adaptively adjusted by Lagrange multipliers, so that reward maximization and constraint satisfaction are considered simultaneously in policy updates.
[0080] S5. Introduce an online training and optimization mechanism, and combine the composite reward quantification index to train and optimize the adaptive decision engine to obtain the joint optimization control command.
[0081] In this embodiment, a two-stage training and online continuous optimization mechanism is constructed as the online training optimization mechanism, including the following: Phase 1: Pre-training in a high-fidelity digital twin simulation environment.
[0082] Simulation Environment Construction: A high-fidelity digital twin simulation environment was built based on physics engines such as PyBullet and MuJoCo. This environment accurately simulates the rigid body dynamics of the robot, the flexible body dynamics of the yarn, the workshop airflow disturbance model, and various random faults (such as spindle position deviation, random initial yarn position, random drift amplitude, environmental wind speed, and spindle position offset). The simulation environment interacts with the reinforcement learning adaptive decision engine through a standard interface, providing low-cost, high-throughput multi-objective evaluation capabilities.
[0083] Large-scale offline training: In a digital twin simulation environment, the adaptive decision engine begins its exploration based on an initial random policy or a simple heuristic policy. The training process is divided into a pre-training phase, which involves millions of rounds of experience collection in the simulation environment, with each round containing thousands of random head-to-head tasks. The adaptive decision engine collects a large number of experience data tuples by interacting with the environment. The Proximal Policy Optimization (PPO) algorithm is used to optimize and update the policy network parameters. The PPO algorithm ensures training stability by limiting the difference between the old and new policies (using a pruning replacement objective function). Its core pruning replacement objective function satisfies the following relationship: in, To replace the objective function with a pruning tool, For the expectation operator, To find the minimum value function, This represents the probability ratio between the old and new strategies. The advantage function is calculated using generalized advantage estimation. For the clipping function, This is the clipping hyperparameter (usually 0.1 or 0.2). The probability output of the new strategy. For at any time Strategic actions, It is a high-dimensional state vector. This is the probability output of the old strategy.
[0084] The above empirical data tuples satisfy the following relationship: in, It is a high-dimensional state vector. The action instructions output by the policy network. As a quantitative indicator for composite rewards, For at any time The high-dimensional state vector, This is the round termination marker.
[0085] Meanwhile, the Critic network is updated by minimizing the mean squared error of the value function. This stage involves millions of steps (e.g., 100,000 to 1 million rounds) of iterative training until the policy can stably achieve high performance metrics in the simulation environment. After pre-training, the policy network possesses basic capabilities.
[0086] Phase Two: Online Security Fine-tuning and Continuous Lifelong Learning.
[0087] Policy Deployment and Secure Startup: Simulating Pre-trained Converged Policy Networks Deployed into the physical robot controller. During the initial operation phase, a "hybrid strategy" is employed: the "simulation-optimal" action output by the current strategy network is executed with a high probability (e.g., 80%), while adaptive exploratory noise (e.g., Ornstein-Uhlenbeck process noise) is added to the output action with a low probability (e.g., 20%). The noise variance is relatively large in the initial stage (e.g., 0.3), linearly decreasing to a lower value (e.g., 0.05) as online learning progresses. This design ensures safe exploration while utilizing simulation knowledge in the initial stage, adapting to the differences between the real world and simulation. Fine-tuning in the actual environment employs a combination of online imitation learning and safe exploration.
[0088] Real-time experience collection and storage: During robot operation in the actual workshop, the complete experience data generated in each control cycle is stored in real time into a priority-based experience playback buffer. This buffer is implemented using a priority-based circular queue data structure, and its basic storage unit is a complete experience trajectory tuple, satisfying the following relationship: in, It is a high-dimensional state vector. The action instructions output by the policy network. As a quantitative indicator for composite rewards, For at any time The high-dimensional state vector, This is the round end marker. This is the initial estimate of the time difference error for this experience.
[0089] Specifically, the experience replay buffer physically allocates a fixed-size contiguous storage space in the GPU's video memory, with a total capacity set to N=500,000 experiences. When a new experience is stored, the oldest data is updated in a circular overwrite manner. Each experience is based on its... The absolute value is assigned sampling priority and stored in a dedicated summation tree data structure for implementation. Prioritize sampling and updating complexity.
[0090] In this embodiment, periodic online parameter updates are implemented: the system triggers a background online learning thread at a fixed interval (e.g., every K=200 environmental interactions). This includes the following steps: First, sampling: Based on the priority distribution of experience, sample a micro-batch of data of size B=256 from the experience playback buffer.
[0091] Second, calculate the loss: using the sampled micro-batch data, calculate the loss functions for the policy network (Actor) and the value network (Critic). For the case of using a deep deterministic policy gradient framework, the loss function of the policy network satisfies the following relationship: in, Let the loss function be the policy network. It is a mean function. Let the action value function be... The current state sampled from the experience replay buffer, This refers to the current action instruction of the current policy network.
[0092] The loss function of the value network satisfies the following relationship: in, The loss function of the value network, It is a mean function. For instant reward value, As a discount factor, The target expected return value, For the next state, For the next action instruction, Let the action value function be... The current state sampled from the experience replay buffer, This refers to the current action instruction of the current policy network.
[0093] Third, parameter update: Perform a gradient descent update on the parameters of both networks. The learning rate during the online learning phase is typically set smaller than that during the pre-training phase (e.g., ...). , This allows for smooth fine-tuning. After computation, a gradient descent update is performed on the parameters of the policy network and value network.
[0094] Fourth, soft update of the target network: To stabilize training, a separate target value network is used to calculate the target value. After each update, the current value network parameters are slowly synchronized to the target network using a soft update method, satisfying the following relationship: in, For the target network parameters, This is the soft update coefficient. These are the current network parameters.
[0095] It should be noted that the soft update coefficient is a very small coefficient (e.g., 0.005). After each update, the target network parameters are synchronized using the soft update method.
[0096] Fifth, priority update: The initial estimate of the time difference error in the sampling experience is corrected to a new update error value after this network parameter update, and the sampling priority of the corresponding experience data is updated according to the update error value; the sampling priority is synchronously written into the summation tree data structure for priority distribution adjustment in subsequent batch sampling.
[0097] The physical robot's real-time control thread executes forward inference and servo control of the currently deployed policy network; network parameter updates are performed periodically in a separate background computing thread; the entire update process is carried out in a background computing thread independent of the environment interaction thread, ensuring that data collection and policy optimization are executed in parallel.
[0098] In this embodiment, convergence monitoring and continuous evolution are implemented: the online learning process continues. The system periodically (e.g., every 1000 joints) evaluates the average success rate and average reward of the current policy on an independent validation set. When these performance metrics no longer show significant improvement (changes less than a threshold) over N consecutive evaluation periods (e.g., N=20), the system proceeds. ,like When the system reaches its optimal or near-optimal state under the current operating conditions, or has reached the preset maximum number of training steps, it can be considered that the strategy has converged to the optimal state or reached the maximum. However, the system does not stop learning. As a "lifelong learning" system, it continuously collects new data and performs regular fine-tuning, thereby adaptively responding to wear and tear after long-term equipment operation, seasonal changes in the workshop environment, or the introduction of new yarn varieties, achieving continuous performance evolution and maintenance.
[0099] S6. Execute the joint optimization control command and obtain the actual joint data in all states, and iterate the adaptive decision engine in a closed loop to complete the self-learning joint control.
[0100] In this embodiment, the motion control card converts the determined motion commands into servo drive signals, precisely driving the three-axis robotic arm, pneumatic splicer, and adaptive yarn guide to collaboratively complete a series of actions such as yarn feeding, yarn threading, and splicing. This is a precision chain that includes command parsing, inverse kinematics, real-time trajectory interpolation, and closed-loop control.
[0101] First, after the adaptive decision engine outputs a set of normalized joint optimization control instructions, the instruction packet is transmitted to the memory buffer of the motion control card. The instruction packet usually contains the discrete target point sequence of the robotic arm end in the task space, the expected velocity and acceleration curves of each point, and the switching timing of auxiliary input / output devices (IOs) such as splicers.
[0102] Secondly, the firmware core of the motion control card then starts real-time analysis. For the robotic arm's motion, the control card calls its built-in inverse kinematics solver library and quickly converts the end-effector pose in the task space into target joint angle values based on the pre-calibrated robot DH parameters.
[0103] Next, the trajectory interpolator performs smooth fitting on these discrete target joint angle values within a microsecond-level control cycle (usually using spline interpolation algorithm), and calculates the instantaneous target values required for each joint in each cycle, including instantaneous target position, target velocity and target acceleration.
[0104] Subsequently, the instantaneous target value is sent to the independent servo control loop corresponding to each joint. The control loop uses a proportional-integral-derivative controller or a more advanced model predictive control algorithm to calculate the torque command required to track the target joint angle value based on the actual joint angle fed back in real time by the photoelectric encoder.
[0105] Then, the torque command is converted from digital to analog signal to analog voltage signal or pulse sequence, and sent to the servo driver of each joint to drive the servo motor to output precise torque and speed.
[0106] Meanwhile, for I / O control commands such as those for the splicer and yarn guide, the multi-functional I / O interface on the motion control card will output switching signals or pulse width modulation signals at the synchronization moment when a specific trajectory point is reached, based on the precise timestamp in the command packet, directly controlling the solenoid valve or micro motor. The entire conversion and execution process strictly follows hardware-based precise timing synchronization, ensuring that the coordination error between the robotic arm movement and the auxiliary I / O actions is within milliseconds, thereby achieving precise, smooth, and synchronized execution.
[0107] Finally, the actual joint data of the full state during the execution process is transmitted back to the experience playback buffer in real time, forming the data basis for control closed-loop feedback and background training updates.
[0108] Please see Figure 2 In an optional embodiment, the present invention provides a self-learning connector control system based on multimodal feedback and reinforcement learning. The system includes an input device, an output device, a processor, and a memory, all interconnected. The memory stores a computer program comprising program instructions, and the processor is configured to invoke the program instructions to execute specific steps as described in the embodiments of the self-learning connector control method based on multimodal feedback and reinforcement learning provided by the present invention. The self-learning connector control system based on multimodal feedback and reinforcement learning provided by the present invention has a complete and stable structure, enhancing the overall applicability and practical application capability of the present invention.
[0109] In summary, the self-learning connector control method and system based on multimodal feedback and reinforcement learning provided by this invention acquires real-time multimodal sensing data by constructing a multimodal synchronous sensing array. After feature extraction and deep fusion, a high-dimensional state vector is formed. An adaptive decision engine is built based on an actuator-evaluator framework. Control commands are generated according to the high-dimensional state vector. A multi-objective composite reward function is designed to accurately guide learning. A hybrid training mechanism combining digital twin simulation pre-training and online optimization of the physical robot is adopted to achieve continuous iterative optimization of the adaptive decision engine. By parsing and executing optimization commands and synchronously collecting real-world data for closed-loop feedback, the system can autonomously and efficiently complete the connector task. This invention is easy to understand, computationally simple, requires less work, and is convenient for engineering applications, providing a theoretical foundation and technical support for the further development of the field of automation control technology.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A self-learning connector control method based on multimodal feedback and reinforcement learning, characterized in that, Includes the following steps: A multimodal synchronous sensing array is constructed, and multimodal real-time sensing data during the joint process is acquired through the multimodal synchronous sensing array; Multi-level feature extraction and deep fusion are performed on the multimodal real-time sensing data to obtain a high-dimensional state vector; An adaptive decision engine is constructed based on deep reinforcement learning. The high-dimensional state vector is input into the adaptive decision engine to output the connector control command. A multi-objective composite reward function is established, and a composite reward quantification index for the joint control command is obtained based on the multi-objective composite reward function; An online training and optimization mechanism is introduced, and the adaptive decision engine is trained and optimized in combination with the composite reward quantification index to obtain the joint optimization control command. The joint optimization control command is executed and the actual joint data in all states is obtained to iterate the adaptive decision engine in a closed loop and complete the self-learning joint control.
2. The self-learning connector control method based on multimodal feedback and reinforcement learning according to claim 1, characterized in that, The construction of the multimodal synchronous sensing array, and the acquisition of real-time multimodal sensing data during the joint process through the multimodal synchronous sensing array, includes: The multimodal synchronous sensing array is constructed by using a visual sensing array, a force sensing array, and a pose sensing array. During the connection process, visual image data, force signal data, and pose state data are acquired through the multimodal synchronous sensing array as the multimodal real-time sensing data.
3. The self-learning connector control method based on multimodal feedback and reinforcement learning according to claim 1, characterized in that, The step of performing multi-level feature extraction and deep fusion on the multimodal real-time sensing data to obtain a high-dimensional state vector includes: Visual feature extraction, force feature extraction, and pose feature extraction are performed on the multimodal real-time sensing data to obtain heterogeneous feature vectors, which include visual state vectors, force state vectors, and pose state vectors. A multimodal feature fusion network is established, and the heterogeneous feature vectors are concatenated in a preset order to form the original feature vector; The original feature vector is input into a multilayer perceptron for fusion mapping to generate the high-dimensional state vector.
4. The self-learning connector control method based on multimodal feedback and reinforcement learning according to claim 1, characterized in that, The adaptive decision engine built based on deep reinforcement learning, which inputs the high-dimensional state vector into the adaptive decision engine to output joint control commands, includes: The adaptive decision engine is constructed based on the executor-evaluator algorithm framework in deep reinforcement learning. The high-dimensional state vector is input into the adaptive decision engine to output incremental adjustment instructions for the control parameters of the underlying actuator as the joint control instructions.
5. The self-learning connector control method based on multimodal feedback and reinforcement learning according to claim 1, characterized in that, The step of establishing a multi-objective composite reward function and obtaining a composite reward quantification index for the joint control command based on the multi-objective composite reward function includes: Obtain multi-dimensional reward quantification indicators, and construct the multi-objective composite reward function based on the multi-dimensional reward quantification indicators; The total reward value of the joint control command is calculated based on the multi-objective composite reward function, and the total reward value is used as the composite reward quantification index.
6. The self-learning connector control method based on multimodal feedback and reinforcement learning according to claim 5, characterized in that, The process of obtaining the multi-dimensional reward quantification index and constructing the multi-objective composite reward function based on the multi-dimensional reward quantification index includes: Perform the joint action according to the joint control command, and determine the joint result to output a success flag. When the success flag indicates that the connector was successfully connected, a basic success reward is allocated. Based on the successful jointing, the strength ratio and diameter error rate at the joint are obtained to calculate the comprehensive joint quality score, and the joint quality reward is allocated accordingly. The total operation time of the aforementioned joint action is calculated, and an efficiency penalty reward is obtained by combining the time penalty coefficient. The smoothness index of the joint action is quantified, and a smoothness penalty reward is obtained based on the smoothness penalty coefficient; An energy-saving reward is set based on the energy consumption of the joint action, and the energy-saving reward is used as other auxiliary rewards. The success flag, the basic success reward, the connector quality reward, the efficiency penalty reward, the smoothing penalty reward, and the other auxiliary rewards are used as the multi-dimensional reward quantification indicators to construct the multi-objective composite reward function.
7. The self-learning connector control method based on multimodal feedback and reinforcement learning according to claim 1, characterized in that, The introduction of an online training and optimization mechanism, combined with the composite reward quantification index, trains and optimizes the adaptive decision engine to obtain joint optimization control instructions, including: A digital twin simulation environment is constructed, and the adaptive decision engine is pre-trained in the digital twin simulation environment to obtain an adaptive decision training engine; The adaptive decision training engine is deployed to the physical robot controller, and the empirical data for each control cycle is acquired and stored in the experience playback buffer. An adaptive decision update engine is obtained by updating the parameters of the adaptive decision training engine online based on the empirical data. The adaptive decision update engine is continuously monitored and evolved to output the joint optimization control command.
8. The self-learning connector control method based on multimodal feedback and reinforcement learning according to claim 7, characterized in that, The process of updating the parameters of the adaptive decision training engine online based on the empirical data to obtain an adaptive decision update engine includes: Based on the priority distribution of the empirical data, micro-batch data is sampled from the empirical replay buffer; The loss function value of the adaptive decision training engine is calculated based on the micro-batch data; The parameters of the adaptive decision training engine are updated by gradient descent based on the loss function value to obtain the network update parameters; Based on the aforementioned network update parameters, a soft update is performed on the target network to synchronize the target network parameters; The update error value corresponding to the empirical data is calculated using the target network parameters; The sampling priority of the empirical data in the empirical replay buffer is updated according to the update error value, and the priority distribution of the empirical data in the summation tree data structure is adjusted synchronously to obtain the adaptive decision update engine.
9. The self-learning connector control method based on multimodal feedback and reinforcement learning according to claim 7, characterized in that, The process of executing the joint optimization control command and acquiring full-state actual joint data to iterate the adaptive decision engine in a closed loop and complete self-learning joint control includes: The target joint angle value is obtained by parsing the joint optimization control command; The instantaneous target value is obtained by smoothly fitting the target joint angle value; Combine the instantaneous target value with the actual joint angle value to calculate the torque command required to achieve the target joint angle value; The torque command is converted into servo drive signals for each joint servo driver; During the execution of the servo drive signal, the actual connector data in the full state is collected and transmitted back to the experience playback buffer in real time to complete closed-loop feedback and parameter update.
10. A self-learning connector control system based on multimodal feedback and reinforcement learning, characterized in that, The system includes an input device, an output device, a processor, and a memory, which are interconnected. The memory stores a computer program, which includes program instructions. The processor is configured to invoke the program instructions to execute the self-learning joint control method based on multimodal feedback and reinforcement learning as described in any one of claims 1-9.
Citation Information
Cited By
Industrial robot adaptive trajectory planning and control system based on reinforcement learning
CN122008262A