A steel ring plate control method and system based on reinforcement learning and error compensation
Patent Information
- Application Number
- CN202610885114.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-06-18
AI Technical Summary
(1)机器人行走底座存在双向定位误差,无针对性补偿手段:纺织车间存在地面不平整、设备长期运行振动、飞花堆积等问题,机器人行走底座在移动至目标工位后,易产生X方向的靠近偏差和水平旋转角度偏差,而激光位移传感器直接固定于行走底座上,无专门的双向误差补偿结构,导致传感器无法精准对准钢领板检测区域(正上方),采集的位置数据存在系统性偏差,该偏差会直接传递至机器人控制器,引发随动控制的位置匹配错误
[0017]第二方面,为能够高效地执行本发明所提供的一种基于强化学习与误差补偿的钢领板控制方法,本发明还提供了一种基于强化学习与误差补偿的钢领板控制系统,包括:输入设备、输出设备、处理器、存储器,所述输入设备、输出设备、处理器、存储器相互连接,所述存储器存储有程序指令,所述程序指令用于基于强化学习与误差补偿的钢领板控制方法。本发明的一种基于强化学习与误差补偿的钢领板控制系统,结构紧凑、性能稳定,能够稳定地执行本发明提供的一种基于强化学习与误差补偿的钢领板控制方法,进一步提升本发明整体适用性和实际应用能力。
Smart Images

Figure CN122449957B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automation control technology, specifically to a steel collar control method and system based on reinforcement learning and error compensation. Background Technology
[0002] In the production operation of ring spinning machines, the automated splicing of yarn after breakage is the core link to improve spinning efficiency. The dual-arm splicing robot needs to follow the real-time movement of the ring rail to complete the yarn threading operation. The accurate position detection of the ring rail and the efficient follow-up control of the robot are the key to successful splicing.
[0003] The core mode of existing ring bar follow-up control technology is direct sensor detection → direct signal transmission → robot execution. The overall system mainly consists of displacement detection sensors, robot controller, jointing robot body, and robot walking base. Its basic workflow is as follows: After the robot walking base moves the robot to the target station of the spinning machine, the laser sensor collects the spatial position (height and horizontal position) data of the ring bar in real time. The sensor directly transmits the collected position signal to the robot controller. The controller calculates the robot follow-up control quantity according to the preset control algorithm, and finally drives the joint movement of the dual-arm robot to achieve follow-up tracking of the ring bar, providing a position matching basis for yarn threading.
[0004] In textile field applications, the ring rail does not perform regular periodic movements. Instead, it undergoes non-periodic up-and-down reciprocating movements in accordance with the operating status and process parameters of the spinning machine. Furthermore, the wire traveler moves in a circular motion along with the ring rail. Therefore, the requirements for the real-time performance and position matching accuracy of the robot's follow-up are extremely high. The position detection accuracy of the laser displacement sensor directly determines the follow-up control effect.
[0005] Existing ring rail servo control technology faces numerous unresolved problems in complex textile applications due to inherent defects in hardware positioning, signal transmission, and control logic. These problems and their causes are as follows: (1) The robot walking base has bidirectional positioning error and no specific compensation method: There are problems such as uneven ground, long-term vibration of equipment, and accumulation of flying flowers in the textile workshop. After the robot walking base moves to the target work position, it is easy to generate X-direction approach deviation and horizontal rotation angle deviation. The laser displacement sensor is directly fixed on the walking base without a special bidirectional error compensation structure, which makes the sensor unable to accurately align with the steel collar detection area (directly above). The collected position data has systematic deviation, which will be directly transmitted to the robot controller, causing position matching error of the follow-up control.
[0006] (2) The sensor signal is directly transmitted to the controller without considering the impact of on-site interference on the detection data, which amplifies the control error: Factors such as fly hair blocking, yarn swing, and equipment electromagnetic interference in the spinning field can cause instantaneous changes and noise interference in the position data collected by the laser sensor. In the existing technology, the sensor directly sends the original detection signal to the robot controller without effective preprocessing and error correction mechanism. The distorted position signal directly participates in the calculation of control quantity, which can easily cause the robot to shake and lag, or even cause the yarn to fail to pass through the loop.
[0007] (3) Insufficient accuracy in sensing the motion state of the steel collar plate and lack of intelligent optimization mechanism: Existing technology only performs simple filtering on the raw data collected by laser, without fusion optimization of hardware compensation results and laser detection data. It cannot accurately capture the real motion state (height, speed, motion trend) of the steel collar plate, nor can it predict the motion trend of the steel collar plate in advance, resulting in inherent data deviation in the subsequent robot follow-up control, making it difficult to achieve high-precision tracking. Summary of the Invention
[0008] To address the shortcomings of existing methods and the needs of practical applications, this invention provides a steel ring control method based on reinforcement learning and error compensation, comprising the following steps: The system acquires the bidirectional positioning error of the moving carrier associated with the target moving part, the bidirectional positioning error including directional proximity deviation and horizontal rotation angle deviation; based on the bidirectional positioning error, it controls the compensation mechanism to perform directional compensation motion and rotation angle compensation motion on the displacement sensor mounted on the moving carrier, so that the displacement sensor is accurately aligned with the detection area of the target moving part, and generates a mechanism compensation result; it acquires the original detection data collected by the displacement sensor after accurate alignment, preprocesses the original detection data to obtain preprocessed laser data, and fuses the mechanism compensation result with the preprocessed laser data to obtain a first fusion vector; based on the first fusion vector, it performs optimization processing through a reinforcement learning model, outputting a first action vector representing the optimized motion state of the target moving part, the reward function of the reinforcement learning model is constructed based on the matching degree between the actual motion state of the target moving part and the first action vector; based on the first action vector, it generates follow-up control commands for the actuator.
[0009] Optionally, obtaining the bidirectional positioning error of the moving carrier associated with the target moving part includes: The actual directional position and actual horizontal rotation angle of the mobile carrier are obtained through the position detection module. The actual directional position is compared with the preset standard workstation directional position to obtain the directional deviation; the actual horizontal rotation angle is compared with the preset standard horizontal rotation angle to obtain the horizontal rotation angle deviation.
[0010] Optionally, the steel collar control method based on reinforcement learning and error compensation further includes, after generating the mechanism compensation result, verifying the alignment accuracy between the displacement sensor and the detection area through the position detection module or the alignment verification sensor. If the alignment accuracy does not reach a preset threshold, a fine-tuning command is generated to control the compensation mechanism to perform fine-tuning until the alignment accuracy reaches the preset threshold.
[0011] Optionally, based on the bidirectional positioning error, the control compensation mechanism includes: The directional approach deviation and the horizontal rotation angle deviation are converted into corresponding compensation control commands; The compensation control command is sent to the drive module of the compensation mechanism, and the drive module executes the corresponding compensation action.
[0012] Optionally, the compensation mechanism includes a shaft precision compensation slide and a rotary fine-tuning motor; The compensation action includes: controlling the axis precision compensation slide to drive the displacement sensor to translate along the axis direction by a first distance, wherein the first distance is equal to the absolute value of the deviation in the direction and opposite in direction; In addition, the rotary fine-tuning motor is controlled to drive the displacement sensor to rotate around the vertical axis by a first angle, the absolute value of which is equal to and opposite to the deviation of the horizontal rotation angle.
[0013] Optionally, the compensation result of the mechanism includes the final compensation displacement value in the direction, the final compensation value in the rotation angle, and a compensation completion indicator.
[0014] Optionally, fusing the mechanism compensation result with the preprocessed laser data to obtain a first fusion vector includes: The final compensated displacement value of the direction, the final compensated value of the rotation angle, and the compensation completion mark in the compensation result of the mechanism are combined with the real-time detection height, the detection height change rate, and the detection height acceleration in the preprocessed laser data to form the first fusion vector.
[0015] Optionally, the output, representing the first motion vector characterizing the optimized motion state of the target moving part, includes: The reinforcement learning model outputs a first motion vector based on the first fusion vector, which includes the optimized real-time height, the optimized rise and fall speed, and the predicted value of the motion trend in the short term.
[0016] Optionally, the target moving part is a ring rail in a spinning machine, the moving carrier is a robot walking base, the displacement sensor is a laser displacement sensor, the actuator is a dual-arm joint robot, and the method is applied to the laser follow-up control of the ring rail.
[0017] Secondly, to efficiently execute the reinforcement learning and error compensation-based ring rail control method provided by this invention, this invention also provides a reinforcement learning and error compensation-based ring rail control system, comprising: an input device, an output device, a processor, and a memory, wherein the input device, output device, processor, and memory are interconnected, and the memory stores program instructions used for the reinforcement learning and error compensation-based ring rail control method. The reinforcement learning and error compensation-based ring rail control system of this invention has a compact structure and stable performance, and can stably execute the reinforcement learning and error compensation-based ring rail control method provided by this invention, further improving the overall applicability and practical application capability of this invention.
[0018] This invention first eliminates the displacement sensor detection area deviation caused by inaccurate positioning of the mobile carrier by acquiring the bidirectional positioning error and controlling the compensation mechanism for precise alignment. This directly eliminates the error in the displacement sensor detection area caused by inaccurate positioning of the mobile carrier at the hardware source, providing a hardware foundation free of systematic errors for subsequent data acquisition. Next, this invention deeply fuses the mechanism compensation result, which characterizes the hardware alignment state, with the preprocessed sensor detection data to construct a first fusion vector containing multi-dimensional information. Based on this, a reinforcement learning model that constructs a reward function based on matching degree is introduced to optimize this fusion vector. This allows the model to intelligently correct random errors and interference in the sensor data through learning, and accurately predict the short-term motion trend of the target moving parts.
[0019] Therefore, the final output first motion vector not only eliminates the deviation introduced by hardware positioning, but also significantly improves the perception accuracy and foresight of the real-time height, speed and motion trend of the target moving parts, thus providing a reliable data core for the high-precision and high-real-time follow-up control of the actuator. Attached Figure Description
[0020] Figure 1 A flowchart of a steel collar control method based on reinforcement learning and error compensation is provided for an embodiment of the present invention; Figure 2 A schematic diagram of a laser-guided steel collar plate structure for a dual-arm joint robot provided in an embodiment of the present invention; Figure 3 This is a framework diagram of a steel collar control system based on reinforcement learning and error compensation, provided for an embodiment of the present invention. Detailed Implementation
[0021] Specific embodiments of the present invention will now be described in detail. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the invention. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, well-known circuits, software, or methods have not been specifically described to avoid obscuring the invention.
[0022] Throughout this specification, references to "an embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with that embodiment or example is included in at least one embodiment of the invention. Therefore, the phrases "in an embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily refer to the same embodiment or example. Furthermore, specific features, structures, or characteristics can be combined in one or more embodiments or examples in any suitable combination and / or sub-combination. Moreover, those skilled in the art will understand that the illustrations provided herein are for illustrative purposes and are not necessarily drawn to scale.
[0023] Please see Figure 1 This invention provides a steel collar control method based on reinforcement learning and error compensation, comprising the following steps: S1. Obtain the bidirectional positioning error of the moving carrier associated with the target moving part, wherein the bidirectional positioning error includes directional proximity deviation and horizontal rotation angle deviation.
[0024] First, the actual directional position and actual horizontal rotation angle of the mobile carrier are obtained through the position detection module. The actual directional position is compared with the preset standard workstation directional position to obtain the directional deviation. Then, the actual horizontal rotation angle is compared with the preset standard horizontal rotation angle to obtain the horizontal rotation angle deviation.
[0025] The position detection module may include an encoder, a laser rangefinder, or a vision sensor. The processor can calculate the bidirectional positioning error by reading its signals and performing a simple subtraction operation.
[0026] In other embodiments, the alignment accuracy between the displacement sensor and the detection area can also be verified by the position detection module or the alignment verification sensor. If the alignment accuracy does not reach a preset threshold, a fine-tuning command is generated to control the compensation mechanism to make fine adjustments until the alignment accuracy reaches the preset threshold.
[0027] The alignment verification sensor can be another set of high-precision vision or proximity sensors. The processor compares the deviation value it feeds back with preset thresholds (e.g., 0.1 mm and 0.1°). If it exceeds the threshold, the fine-tuning amount is recalculated and sent to the compensation mechanism.
[0028] S2. Based on the bidirectional positioning error, the control compensation mechanism performs directional compensation motion and rotation angle compensation motion on the displacement sensor installed on the mobile carrier, so that the displacement sensor is accurately aligned with the detection area of the target moving part, and generates the mechanism compensation result.
[0029] In this embodiment, the directional approach deviation and the horizontal rotation angle deviation are converted into corresponding compensation control commands, and then the compensation control commands are sent to the drive module of the compensation mechanism, which executes the corresponding compensation action.
[0030] Based on the error value, the processor generates digital or pulse commands containing the target displacement and angle through a preset mapping relationship or scaling factor, and sends them to the drive module through a communication interface (such as CAN bus or EtherCAT).
[0031] Specifically, the compensation mechanism includes a shaft precision compensation slide and a rotary fine-tuning motor.
[0032] The compensation action includes: controlling the axis precision compensation slide to move the displacement sensor along the axis direction by a first distance, the first distance being equal to the absolute value of the deviation in the direction and opposite in direction; and controlling the rotary fine-tuning motor to move the displacement sensor around the vertical axis by a first angle, the first angle being equal to the absolute value of the deviation in the horizontal rotation angle and opposite in direction.
[0033] The drive module (such as a servo driver) parses the received instructions and drives the linear motor and rotary motor to perform precise displacement and rotation movements respectively.
[0034] The compensation result of the mechanism includes the final compensation displacement value in the direction. Final compensation value of rotation angle and compensation completion mark .in, The value is set to 1 after the compensation mechanism completes its action and passes verification; otherwise, it is set to 0.
[0035] S3. Obtain the original detection data collected by the displacement sensor after precise alignment, preprocess the original detection data to obtain preprocessed laser data, and fuse the mechanism compensation result with the preprocessed laser data to obtain a first fusion vector.
[0036] In one embodiment, preprocessing is typically performed by a processor, including but not limited to digital filtering (such as Kalman filtering, moving average filtering) to remove noise, and differential calculation of discrete height samples to obtain rate of change and acceleration.
[0037] Furthermore, fusing the mechanism compensation result with the preprocessed laser data to obtain a first fusion vector includes combining the final directional compensation displacement value, the final rotation angle compensation value, and the compensation completion indicator in the mechanism compensation result with the real-time detection height, detection height change rate, and detection height acceleration in the preprocessed laser data to form the first fusion vector.
[0038] The first fusion vector The expression is: in, At the current sampling time, The final compensation displacement value in the stated direction. This is the final compensation value for the rotation angle. This is the indicator for the completion of the compensation. The real-time detection height, The rate of change of the detected height, For the detected height acceleration, each component is normalized and mapped to the same dimensionless interval. .
[0039] S4. Based on the first fusion vector, the optimization process is performed through a reinforcement learning model to output a first action vector representing the optimized motion state of the target moving part. The reward function of the reinforcement learning model is constructed based on the matching degree between the actual motion state of the target moving part and the first action vector.
[0040] In this embodiment, based on the first fusion vector The optimization process is performed using a reinforcement learning model, and the first action vector representing the optimized motion state of the target moving part is output. The reinforcement learning model, based on the first fusion vector, outputs a first action vector containing optimized real-time height, optimized ascent and descent speed, and predicted short-term motion trends. The first action vector... The expression is: in, for The optimized real-time height at the specified moment. for The optimized acceleration and deceleration speed at that moment. for The predicted future short-term motion trend values at time t is normalized and mapped to the same dimensionless interval for each component. .
[0041] The reinforcement learning model may include a state input module, a motion perception / prediction output module, a reward evaluation module, and a policy optimization module.
[0042] The reward function of the reinforcement learning model The matching degree between the actual motion state of the target moving part and the first action vector is constructed (the components of various reward functions are normalized and mapped to the same dimensionless interval). ).
[0043] Specifically, the reward function Accurate rewards based on detection correction Sports trend-aligned rewards and data stability rewards The construction, expressed as: in, , , The preset weighting coefficients, and .
[0044] The detection and correction accuracy reward The calculation expression is: in, For height deviation, for The actual height of the target moving part at that moment; For speed deviation, for The actual velocity of the target moving part at that moment; This is the preset optimal correction accuracy threshold; This is the preset maximum tolerance deviation threshold; The maximum reward value is adjusted to the preset detection accuracy.
[0045] The movement trend fits the reward The calculation expression is: in, For symbolic functions, for The actual acceleration of the target moving part at that moment. The maximum reward value is matched to the preset movement trend.
[0046] The data stability reward To encourage the stationarity of output data, its calculation can be based on and The judgment is made based on whether the rate of change exceeds a preset stability threshold. The model can adopt an online learning approach, with the processor based on the real-time collected actual state. , , The reward is calculated, and the policy network parameters are updated using algorithms such as Q-learning and DDPG to continuously optimize the output.
[0047] S5. Based on the first motion vector, generate follow-up control commands for the actuator.
[0048] The first action vector and As input, the servo motor control quantities of each joint of the actuator are calculated through preset inverse kinematics and PID control algorithms, and commands are issued to achieve high-precision follow-up tracking.
[0049] In a specific application scenario, such as Figure 2 As shown, the target moving part is the ring rail 1 in the spinning equipment, the moving carrier is the robot walking base 6, the displacement sensor is the laser displacement sensor 2, the actuator is the dual-arm joint robot 7, and the compensation mechanism includes the axis precision compensation slide 4 and the rotary fine-tuning motor 3. The method is applied to the laser follow-up control of the ring rail.
[0050] After the robot's walking base 6 moves the entire system to the workstation, the position detection module 5 detects the bidirectional positioning error, and the compensation mechanism drives the laser displacement sensor 2 to complete precise alignment.
[0051] The laser displacement sensor 2 collects the height data of the steel collar plate 1. After being fused and optimized by the processor, it generates instructions to control the dual-arm joint robot 7 to follow the movement of the steel collar plate 1.
[0052] It should be noted that the specific implementation methods described above, such as image processing, numerical simulation, and the construction and training of machine learning models, can all be accomplished by the processor by calling the corresponding computer program instructions stored in memory. Those skilled in the art can implement the above functions using algorithms and tools known in the prior art, according to actual needs.
[0053] Please see Figure 3In an embodiment, to efficiently execute the reinforcement learning and error compensation-based ring rail control method provided by this invention, this invention also provides a reinforcement learning and error compensation-based ring rail control system, including: an input device 10, an output device 20, a processor 30, and a memory 40. The input device 10, output device 20, processor 30, and memory 40 are interconnected. The memory 40 stores program instructions used to execute the steps of the reinforcement learning and error compensation-based ring rail control method. The reinforcement learning and error compensation-based ring rail control system of this invention has a compact structure and stable performance, and can stably execute the reinforcement learning and error compensation-based ring rail control method of this invention, further improving the overall applicability and practical application capability of this invention.
[0054] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the present invention.
Claims
1. A method for controlling a steel ring rail based on reinforcement learning and error compensation, characterized in that, Includes the following steps: The bidirectional positioning error of the moving carrier associated with the target moving part is obtained, and the bidirectional positioning error includes directional proximity deviation and horizontal rotation angle deviation. Based on the bidirectional positioning error, the control compensation mechanism performs directional compensation motion and rotation angle compensation motion on the displacement sensor installed on the mobile carrier, so that the displacement sensor is accurately aligned with the detection area of the target moving part, and generates mechanism compensation results. The original detection data collected by the displacement sensor after precise alignment is obtained, the original detection data is preprocessed to obtain preprocessed laser data, and the mechanism compensation result is fused with the preprocessed laser data to obtain a first fusion vector; Based on the first fusion vector, the optimization process is performed through a reinforcement learning model to output a first action vector representing the optimized motion state of the target moving part. The reward function of the reinforcement learning model is constructed based on the matching degree between the actual motion state of the target moving part and the first action vector. Based on the first motion vector, a follow-up control command for the actuator is generated.
2. The steel ring control method based on reinforcement learning and error compensation according to claim 1, characterized in that, The acquisition of the bidirectional positioning error of the moving carrier associated with the target moving part includes: The actual directional position and actual horizontal rotation angle of the mobile carrier are obtained through the position detection module. The actual directional position is compared with the preset standard workstation directional position to obtain the directional deviation; the actual horizontal rotation angle is compared with the preset standard horizontal rotation angle to obtain the horizontal rotation angle deviation.
3. The steel ring control method based on reinforcement learning and error compensation according to claim 2, characterized in that, It also includes verifying the alignment accuracy between the displacement sensor and the detection area through the position detection module or the alignment verification sensor after generating the mechanism compensation result. If the alignment accuracy does not reach the preset threshold, a fine-tuning command is generated to control the compensation mechanism to make fine adjustments until the alignment accuracy reaches the preset threshold.
4. The steel ring control method based on reinforcement learning and error compensation according to claim 1, characterized in that, Based on the bidirectional positioning error, the control compensation mechanism includes: The directional approach deviation and the horizontal rotation angle deviation are converted into corresponding compensation control commands; The compensation control command is sent to the drive module of the compensation mechanism, and the drive module executes the corresponding compensation action.
5. The steel ring control method based on reinforcement learning and error compensation according to claim 4, characterized in that, The compensation mechanism includes a shaft precision compensation slide and a rotary fine-tuning motor; The compensation action includes: controlling the axis precision compensation slide to drive the displacement sensor to translate along the axis direction by a first distance, wherein the first distance is equal to the absolute value of the deviation in the direction and opposite in direction; In addition, the rotary fine-tuning motor is controlled to drive the displacement sensor to rotate around the vertical axis by a first angle, the absolute value of which is equal to and opposite to the deviation of the horizontal rotation angle.
6. The steel ring control method based on reinforcement learning and error compensation according to claim 1, characterized in that, The compensation results of the mechanism include the final compensated displacement value in the direction, the final compensated value in the rotation angle, and a compensation completion indicator.
7. The steel ring control method based on reinforcement learning and error compensation according to claim 1, characterized in that, The step of fusing the mechanism compensation result with the preprocessed laser data to obtain a first fusion vector includes: The final compensated displacement value of the direction, the final compensated value of the rotation angle, and the compensation completion mark in the compensation result of the mechanism are combined with the real-time detection height, the detection height change rate, and the detection height acceleration in the preprocessed laser data to form the first fusion vector.
8. The steel ring control method based on reinforcement learning and error compensation according to claim 1, characterized in that, The output, representing the first motion vector of the optimized motion state of the target moving part, includes: The reinforcement learning model outputs a first motion vector based on the first fusion vector, which includes the optimized real-time height, the optimized rise and fall speed, and the predicted value of the motion trend in the short term.
9. The steel ring control method based on reinforcement learning and error compensation according to claim 1, characterized in that, The target moving part is the ring rail in the spinning equipment, the moving carrier is the robot walking base, the displacement sensor is a laser displacement sensor, the actuator is a dual-arm joint robot, and the method is applied to the laser follow-up control of the ring rail.
10. A steel ring control system based on reinforcement learning and error compensation, characterized in that, The collar control system based on reinforcement learning and error compensation includes: an input device, an output device, a processor, and a memory. The input device, output device, processor, and memory are interconnected. The memory stores program instructions, which are used to execute the collar control method based on reinforcement learning and error compensation as described in any one of claims 1-9.
Citation Information
Patent Citations
Multi-robot formation control method based on physical model feedforward and deep reinforcement learning
CN121806890A
Self-learning joint control method and system based on multi-modal feedback and reinforcement learning
CN121859945A