Robot high-precision compliant control method and system based on man-machine cooperation teaching track path

By constructing a structured trajectory library through human-machine collaborative teaching and combining it with reinforcement learning to optimize control parameters, the problem of insufficient accuracy of traditional robot compliant control methods when facing workpiece pose deviations and material elastic deformation is solved, achieving high-precision compliant control and reducing training costs and safety risks.

CN121821389APending Publication Date: 2026-04-10SHENZHEN MOYING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-12
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional robot compliant control methods struggle to achieve high-precision operation when faced with workpiece pose deviations and material elastic deformation. Furthermore, existing methods are highly dependent on operator experience, have high reinforcement learning costs and significant safety risks, and the teach trajectory library fails to effectively support intelligent control.

Method used

A structured trajectory library is constructed by teaching the trajectory through human-machine collaboration, which serves as prior knowledge for reinforcement learning. An initial range of control parameters is set, and a joint state space of contact force and pose is constructed during training. A reward function is designed for optimization training, generating compliant control parameters. The control parameters are collected in real time and dynamically adjusted to achieve high-precision control.

Benefits of technology

It improves the control accuracy of robots under complex working conditions, reduces the reliance on manual parameter tuning, reduces training costs and safety risks, and achieves the coordinated optimization of sub-millimeter pose and millinews force control to meet the needs of precision assembly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121821389A_ABST
    Figure CN121821389A_ABST
Patent Text Reader

Abstract

The invention provides a robot high-precision compliance control method and system based on a man-machine cooperation teaching track path. The method belongs to the cross technical field of robot intelligent control, man-machine cooperation and machine learning. The method comprises the steps of collecting teaching data through man-machine cooperation teaching operation, and performing structured processing to generate a structured teaching track library; according to task labels in the structured teaching track library, tracks are classified and stored, and track sub-libraries of different task types are formed; a structured trajectory library is constructed on the basis of a man-machine cooperation teaching trajectory path and serves as reinforcement learning priori knowledge, compliant control parameters can be generated by self-adaption to environment disturbance, the control precision of the robot under the complex working condition is greatly improved, submillimeter-level precision and millinewton-level force control collaborative optimization is achieved, and the control precision of the robot under the complex working condition is improved. And high-precision task requirements such as precision assembly are met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application provides a robot high-precision compliant control method and system based on human-robot collaborative teaching trajectory, and belongs to the technical field of robot intelligent control, human-robot collaboration and machine learning. BACKGROUND

[0002] In the process of industrial production towards intelligence and precision, robots are undertaking increasingly complex and delicate tasks such as precision assembly and flexible operation, which puts high requirements on the control precision and compliance of robots.

[0003] Traditional mechanical arm compliant control and skill reuse technology has many deficiencies. Traditional methods such as drag teaching or programming teaching can only record fixed trajectories, and when facing actual disturbances such as workpiece pose deviation and material elastic deformation, it is easy to cause problems such as insertion failure or surface scratching, which is difficult to meet the demand of high-precision operation. Impedance / admittance control can achieve compliant control, but it needs to manually set the stiffness and damping matrix, and is highly dependent on the experience of the operator, and it is difficult to find a perfect balance between precision and compliance. Reinforcement learning provides a new way to solve the compliant control problem, but training a compliant strategy from scratch requires a lot of real interaction, which is costly, time-consuming and has safety risks. In addition, the existing "teaching trajectory library" is mainly used for simple playback or interpolation, and is not deeply integrated with the intelligent control strategy generation process. The related concepts in the public data are relatively isolated, and high-quality human-machine teaching trajectory library is not used as an effective support for reinforcement learning, making it difficult to achieve intelligent compliant control driven by human experience. SUMMARY

[0004] The application provides a robot high-precision compliant control method and system based on human-robot collaborative teaching trajectory, which solves the problems mentioned in the background technology:

[0005] The application provides a robot high-precision compliant control method based on human-robot collaborative teaching trajectory, which comprises the following steps:

[0006] S1, through human-robot collaborative teaching operation, collecting teaching data and performing structured processing to generate a structured teaching trajectory library; according to the task label in the structured teaching trajectory library, the trajectory is classified and stored to form a trajectory sub-library of different task types;

[0007] S2, using the structured teaching trajectory library as prior knowledge, initializing the reinforcement learning strategy and setting the initial control parameter range; at the same time, using the trajectory data in the trajectory sub-library to construct constraint conditions to constrain the training process of the reinforcement learning strategy;

[0008] S3, during the reinforcement learning training process, a contact force and pose joint state space is constructed, and a reward function is designed based on the state space; the reinforcement learning strategy is optimized and trained according to the reward function, and the compliant control parameter is generated;

[0009] S4, the generated compliant control parameter is used to control the robot, and the actual pose information and contact force information of the robot are collected in real time during the execution of the robot task; the actual pose information and contact force information are compared and analyzed with the corresponding data in the structured teaching trajectory library, and pose deviation data and force deviation data are obtained;

[0010] S5, the compliant control parameter is dynamically adjusted according to the pose deviation data and force deviation data, the high-precision compliant control instruction of the robot is generated based on the adjusted compliant control parameter, and the robot is accurately controlled according to the high-precision compliant control instruction.

[0011] The system for implementing the robot high-precision compliant control method based on the human-machine cooperative teaching trajectory path as described above is provided, and the system comprises:

[0012] The sub-library forming module: through human-machine cooperative teaching operation, teaching data are collected and structured, and a structured teaching trajectory library is generated; the trajectories are classified and stored according to the task labels in the structured teaching trajectory library, and trajectory sub-libraries of different task types are formed;

[0013] The process constraint module: the structured teaching trajectory library is used as prior knowledge to initialize the reinforcement learning strategy, and an initial control parameter range is set; at the same time, the trajectory data in the trajectory sub-library are used to construct constraint conditions, and the training process of the reinforcement learning strategy is constrained;

[0014] The parameter generation module: during the reinforcement learning training process, a contact force and pose joint state space is constructed, and a reward function is designed based on the state space; the reinforcement learning strategy is optimized and trained according to the reward function, and the compliant control parameter is generated;

[0015] The data acquisition module: the generated compliant control parameter is used to control the robot, and the actual pose information and contact force information of the robot are collected in real time during the execution of the robot task; the actual pose information and contact force information are compared and analyzed with the corresponding data in the structured teaching trajectory library, and pose deviation data and force deviation data are obtained;

[0016] The accurate control module: the compliant control parameter is dynamically adjusted according to the pose deviation data and force deviation data, the high-precision compliant control instruction of the robot is generated based on the adjusted compliant control parameter, and the robot is accurately controlled according to the high-precision compliant control instruction.

[0017] The application has the advantages that: by constructing a structured trajectory library based on human-robot collaborative teaching trajectory, and taking it as prior knowledge of reinforcement learning, adaptive environment disturbance can be generated to generate compliant control parameters, which greatly improves the control accuracy of the robot under complex working conditions, realizes sub-millimeter level precision and millinewton level force control collaborative optimization, and meets the high-precision task demand of precision assembly; at the same time, the method reduces the dependence on manual parameter adjustment, and no longer needs the operator to manually set the stiffness and damping matrix according to experience, thereby reducing the problem of poor control effect caused by improper manual parameter adjustment. The adaptability of the robot to actual disturbances such as workpiece pose deviation and material elastic deformation is enhanced, so that the robot can more flexibly complete the task; the number of real interactions required for zero-start training of reinforcement learning is reduced, the training cost and period are reduced, and the safety risk brought by a large number of interactions is also avoided. The method can not only quickly initialize and constrain the reinforcement learning strategy by using human teaching experience, but also realize autonomous optimization of compliant control parameters through reinforcement learning, thereby avoiding the problems of lack of generalization of traditional teaching and low efficiency of pure reinforcement learning, and providing reliable protection for precise operation of the robot. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 The method steps of the application are shown in the following figure;

[0019] Figure 2 The system module diagram of the application is shown in the following figure. DETAILED DESCRIPTION

[0020] The preferred embodiments of the application are described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the application, and are not used to limit the application.

[0021] One embodiment of the application is shown in the following figure, a robot high-precision compliant control method based on human-robot collaborative teaching trajectory, the method comprises: Figure 1

[0022] S1, through human-robot collaborative teaching operation, collect teaching data containing pose information, force information and task label, and generate a structured teaching trajectory library by structuring the teaching data; according to the task label in the structured teaching trajectory library, the trajectory is classified and stored to form a trajectory sub-library of different task types;

[0023] S2, taking the structured teaching trajectory library as prior knowledge, initializing the reinforcement learning strategy and setting the initial control parameter range; at the same time, using the trajectory data in the trajectory sub-library to construct constraint conditions, and constraining the training process of the reinforcement learning strategy to ensure that the compliant control parameters generated in the training process meet the actual working condition requirements;

[0024] ​S3, during the reinforcement learning training process, a contact force and pose joint state space is constructed, and a reward function is designed based on the state space; the reinforcement learning strategy is optimized and trained according to the reward function, to generate a compliant control parameter that can adapt to environmental disturbance, the compliant control parameter including a stiffness parameter and a damping parameter;

[0025] S4, the generated compliant control parameter is used to control the robot, and during the execution of the task by the robot, actual pose information and contact force information of the robot are collected in real time; the actual pose information and contact force information are compared and analyzed with corresponding data in the structured teaching trajectory library, to obtain pose deviation data and force deviation data;

[0026] S5, the compliant control parameter is dynamically adjusted according to the pose deviation data and the force deviation data, a high-precision compliant control instruction of the robot is generated based on the adjusted compliant control parameter, the robot is accurately controlled according to the high-precision compliant control instruction, and a high-precision compliant operation task is completed.

[0027] The working principle and effects of the above technical solution are as follows: the structured trajectory library is constructed through human-machine cooperative teaching, which greatly improves the utilization rate of teaching data and reduces the time loss caused by invalid teaching operation; the reinforcement learning training is constrained based on the trajectory library, which significantly enhances the adaptability of the compliant control parameter to the actual working condition and avoids control failure caused by the parameter deviating from the working scene; the deviation data is collected in real time and the parameter is dynamically adjusted, which greatly improves the control precision of the robot, realizes the collaborative optimization of sub-millimeter level pose control and millinewton level force control, avoids the work defects or failures caused by insufficient precision, effectively reduces the difficulty and cost of manual debugging, reduces the parameter optimization period, stably copes with environmental disturbance, guarantees the execution quality of complex tasks, and significantly improves the reliability and practicality of the robot compliant operation.

[0028] In an embodiment of the present application, the S1 comprises:

[0029] S11, a teaching action is performed through a force feedback type human-machine cooperative operation device, three-dimensional pose coordinates of a robot end, six-dimensional contact force / torque signals and a task type label are collected in real time, and a multi-dimensional original teaching data set is generated;

[0030] S12, based on the original teaching data set, a time sequence alignment algorithm is used to synchronize and calibrate discrete pose and force signals, and an outlier detection algorithm is used to eliminate impact noise data, to generate clean teaching data;

[0031] S13, the clean teaching data is subjected to structured coding processing, the pose information, the force information and the task label are associated and mapped, a structured teaching trajectory library is generated, and the structured teaching trajectory library includes a trajectory time sequence index;

[0032] S14, based on the task label in the structured teaching trajectory library, a hierarchical clustering algorithm is used to classify the trajectory data, and a trajectory sub-library of different task types is constructed, the different task types include assembly, polishing and grabbing, and an association index mechanism between the sub-libraries is established;

[0033] S15, the trajectory data in each trajectory sub-library is preprocessed, the missing time points are completed by B-spline interpolation algorithm, and a high-integrity task-specific trajectory sub-library is generated.

[0034] The working principle and effect of the above technical scheme are as follows: the multi-dimensional teaching data is accurately collected by the force feedback device, the time alignment and outlier elimination processing are combined, the purity of the teaching data is greatly improved, and the subsequent analysis is avoided. Interference of noise data; a plurality of information is associated through structured coding, the correlation and availability of the data are enhanced, and the proportion of invalid data is reduced; the hierarchical clustering constructs the task sub-library and establishes the association index, the classification specification of the trajectory data is improved, and the time cost of subsequent retrieval and calling is reduced; the missing time points are completed through smoothing preprocessing, the integrity of the trajectory data is improved, and the deviation of subsequent training or control caused by incomplete data is avoided. Provide a reliable data basis for subsequent accurate control.

[0035] In an embodiment of the present application, the S14 comprises:

[0036] The task label field and corresponding trajectory data in the structured teaching trajectory library are extracted, a one-to-one mapping table of the label and the trajectory is established, and a label-associated trajectory data set is generated;

[0037] Based on the label-associated trajectory data set, the hierarchical clustering level division rule is determined, the first level is the core task type, and the second level is the subdivided work scene; and the trajectory similarity threshold is set as the clustering determination index, and a clustering parameter configuration scheme is generated;

[0038] Based on the hierarchical clustering algorithm, the label-associated trajectory data is iteratively clustered, three core task clustering clusters of assembly, polishing and grabbing are obtained by aggregation, and the data in each cluster is subdivided according to the work scene, and a multi-level task trajectory clustering result is generated;

[0039] Based on the clustering result, an assembly task trajectory sub-library, a polishing task trajectory sub-library and a grabbing task trajectory sub-library are constructed, the trajectory data of the corresponding clustering cluster is imported into the sub-library and time sequence sorting is completed, and an initial task sub-library set is generated;

[0040] The trajectory feature correlation of each task sub-library is analyzed, the index association rule is designed, the trajectory cross index table between the sub-libraries is established, the construction of the association index mechanism between the sub-libraries is completed, and the trajectory feature correlation includes the connection trajectory correlation of assembly and grabbing.

[0041] The working principle and effect of the above technical solution are that: by first establishing accurate mapping of labels and trajectories, the accuracy of data association is greatly improved, and classification deviation caused by mismatch of labels and trajectories is avoided;Through hierarchical clustering, multi-level division of tasks is realized, the logicality and standardization of trajectory classification are improved, and the filtering cost in subsequent data calling is reduced;According to the clustering results, a dedicated task sub-library is constructed and time-sequentially sorted, further improving the efficiency of data retrieval and reducing the time-consuming of invalid search;The association index between the sub-libraries is built, the linkage of different task trajectories is enhanced, the data gap in the connection of multiple tasks is avoided, and coherent and reliable trajectory data support is provided for subsequent cross-task collaborative control.

[0042] In one embodiment of the present application, the S2 comprises:

[0043] S21, extract the trajectory feature parameters in the structured teaching trajectory library, the trajectory feature parameters include trajectory curvature, force control threshold and motion rhythm, embed the trajectory feature parameters as prior knowledge into the reinforcement learning framework, initialize the DQN (deep Q network) reinforcement learning strategy, and generate an initial strategy model;

[0044] S22, set an initial control parameter range, the initial control parameter range includes a stiffness coefficient interval, a damping coefficient interval and a response speed threshold, and combine constraint conditions, the constraint conditions include robot joint physical limits and operation space boundaries, and a control parameter constraint set is constructed;

[0045] S23, randomly extract typical trajectory data from each task dedicated trajectory sub-library, simulate the scene in the actual working condition, and construct a dynamic constraint scene library, the scene includes load change and path deviation;

[0046] S24, integrate the control parameter constraint set and the dynamic constraint scene library into the reinforcement learning training process, design a constraint checking module, and real-time check each set of compliant control parameters generated in the training process;

[0047] S25, based on the initial parameters that pass the verification, generate an initial configuration scheme for reinforcement learning training.

[0048] The working principle and effect of the above technical solution are that: by embedding the teaching trajectory features as prior knowledge into the reinforcement learning, the working condition adaptability of the initial strategy model is greatly improved, and the training inefficiency caused by too large initial deviation of the model is avoided;The control parameter constraint set and the dynamic scene library are constructed, the pertinence of the training process is enhanced, and the generated parameters are prevented from exceeding the physical limits of the robot or deviating from the actual operation scene;Invalid parameters are real-time checked and removed, resource consumption of invalid training is reduced, and training efficiency is improved;A standard initial configuration scheme is generated, the training process is avoided from being disordered and chaotic, the training direction is ensured to always fit the actual operation demand, and a reliable foundation is laid for subsequent optimization training.

[0049] In one embodiment of the present application, the S3 comprises:

[0050] S31, based on the initial configuration scheme, the contact force and pose time sequence data of the robot in different working conditions are collected, the dimensionality reduction processing is performed through the principal component analysis algorithm, the high-dimensional contact force and pose joint state space is constructed, and efficient representation of state information is realized;

[0051] S32, in the joint state space, a multi-objective reward function is designed, the function includes a pose accuracy reward item, a force control compliance reward item and an environmental adaptability reward item, wherein the pose accuracy reward is associated with sub-millimeter level deviation, and the force control compliance reward is associated with millinewton level force fluctuation;

[0052] S33, the scene data in the dynamic constraint scene library is input as a training sample into a reinforcement learning framework for iterative training, and the feedback value is calculated through the reward function after each round of training to adjust the weight parameters of the strategy network;

[0053] S34, an adaptive exploration mechanism is introduced, the exploration rate is dynamically adjusted in the training process, the globality and convergence speed of parameter optimization are balanced, a preliminary compliant control parameter set is generated, and the preliminary compliant control parameter set includes dynamic stiffness parameters, variable damping parameters and response delay compensation parameters;

[0054] S35, the preliminary compliant control parameter set is tested offline, the adaptability of the parameters in the complex scene is verified through a robot dynamics simulation platform, the optimal parameter combination is selected, and the final compliant control parameter is generated, which can adapt to environmental disturbance.

[0055] The working principle and effect of the above technical solution are as follows: the joint state space is constructed through principal component analysis dimensionality reduction, which greatly improves the representation efficiency of state information and reduces the interference of redundant data on training; the multi-objective reward function considers the requirements of pose accuracy and force control compliance, enhances the pertinence of parameter optimization, and avoids control imbalance caused by optimization bias to a single target; the adaptive exploration mechanism dynamically adjusts the exploration rate, balances the globality and convergence speed of parameter optimization, reduces the training period, and avoids the problems of incomplete optimization or slow convergence; offline simulation test verifies the adaptability of the parameters and selects the optimal combination, enhances the environmental adaptation ability of the final parameters, avoids control fluctuation caused by parameter mismatch in actual operation, and provides high-quality parameter support for precise compliant control.

[0056] In one embodiment of the present application, the S34 comprises:

[0057] S341, extract the iteration progress, parameter optimization range and reward function convergence trend data of the current training stage of reinforcement learning, determine the core control parameters of the adaptive exploration mechanism, generate an exploration mechanism parameter initialization scheme, and the core control parameters include an initial exploration rate, a decay threshold and a convergence determination coefficient;

[0058] S342, based on the parameter initialization scheme, build an exploration rate dynamic adjustment model, design an exploration rate adjustment rule driven by iteration process and reward feedback, and generate an exploration rate adaptive adjustment logic;

[0059] S343, embed the adjustment logic into the reinforcement learning training process, collect the parameter optimization coverage range and convergence speed data of each round of training in real time, dynamically correct the exploration rate through the adjustment logic, balance the globality and convergence efficiency of parameter optimization, and generate an optimized exploration strategy;

[0060] S344, based on the optimized exploration strategy, drive the reinforcement learning framework to perform compliant control parameter optimization, output multiple candidate parameters, and generate a candidate compliant control parameter set, wherein the multiple candidate parameters include dynamic stiffness parameters, variable damping parameters and response delay compensation parameters;

[0061] S345, perform preliminary effectiveness verification on the candidate parameter set, eliminate invalid parameters that exceed the physical constraint range of the robot joint, retain parameter combinations that meet the basic requirements of the working condition, and generate a preliminary compliant control parameter set.

[0062] The working principle and effect of the above technical solution are as follows: by extracting key training data to determine core parameters, the accuracy of the exploration mechanism initialization is greatly improved, and the adjustment error caused by blind parameter setting is avoided; the adjustment logic constructed by the double-driven adjustment rule enhances the pertinence of the exploration rate correction and reduces the adjustment deviation; real-time data collection dynamically corrects the exploration rate, effectively balances the globality and convergence efficiency of parameter optimization, shortens the training cycle, avoids the problems of incomplete optimization or slow convergence, generates multiple candidate parameters and performs effectiveness verification, eliminates invalid parameters, retains combinations that adapt to the working condition, reduces the subsequent screening cost, avoids the interference of invalid parameters on the subsequent process, and provides protection for generating a high-quality preliminary compliant control parameter set.

[0063] In an embodiment of the application, the S343 comprises:

[0064] Analyze the parameter interaction interface protocol of the reinforcement learning training framework, determine the embedding node of the adjustment logic (before the parameter update link), generate an interface adaptation scheme and a data interaction format specification;

[0065] Based on the interface adaptation scheme, the exploration rate adaptive adjustment logic is integrated into the reinforcement learning training process, a real-time data acquisition link is built, and the data on the coverage of parameter optimization and convergence speed in each round of training are collected synchronously to generate a training and acquisition linkage module.

[0066] The linkage module obtains the optimization coverage data and convergence speed index of the current training round in real time, inputs them into the adjustment logic, calculates the exploration rate correction coefficient, and generates a dynamic adjustment instruction for the exploration rate.

[0067] Execute adjustment instructions to update reinforcement learning exploration rate parameters, simultaneously collect optimization coverage and convergence speed data for the next round of training after adjustment, and generate adjustment effect verification dataset;

[0068] The dataset is compared with a preset balance threshold, which is an optimization coverage of ≥85% and a convergence speed fluctuation of ≤10%. It is determined whether a balance between global optimization and convergence efficiency has been achieved. If not, the above steps are repeated. If the balance is achieved, the current adjustment rule is locked and an optimized exploration strategy is generated.

[0069] The working principle and effects of the above technical solution are as follows: By analyzing the interface protocol to determine the embedded nodes, the adaptability of the adjustment logic and the training framework is greatly improved, avoiding the problem of data interaction disorder during the integration process; a training and acquisition linkage module is built to enhance the real-time and synchronization of data acquisition and reduce the adjustment deviation caused by data latency; the exploration rate correction coefficient is calculated in real time to generate adjustment instructions, making the exploration rate adjustment more accurate and avoiding the optimization imbalance caused by blind adjustment; by comparing and iterating the verification dataset with the balance threshold, it is ensured that a balance between global optimization and convergence efficiency can be achieved in the end, reducing the situation of insufficient optimization; the optimized adjustment rules are locked, which improves the stability of subsequent training, avoids the interference of repeated fluctuations in the exploration strategy on the training effect, and provides reliable support for efficient parameter optimization.

[0070] In one embodiment of the present invention, step S4 includes:

[0071] S41. Load the final compliant control parameters into the robot controller, generate initial control commands, and drive the robot to perform the target task.

[0072] S42. Using a laser tracker (pose acquisition accuracy ±0.1mm) and a high-precision six-dimensional force sensor (force control acquisition accuracy ±0.1mN), the actual pose data and contact force data of the robot end are acquired in real time, generating a millisecond-level real-time status data stream;

[0073] S43. Use a timestamp synchronization algorithm to perform time-series calibration on the real-time status data stream, compare it frame by frame with the standard trajectory data and standard force data of the corresponding task in the structured teaching trajectory library, and calculate the pose deviation value and force deviation value.

[0074] S44. Use the Kalman filter algorithm to suppress noise in the deviation value, eliminate interference from measurement noise, and generate high-precision pose deviation data and force deviation data.

[0075] S45. Construct a deviation time series change curve based on the deviation data, analyze the trend of deviation change, the trend of change includes increasing / decreasing / fluctuating, and generate a deviation analysis report, the deviation analysis report includes the deviation magnitude, rate of change, and influencing factors.

[0076] The working principle and effects of the above technical solution are as follows: By loading the final compliant control parameters to drive the robot's operation, and combining high-precision sensors to achieve millisecond-level data acquisition, the accuracy and real-time performance of the actual state data acquisition are greatly improved, avoiding data lag or distortion problems; the timing calibration algorithm makes the comparison between actual data and standard data more accurate, reducing the deviation calculation error caused by timing misalignment; Kalman filtering effectively removes measurement noise, enhances the reliability of deviation data, and avoids misjudgment of deviation caused by noise interference; by constructing deviation timing curves and generating analysis reports, the trend of deviation changes and key information are clearly presented, reducing the blindness of subsequent parameter adjustments and avoiding improper adjustments due to a lack of understanding of deviation patterns, providing accurate and comprehensive decision-making basis for the dynamic optimization of subsequent compliant control parameters.

[0077] In one embodiment of the present invention, step S5 includes:

[0078] S51. Based on the deviation analysis report, establish a deviation and parameter mapping model, train the model through the BP neural network algorithm, determine the quantitative mapping relationship between pose deviation, force deviation and compliance control parameters, and determine the parameter adjustment direction and step size.

[0079] S52. Using model predictive control algorithms, the subsequent deviation development is predicted based on the deviation change trend, and the compliant control parameters are dynamically corrected in advance to achieve synergistic optimization of sub-millimeter pose accuracy and millinews force control accuracy.

[0080] S53. The corrected compliant control parameters are fused with the robot joint dynamics model to generate a high-precision compliant control instruction set, which includes position instructions, velocity instructions, and force control instructions, with an instruction update frequency of 1kHz.

[0081] S54. The control instruction set is sent to the robot's joint drivers through the EtherCAT real-time industrial bus. At the same time, a vision servo system is introduced for real-time feedback calibration to dynamically compensate for dynamic errors during the robot's execution.

[0082] S54. Continuously collect pose and force data during robot execution, fine-tune control commands through a closed-loop feedback mechanism to ensure accuracy and stability throughout the task execution process, complete high-precision compliant operation tasks, and generate a task execution accuracy report, which includes indicators such as pose accuracy error, force control accuracy error, and task completion efficiency.

[0083] The working principle and effects of the above technical solution are as follows: A mapping model is established based on the deviation analysis report to accurately determine the quantitative relationship between deviation and control parameters, significantly improving the targeting of parameter adjustments and avoiding precision imbalance caused by blind adjustments; Model predictive control corrects parameters in advance, achieving synergistic optimization of sub-millimeter pose accuracy and millinews force control accuracy, enhancing control precision; The corrected parameters are integrated with the dynamic model to generate high-frequency commands, improving the adaptability of commands to robot motion characteristics and reducing execution deviations; Real-time bus command issuance combined with visual servo calibration dynamically compensates for dynamic errors, enhancing the stability of task execution; Closed-loop feedback continuously fine-tunes commands, avoiding precision fluctuations throughout the operation process, reducing operation failures due to insufficient precision, and the generated precision report provides a reliable basis for subsequent optimization, ensuring the stable completion of high-precision compliant operation.

[0084] In one embodiment of the present invention, S52 includes:

[0085] Extract the time-series change sequences of pose deviation and force deviation, as well as the deviation growth rate data, from the deviation analysis report to generate a deviation trend feature dataset;

[0086] Based on the deviation trend feature dataset, a rolling optimization model for model predictive control is built. The prediction time domain and the control time domain are set, and a model structure configuration scheme is generated. The prediction time domain is the next 5 control cycles, and the control time domain is the current 2 control cycles.

[0087] Input the deviation trend characteristic data into the rolling optimization model, and use the recursive algorithm to predict the peak value of the pose deviation and the fluctuation range of the force deviation in the subsequent control cycle, and generate a deviation prediction result set.

[0088] With sub-millimeter pose accuracy (≤0.1mm) and millinews force control accuracy (≤5mN) as the synergistic optimization objectives, an error loss function between the deviation prediction value and the accuracy target is constructed, and the correction amount of the compliant control parameters is calculated.

[0089] The compliance control parameters are dynamically corrected in advance based on the correction amount. The parameter response data after correction is collected synchronously to verify whether the accuracy index meets the standard. The parameter correction and optimization verification results are generated to complete the collaborative optimization. The compliance control parameters include stiffness and damping parameters.

[0090] The working principle and effects of the above technical solution are as follows: By extracting the time-series changes and growth rate data of deviations to generate a feature set, the accuracy of deviation trend capture is greatly improved, avoiding the omission of key deviation change information; by building a rolling optimization model and reasonably setting the prediction and control time domains, the pertinence and reliability of deviation prediction are enhanced, and the problem of excessive prediction deviation is reduced; by using a recursive algorithm to accurately predict the peak value and fluctuation range of deviations, parameter correction has a clear direction and avoids blind adjustment; by calculating the correction amount with sub-millimeter and millinewton precision as the target, the accuracy of parameter correction is improved, and the synergistic optimization of pose and force control precision is achieved; by dynamically correcting parameters in advance and verifying compliance, the foresight of control is enhanced, the precision imbalance caused by the continuous expansion of deviation is avoided, the precision fluctuation during operation is reduced, and a solid foundation is laid for subsequent high-precision control.

[0091] In one embodiment of the present invention, S53 includes:

[0092] Extract the corrected compliant control parameters and the core parameters of the robot joint dynamics model to generate a fusion calculation input parameter set. The compliant control parameters include dynamic stiffness and variable damping parameters; the core parameters of the robot joint dynamics model include joint inertia, transmission ratio, and friction coefficient.

[0093] Based on the Lagrange dynamics equations, a fusion calculation framework for compliant control parameters and joint dynamics models is built, a parameter coupling mapping algorithm is designed, and fusion calculation logic rules are generated.

[0094] Substitute the input parameter set into the fusion computing framework, calculate the target torque and motion acceleration of each joint through numerical solution algorithms, and further derive the basic values ​​of position command and velocity command to generate preliminary command parameters.

[0095] A real-time optimization algorithm is introduced to perform time-series discretization on the initial instruction parameters. The instruction timing points are split according to the update frequency of 1kHz to ensure the real-time performance and continuity of instruction issuance and generate a time-series instruction set.

[0096] Integrate time-sequential position commands, speed commands, and force control commands, add command check codes and execution priority identifiers, generate a complete high-precision compliant control command set, and simultaneously verify the compliance of command format.

[0097] The working principle and effects of the above technical solution are as follows: By extracting and modifying the compliant control parameters and generating the input set with the core parameters of robot dynamics, the integrity and matching degree of the parameters are greatly improved, avoiding calculation deviations caused by missing or mismatched parameters; a fusion calculation framework is built based on the Lagrange equation, which enhances the adaptability of instruction calculation to robot motion characteristics and reduces the problem of inconsistency between instructions and actual dynamic states; numerical solution derives preliminary instruction parameters, which improves the accuracy of the calculation of basic instruction values ​​and avoids the accumulation of basic data errors; 1kHz time-series discretization processing improves the real-time performance and continuity of instruction issuance and reduces control lag; integrating instructions and verifying compliance enhances the reliability and integrity of the instruction set, avoids instruction execution failures caused by format errors, and provides accurate and reliable instruction support for the high-precision execution of robot movements.

[0098] One embodiment of the present invention, such as Figure 2 As shown, a system for implementing the high-precision compliant control method for a robot based on a human-machine collaborative teaching trajectory path as described above is provided, the system comprising:

[0099] Sub-library formation module: Through human-computer collaborative teaching operations, teaching data is collected and structured to generate a structured teaching trajectory library; the trajectories are classified and stored according to the task tags in the structured teaching trajectory library to form trajectory sub-libraries of different task types;

[0100] Process constraint module: Using the structured teaching trajectory library as prior knowledge, the reinforcement learning strategy is initialized and the initial control parameter range is set; at the same time, the trajectory data in the trajectory sub-library is used to construct constraint conditions to constrain the training process of the reinforcement learning strategy.

[0101] Parameter generation module: During reinforcement learning training, a joint state space of contact force and pose is constructed, and a reward function is designed based on this state space; the reinforcement learning strategy is optimized and trained according to the reward function to generate compliant control parameters;

[0102] Data acquisition module: Controls the robot using the generated compliant control parameters. During the robot's task execution, it collects the robot's actual pose information and contact force information in real time. It compares and analyzes the actual pose information and contact force information with the corresponding data in the structured teaching trajectory library to obtain pose deviation data and force deviation data.

[0103] Precision control module: Based on the posture deviation data and force deviation data, dynamically adjusts the compliance control parameters, generates high-precision compliance control commands for the robot based on the adjusted compliance control parameters, and performs precise control of the robot according to the high-precision compliance control commands.

[0104] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A high-precision compliant control method for a robot based on a human-machine collaborative teaching trajectory path, characterized in that, The method includes: S1. Through human-computer collaborative teaching operations, teaching data is collected and structured to generate a structured teaching trajectory library; the trajectories are classified and stored according to the task tags in the structured teaching trajectory library to form trajectory sub-libraries of different task types. S2. Using the structured teaching trajectory library as prior knowledge, the reinforcement learning strategy is initialized, and an initial range of control parameters is set. At the same time, the trajectory data in the trajectory sub-library is used to construct constraints to constrain the training process of the reinforcement learning strategy. S3. During the reinforcement learning training process, a joint state space of contact force and pose is constructed, and a reward function is designed based on this state space; the reinforcement learning strategy is optimized and trained according to the reward function to generate compliant control parameters. S4. Control the robot using the generated compliant control parameters. During the robot's task execution, collect the robot's actual pose information and contact force information in real time. Compare and analyze the actual pose information and contact force information with the corresponding data in the structured teaching trajectory library to obtain pose deviation data and force deviation data. S5. Based on the posture deviation data and force deviation data, dynamically adjust the compliance control parameters, generate high-precision compliance control commands for the robot based on the adjusted compliance control parameters, and perform precise control of the robot according to the high-precision compliance control commands.

2. The high-precision compliant robot control method based on human-machine collaborative teaching trajectory path according to claim 1, characterized in that, S1 includes: S11. The teaching action is executed through the force feedback human-machine collaborative operation device, and the three-dimensional pose coordinates, six-dimensional contact force / torque signals and task type labels of the robot end are collected in real time to generate a multi-dimensional original teaching dataset. S12. Based on the original teaching dataset, a time-series alignment algorithm is used to synchronously calibrate the discrete pose and force signals, and an outlier detection algorithm is used to remove impact noise data to generate clean teaching data. S13. Perform structured coding on the clean teaching data, associate and map the pose information, force information and task labels to generate a structured teaching trajectory library. S14. Based on the task labels in the structured teaching trajectory library, a hierarchical clustering algorithm is used to classify the trajectory data, construct trajectory sub-libraries for different task types, and establish an association index mechanism between sub-libraries. S15. Perform smoothing preprocessing on the trajectory data in each trajectory sub-library, and use the B-spline interpolation algorithm to fill in the missing time series points to generate a highly complete task-specific trajectory sub-library.

3. The high-precision compliant robot control method based on human-machine collaborative teaching trajectory path according to claim 1, characterized in that, The S2 includes: S21. Extract trajectory feature parameters from the structured teaching trajectory library, embed them as prior knowledge into the reinforcement learning framework, initialize the DQN reinforcement learning strategy, and generate an initial policy model. S22. Set the initial control parameter range and construct a set of control parameter constraints in conjunction with the constraints. S23. Randomly extract typical trajectory data from the trajectory sub-library of each task, simulate the scenario in the actual working condition, and build a dynamic constraint scenario library; S24. Integrate the set of control parameter constraints and the dynamic constraint scenario library into the reinforcement learning training process, and design a constraint verification module to verify each set of compliant control parameters generated during the training process in real time. S25. Based on the validated initial parameters, generate an initial configuration scheme for reinforcement learning training.

4. The high-precision compliant robot control method based on human-machine collaborative teaching trajectory path according to claim 1, characterized in that, The S3 includes: S31. Based on the initial configuration scheme, collect the contact force and pose time sequence data of the robot under different working conditions, and construct a high-dimensional joint state space of contact force and pose by dimensionality reduction processing through principal component analysis algorithm. S32. Design a multi-objective reward function in the joint state space; S33. Use the scene data in the dynamic constraint scene library as training samples, input them into the reinforcement learning framework for iterative training, calculate the feedback value through the reward function after each round of training, and adjust the weight parameters of the policy network. S34. An adaptive exploration mechanism is introduced to dynamically adjust the exploration rate during training, balancing the globality of parameter optimization with the convergence speed, and generating a preliminary set of compliant control parameters. S35. Conduct offline simulation tests on the preliminary compliant control parameter set, verify the adaptability of the parameters in complex scenarios through a robot dynamics simulation platform, select the optimal parameter combination, and generate the final compliant control parameters that can adapt to environmental disturbances.

5. The high-precision compliant robot control method based on human-machine collaborative teaching trajectory path according to claim 4, characterized in that, S34 includes: S341. Extract the iteration progress, parameter optimization range, and reward function convergence trend data of the current training stage of reinforcement learning, determine the core control parameters of the adaptive exploration mechanism, and generate an exploration mechanism parameter initialization scheme. S342. Based on the parameter initialization scheme, build a dynamic adjustment model for the exploration rate, design exploration rate adjustment rules driven by both iterative process and reward feedback, and generate adaptive adjustment logic for the exploration rate. S343. Embed the adjustment logic into the reinforcement learning training process, collect the parameter optimization coverage and convergence speed data of each training round in real time, dynamically adjust the exploration rate through the adjustment logic, balance the globality of parameter optimization and convergence efficiency, and generate an optimized exploration strategy. S344. Based on the optimized exploration strategy, drive the reinforcement learning framework to optimize the compliant control parameters, output multiple sets of candidate parameters, and generate a candidate compliant control parameter set. S345. Perform preliminary validity checks on the candidate parameter set to generate a preliminary compliant control parameter set.

6. The high-precision compliant robot control method based on human-machine collaborative teaching trajectory path according to claim 5, characterized in that, S343 includes: Analyze the parameter interaction interface protocol of the reinforcement learning training framework, determine the embedding nodes of the adjustment logic, and generate interface adaptation schemes and data interaction format specifications. Based on the interface adaptation scheme, the exploration rate adaptive adjustment logic is integrated into the reinforcement learning training process, a real-time data acquisition link is built, and the data on the coverage of parameter optimization and convergence speed in each round of training are collected synchronously to generate a training and acquisition linkage module. The linkage module obtains the optimization coverage data and convergence speed index of the current training round in real time, inputs them into the adjustment logic, calculates the exploration rate correction coefficient, and generates a dynamic adjustment instruction for the exploration rate. Execute adjustment instructions to update reinforcement learning exploration rate parameters, simultaneously collect optimization coverage and convergence speed data for the next round of training after adjustment, and generate adjustment effect verification dataset; Compare the validation dataset with the preset balance threshold to determine whether a balance between global optimization and convergence efficiency has been achieved. If not, repeat the above steps. If so, lock the current adjustment rule and generate an optimized exploration strategy.

7. The high-precision compliant robot control method based on human-machine collaborative teaching trajectory path according to claim 1, characterized in that, The S4 includes: S41. Load the final compliant control parameters into the robot controller, generate initial control commands, and drive the robot to perform the target task. S42. Using a laser tracker and a high-precision six-dimensional force sensor, the actual pose data and contact force data of the robot end are collected in real time, generating a millisecond-level real-time status data stream. S43. Use a timestamp synchronization algorithm to perform time-series calibration on the real-time status data stream, compare it frame by frame with the standard trajectory data and standard force data of the corresponding task in the structured teaching trajectory library, and calculate the pose deviation value and force deviation value. S44. Use the Kalman filter algorithm to suppress noise in the deviation value, eliminate interference from measurement noise, and generate high-precision pose deviation data and force deviation data. S45. Construct a deviation time series change curve based on the deviation data, analyze the trend of deviation change, and generate a deviation analysis report.

8. The high-precision compliant robot control method based on human-machine collaborative teaching trajectory path according to claim 1, characterized in that, The S5 includes: S51. Based on the deviation analysis report, establish a deviation and parameter mapping model, train the model through the BP neural network algorithm, determine the quantitative mapping relationship between pose deviation, force deviation and compliance control parameters, and determine the parameter adjustment direction and step size. S52. Using model predictive control algorithms, predict the subsequent development of deviations based on the trend of deviation changes, and dynamically correct the compliance control parameters in advance. S53. The corrected compliant control parameters are fused with the robot joint dynamics model to generate a high-precision compliant control instruction set. S54. The control instruction set is sent to the robot's joint drivers through the EtherCAT real-time industrial bus. At the same time, a vision servo system is introduced for real-time feedback calibration to dynamically compensate for dynamic errors during the robot's execution. S54. Continuously collect pose and force data during robot execution, fine-tune control commands through a closed-loop feedback mechanism, and generate a task execution accuracy report.

9. The high-precision compliant robot control method based on human-machine collaborative teaching trajectory path according to claim 8, characterized in that, S52 includes: Extract the time-series change sequences of pose deviation and force deviation, as well as the deviation growth rate data, from the deviation analysis report to generate a deviation trend feature dataset; Based on the deviation trend feature dataset, a rolling optimization model for model predictive control is built, the prediction time domain and control time domain are set, and a model structure configuration scheme is generated. Input the deviation trend characteristic data into the rolling optimization model, and use the recursive algorithm to predict the peak value of the pose deviation and the fluctuation range of the force deviation in the subsequent control cycle, and generate a deviation prediction result set. With sub-millimeter pose accuracy and millinews force control accuracy as the synergistic optimization objectives, an error loss function is constructed between the predicted deviation value and the accuracy target, and the correction amount of the compliant control parameters is calculated. Based on the correction amount, the compliance control parameters are dynamically corrected in advance, the parameter response data after correction is collected synchronously, the accuracy index is verified, the parameter correction and optimization verification results are generated, and the collaborative optimization is completed.

10. A system for implementing the high-precision compliant robot control method based on human-machine collaborative teaching trajectory path as described in claim 1, characterized in that, The system includes: Sub-library formation module: Through human-computer collaborative teaching operations, teaching data is collected and structured to generate a structured teaching trajectory library; the trajectories are classified and stored according to the task tags in the structured teaching trajectory library to form trajectory sub-libraries of different task types; Process constraint module: Using the structured teaching trajectory library as prior knowledge, the reinforcement learning strategy is initialized and the initial control parameter range is set; at the same time, the trajectory data in the trajectory sub-library is used to construct constraint conditions to constrain the training process of the reinforcement learning strategy. Parameter generation module: During reinforcement learning training, a joint state space of contact force and pose is constructed, and a reward function is designed based on this state space; the reinforcement learning strategy is optimized and trained according to the reward function to generate compliant control parameters; Data acquisition module: Controls the robot using the generated compliant control parameters. During the robot's task execution, it collects the robot's actual pose information and contact force information in real time. It compares and analyzes the actual pose information and contact force information with the corresponding data in the structured teaching trajectory library to obtain pose deviation data and force deviation data. Precision control module: Based on the posture deviation data and force deviation data, dynamically adjusts the compliance control parameters, generates high-precision compliance control commands for the robot based on the adjusted compliance control parameters, and performs precise control of the robot according to the high-precision compliance control commands.